Module Module 4. Parallelism via Multiprocessing
Objective Build a parallel program using multiprocessing.Pool to accelerate a CPU-bound task and compare it with sequential execution.
- Python 3 installed.
- Git installed and basic familiarity with
clone,add,commit,push. - Concepts from Module 4:
- Why multiprocessing can help CPU-bound work.
- Process creation (
multiprocessing.Process). - Process pools (
multiprocessing.Pool) andmap. - Pool lifecycle and the
withcontext manager. - Pickling and inter-process communication.
- The need for an
if __name__ == "__main__":guard in multiprocessing programs.
README.mdthis filelab3_parallel_map.pystarter with function skeletons and amainblockanalysis.mdwhere you record timings and answer the analysis questions
General instructions
- Clone your GitHub Classroom repository.
- Modify
lab3_parallel_map.pyto complete the tasks. - Keep the supplied benchmark structure so sequential and parallel runs process the same data.
- Record results in
analysis.md. - Commit frequently and push before the deadline.
- Open
lab3_parallel_map.py. - Complete
cpu_intensive_task(n)with a deterministic CPU-bound calculation. - A suitable implementation is
math.factorial(n). - Keep the worker function at module scope. Functions passed to a process pool must be serializable by the multiprocessing machinery.
Do not add sleeps or I/O to this function. The purpose of this lab is to measure CPU-bound parallelism.
In run_sequential(data):
- Process every value in
datawithcpu_intensive_task. - Preserve the input order.
- Return the complete list of results.
The sequential run and the pool run must perform the same work.
In run_parallel_map(data, pool_size):
- Create a
multiprocessing.Poolwithpool_sizeworkers using awithblock. - Apply
pool.map(cpu_intensive_task, data). - Return the complete result list.
Do not create one process manually for every input value. This task is specifically about process pools.
- Run
python lab3_parallel_map.py. - The script reports:
- Python version;
- multiprocessing start method;
- number of tasks;
- pool size;
- sequential time;
- parallel time;
- speedup.
- The script also verifies that the sequential and parallel result lists are identical.
- If your machine is unusually slow or fast, you may adjust
VALUES_UPPER_BOUNDmodestly. KeepNUMBER_LIST_SIZElarge enough to provide multiple tasks to the pool.
The supplied defaults are intended to make process-pool overhead small enough for the experiment to be meaningful while keeping runtime practical on typical student hardware. A speedup is not guaranteed on every machine.
Answer the following:
- Speedup Compute
Sequential / Parallel. Interpret the result, including the possibility of no speedup. - Why processes can help Explain how separate Python processes can execute CPU-bound work concurrently and how this differs from normal GIL-enabled threading.
Pool.mapExplain whyPool.mapis convenient for data-parallel workloads and what ordering guarantee it provides.- Overheads Identify process startup, task scheduling, serialization/pickling, inter-process data transfer, result collection, and load imbalance.
- Pool size Explain why creating more worker processes than useful CPU resources or tasks can reduce performance.
- Correctness Explain why timing alone is insufficient and why the sequential and parallel outputs must be checked for equality.
- Ensure
lab3_parallel_map.pyruns successfully and the result verification passes. - Ensure
analysis.mdincludes your recorded environment, timings, speedup, and answers. - Stage:
git add lab3_parallel_map.py analysis.md(orgit add .) - Commit:
git commit -m "Complete Lab 3 Parallel Map" - Push:
git push origin main(or your default branch) - Verify on GitHub that
lab3_parallel_map.pyandanalysis.mdare updated.