Futures & Executors Proficient¶
When you'd use this
Offload work to thread and process pools with concurrent.futures.
Submit work to thread/process pools and collect results via futures — a simple API for parallelism without managing threads directly.
What you'll learn¶
- What a future represents
- Thread vs process pools
-
mapandsubmit/as_completed(tested) - Choosing pool type by workload
- Handling results and errors
concurrent.futures is the standard-library's high-level interface for running work concurrently — you hand it functions, it runs them in a pool of threads or processes, and gives you futures (placeholders for results). It's the easiest way to parallelize. Examples here are run-verified.
A future is a promise of a result¶
A future is a promise of a result — a key concept in Futures & Executors.
A future is an object representing a computation that may not be done yet. You get it immediately when you submit work; later you call .result() to get the value (blocking until ready).
Two executors manage the pool: - ThreadPoolExecutor — pool of threads. Best for I/O-bound work (network, disk) where the GIL is released while waiting. - ProcessPoolExecutor — pool of processes. Best for CPU-bound work, sidestepping the GIL for true parallelism (see Threading on the GIL).
map — parallel over an iterable (tested)¶
map — parallel over an iterable, part of Futures & Executors.
The simplest pattern: apply a function to every item, in parallel:
from concurrent.futures import ThreadPoolExecutor
def square(x):
return x * x
with ThreadPoolExecutor(max_workers=4) as ex:
results = list(ex.map(square, range(6)))
print(results)
Output:
ex.map runs square on each input across the pool and returns results in input order. The with block cleanly shuts the pool down when done. Swap ThreadPoolExecutor for ProcessPoolExecutor and CPU-bound work runs on multiple cores.
submit + as_completed — results as they finish (tested)¶
submit + as_completed — results as they finish, part of Futures & Executors.
When you want each result the moment it's ready (not in submission order):
from concurrent.futures import ThreadPoolExecutor, as_completed
def square(x):
return x * x
with ThreadPoolExecutor(max_workers=3) as ex:
futures = [ex.submit(square, n) for n in [2, 3, 4]]
results = sorted(f.result() for f in as_completed(futures))
print(results)
Output:
submit returns a future immediately; as_completed yields each future as it finishes. This is ideal when tasks take varying times and you want to process fast ones without waiting for slow ones. (.result() also re-raises any exception the task hit — so errors surface when you collect results.)
Choosing the pool¶
ThreadPool for I/O-bound, ProcessPool for CPU-bound.
| Workload | Executor | Why |
|---|---|---|
| I/O-bound (HTTP, files, DB) | ThreadPoolExecutor | Threads wait efficiently; GIL released during I/O |
| CPU-bound (math, parsing, compression) | ProcessPoolExecutor | True parallelism across cores, bypassing the GIL |
Getting this wrong is the classic mistake: threads won't speed up CPU-bound work (the GIL serializes it), and processes add overhead for I/O-bound work.
Start here for parallelism
concurrent.futures is usually the right first tool — higher-level and safer than managing raw threads/processes. Reach for lower-level threading/multiprocessing only when you need fine control. For async I/O at large scale, see Asyncio.
Error handling¶
Exceptions surface when you read a future's result — handle them there.
Exceptions in a task don't crash the pool — they're stored in the future and raised when you call .result():
from concurrent.futures import ThreadPoolExecutor
def risky(x):
if x == 0:
raise ValueError("cannot process zero")
return 10 / x
with ThreadPoolExecutor() as ex:
fut = ex.submit(risky, 0)
try:
fut.result() # re-raises the ValueError here
except ValueError as e:
print("caught:", e) # caught: cannot process zero
Always retrieve .result() (or check .exception()) so failures aren't silently swallowed.
Practice exercises¶
- Use
ThreadPoolExecutorto fetch several URLs "concurrently" (simulate withtime.sleep) and time it vs sequential. - Switch a CPU-bound task (e.g. summing squares to a big N) from thread to process pool and compare timing.
- Use
as_completedwith tasks of different durations and print results in completion order. - Add a per-task timeout with
future.result(timeout=...)and handleTimeoutError. - Explain why a
ThreadPoolExecutorwon't speed up a pure-Python CPU-bound loop.
💬 Discussion
Have a question about this topic? Found an error? Share your thoughts below.