
In Part 53, you made your code clear and type-safe. Now we tackle a different kind of slowness — the kind that comes from waiting on the network or disk, and the kind that comes from using only one CPU core while others sit idle. This one part covers all three concurrency tools: async, threads, and processes.
Hands-on lab:
cpu-demo.py— see the GIL in Activity Monitor with your own eyes.
Read order (do it exactly like this, once):
Part-54/ and run python cpu-demo.py 1. Open Activity Monitor (Mac) / Task Manager (Windows). Look at your CPU cores. Watch just 1 of them max out. Leave that terminal open — it plants the question this whole part answers.python cpu-demo.py 3. Watch every core light up. You will now be able to explain exactly why, in one paragraph.The concurrency-visual-guide.html file has the full diagrams. Use it for the classroom walkthrough; use this file for reading later.
The problem: Sequential code either (A) sits idle while waiting on network/disk, or (B) uses one CPU core while the rest do nothing.
Concurrency = many tasks in progress (structure). Parallelism = many tasks running at the same instant (execution). Concurrency does not require parallelism — async on one thread is concurrent but not parallel.
Python's three tools:
| Tool | Solves | Concurrent? | Parallel on multiple cores? |
|---|---|---|---|
asyncio | IO-bound waiting | Yes | No — one thread, smarter switching |
threading | IO-bound with blocking libs (requests) | Yes | No for pure Python; yes for C extensions that release the GIL |
multiprocessing | CPU-bound compute | Yes | Yes — each process has its own GIL |
The one-line summary: "asyncio and threading help Python deal with many things at once. Only multiprocessing lets Python actually do many things at once."
GIL: Only one thread runs Python bytecode at a time. Threads help IO, not CPU math. Processes bypass this. Python 3.14 (free-threaded build) makes it optional, but the standard build still has it.
Do not confuse these two. Getting this straight saves a lot of interview grief.
threading module was added by Guido in 1992 — 13 years before consumer multi-core existed. So no, concurrency was not born in 2005.| Era | Consumer chip clock | Cores per chip | What concurrency was used for |
|---|---|---|---|
| 1960s–70s | kHz → few MHz | 1 | Time-sharing on mainframes (many users, one CPU) |
| 1980s–90s | MHz → hundreds of MHz | 1 | Preemptive multitasking on desktop OSes (many programs on one core, IO overlap) |
| 1993–2005 | ~200 MHz → ~3.8 GHz | 1 (Intel HT gave 1 core → 2 threads from 2002) | Same, plus GUIs staying responsive while background work happened |
| 2005 → today | stuck at ~3–5 GHz | 2 → 4 → 8 → 18+ | Parallel compute in application code — real CPU speedups from using many cores |
From 1971 (Intel 4004, 740 kHz clock, ~92,000 instructions/sec) to 2005 (Pentium 4 Prescott, ~3.8 GHz), consumer clock speeds rose ~5,000×. Existing code ran faster every year — programmers got a free lunch. In 2005, Intel's Pentium 4 hit the wall — physics stopped clock speeds from rising (heat, power, and the speed of light itself). Chipmakers stopped racing on frequency and started adding cores.
So what changed for us as Python developers?
multiprocessing. Python's threading still doesn't give it to you because of the GIL. That's exactly why multiprocessing was added in Python 2.6 (2008) — to let application code finally benefit from multi-core hardware.Your laptop today has 10–18 cores. Your Python script uses one. Threads take turns because of the GIL. Only processes truly use all your cores. That's the problem this whole part exists to solve.
Chip design and manufacturing split into two industries:
| Role | Companies (2026) |
|---|---|
| Designers (draw the blueprint) | Apple, AMD, Intel, Qualcomm, NVIDIA |
| Foundries (actually etch the silicon) | TSMC (Taiwan) ~70% share, Samsung, Intel Foundry |
| IP licensors (own the CPU instruction set) | ARM Holdings (UK), Intel (x86), RISC-V (open) |
| Toolmakers (the machines that make chips) | ASML (Netherlands, sole EUV lithography source) |
When you land on python.org/downloads, the site asks you to pick a build. That choice IS the ARM vs x86 story from above.
| Your machine | Chip | ISA | What to download |
|---|---|---|---|
| This MacBook (M4 Pro) or any Mac since 2020 | Apple M1/M2/M3/M4 | ARM64 | macOS 64-bit universal2 installer — one .pkg, both arm64 and x86_64 inside, macOS picks automatically |
| Intel Mac (2006–2020) | Intel Core i5/i7/i9 | x86-64 | Same universal2 installer works |
| Windows Intel/AMD PC | Intel Core / AMD Ryzen | x86-64 | Windows installer (64-bit) |
| Windows on ARM — Copilot+ PC | Snapdragon X / X2 Elite | ARM64 | Windows installer (ARM64) — different file from the 64-bit x86 one |
| Legacy 32-bit Windows | Old Intel/AMD | x86 (32-bit) | Windows installer (32-bit) — only for very old machines |
Verify what you're actually running:
python3 -c "import platform, sys; print('Machine:', platform.machine(), '| Python:', sys.version.split()[0])"
On this M4 Pro laptop it prints Machine: arm64. On Intel/AMD it prints x86_64. On Snapdragon it prints ARM64.
The gotcha: if you accidentally install the x86_64 Python on an Apple Silicon Mac, it will still run — through Rosetta 2, Apple's on-the-fly Intel-to-ARM translator — but ~20–40% slower, and every C extension you pip install will be x86 too. Always use the universal2 installer on Mac and the ARM64 installer on Snapdragon.
Why the choice exists at all: a compiled program contains native machine code for one specific instruction set. ARM64 code cannot run directly on x86, and vice versa. Every .so/.pyd file in your Python install is native — no magic. That is exactly why ARM Holdings collects a royalty every time an ARM binary runs.
An SoC (System-on-Chip) puts CPU cores + GPU + Neural Engine + memory controller + I/O all on one silicon die. Apple's M-series is the poster child:
When you run python cpu-demo.py 1, exactly one P-core on that giant SoC lights up. The other P-core, all E-cores, the whole GPU, and the whole NPU sit idle. That's the waste this course teaches you to fix.
Every "14 cores / 20 GPU / 48 GB" reference below is the real laptop this course is filmed on. Verified from system_profiler SPHardwareDataType:
Model Name: MacBook Pro
Model Identifier: Mac16,7 (16-inch, late 2024)
Chip: Apple M4 Pro
Total Number of Cores: 14 (10 Performance and 4 Efficiency)
GPU: 20-core (Metal 4)
Memory: 48 GB unified (LPDDR5X, 273 GB/s)
Neural Engine: 16 cores, 38 TOPS
OS: macOS 26.6.2
Full spec sheet:
| Spec | This MacBook (M4 Pro, upgraded tier) |
|---|---|
| Model | MacBook Pro 16" (Mac16,7) |
| CPU cores | 14 (10 P-cores + 4 E-cores) |
| CPU threads | 14 (Apple has no Hyper-Threading — 1 thread per core) |
| GPU cores | 20 |
| Neural Engine | 16 cores, 38 TOPS |
| P-core max clock | 4.51 GHz |
| E-core max clock | 2.6 GHz |
| Unified memory | 48 GB (LPDDR5X) |
| Memory bandwidth | 273 GB/s |
| Transistors | ~55 billion |
| Die size | 325 mm² |
| Process | TSMC N3E (3 nm, 2nd-gen) |
| Released | Oct 30, 2024 |
Where the M4 Pro sits in Apple's lineup (Oct 2024 → today):
If your machine is different, the math below scales — just substitute your core count wherever "14" appears.
None of this course requires you to understand chip manufacturing. But if a student asks, here is the honest answer with links to the definitive sources.
The silicon in your MacBook did not come from "beach sand" — that's a common oversimplification. It started as a specific rock and went through four industrial stages to reach the purity a chip demands.
| # | Stage | Material | Purity | How it's done |
|---|---|---|---|---|
| 1 | Mining | Quartzite rock (80–99% SiO₂) | ~99% | Quarried from specific mines: Spruce Pine (USA), Kyshtym (Russia), Norway, Brazil, China |
| 2 | Arc-furnace reduction | Metallurgical-grade silicon (MGS) | ~98–99% | Quartzite + carbon (coke) heated to 1,800–2,000 °C in an arc furnace: SiO₂ + 2 C → Si + 2 CO |
| 3 | Siemens process | Electronic-grade polysilicon rods | 99.9999999% (9N–11N) | MGS + HCl → trichlorosilane (SiHCl₃) liquid → fractionally distilled → CVD-deposited on hot silicon rods at 1,100 °C for 200–300 hours |
| 4 | Czochralski method | Monocrystalline silicon ingot (boule) | Same purity, now a single crystal | Polysilicon melted at 1,420 °C in a quartz crucible → seed crystal dipped in → rotated and pulled up at 1–2 cm/hour → 300 mm × 2 m ingot, 200–300 kg, hangs from a "necked" thread just a few mm wide |
| 5 | Wire sawing + CMP | 300 mm silicon wafer, ~0.775 mm thick | Same | Ingot sliced by diamond wire saw, polished by Chemical Mechanical Polishing to atomic flatness |
Key facts worth memorising:
Refinement pipeline:
Transistor physics:
Videos:
Industry:
Get this straight before any code. The word "thread" is used in three different layers and mixing them up is the #1 source of confusion.
| Layer | What it is | Example on this course's laptop (14-core Apple M4 Pro) |
|---|---|---|
| Physical core | A real block of silicon that executes instructions | 14 (10 P-cores + 4 E-cores) |
| Hardware thread / logical processor | An execution slot on a core (Intel/AMD Hyper-Threading gives 2 per P-core; Apple gives 1 per core) | 14 (Apple has no HT). A 14-core Intel i9 with HT on P-cores would report 24+ |
| OS thread | A software execution path scheduled by the OS onto hardware threads | Thousands at once (macOS/Windows/Linux time-slice them) |
Python thread (threading.Thread) | A real OS thread that runs Python bytecode | Same as OS thread — but only 1 holds the GIL at a time |
Intel calls it Hyper-Threading (HT); AMD calls it SMT (Simultaneous Multi-Threading). Same idea: a single physical core keeps two thread contexts warm and switches between them when one stalls waiting for memory. Result: ~20–30% more throughput per core, not 2×.
That's why sysctl -n hw.ncpu on this 14-core M4 Pro reports 14, but on a 14-core Intel Core i9 (with HT on P-cores) would report 24 or higher.
Python thread (threading.Thread)
└── OS thread (kernel-scheduled)
└── Hardware thread (a slot on a physical core)
└── Physical core (silicon that executes bytecode)
Every Python thread is a real OS thread — that's not the problem. The problem is the GIL at the top: only one Python thread may execute bytecode at a time, no matter how many cores you have. We get to that in Part VIII.
Before writing concurrent code, diagnose which kind of slow you have. Wrong diagnosis = wrong tool = wasted effort.
| Symptom | Kind | Example | Right tool |
|---|---|---|---|
| One core pinned at 100%, others idle | CPU-bound | Number crunching, image resizing, ML training | multiprocessing |
| All cores idle, program still slow | IO-bound | API calls, DB queries, file reads | asyncio or threading |
Rule of thumb: look at Activity Monitor before you code. If your python process shows ~100% CPU, you're CPU-bound. If it shows ~2% CPU while the program still waits, you're IO-bound.
Rob Pike (co-creator of Go) gave the definitive talk in 2012 — Concurrency Is Not Parallelism. His one-sentence version:
Pike's gopher analogy: one gopher moving books from one pile to another is sequential. Two gophers carrying two books at once is parallelism. One gopher who cleverly interleaves loading, walking, and unloading so books flow smoothly is concurrency — even alone.
Concurrency enables parallelism, but they are not the same. asyncio is 100% concurrent, 0% parallel — proof that the two ideas are separate.
Mapping to Python:
| Tool | Concurrency? | Parallelism? |
|---|---|---|
asyncio | Yes (one thread, cooperative) | No |
threading | Yes | No for pure Python (GIL); yes when C code releases the GIL |
multiprocessing | Yes | Yes — separate processes, separate GILs |
This is the concept behind how the switching between tasks actually happens. It's the difference between threading and asyncio.
| Style | Who decides when to switch? | Used by |
|---|---|---|
| Preemptive | The OS forcibly pauses a thread mid-instruction and hands the core to another thread | threading, multiprocessing, every OS-level scheduler |
| Cooperative | Tasks voluntarily yield control at specific points; nothing else can run until they do | asyncio (yields at every await) |
Consequences:
counter += 1. That's why threads need Locks to guard shared data (races are real).await. Between two awaits the code is atomic for other coroutines in the same loop. Fewer races, but one runaway coroutine that never awaits will freeze the whole event loop.Restaurant analogy:
await.asyncioimport time
import requests
def fetch_user(user_id: int) -> dict:
return requests.get(
f"https://jsonplaceholder.typicode.com/users/{user_id}",
timeout=10,
).json()
start = time.time()
fetch_user(1)
fetch_user(2)
fetch_user(3)
print(f"Total: {time.time() - start:.2f}s") # ~0.9s — all sequential
Three requests, ~0.9s total — but the CPU was idle 99% of that time, waiting on the network. What if you could start all three at once and wait together? That's async.
asyncio runsA coroutine is a function that can pause mid-execution and hand control back to the event loop, then resume later exactly where it left off. It is not a thread — it's a plain Python object that lives in one thread. Thousands of coroutines can share one thread.
async and await are language keywords — async def marks a function as a coroutine; await is where it can pause.asyncio is the standard library that provides the event loop, primitives (gather, Lock, Queue, TaskGroup), and network integration.async def creates. Calling it doesn't run it — it returns a coroutine object that the event loop must schedule.import asyncio
async def greet(name: str) -> str:
print(f"Starting {name}")
await asyncio.sleep(1) # pauses THIS coroutine, lets others run
return f"Hello, {name}!"
async def main():
print(await greet("Alice"))
asyncio.run(main())
async def defines a coroutine; await pauses it while waiting and lets other coroutines run.asyncio.run(main()) creates the loop, runs main(), and closes the loop. One thread, one event loop. The loop rotates between coroutines whenever one hits an await. That's cooperative multitasking — no OS preemption, no GIL contention between coroutines, no thread overhead.
time.sleep(2) # BLOCKS the whole event loop — nothing else runs
await asyncio.sleep(2) # YIELDS control — other coroutines run meanwhile
Using time.sleep() (or any blocking call) inside async code defeats the purpose.
asyncio.gatherimport asyncio
import time
async def fetch(source: str, delay: float) -> str:
await asyncio.sleep(delay)
return f"Data from {source}"
async def main():
start = time.time()
results = await asyncio.gather(
fetch("API-1", 2),
fetch("API-2", 1),
fetch("API-3", 1.5),
)
print(results, f"{time.time() - start:.2f}s") # ~2s, not 4.5s
asyncio.run(main())
asyncio.gather() starts all coroutines at once; the slowest one determines total time. The work didn't get faster — the waiting got smarter.
asyncio.TaskGroup — the modern gather (Python 3.11+)TaskGroup is what you should reach for in new code. It gives you structured concurrency — if one task fails, the group cleanly cancels the others.
import asyncio
async def fetch(url: str) -> str:
await asyncio.sleep(0.1)
return f"body of {url}"
async def main():
async with asyncio.TaskGroup() as tg:
t1 = tg.create_task(fetch("https://a.example"))
t2 = tg.create_task(fetch("https://b.example"))
t3 = tg.create_task(fetch("https://c.example"))
print(t1.result(), t2.result(), t3.result())
asyncio.run(main())
Prefer TaskGroup over gather when you want all-or-nothing behavior.
asyncio.as_completed — stream results as they arriveimport asyncio
async def fetch(url: str) -> str:
await asyncio.sleep(0.1)
return f"body of {url}"
async def main():
urls = ["https://a.example", "https://b.example", "https://c.example"]
tasks = [fetch(u) for u in urls]
for coro in asyncio.as_completed(tasks):
result = await coro
print("first back:", result) # in whatever order they finish
asyncio.run(main())
Great when you want to react as soon as any result is ready (e.g. return the fastest of several redundant API calls).
asyncio.timeout — Python 3.11+ (the good timeout)import asyncio
async def slow_operation():
await asyncio.sleep(10)
async def main():
try:
async with asyncio.timeout(5):
await slow_operation()
except TimeoutError:
print("Timed out after 5s")
asyncio.run(main())
Older asyncio.wait_for(coro, timeout=5) still works, but asyncio.timeout() composes better and is preferred for new code.
asyncio.Semaphore — rate limitingLimit how many things happen concurrently (e.g. don't hammer an API with 1000 requests at once).
import asyncio
sem = asyncio.Semaphore(10) # at most 10 concurrent
async def fetch(url):
async with sem:
await asyncio.sleep(0.1)
return f"body of {url}"
asyncio.to_thread — bridge to blocking librariesYou want to call a blocking library (requests, DB drivers, disk IO) inside async code without blocking the event loop. Use to_thread:
import asyncio
import requests
def blocking_get(url):
return requests.get(url, timeout=10).text
async def main():
text = await asyncio.to_thread(blocking_get, "https://example.com")
print(text[:100])
asyncio.run(main())
Under the hood, to_thread runs the function in a ThreadPoolExecutor and awaits its result. Uses ~1 real OS thread per call.
loop.run_in_executor — CPU work inside async codeSame idea as to_thread, but you supply your own pool — and it can be a ProcessPoolExecutor for CPU-bound work:
import asyncio
from concurrent.futures import ProcessPoolExecutor
def heavy_compute(n):
return sum(i * i for i in range(n))
async def main():
loop = asyncio.get_running_loop()
with ProcessPoolExecutor() as pool:
result = await loop.run_in_executor(pool, heavy_compute, 10_000_000)
print(result)
if __name__ == "__main__":
asyncio.run(main())
httpximport asyncio
import httpx
async def fetch_user(client: httpx.AsyncClient, uid: int) -> dict:
r = await client.get(f"https://jsonplaceholder.typicode.com/users/{uid}")
return r.json()
async def main():
async with httpx.AsyncClient(timeout=10) as client:
users = await asyncio.gather(
*(fetch_user(client, i) for i in range(1, 6))
)
for u in users:
print(u["name"])
asyncio.run(main()) # pip install httpx
httpx.AsyncClient is the async version of requests.Session; async with closes it automatically. All 5 requests run concurrently.
The event loop runs in one thread. Anything that blocks that thread stalls every coroutine.
time.sleep(), requests.get(), DB queries, or CPU-heavy loops directly in async code.await asyncio.sleep, httpx.AsyncClient, async DB drivers).asyncio.to_thread(...).threadingimport threading
import time
def download(url: str) -> None:
time.sleep(1) # simulate network delay
print(f"got {url}")
start = time.time()
threads = [
threading.Thread(target=download, args=(f"url/{i}",))
for i in range(5)
]
for t in threads: t.start()
for t in threads: t.join()
print(f"Total: {time.time() - start:.2f}s") # ~1s, not 5s
Great for IO-bound work with blocking libraries. No speedup for pure-Python CPU work (see GIL).
Even with the GIL, threads share memory. This code is broken:
import threading
counter = 0
def increment():
global counter
for _ in range(1_000_000):
counter += 1 # NOT atomic — read, add, write
threads = [threading.Thread(target=increment) for _ in range(4)]
for t in threads: t.start()
for t in threads: t.join()
print(counter) # NOT 4_000_000. Something like 1_842_193.
counter += 1 compiles to 3 bytecodes (LOAD, ADD, STORE) — a thread can be swapped out between them (preemptive multitasking, remember?).
threading.Lockimport threading
counter = 0
lock = threading.Lock()
def increment():
global counter
for _ in range(1_000_000):
with lock:
counter += 1
threads = [threading.Thread(target=increment) for _ in range(4)]
for t in threads: t.start()
for t in threads: t.join()
print(counter) # 4_000_000
Now the read-modify-write is atomic. Correct result, ~10× slower.
| Primitive | Use it when |
|---|---|
Lock | Only one thread in the critical section at a time |
RLock | Same thread may acquire the lock more than once (reentrant) |
Semaphore(n) | At most N threads inside a section |
Event | One thread signals, others wait |
Condition | Wait for a state change (producer/consumer) |
Barrier(n) | N threads must all reach a point before any proceeds |
queue.Queue | Thread-safe producer/consumer — often the simplest answer |
import threading
lock_a = threading.Lock()
lock_b = threading.Lock()
def worker_1():
with lock_a:
with lock_b: # waiting for B
...
def worker_2():
with lock_b:
with lock_a: # waiting for A
...
If worker_1 gets A while worker_2 gets B, both wait forever. That's a deadlock.
Four conditions must all hold (Coffman conditions):
Break any one to prevent deadlock. Practical rules:
queue.Queue) over multiple locks.Lock.acquire(timeout=…) and give up cleanly if it times out.ThreadPoolExecutor — the simple APIfrom concurrent.futures import ThreadPoolExecutor
import requests
def fetch(url):
return requests.get(url, timeout=10).json()
urls = [
f"https://jsonplaceholder.typicode.com/users/{i}"
for i in range(1, 6)
]
with ThreadPoolExecutor(max_workers=5) as ex:
results = list(ex.map(fetch, urls))
for r in results:
print(r["name"])
multiprocessingEach process has its own interpreter and its own GIL, so processes run truly in parallel across CPU cores:
import multiprocessing
def cpu_heavy(n: int) -> int:
return sum(range(n))
if __name__ == "__main__": # required — without it, children re-import and spawn forever
procs = [
multiprocessing.Process(target=cpu_heavy, args=(50_000_000,))
for _ in range(4)
]
for p in procs: p.start()
for p in procs: p.join()
If you launch 20 processes on a 10-core machine, the OS scheduler time-slices them onto the 10 hardware threads. You don't get 20-way parallelism; you get context-switching overhead and cache thrashing. Rule of thumb: ProcessPoolExecutor() defaults to os.cpu_count() workers for a reason.
| Method | Default on | Behavior |
|---|---|---|
fork | Linux (until Python 3.14) | Child copies parent memory. Fast start. Dangerous with threads. |
spawn | Windows, macOS, Linux from Python 3.14+ | Child starts a fresh interpreter and re-imports your module. Slower start. Safe. |
forkserver | Optional on Linux | A dedicated server process forks children. |
Set explicitly for portability:
import multiprocessing as mp
mp.set_start_method("spawn", force=True)
Processes have separate memory. To share data, you pick one of:
| Primitive | What it is | Best for |
|---|---|---|
multiprocessing.Queue | Thread- and process-safe FIFO | Message passing between workers |
multiprocessing.Pipe | Two-endpoint connection | Simple parent ↔ child channel |
multiprocessing.shared_memory.SharedMemory | Raw bytes / numpy arrays with zero-copy | Big numeric arrays without pickling |
multiprocessing.Manager | Server process hosting shared dicts/lists | Convenient, but slower (all access is RPC) |
ProcessPoolExecutor — the simple APIfrom concurrent.futures import ProcessPoolExecutor
def fib(n):
a, b = 0, 1
for _ in range(n):
a, b = b, a + b
return b
if __name__ == "__main__":
with ProcessPoolExecutor() as ex: # CPU-bound
results = list(ex.map(fib, [300_000, 350_000, 400_000]))
for r in results:
print(str(r)[:15], "...")
Same interface as ThreadPoolExecutor — swap the executor to match the workload.
The Global Interpreter Lock is a mutex in CPython that lets only one thread execute Python bytecode at a time. It exists because CPython's memory management (reference counting) is not thread-safe.
If you ran Lab 0 Demo 2, you already saw this: N threads, one core. The monitor doesn't lie.
sleep), it releases the GIL, so other threads run.numpy, pandas, PyTorch, Pillow, and most scientific libraries drop the GIL when they enter C code. That's why:
# threading + numpy = actual parallel speedup
import numpy
from concurrent.futures import ThreadPoolExecutor
big_matrices = [numpy.random.rand(1000, 1000) for _ in range(4)]
with ThreadPoolExecutor() as ex:
results = list(ex.map(numpy.linalg.svd, big_matrices))
can saturate multiple cores — the C code isn't executing Python bytecode.
David Beazley's Understanding the Python GIL is the canonical GIL talk. Key findings:
The one-sentence takeaway for a Python developer: the GIL punishes you for using threads for CPU work — always has, always will (until 3.14t).
Python 3.14 ships two builds:
python runs by default.python3.14t) — GIL removed. Threads actually parallelise pure Python code.Caveats: free-threaded Python is ~10% slower single-threaded, and many C extensions haven't been ported yet. It's officially supported (Phase II, PEP 779) but not the default. Treat it as "the future," not "today's production."
New in 3.14: run multiple Python interpreters in one OS process, each with its own GIL. Access via concurrent.futures.InterpreterPoolExecutor. Cheaper than multiprocessing (no separate process, faster startup, easier data transfer than pickling) but with strict isolation. Another future path to parallelism.
The single most important concurrency decision:
| Task type | Examples | Solution |
|---|---|---|
| CPU-bound | Number crunching, image processing, ML | multiprocessing, ProcessPoolExecutor |
| IO-bound | API calls, DB queries, file reading | asyncio, threading, ThreadPoolExecutor |
Need concurrency?
├── IO-bound (network, disk, DB)?
│ ├── Whole codebase is async? → asyncio (TaskGroup, gather)
│ ├── Blocking libraries (requests, DB drivers)? → ThreadPoolExecutor
│ └── Mostly async but one blocking call? → asyncio.to_thread(...)
│
└── CPU-bound (computation)?
├── Uses numpy/pandas/PyTorch heavily? → threading MAY work (GIL released)
├── Pure Python compute? → ProcessPoolExecutor / multiprocessing
└── Very short tasks? → often faster to stay sequential (fork/spawn overhead)
threading). Cooperative = tasks yield voluntarily at await (asyncio).async def that can pause at await and resume later. Lives inside one thread. Thousands can share one thread.asyncio parallel? No — one thread, one core. It's concurrent only.gather vs TaskGroup? TaskGroup (3.11+) gives structured concurrency: if one task fails, the group cancels the others cleanly.fork vs spawn? fork copies parent memory (fast, unsafe with threads, Linux default until 3.14). spawn restarts a fresh interpreter (slow, safe, Windows/macOS default and Linux from 3.14+).Queue, Pipe, shared_memory.SharedMemory, or Manager.time.sleep() in async code? You block the whole event loop; every other coroutine stalls. Use await asyncio.sleep(...).counter += 1? Yes — it's 3 bytecodes. Protect with threading.Lock.python3.14t) that runs without the GIL. The default python still has it.async def endpoints handle thousands of concurrent requests without blocking.asyncio.gather / TaskGroup).ThreadPoolExecutor.ProcessPoolExecutor / multiprocessing.asyncio.Semaphore(N) to cap in-flight requests.asyncio.timeout(...) to guarantee bounded latency.numpy / PyTorch release the GIL, so ThreadPoolExecutor can genuinely parallelise matrix ops.Save as concurrency.py (or a folder with one file per task). Every snippet in this file is copy-paste runnable — start from those.
cpu-demo.py demos 1–3. Screenshot or note the CPU % for each. Write one sentence explaining why Demo 2 and Demo 3 look different.simulate_api_call(name, delay) (uses asyncio.sleep), then a main() that calls it three times sequentially vs. concurrently with asyncio.gather, printing both durations.asyncio.TaskGroup (3.11+) instead of gather. Add a task that raises an exception and observe how TaskGroup cancels the others.ThreadPoolExecutor to fetch /users/1–/users/5 from JSONPlaceholder; compare to sequential requests.get.counter += 1 example from above. Run it. Then fix it with threading.Lock.ProcessPoolExecutor to compute Fibonacci for [100_000, 200_000, 300_000, 400_000]; compare to sequential.asyncio.to_thread: call requests.get from inside an async function without blocking the event loop.