/ Single-Thread vs Multi-Thread in Python
Single-Thread vs Multi-Thread in Python¶
This notebook demonstrates the difference between single-threaded and multi-threaded execution in Python.
We use an I/O-bound workload — simulated network requests with time.sleep — which is the ideal scenario where threading shines.
Note: Python's Global Interpreter Lock (GIL) prevents true parallel execution of CPU-bound tasks using threads. For CPU-bound work, prefer
multiprocessing. For I/O-bound work,threading(orasyncio) gives real speedups.
import time
import threading
from concurrent.futures import ThreadPoolExecutor
import random
The Task¶
We simulate fetching data from a list of remote endpoints. Each "fetch" sleeps for a random duration between 0.5 and 1.5 seconds, mimicking network latency.
ENDPOINTS = [
"https://api.example.com/users",
"https://api.example.com/products",
"https://api.example.com/orders",
"https://api.example.com/inventory",
"https://api.example.com/reports",
"https://api.example.com/analytics",
"https://api.example.com/notifications",
"https://api.example.com/settings",
]
LATENCY_SEED = 42 # reproducible latencies
def get_latencies():
rng = random.Random(LATENCY_SEED)
return {url: round(rng.uniform(0.5, 1.5), 2) for url in ENDPOINTS}
LATENCIES = get_latencies()
def fetch(url: str) -> dict:
"""Simulate an HTTP GET request with network latency."""
delay = LATENCIES[url]
time.sleep(delay)
return {"url": url, "status": 200, "latency_s": delay}
print("Simulated latencies (seconds):")
for url, latency in LATENCIES.items():
print(f" {url:<45} {latency:.2f}s")
print(f"\nTotal if sequential: {sum(LATENCIES.values()):.2f}s")
Single-Threaded Execution¶
In single-threaded execution, each request is made one after another. The total time is the sum of all individual latencies.
def run_single_threaded(urls):
results = []
for url in urls:
result = fetch(url)
results.append(result)
print(f" Fetched {result['url'].split('/')[-1]:<15} in {result['latency_s']:.2f}s")
return results
print("Single-threaded run:")
t0 = time.perf_counter()
single_results = run_single_threaded(ENDPOINTS)
single_elapsed = time.perf_counter() - t0
print(f"\nTotal elapsed: {single_elapsed:.2f}s")
Multi-Threaded Execution¶
With ThreadPoolExecutor, all requests are dispatched concurrently.
While one thread waits on I/O, others can proceed — so the total time approaches the maximum individual latency rather than the sum.
def run_multi_threaded(urls, max_workers=8):
results = []
lock = threading.Lock()
def fetch_and_record(url):
result = fetch(url)
with lock:
results.append(result)
print(f" Fetched {result['url'].split('/')[-1]:<15} in {result['latency_s']:.2f}s")
with ThreadPoolExecutor(max_workers=max_workers) as executor:
executor.map(fetch_and_record, urls)
return results
print("Multi-threaded run:")
t0 = time.perf_counter()
multi_results = run_multi_threaded(ENDPOINTS)
multi_elapsed = time.perf_counter() - t0
print(f"\nTotal elapsed: {multi_elapsed:.2f}s")
Performance Comparison¶
speedup = single_elapsed / multi_elapsed
theoretical_max = single_elapsed / max(LATENCIES.values())
print(f"{'Metric':<35} {'Value':>10}")
print("-" * 46)
print(f"{'Single-threaded time':<35} {single_elapsed:>9.2f}s")
print(f"{'Multi-threaded time':<35} {multi_elapsed:>9.2f}s")
print(f"{'Speedup':<35} {speedup:>9.2f}x")
print(f"{'Theoretical max speedup':<35} {theoretical_max:>9.2f}x")
print(f"{'Endpoints fetched':<35} {len(ENDPOINTS):>10}")
The GIL: Why Threading Doesn't Help CPU-Bound Work¶
Python's Global Interpreter Lock (GIL) allows only one thread to execute Python bytecode at a time. For CPU-bound tasks (math, compression, image processing), threads take turns rather than running in parallel — yielding little or no speedup.
| Workload type | Best tool |
|---|---|
| I/O-bound (network, disk) | threading or asyncio |
| CPU-bound (computation) | multiprocessing or concurrent.futures.ProcessPoolExecutor |
The example below shows a CPU-bound task where threading gives no speedup.
def cpu_task(n: int) -> int:
"""Sum squares up to n — a pure CPU workload."""
return sum(i * i for i in range(n))
TASKS = [5_000_000] * 8
# Single-threaded
t0 = time.perf_counter()
cpu_single = [cpu_task(n) for n in TASKS]
cpu_single_elapsed = time.perf_counter() - t0
# Multi-threaded (GIL prevents true parallelism)
t0 = time.perf_counter()
with ThreadPoolExecutor(max_workers=8) as executor:
cpu_multi = list(executor.map(cpu_task, TASKS))
cpu_multi_elapsed = time.perf_counter() - t0
print("CPU-bound task (sum of squares, 8 tasks of 5M iterations each):")
print(f" Single-threaded: {cpu_single_elapsed:.2f}s")
print(f" Multi-threaded: {cpu_multi_elapsed:.2f}s")
print(f" Speedup: {cpu_single_elapsed / cpu_multi_elapsed:.2f}x (near 1.0 due to GIL)")
Summary¶
- I/O-bound tasks: threading delivers real concurrency because threads yield the GIL while waiting on I/O.
- CPU-bound tasks: threads compete for the GIL — use
multiprocessinginstead. ThreadPoolExecutoris the idiomatic high-level API for thread pools in modern Python.- For even higher concurrency on I/O-bound work, consider
asynciowithasync/await.
© 2023 Ivan Cao-Berg Pittsburgh Supercomputing Center, Carnegie Mellon University
Licensed under the GNU General Public License v2.0 (GPL-2).
You may redistribute and/or modify this work under the terms of GPL-2.
This work is distributed WITHOUT ANY WARRANTY.
Happy computing.