Skip to content

Performance Anti-Patterns Advanced

⚙️ Performance & Systems
⏱️ ~2 days 📚 Prerequisites: Profiling, Data Structures

When you'd use this

Common slow patterns in Python and the faster alternatives.

Recognize and avoid the slow patterns (string += in loops, needless copies, N+1 queries) that quietly kill performance.

What you'll learn

  • The most common Python slow-downs
  • String building the wrong way (tested)
  • Wrong data structure for the job (tested)
  • Repeated work and needless allocation
  • Measure before you optimize

Most Python performance problems come from a handful of recurring mistakes. Recognizing them saves you from most slow code. The comparisons here are run-verified.

Profile first

Don't guess where the slowness is — measure it (see Profiling). "Premature optimization is the root of all evil"; but known anti-patterns are worth avoiding from the start.


Anti-pattern 1: string concatenation in a loop

s += x in a loop is O(n²); build a list and join instead.

Building a string with += in a loop is O(n²) — each += creates a whole new string (strings are immutable). Use join instead. Tested:

n = 20000

# SLOW — O(n²): a new string built every iteration
s = ""
for i in range(n):
    s += "x"

# FAST — O(n): collect parts, join once
parts = []
for i in range(n):
    parts.append("x")
result = "".join(parts)

print(len(s), len(result))

Output:

20000 20000

Both produce the same 20,000-character string, but join is dramatically faster at scale because it allocates once instead of n times. Rule: build a list, then "".join(list) — never += strings in a loop.


Anti-pattern 2: wrong data structure (list vs set for membership)

x in list is O(n); use a set for O(1) membership.

Checking x in list is O(n) — it scans every element. x in set is O(1). If you test membership repeatedly, use a set. Tested:

import time

data = list(range(10000))
data_set = set(data)
target = 9999

t0 = time.perf_counter()
for _ in range(1000):
    _ = target in data           # O(n) scan each time
list_time = time.perf_counter() - t0

t0 = time.perf_counter()
for _ in range(1000):
    _ = target in data_set       # O(1) hash lookup
set_time = time.perf_counter() - t0

print("set faster than list:", set_time < list_time)

Output:

set faster than list: True

The set is far faster for repeated membership tests. Rule: if you check membership a lot, store it in a set (or dict), not a list. Choosing the right data structure (Data Structures) is often the single biggest performance lever.


Anti-pattern 3: recomputing inside a loop

Hoist invariant work out of the loop instead of redoing it each pass.

Computing the same value every iteration wastes work. Hoist invariants out:

# SLOW — len(data) and the lookup recomputed every iteration
for i in range(len(data)):
    if data[i] > some_object.threshold.value:
        ...

# FAST — compute once, before the loop
n = len(data)
threshold = some_object.threshold.value
for i in range(n):
    if data[i] > threshold:
        ...

Attribute lookups (some_object.threshold.value) and function calls have real cost in Python. Pull anything constant out of the loop body.


Anti-pattern 4: needless allocation & copying

Avoid building throwaway intermediate lists; stream or reuse buffers.

Creating throwaway lists, or copying data you could iterate lazily, wastes memory and time:

# SLOW — builds a full list just to sum it
total = sum([x * x for x in range(100000)])   # list comprehension allocates a list

# FAST — generator expression, no intermediate list
total = sum(x * x for x in range(100000))     # streams values, no allocation

Dropping the brackets ([...] → (...)) turns a list comprehension into a generator expression — same result, no intermediate list allocated. Use generators when you only iterate once.


Anti-pattern 5: looping in Python over big numeric data

Vectorize with NumPy instead of per-element Python loops.

For heavy numeric work, a Python for loop is slow because each iteration runs interpreted bytecode. Push the loop into C — via builtins (sum, map) or, for real numeric arrays, NumPy (see Vectorization):

# Pure-Python loop (interpreted each step)
total = 0
for x in range(100000):
    total += x * x

# Builtin pushes iteration into C — faster
total = sum(x * x for x in range(100000))

Both give the same answer; the builtin version does the looping in optimized C. For large arrays, NumPy vectorization is orders of magnitude faster still.


The meta-lesson

Measure first, fix the real bottleneck, and prefer the right data structure.

Anti-pattern Fix
+= strings in a loop "".join(list)
in list repeatedly use a set/dict
Recompute in loop hoist invariants out
Needless list allocation generator expressions
Python loop over numeric data builtins / NumPy

But measure — don't cargo-cult

These are known patterns worth avoiding, but the golden rule stands: profile before optimizing (Profiling). Optimizing code that isn't the bottleneck wastes effort and can hurt readability. Fix the anti-patterns in hot paths the profiler identifies.


Practice exercises

  1. Time string += vs join for n = 100,000 and report the ratio.
  2. Find a place in your own code testing membership on a list and switch it to a set.
  3. Rewrite a loop that recomputes len() or an attribute each iteration to hoist it out.
  4. Convert a list comprehension used only for summing into a generator expression.
  5. Profile a slow function you have and identify which anti-pattern (if any) it hits.

💬 Discussion

Have a question about this topic? Found an error? Share your thoughts below.