Memory Model Advanced¶
Variables are references, not boxes¶
A Python variable is a name bound to an object, not a container holding a value — grasping this explains aliasing, in-place mutation, and why two names can see each other's changes.
In Python, variables don't "contain" values — they're name tags pointing to objects:
a = [1, 2, 3]
b = a # b points to the SAME object
b.append(4)
print(a) # [1, 2, 3, 4] — both names see the change
print(id(a) == id(b)) # True — same memory address
print(a is b) # True — identity check
Assignment rebinds the name, it doesn't copy:¶
Assigning to a variable points the name at a different object; it never mutates what the name previously referred to — the key distinction behind "why didn't my other variable change?"
id() — object identity¶
The identity of an object — what is compares, distinct from ==.
id(obj) returns the memory address of the object (in CPython):
x = "hello"
print(id(x)) # e.g. 140234567890 (integer address)
print(hex(id(x))) # 0x7f5a3b2c1234
# Same object → same id
y = x
print(id(x) == id(y)) # True
# Different object → different id (usually)
z = "hello" # may be interned to same object
w = list(x) # definitely new object
print(id(x) == id(w)) # False
id() is only unique while the object exists
After an object is garbage collected, its id can be reused by a new object.
Mutable vs Immutable¶
Whether an object can change in place drives Python's trickiest bugs — shared mutable defaults, aliasing, and surprising == vs is results. Immutables (int, str, tuple) are safe to share; mutables aren't.
| Immutable | Mutable |
|---|---|
int, float, bool | list |
str, bytes | dict |
tuple, frozenset | set |
None | Custom objects (usually) |
# Immutable: operations create NEW objects
a = "hello"
b = a.upper() # new string — a is untouched
print(id(a) == id(b)) # False
# Mutable: operations modify IN PLACE
c = [1, 2, 3]
d = c
c.append(4) # mutates the same object
print(d) # [1, 2, 3, 4]
Shallow vs Deep Copy¶
Shallow copies share nested objects; deep copies duplicate them fully.
import copy
original = [[1, 2, 3], [4, 5, 6], {"key": "value"}]
# Shallow copy: new outer container, same inner objects
shallow = copy.copy(original)
# Also: list(original), original[:], original.copy()
shallow.append([7, 8, 9]) # doesn't affect original
print(len(original)) # 3 (unaffected)
shallow[0].append(99) # DOES affect original!
print(original[0]) # [1, 2, 3, 99] ← shared inner list
# Deep copy: recursively copies everything
deep = copy.deepcopy(original)
deep[0].append(888)
print(original[0]) # [1, 2, 3, 99] ← unaffected
When to use which:¶
| Scenario | Method |
|---|---|
| Top-level container only | copy.copy() or list() |
| Nested mutable objects | copy.deepcopy() |
| Immutable contents | No copy needed (share safely) |
| Performance-critical | Avoid deep copy, design with immutables |
Object sizes and sys.getsizeof¶
Measure an object's memory footprint.
import sys
# Basic objects
print(sys.getsizeof(None)) # 16
print(sys.getsizeof(True)) # 28
print(sys.getsizeof(0)) # 28
print(sys.getsizeof(2**30)) # 32
print(sys.getsizeof(2**60)) # 36
# Containers (shallow — doesn't count contents)
print(sys.getsizeof([])) # 56
print(sys.getsizeof([1])) # 64
print(sys.getsizeof([1,2,3,4,5])) # 96
print(sys.getsizeof({})) # 64
print(sys.getsizeof({"a": 1})) # 184
print(sys.getsizeof("")) # 49
print(sys.getsizeof("a")) # 50
print(sys.getsizeof("hello")) # 54
True deep size:¶
sys.getsizeof only counts the container itself, not its contents; this recursive helper walks the whole object graph (guarding against cycles) to get the real total — what you want when profiling a large data structure.
def deep_getsizeof(obj, seen=None):
"""Recursively compute total memory of an object graph."""
if seen is None:
seen = set()
obj_id = id(obj)
if obj_id in seen:
return 0
seen.add(obj_id)
size = sys.getsizeof(obj)
if isinstance(obj, dict):
size += sum(deep_getsizeof(k, seen) + deep_getsizeof(v, seen)
for k, v in obj.items())
elif isinstance(obj, (list, tuple, set, frozenset)):
size += sum(deep_getsizeof(i, seen) for i in obj)
elif hasattr(obj, '__dict__'):
size += deep_getsizeof(obj.__dict__, seen)
return size
data = {"users": [{"name": "Alice", "scores": [95, 87, 92]}] * 100}
print(f"Shallow: {sys.getsizeof(data)} bytes")
print(f"Deep: {deep_getsizeof(data)} bytes")
tracemalloc — tracking memory allocations¶
Trace where memory is allocated to find growth and leaks.
import tracemalloc
tracemalloc.start()
# Code to measure
data = [list(range(1000)) for _ in range(1000)]
snapshot = tracemalloc.take_snapshot()
stats = snapshot.statistics("lineno")
print("Top 5 memory consumers:")
for stat in stats[:5]:
print(f" {stat}")
Output:
Memory-efficient patterns¶
Generators, __slots__, and smaller types to cut memory use.
Use __slots__ for data-heavy classes:¶
Declaring __slots__ drops each instance's per-object __dict__, cutting memory roughly 3× — a big win when you create millions of small objects (points, records, nodes).
class PointSlots:
__slots__ = ('x', 'y', 'z')
def __init__(self, x, y, z):
self.x, self.y, self.z = x, y, z
class PointDict:
def __init__(self, x, y, z):
self.x, self.y, self.z = x, y, z
# Create 1 million points
import sys
slots_list = [PointSlots(i, i, i) for i in range(100_000)]
dict_list = [PointDict(i, i, i) for i in range(100_000)]
print(f"Slots: {sys.getsizeof(slots_list[0])} bytes each") # ~56
print(f"Dict: {sys.getsizeof(dict_list[0]) + sys.getsizeof(dict_list[0].__dict__)} bytes each") # ~168
Use generators instead of lists for streaming:¶
Yield items one at a time instead of building a full list, so memory stays flat no matter how large the input — the standard fix for processing huge files or streams.
# Bad — stores all in memory
all_lines = [line.strip() for line in open("huge.txt")]
# Good — streams one line at a time
def stripped_lines(path):
with open(path) as f:
for line in f:
yield line.strip()
Use array.array for homogeneous numeric data:¶
Store millions of same-type numbers in a compact C array instead of a list of boxed Python ints — about half the memory, and the gateway to NumPy for even more.
import array, sys
py_list = list(range(1_000_000))
c_array = array.array('i', range(1_000_000))
print(f"List: {sys.getsizeof(py_list):>10,} bytes") # ~8.4 MB
print(f"Array: {sys.getsizeof(c_array):>10,} bytes") # ~4.0 MB
# NumPy array would be even smaller + faster
Weak References¶
Reference an object without keeping it alive — for caches and back-references.
Normal references keep objects alive. Weak references allow the object to be garbage collected:
import weakref
class ExpensiveObject:
def __init__(self, name):
self.name = name
def __del__(self):
print(f" {self.name} deleted")
obj = ExpensiveObject("resource")
weak = weakref.ref(obj)
print(weak()) # <ExpensiveObject object>
print(weak().name) # resource
del obj # Output: resource deleted
print(weak()) # None — object was collected
WeakValueDictionary — cache that doesn't prevent GC:¶
A cache whose entries vanish once nothing else references the value — gives you memoization without the memory leak of a dict that pins every object it ever stored.
cache = weakref.WeakValueDictionary()
def get_data(key):
if key in cache:
return cache[key]
data = ExpensiveObject(key) # expensive creation
cache[key] = data
return data
Practice Exercises¶
- Create a reference cycle and show that
deldoesn't free the memory (needgc.collect()). - Implement
deep_getsizeofand measure the true memory of a nested data structure. - Compare memory usage of 1M objects with
__slots__vs without. - Use
tracemallocto find the top 3 memory allocations in a complex script. - Build a weak-reference cache that expires entries when no strong references exist.
- Demonstrate that tuple reuse happens for small tuples:
() is ()is True.
💬 Discussion
Have a question about this topic? Found an error? Share your thoughts below.