Skip to content

Kernel Tracing Advanced

🔧 Embedded & Systems
⏱️ ~1 week 📚 Prerequisites: Profiling, eBPF

When you'd use this

Trace syscalls and kernel events with perf, ftrace and eBPF from Python.

Trace OS/kernel behavior to diagnose latency, scheduling, and I/O issues on embedded and server systems.

What you'll learn

  • Why trace the kernel at all
  • The main Linux tracing tools (strace, perf, ftrace, eBPF)
  • Trace system calls from and about a program
  • Use eBPF/BCC from Python
  • Read tracing output to debug performance

Linux-only, requires elevated privileges

Kernel tracing is a Linux capability and needs root/CAP_SYS_ADMIN. The tools here (perf, ftrace, eBPF/BCC) don't exist on this Windows host, so the commands and BCC snippets aren't run-verified — they follow the documented interfaces. To use them you need a Linux system with the tracing tools and kernel support installed.


Why trace the kernel

A core question explored in Kernel Tracing: Why trace the kernel.

Application profilers (see Profiling) show where your Python code spends time. But when a program is slow because of what it asks the operating system to do — reading files, waiting on the network, spawning processes, contending for locks — you need to see below your code, into the kernel. Kernel tracing reveals the system calls a program makes, how long they take, and what the kernel does in response.

   your Python code          ← profilers see here
   ─────────────────────
   system calls (read,       ← kernel tracing sees here
   write, open, futex...)
   ─────────────────────
   kernel: schedulers, I/O,
   network stack, filesystems

The Linux tracing toolbox

strace, perf, ftrace, and eBPF — the layers of Linux observability.

Tool What it does Overhead
strace Log every syscall a process makes High (good for debugging, not production)
perf Sample CPU, count events, profile system-wide Low–medium
ftrace Built-in kernel function tracer (via /sys/kernel/debug/tracing) Low
eBPF Run safe custom programs inside the kernel to trace anything Very low (production-grade)

The trend is toward eBPF — it's programmable, low-overhead, and safe enough to run on production systems.


strace: see the syscalls

Watch every system call a process makes to diagnose I/O and permission issues.

The quickest way to see what a program asks the kernel to do:

strace -c python3 myscript.py     # -c summarizes syscall counts and time

Example summary (illustrative):

% time     seconds  usecs/call     calls    errors syscall
------ ----------- ----------- --------- --------- ----------------
 42.11    0.001234          12       100           read
 28.05    0.000822           8       100           write
 ...

This instantly answers "is my program spending its time in read, futex, poll…?" — often revealing that a "slow" program is actually waiting on I/O or locks, not computing.


perf: profile the whole system

Sample CPU across the whole machine to find hotspots.

perf samples what's running across the system, including kernel time:

perf record -g python3 myscript.py    # -g captures call graphs
perf report                            # interactive breakdown

It attributes CPU time down into kernel functions, so you can see, say, that time is going into the network stack or the filesystem — invisible to a pure-Python profiler.


eBPF from Python (BCC)

Attach safe programs to kernel events and read them from Python.

eBPF lets you load small, verified programs into the kernel that fire on events (a syscall, a function entry, a network packet). The BCC toolkit exposes this from Python: you write the in-kernel probe in a C snippet and the orchestration in Python.

from bcc import BPF     # requires bcc installed on Linux

# In-kernel probe: fire on every open() syscall
program = r"""
int trace_open(void *ctx) {
    bpf_trace_printk("open() called\\n");
    return 0;
}
"""

b = BPF(text=program)
b.attach_kprobe(event=b.get_syscall_fnname("openat"), fn_name="trace_open")

print("tracing openat()... Ctrl-C to stop")
b.trace_print()      # stream events as they fire

This attaches a kprobe (kernel probe) to the openat syscall; every time any process opens a file, the probe fires and prints. Real eBPF tools (like the bcc and bpftrace suites — execsnoop, opensnoop, biolatency) do exactly this to trace process launches, file opens, disk latency, and more with negligible overhead.

Reach for prebuilt eBPF tools first

Before writing your own probe, check the bcc-tools/bpftrace collection — opensnoop, execsnoop, tcpconnect, biolatency, and dozens more already cover common questions. Write custom eBPF only when nothing off-the-shelf answers your question. See the eBPF topic for depth.


Reading the output to debug

Interpret traces to find the actual cause of latency or failures.

The workflow for a "mysteriously slow" program:

  1. strace -c — is it syscall-bound? Which syscall dominates?
  2. If it's waiting (lots of time in read, poll, futex) → it's I/O or lock contention, not CPU.
  3. perf — where does CPU time go, including kernel?
  4. eBPF tool — drill into the specific subsystem (disk latency with biolatency, slow opens with opensnoop).

This top-down path moves you from "the program is slow" to a specific kernel-level cause — the kind of insight application profilers can't give.


Practice exercises

  1. On a Linux machine, run strace -c on a simple Python script that reads a file in a loop, and identify the dominant syscall.
  2. Use strace -e trace=network on a script that makes an HTTP request and observe the connect/send/recv calls.
  3. Run the opensnoop bcc tool (if available) and watch which files programs open system-wide.
  4. Explain the difference between what cProfile shows and what strace shows for the same program.
  5. Describe a performance problem that a kernel tracer would reveal but a Python profiler would completely miss.

💬 Discussion

Have a question about this topic? Found an error? Share your thoughts below.