Skip to content

System Design Notes

Working notes on systems engineering, written while learning each topic from primary sources rather than summaries. They go from CPU and OS primitives up through storage engines, distributed systems, Kubernetes, observability, and retrieval-augmented generation.

Most chapters end with lab exercises, and several have runnable code in the repository alongside them.

How to read these

Each topic directory is ordered numerically and meant to be read in sequence -- later chapters assume the earlier ones. Use the navigation on the left to follow a track, or the search box (press /) to jump straight to a concept.

Tracks

  • :material-language-python:{ .lg .middle } Python & Systems Internals


    CPU execution model, caches and memory ordering, virtual memory, allocators, syscalls and IO, then CPython itself: refcounting, the eval loop, GC, the GIL, free-threading, and asyncio internals.

    :octicons-arrow-right-24: 30 chapters

  • :material-database:{ .lg .middle } Databases & Storage


    Storage engine fundamentals, encoding formats, access methods, query engines, transactions and concurrency control, B-trees and LSM-trees, write-ahead logging, and vector search internals.

    :octicons-arrow-right-24: 25 chapters

  • :material-lan:{ .lg .middle } Distributed Systems


    Consensus, replication, failure detection, and coordination -- the staff-level roadmap and its reading map.

    :octicons-arrow-right-24: Roadmap

  • :material-kubernetes:{ .lg .middle } Kubernetes & Containers


    From Linux namespaces and cgroups up through etcd, the API server, scheduler, and kubelet internals, then CNI, Cilium and eBPF, CSI, operators, multi-tenancy, and supply-chain security.

    :octicons-arrow-right-24: 46 chapters

  • :material-chart-timeline-variant:{ .lg .middle } SRE & Observability


    OpenTelemetry, instrumentation, collection and transport, storage for metrics/logs/traces, query layers, SLO engineering, on-call and incident response, cardinality and cost control.

    :octicons-arrow-right-24: 47 chapters

  • :material-expansion-card-variant:{ .lg .middle } GPU Observability


    DCGM exporter internals, GPU cluster telemetry on Kubernetes, allocation and utilization efficiency, hardware failure detection, and observability for LLM inference and distributed training.

    :octicons-arrow-right-24: 23 chapters

  • :material-vector-triangle:{ .lg .middle } AI & RAG


    Embeddings and representation, chunking and document processing, vector indexes, hybrid retrieval and reranking, and evaluation methodology -- with document-processing and golden-set labs.

    :octicons-arrow-right-24: Reading map

  • :material-hammer-wrench:{ .lg .middle } Design Practice


    Design tasks stated at four scale tiers (10k → 10m), worked solutions, and reference implementations for each tier: Twitter search, Instagram feed, a distributed counter, and FastAPI RBAC.

    :octicons-arrow-right-24: Start with the tasks

Also here

  • System Design Guide — the cross-cutting reference that ties the tracks together.
  • Kubernetes Labs — hands-on task sheets with manifests, separate from the Kubernetes theory track.

The source lives at github.com/Harut8/system-design. Corrections are welcome via issues or pull requests.