Skip to content

System Design Tasks

Twenty-one interview-style design problems. Each has a full reference solution in ../solutions/ and ends with an Interview Kit: what to read first, curveballs an interviewer would throw, must-answer security and privacy questions, and a phased delivery plan.

Index

Task Level Time box Core topics Solution
Distributed counter Mid 45 min Hot keys, idempotency, write aggregation, caching solution
Instagram feed Mid 45 min Fan-out on write vs. read, caching, ranking solution
Twitter search Mid–Senior 45 min Inverted index, real-time ingestion, visibility filtering solution
API gateway + rate limiter Mid–Senior 45 min Token bucket / GCRA, distributed limits, config safety solution
FastAPI RBAC Mid 45 min, plus code Roles, tenancy, permission caching, audit solution
Job scheduler on Postgres Senior 60 min SKIP LOCKED, leases, vacuum, fairness solution
Key-value store Senior–Staff 60 min Raft, partitioning, linearizability, leases solution
IBKR-style trading platform Staff 60 min OMS, risk, market data, compliance, idempotency solution
Feature store Senior 60 min Point-in-time joins, streaming features, skew solution
Recommendation system Senior–Staff 60 min Two-tower retrieval, ANN, multi-stage ranking solution
ML inference platform Senior–Staff 60 min Batching, autoscaling GPUs, LLM serving, registry solution
Parallel ML training Staff 60 min 3D parallelism, gang scheduling, checkpointing solution
DAG pipeline orchestration Senior 60 min Scheduler loop, executors, backfills, metadata DB solution
Workflow orchestration Staff 60 min Event sourcing, deterministic replay, sharding solution
RAG platform Senior–Staff 60 min Hybrid retrieval, ACLs, evaluation, injection solution
AI search engine Staff 60 min Crawling, index tiers, ranking, answer generation solution
LLM gateway Senior 45 min Provider routing, fallbacks, token budgets solution
Agent orchestration Staff 60 min Planning, durable runs, human oversight solution
AI agent platform Staff 60 min Multi-team platform, credentials, evals solution
Tool platform Staff 60 min Tool registry, sandboxes, on-behalf-of auth solution
AI observability & evaluation Senior–Staff 60 min Tracing, LLM judges, sampling, privacy solution

Suggested order: counter → feed → rate limiter → job scheduler → key-value store for the distributed-systems core. Then feature store → recommendation → inference for ML. Then LLM gateway → RAG → tool platform → agent orchestration for AI systems.

How to practice

  1. Time-box it (45 or 60 minutes, see the index) and answer out loud or on paper before opening the solution.
  2. Follow this shape. Interviewers score structure as much as content:
Minutes (of 45) Step Output
0–5 Clarify Functional scope, non-goals, the two or three numbers that drive the design
5–10 Estimate QPS (peak, not average), storage growth, the read/write ratio, one cost figure
10–20 High-level design Boxes and arrows, the data model, the API for the main call
20–35 Deep dives The two hardest parts, chosen by you, with trade-offs stated
35–40 Failure and security What breaks first, what happens then, who can see which data
40–45 Evolution MVP → Growth → Scale, and what you would not build yet
  1. Then work through the task's Interview Kit curveballs without looking at the hints.
  2. Read the solution and score yourself with the rubric below. Redo the task a week later.

Scoring rubric

Score each row 1–4. Around 20 of 28 is a pass at senior level. Staff level needs no row below 3.

Dimension 1: Weak 2: Developing 3: Strong 4: Exceptional
Requirements Starts drawing immediately Lists features, no numbers Scope, non-goals and the key numbers stated up front Finds the requirement that changes the design (e.g. "exact counts only below 1K")
Estimation None Numbers without using them Peak QPS, storage and cost estimated and used to pick components Estimates expose the bottleneck before the design does
Architecture Box soup, no data flow Reasonable boxes, vague data model Clear data model, API, read and write paths Simplest design that meets the numbers, with a named reason for each component
Deep dives Stays at the surface Explains one component Two hard parts in depth, with alternatives compared Quantifies the trade-off (latency, cost, consistency) and names when the choice flips
Failure handling "We'll add retries" Lists failures Behaviour under node, zone and dependency failure. Retries have budgets, overload is shed Recovery is designed too: no thundering herds, backlog handled, blast radius bounded (cells, shuffle sharding)
Security & privacy Not mentioned "Use HTTPS and auth" Authn/z at every boundary, tenant isolation, secrets, PII handling and deletion Threat model specific to this system (IDOR, injection, abuse), with GDPR/PCI trade-offs resolved
Evolution & operations Only the end state Mentions monitoring MVP → Scale phases, SLOs, key alerts, cost awareness Says what not to build yet and what signal would trigger the next phase

Reference material

The chapters these tasks draw on: ../distributed-systems/, ../databases/, ../ai-rag/ and ../sre-observability/. Pattern references with no matching task: API design patterns, API message patterns, Big-tech API standards, Database design best practices, Write-ahead log deep dive and Zero-GC Python.