System Design Tasks¶
Twenty-one interview-style design problems. Each has a full reference solution in
../solutions/ and ends with an Interview Kit: what to read first,
curveballs an interviewer would throw, must-answer security and privacy questions, and a
phased delivery plan.
Index¶
| Task | Level | Time box | Core topics | Solution |
|---|---|---|---|---|
| Distributed counter | Mid | 45 min | Hot keys, idempotency, write aggregation, caching | solution |
| Instagram feed | Mid | 45 min | Fan-out on write vs. read, caching, ranking | solution |
| Twitter search | Mid–Senior | 45 min | Inverted index, real-time ingestion, visibility filtering | solution |
| API gateway + rate limiter | Mid–Senior | 45 min | Token bucket / GCRA, distributed limits, config safety | solution |
| FastAPI RBAC | Mid | 45 min, plus code | Roles, tenancy, permission caching, audit | solution |
| Job scheduler on Postgres | Senior | 60 min | SKIP LOCKED, leases, vacuum, fairness |
solution |
| Key-value store | Senior–Staff | 60 min | Raft, partitioning, linearizability, leases | solution |
| IBKR-style trading platform | Staff | 60 min | OMS, risk, market data, compliance, idempotency | solution |
| Feature store | Senior | 60 min | Point-in-time joins, streaming features, skew | solution |
| Recommendation system | Senior–Staff | 60 min | Two-tower retrieval, ANN, multi-stage ranking | solution |
| ML inference platform | Senior–Staff | 60 min | Batching, autoscaling GPUs, LLM serving, registry | solution |
| Parallel ML training | Staff | 60 min | 3D parallelism, gang scheduling, checkpointing | solution |
| DAG pipeline orchestration | Senior | 60 min | Scheduler loop, executors, backfills, metadata DB | solution |
| Workflow orchestration | Staff | 60 min | Event sourcing, deterministic replay, sharding | solution |
| RAG platform | Senior–Staff | 60 min | Hybrid retrieval, ACLs, evaluation, injection | solution |
| AI search engine | Staff | 60 min | Crawling, index tiers, ranking, answer generation | solution |
| LLM gateway | Senior | 45 min | Provider routing, fallbacks, token budgets | solution |
| Agent orchestration | Staff | 60 min | Planning, durable runs, human oversight | solution |
| AI agent platform | Staff | 60 min | Multi-team platform, credentials, evals | solution |
| Tool platform | Staff | 60 min | Tool registry, sandboxes, on-behalf-of auth | solution |
| AI observability & evaluation | Senior–Staff | 60 min | Tracing, LLM judges, sampling, privacy | solution |
Suggested order: counter → feed → rate limiter → job scheduler → key-value store for the distributed-systems core. Then feature store → recommendation → inference for ML. Then LLM gateway → RAG → tool platform → agent orchestration for AI systems.
How to practice¶
- Time-box it (45 or 60 minutes, see the index) and answer out loud or on paper before opening the solution.
- Follow this shape. Interviewers score structure as much as content:
| Minutes (of 45) | Step | Output |
|---|---|---|
| 0–5 | Clarify | Functional scope, non-goals, the two or three numbers that drive the design |
| 5–10 | Estimate | QPS (peak, not average), storage growth, the read/write ratio, one cost figure |
| 10–20 | High-level design | Boxes and arrows, the data model, the API for the main call |
| 20–35 | Deep dives | The two hardest parts, chosen by you, with trade-offs stated |
| 35–40 | Failure and security | What breaks first, what happens then, who can see which data |
| 40–45 | Evolution | MVP → Growth → Scale, and what you would not build yet |
- Then work through the task's Interview Kit curveballs without looking at the hints.
- Read the solution and score yourself with the rubric below. Redo the task a week later.
Scoring rubric¶
Score each row 1–4. Around 20 of 28 is a pass at senior level. Staff level needs no row below 3.
| Dimension | 1: Weak | 2: Developing | 3: Strong | 4: Exceptional |
|---|---|---|---|---|
| Requirements | Starts drawing immediately | Lists features, no numbers | Scope, non-goals and the key numbers stated up front | Finds the requirement that changes the design (e.g. "exact counts only below 1K") |
| Estimation | None | Numbers without using them | Peak QPS, storage and cost estimated and used to pick components | Estimates expose the bottleneck before the design does |
| Architecture | Box soup, no data flow | Reasonable boxes, vague data model | Clear data model, API, read and write paths | Simplest design that meets the numbers, with a named reason for each component |
| Deep dives | Stays at the surface | Explains one component | Two hard parts in depth, with alternatives compared | Quantifies the trade-off (latency, cost, consistency) and names when the choice flips |
| Failure handling | "We'll add retries" | Lists failures | Behaviour under node, zone and dependency failure. Retries have budgets, overload is shed | Recovery is designed too: no thundering herds, backlog handled, blast radius bounded (cells, shuffle sharding) |
| Security & privacy | Not mentioned | "Use HTTPS and auth" | Authn/z at every boundary, tenant isolation, secrets, PII handling and deletion | Threat model specific to this system (IDOR, injection, abuse), with GDPR/PCI trade-offs resolved |
| Evolution & operations | Only the end state | Mentions monitoring | MVP → Scale phases, SLOs, key alerts, cost awareness | Says what not to build yet and what signal would trigger the next phase |
Reference material¶
The chapters these tasks draw on: ../distributed-systems/,
../databases/, ../ai-rag/ and
../sre-observability/. Pattern references with no matching task:
API design patterns,
API message patterns,
Big-tech API standards,
Database design best practices,
Write-ahead log deep dive and
Zero-GC Python.