Twitter search
System Design Task: Twitter-Scale Real-Time Search¶
Problem Statement¶
Design a scalable, low-latency system to index and search billions of tweets in near real time for a global user base. The system must support search, trending topics, and autocomplete, while handling massive write and read traffic with high availability.
You are expected to design this as if it were going into production at Twitter scale.
Functional Requirements¶
Your design must support:
-
Tweet Search
-
Search by:
- Keywords and phrases
- Hashtags
- User IDs
- Filters: time range, geo/location, language
- Trending Topics
-
Identify and rank trending hashtags and phrases:
- Globally
- Per region
- Trends must reflect near real-time activity
- Autocomplete
-
Provide autocomplete suggestions while users type
-
Suggestions should be based on:
- Popular queries
- Recent searches
- Trending terms
- APIs
-
Search API
- Trending topics API
- Autocomplete API
-
Ranking
-
Rank search results using:
- Recency
- Popularity (engagement)
- Optional personalization signals
Non-Functional Requirements¶
Your system must meet the following constraints:
-
Scale
-
Billions of tweets stored
- Thousands to tens of thousands of queries per second globally
-
Latency
-
P99 latency ≤ 200 ms for search and autocomplete
-
Freshness
-
Trending topics freshness < 1 minute
- Newly created tweets should appear in search within seconds
-
Availability
-
≥ 99.99% uptime
- Tolerant to regional and data-center failures
-
Consistency
-
Eventual consistency is acceptable across shards and regions
What You Should Deliver¶
Provide a practical, production-oriented design that includes:
- Requirement clarification & assumptions
-
High-level architecture
-
Core services
- Data flow (tweet ingestion → indexing → search)
-
Data storage choices
-
Indexing strategy
- Sharding and partitioning
-
Search architecture
-
How queries are executed and ranked
- How latency targets are met
-
Trending topics computation
-
Windowing strategy
- Real-time vs batch trade-offs
-
Autocomplete design
-
Data sources
- Update frequency
-
Scalability strategy
-
Horizontal scaling
- Hot partition mitigation
-
Failure handling
-
Regional failover
- Backpressure, retries
-
Rough capacity estimates
-
Tweets/day
- Index size
- QPS assumptions
-
Trade-offs
- Explicitly explain what you choose not to optimize and why
Expectations¶
- Be concrete (name technologies or categories: inverted index, stream processor, cache, etc.)
- Avoid academic theory unless it directly impacts production behavior
- Prefer simple, reliable designs over clever ones
- Assume this system will be maintained by hundreds of engineers over many years
Interview Kit¶
Read first: Solution · Sharding §9 secondary indexes · Stream processing · Indexing §8 full-text
Curveballs. The interviewer changes one thing mid-design. The hint in italics is what a strong answer reaches for:
- A user blocks someone. Within how many seconds must the blocked person's tweets disappear from their results, and where is that enforced? (At query time, per viewer, not at index time.)
- An election night sends query volume 20× for one hour, all for five terms. (Result cache with short TTL, request coalescing, and shedding of expensive queries.)
- Legal requires a tweet to be withheld in one country only. (A visibility class evaluated with the viewer's region.)
- Autocomplete suggests a private medical query typed by one user. How did that happen, and what threshold prevents it?
Must answer (security, privacy, operations):
- Deleted tweets: the SLO for disappearing from results vs. from the index, and how tombstones bridge the gap
- Query-log retention and anonymization. Why user text never reaches
query_string
Phase it (MVP → Growth → Scale): MVP: Postgres full-text search plus Redis for trends. Growth: Kafka to Elasticsearch with time-sharded indexes. Scale: earlybird-style in-memory real-time tier plus archive tier, multi-region.
Score yourself with the rubric.