System design fundamentals cheat sheet¶
For anyone with an HLD round, and useful for LLD too. You end with 51 concepts in your own words, the latency numbers, and an estimation sheet you can redo from memory.
How to use this page¶
- Work through one table per sitting. For each row, open the free link and read only the section on that concept.
- Close the tab and write the one-line meaning in your own words. If you cannot, read it again.
- Turn rows into flashcards, or import the System Design Primer's ready-made Anki decks. Review 10 minutes a day.
- Copy the latency numbers and the estimation cheat sheet onto one page of notes.
- Before every mock, re-read the trade-off table.
Tip: A new grad loop does not need all 51. If your loop is LLD-only, read the caching, rate limiting, idempotency and isolation-level rows, then move to Low-level design.
Scale and performance¶
| Concept | What it means | Best free link | Where it comes up |
|---|---|---|---|
| Vertical vs horizontal scaling | Vertical is a bigger machine. Horizontal is more machines behind a load balancer, which needs stateless servers. | ByteByteGo: Scale from zero | The first growth step in every design |
| Latency vs throughput | Latency is time per request. Throughput is requests per second. Aim for the most throughput at an acceptable latency. | System Design Primer | Writing non-functional requirements |
| Availability and the nines | The share of time the system works. 99.99% allows about 52.6 minutes down a year. Parts in series lower it; parts in parallel raise it. | ByteByteGo: Estimation | Non-functional requirements, replicas |
| SLI, SLO, SLA | SLI is what you measure (p99 latency). SLO is your target for it. SLA is the promise to customers, with penalties. | Karan Pratap Singh: SLA, SLO, SLI | Wrap-up, monitoring |
| Back-of-the-envelope estimation | Rough QPS, storage and bandwidth to decide if one machine is enough. | ByteByteGo: Estimation | Only when it changes a decision |
Networking and APIs¶
| Concept | What it means | Best free link | Where it comes up |
|---|---|---|---|
| DNS | Turns names into IP addresses. Can also route by geography or weight. | Cloudflare: What is DNS | CDNs, multi-region |
| TCP vs UDP | TCP gives ordered, reliable delivery over a connection. UDP sends packets with no delivery guarantee and less overhead. | Karan Pratap Singh: TCP and UDP | Video calls, games, trading systems |
| HTTP | A stateless request and response protocol: methods, status codes, headers. | MDN: HTTP overview | API design |
| Load balancer | Spreads traffic across servers. L4 routes by IP and port, L7 by HTTP content. Algorithms: round robin, least connections, hashing. | Hello Interview: Networking essentials | Every design |
| Reverse proxy and API gateway | A front door that handles TLS, auth, routing and rate limits before your services. | Hello Interview: API gateway | Microservice entry point |
| CDN | Edge servers cache static content close to users. Push CDNs get files uploaded; pull CDNs fetch on first request. | Cloudflare: What is a CDN | Images, video, static files |
| REST, gRPC, GraphQL | REST for public resources. gRPC for fast internal service calls. GraphQL when clients need flexible fields. Paginate with cursors. | Hello Interview: API design | The API step of every design |
| Long polling, SSE, WebSockets | Long polling holds a request open. SSE pushes from server to client over HTTP. WebSockets are two-way. Pick the simplest that meets the need. | Hello Interview: Networking essentials | Chat, live comments, notifications |
| TLS and mTLS | TLS encrypts traffic. mTLS also proves the client's identity, often between services. | Karan Pratap Singh: SSL, TLS, mTLS | Security requirements |
| OAuth 2.0 and OpenID Connect | OAuth gives an app limited access with a token. OIDC adds login identity on top. | Karan Pratap Singh: OAuth 2.0 and OIDC | "Sign in with Google", third-party APIs |
Data storage¶
| Concept | What it means | Best free link | Where it comes up |
|---|---|---|---|
| SQL vs NoSQL | Relational stores give joins and transactions. Key-value, document, wide-column and graph stores give flexible schemas and easier horizontal scale. Default to Postgres unless you have a reason. | Hello Interview: Key technologies | Choosing the database |
| Data modeling | Pick entities, keys and access patterns first, then the schema. Denormalize read-heavy paths. | Hello Interview: Data modeling | Every design |
| Indexing | B-tree and hash indexes turn full scans into lookups. Writes get slower. Column order in a composite index matters. | Use The Index, Luke | "This query is slow" follow-ups |
| ACID vs BASE | ACID: all-or-nothing, consistent, isolated, durable. BASE: basically available, eventually consistent. | Karan Pratap Singh: ACID and BASE | Booking, payments, inventory |
| Isolation levels | Decide which anomalies concurrent transactions may see (dirty reads, non-repeatable reads, phantoms). | PostgreSQL: Transaction isolation | Two users booking the same seat |
| Replication | Copies of data for availability and read scale: leader-follower, multi-leader, leaderless. The cost is replication lag. | Karan Pratap Singh: Replication | Read-heavy systems, failover |
| Sharding (partitioning) | Split data across machines by key, by range or by hash. Watch for hot keys and cross-shard queries. | Hello Interview: Sharding | Write scale, huge tables |
| Consistent hashing | Servers and keys sit on a ring, so adding a node moves only about 1/N of the keys. Virtual nodes even out the load. | Hello Interview: Consistent hashing | Caches, key-value stores |
| Federation | Split databases by function (users, orders, products) instead of by row. | Karan Pratap Singh: Federation | An early scaling step |
| Blob storage | Keep large files in object storage (S3) and metadata in a database. Clients upload and download directly with presigned URLs. | AWS: Presigned URLs | Dropbox, YouTube, image upload |
| Search (inverted index) | Maps each word to the documents that contain it (Elasticsearch). Keep it in sync from the main database through change data capture or a queue. | Hello Interview: Elasticsearch | Post search, product search |
| Geospatial index | Geohash, quadtree or H3 to find nearby drivers or places fast. | Hello Interview: Proximity search | Uber, Yelp, delivery |
Caching¶
| Concept | What it means | Best free link | Where it comes up |
|---|---|---|---|
| Caching strategies | Cache-aside, write-through, write-behind, refresh-ahead. Caches can sit at the client, CDN, server or database. Main risks: stale data and a thundering herd on expiry. | Hello Interview: Caching | Any read-heavy path |
| Cache eviction | What to drop when memory is full: LRU, LFU, TTL, random. | Redis: Eviction policies | Cache sizing, the LRU LLD question |
Async processing and messaging¶
| Concept | What it means | Best free link | Where it comes up |
|---|---|---|---|
| Message queue | Decouples producers from consumers, absorbs spikes and retries failed work (SQS, RabbitMQ). | AWS: Message queues | Notifications, background jobs, uploads |
| Publish-subscribe | One message goes to every subscriber of a topic. | Karan Pratap Singh: Pub-sub | Notifications, feed fan-out |
| Streams (Kafka) | An append-only, partitioned log. Consumers track offsets, so you get replay and order within a partition. | Hello Interview: Kafka | Analytics, click aggregation |
| Event sourcing and CQRS | Store every change as an event; keep separate models for writes and reads. Rarely needed in a 45-minute round. | Karan Pratap Singh: Event sourcing, CQRS | Ledgers, audit trails |
Consistency and distributed systems¶
| Concept | What it means | Best free link | Where it comes up |
|---|---|---|---|
| CAP theorem | During a network partition you choose consistency or availability. Partition tolerance is not optional. | Hello Interview: CAP | Non-functional requirements, per feature |
| PACELC | If there is a Partition, pick A or C. Else, pick Latency or Consistency. | Karan Pratap Singh: PACELC | Why eventually consistent stores are fast |
| Consistency models | Strong (linearizable), sequential, causal, read-your-writes, eventual: what a reader may see after a write. | Jepsen: Consistency models | Feeds vs bank balances |
| Consensus (Raft) | How replicas agree on one leader and one log despite failures. | Raft and its visual walkthrough | Leader election, config stores |
| Quorum (N, W, R) | With N replicas, if writes wait for W and reads ask R, and W + R > N, every read overlaps the latest write. | ByteByteGo: Key-value store | Key-value stores |
| Distributed transactions | Two-phase commit locks every participant until a coordinator commits. A saga runs local transactions and undoes them with compensating steps on failure. | Karan Pratap Singh: Distributed transactions, microservices.io: Saga | Orders and payments across services |
| Idempotency | Repeating a request has the same effect as doing it once. The client sends an idempotency key; the server stores the result. | Stripe: Idempotency | Payments, any retry |
| Unique ID generation | Options: database auto-increment, UUID, ticket server, Snowflake (timestamp plus machine ID plus sequence, a sortable 64-bit ID). | Twitter: Announcing Snowflake | URL shortener, messages |
| Write-ahead log | Write each change to an append-only log before applying it, so you can recover after a crash. | Wikipedia: Write-ahead logging | Durability questions |
Reliability and operations¶
| Concept | What it means | Best free link | Where it comes up |
|---|---|---|---|
| Rate limiting | Caps requests per client. Algorithms: token bucket, leaking bucket, fixed window, sliding window log or counter. | ByteByteGo: Rate limiter | API gateways, abuse |
| Retries with backoff and jitter | Wait longer after each failed try and add randomness, so clients do not retry in lockstep. | AWS: Exponential backoff and jitter | Any external call |
| Circuit breaker | Stop calling a failing dependency and fail fast; probe again later. | Karan Pratap Singh: Circuit breaker | Payment providers, third-party APIs |
| Service discovery | How services find healthy instances of each other (a registry or DNS). | Karan Pratap Singh: Service discovery | Microservices |
| Observability | Metrics, logs and traces. Watch the four golden signals: latency, traffic, errors, saturation. | Google SRE: Monitoring distributed systems | Wrap-up of every design |
| Disaster recovery | Backups and failover plans. RTO is how long recovery may take; RPO is how much data you can afford to lose. | Karan Pratap Singh: Disaster recovery | Multi-region, durability |
| Monolith vs microservices | Separate deployable services trade simplicity for team and scaling independence. Default to fewer services in an interview. | Martin Fowler: Microservices | "Why not one service?" |
Probabilistic data structures¶
| Concept | What it means | Best free link | Where it comes up |
|---|---|---|---|
| Bloom filter | Answers "definitely not seen" or "maybe seen" in very little memory. No false negatives. | Wikipedia: Bloom filter | Web crawler URL dedup |
| Count-min sketch | Approximate counts for items in a stream in fixed memory. It can overcount, never undercount. | Wikipedia: Count-min sketch | Top-K, heavy hitters |
Building blocks to name in an interview¶
Name a real technology only if you can explain how it works inside. OpenAI interviewers in particular drill into internals (company page).
| Technology | Reach for it when | Free link |
|---|---|---|
| PostgreSQL | You need transactions, joins and constraints. The default choice. | Isolation levels |
| Redis | You need a cache, counters, rate limiting, leaderboards (sorted sets) or short-lived locks. | Hello Interview: Redis, sorted sets |
| Kafka | You need a durable event log with replay, many consumers, and order per key. | Hello Interview: Kafka, Kafka intro |
| Cassandra | Very high write volume with a known query pattern per table. | Hello Interview: Cassandra |
| DynamoDB | A managed key-value store with predictable latency on AWS. | Hello Interview: DynamoDB |
| Elasticsearch | Full-text search, filters and relevance ranking. | Hello Interview: Elasticsearch |
| S3 (object storage) | Files, images, video, backups. | AWS: S3 user guide |
| SQS or RabbitMQ | A work queue where each message is handled once by one worker. | AWS: Message queues |
| API gateway | Auth, rate limits and routing in one place. | Hello Interview: API gateway |
Trade-offs you will be asked to defend¶
| Decision | Pick the first when | Pick the second when |
|---|---|---|
| SQL vs NoSQL | You need transactions, joins or strict constraints (payments, bookings) | You need huge write scale on a simple, known access pattern |
| Cache-aside vs write-through | Reads dominate and slightly stale data is fine | Reads must see the latest write right after it happens |
| Queue (SQS) vs stream (Kafka) | Each job is done once by one worker | Many consumers need the same events, or you need replay |
| Fan-out on write vs on read (feeds) | Most users have few followers: precompute each feed | Celebrity accounts: merge their posts at read time |
| Long polling vs SSE vs WebSockets | Rare updates, simplest client | Server-to-client updates only / two-way, low-latency messages |
| Strong vs eventual consistency | Money, inventory, seat holds | Likes, view counts, feeds |
| Hash vs range sharding | Even spread of load | Range scans (time ranges, sorted IDs) |
| REST vs gRPC | Public or browser-facing APIs | Internal service-to-service calls |
| Monolith vs microservices | Small team, early product, interview default | Teams that must deploy and scale parts independently |
| Retry vs fail fast | The call is idempotent and the failure looks temporary | The dependency is down: open the circuit breaker |
For a URL shortener, know why Hello Interview's breakdown picks a 302 redirect over a 301. Browsers cache a 301, so later clicks skip your server. A 302 sends every click through you, so you can count clicks and change or expire links.
Latency numbers to know¶
From the System Design Primer. Its sources include Jeff Dean's 2009 talk. Hardware is faster now; this interactive chart shows the numbers by year. Learn the orders of magnitude, not the exact values.
| Operation | Time | What it tells you |
|---|---|---|
| L1 cache reference | 0.5 ns | |
| L2 cache reference | 7 ns | |
| Main memory reference | 100 ns | Memory is fast: cache hot data in RAM |
| Compress 1 KB with Zippy | 10 us | Compress before sending over the network |
| Send 1 KB over a 1 Gbps network | 10 us | |
| Read 4 KB randomly from SSD | 150 us | |
| Read 1 MB sequentially from memory | 250 us | |
| Round trip within one datacenter | 500 us | Each extra service hop costs real time |
| Read 1 MB sequentially from SSD | 1 ms | |
| HDD seek | 10 ms | Avoid random disk reads |
| Read 1 MB over a 1 Gbps network | 10 ms | |
| Read 1 MB sequentially from HDD | 30 ms | Read sequentially, not randomly |
| Packet California to Netherlands and back | 150 ms | Put servers and CDNs near users |
Handy throughput figures from the same table: HDD 30 MB/s, 1 Gbps Ethernet 100 MB/s, SSD 1 GB/s, main memory 4 GB/s. You get about 2,000 round trips a second inside one datacenter and 6 to 7 round trips a second across the world.
Estimation cheat sheet¶
Estimate only when the number changes a decision, for example whether a top-K heap fits on one machine. Hello Interview recommends this, and Meta interviewers may withhold numbers so you have to set your own. Say which one you are doing out loud.
Powers of two
| Power | Approximate value | Size |
|---|---|---|
| 2^10 | 1 thousand | 1 KB |
| 2^20 | 1 million | 1 MB |
| 2^30 | 1 billion | 1 GB |
| 2^40 | 1 trillion | 1 TB |
| 2^50 | 1 quadrillion | 1 PB |
Time and traffic conversions (arithmetic)
| Per day | Average per second |
|---|---|
| 1 day | 86,400 seconds, call it 10^5 |
| 1 million requests a day | about 12 QPS |
| 10 million requests a day | about 120 QPS |
| 100 million requests a day | about 1,200 QPS |
| 1 billion requests a day | about 12,000 QPS |
One month is about 2.6 million seconds and one year about 31.5 million seconds. ByteByteGo's worked example uses peak QPS = 2 x average.
Availability budget (arithmetic)
| Availability | Downtime per year | Downtime per month |
|---|---|---|
| 99% | 3.65 days | 7.3 hours |
| 99.9% | 8.8 hours | 43.8 minutes |
| 99.99% | 52.6 minutes | 4.4 minutes |
| 99.999% | 5.26 minutes | 26 seconds |
Formulas
- Storage per day = writes per day x size per item. Multiply by 365, by years kept, and by the number of replicas.
- Bandwidth = QPS x bytes per request (do reads and writes separately).
- Servers = peak QPS / QPS one server handles. State your per-server number as an assumption.
- Cache size = hot items x item size. A common assumption is that a small share of items gets most reads; say the share you assume.
ByteByteGo's worked example (illustrative numbers, not real Twitter data): 300 million monthly users, 50% daily, so 150 million daily users posting 2 tweets a day. Tweet QPS = 150M x 2 / 86,400, about 3,500, with peak about 7,000. If 10% of tweets carry 1 MB of media, that is 30 TB a day, about 55 PB over 5 years (chapter).
Copy this template into your notes and fill it for every practice problem:
ASSUMPTIONS (say each one out loud)
Daily active users: [ ]
Actions per user per day: [ ] writes, [ ] reads
Size per item: [ ] bytes
Retention: [ ] years, replicas: [ ]
TRAFFIC
Write QPS = DAU x writes / 86,400 = [ ] peak (x2) = [ ]
Read QPS = DAU x reads / 86,400 = [ ] peak (x2) = [ ]
Read:write ratio = [ ]
STORAGE
Per day = writes per day x size = [ ]
Total = per day x 365 x years x replicas = [ ]
DECISION THIS CHANGES
[ ] fits on one machine? [ ] needs sharding? [ ] cache worth it?
Rules that keep estimation short:
- Round hard: 86,400 becomes 100,000 and 2.6 million becomes 2.5 million.
- Write every assumption on the board with units.
- Stop as soon as you have the number that drives the decision.
Week 1 checklist¶
- I can explain each concept in the scale, networking and data storage tables in one sentence.
- I can draw cache-aside and explain when the cache goes stale.
- I can explain sharding vs replication and consistent hashing on a whiteboard.
- I can recite the latency orders of magnitude (memory, SSD, datacenter round trip, cross-continent).
- I can convert requests a day to QPS and estimate storage for 5 years without a calculator.
- I can defend each row of the trade-off table with one example.