System design interview ยท Platform

Design a notification system

"Just call Twilio" โ€” famous last words. The API call is 1% of the system. The other 99% is deciding who gets what, in which order, how often โ€” and what you do when the carrier says "slow down or you're banned."

๐ŸŽฏ The takeaway, first

Priorities plus per-channel rate limits are the whole game. A password reset and a promo blast must never share a queue. The design: priority queues feeding token-bucket rate limiters per channel, with retries on a dead-letter loop. Critical alerts jump the queue; marketing emails wait their turn. Think of it like an ER triage desk in front of three operating rooms (SMS, push, email) โ€” each room only admits patients so fast, and the gunshot wound goes before the checkup. If your design treats every notification identically, it's wrong.

๐Ÿšซ Misconception, busted

"A notification system is just API calls to Twilio/SES." The send call is the easy part. The other 90% is templates (rendered, localized, versioned), user preferences + quiet hours (never text someone at 3am about a sale), dedup (don't send the same alert twice), priority ordering, per-carrier rate limits, and retry with backoff. Twilio is a doorknob; you're being asked to design the building.

Requirements

Say these out loud before drawing a single box.

Functional

  • Send via SMS, push, email
  • Templates with variables, per-locale
  • Priorities: critical / high / normal / low
  • User preferences + quiet hours
  • Dedup via idempotency keys; delivery status tracking

Non-functional

  • 10M+/day (~116/s avg, ~10ร— peak)
  • Critical delivered in seconds; low may wait minutes
  • Never breach carrier rate limits (ban = outage)
  • At-least-once delivery with idempotent sends
  • 99.9% availability; provider failover

Back-of-the-envelope math

Assumptions labeled.

WhatAssumptionMath
Throughput10M/dayโ‰ˆ 116/s avg, โ‰ˆ 1.2K/s peak
SMS bottleneckcarriers allow ~10โ€“50 msg/s per long codepeak SMS needs a pool of sending numbers, not one
PushFCM/APNs absorb fan-outour bottleneck is our throughput, not theirs
EmailSES default ~14/srequest a limit raise on day one; batch low-priority email
Queue depthpeak 1.2K/s ร— 30s of retry backlogโ‰ˆ 36K in flight โ€” trivial for Kafka
Delivery log10M/day ร— 500 B, 90-day TTLโ‰ˆ 450 GB โ€” small; keep it, debugging needs it
Go deeper: why is the carrier ban the scariest failure?

Rate limits here aren't your policy โ€” they're the carrier's. Blow past them and the carrier throttles or blacklists your sending numbers, which is an outage you can't fix by scaling: you have to wait out the ban or rotate numbers. That's why the token bucket sits before the provider call, sized conservatively (e.g. 80% of the carrier's stated limit), per sending number. Throughput you leave on the table is the price of never being banned.

Architecture

The centerpiece: priority queues โ†’ per-channel limiters โ†’ workers.

flowchart LR
    API[Notification API] --> VAL[Validation
prefs, quiet hours, dedup] VAL --> TPL[Template Service
render + localize] TPL --> QC[Priority Queues
Kafka: critical/high/normal/low] QC --> DISP[Dispatcher
strict priority drain] DISP --> RL1[Rate Limiter
SMS token bucket] DISP --> RL2[Rate Limiter
Push token bucket] DISP --> RL3[Rate Limiter
Email token bucket] RL1 --> WS[SMS Workers
Twilio] RL2 --> WP[Push Workers
FCM + APNs] RL3 --> WE[Email Workers
SES] WS --> LOG[(Delivery Log)] WP --> LOG WE --> LOG LOG --> RETRY[Retry with backoff
โ†’ DLQ after N tries] RETRY --> QC

Requests are validated (preferences, quiet hours, idempotency-key dedup), rendered from templates, and enqueued by priority. The dispatcher always drains critical before high before normal before low. Each channel has its own token bucket in Redis sized under the provider's limit โ€” a notification only leaves when its channel has a token. Failures go to a retry loop with exponential backoff, then a dead-letter queue for humans.

Go deeper: strict priority vs weighted fair queuing?

Strict priority can starve low-priority forever during a sustained critical flood (e.g. an outage paging everyone). Production systems usually run weighted draining โ€” critical gets 70% of dispatcher slots, high 20%, normal 8%, low 2% โ€” so the promo email eventually goes out even during an incident. Mention this as the v2; strict priority is the correct v1 because it's explainable in an interview.

Component deep-dives

Send โ€” validate, render, enqueue

sequenceDiagram
    participant C as Caller service
    participant A as Notification API
    participant P as Prefs + Dedup (Redis)
    participant T as Template Service
    participant Q as Kafka priority topic
    C->>A: POST /v1/notifications {user, template, vars, priority, channels}
    A->>P: idempotency key seen? prefs allow? quiet hours?
    alt duplicate or opted out or quiet hours
        P-->>A: reject with reason
        A-->>C: 200 {status: skipped, reason}
    else ok
        A->>T: render template + locale
        T-->>A: rendered bodies
        A->>Q: produce to {priority} topic
        A-->>C: 202 {notification_id, status: queued}
    end

Dispatch โ€” the token bucket gate

sequenceDiagram
    participant D as Dispatcher
    participant B as Token Bucket (Redis)
    participant W as Channel Worker
    participant V as Provider (Twilio)
    D->>D: pick highest-priority non-empty queue
    D->>B: take token for sms:+1555 pool?
    alt token available
        B-->>D: granted
        D->>W: send job
        W->>V: API call
        alt success
            V-->>W: delivered
        else 429 / 5xx
            W->>D: retry with backoff + jitter
        end
    else bucket empty
        B-->>D: denied โ€” requeue, try next channel
        Note over D: low-priority waits here.
This is throttling working as designed. end

API + data model

POST /v1/notifications
  { "user_id": "u_123", "template_id": "order_shipped",
    "vars": { "order": "A-9918" }, "priority": "normal",
    "channels": ["push", "email"],
    "idempotency_key": "ord-9918-shipped" }
  โ†’ 202 { "notification_id": "n_...", "status": "queued" }
  โ†’ 200 { "status": "skipped", "reason": "quiet_hours" }

GET  /v1/notifications/{id}   โ†’ delivery status + attempts
templates(user-facing) { template_id, channel, locale, body, version }
user_prefs             { user_id PK, sms/push/email_enabled, quiet_hours_start/end }
notification_log       { notification_id PK, user_id, priority, channel, status, attempts, created_at }
idempotency_keys       { key PK, notification_id }  -- 24h TTL in Redis

Trade-offs

DecisionOption AOption BPick
Queue shapeTopic per priority (4 topics)Single queue, priority fieldPer-priority topics โ€” the dispatcher logic stays dumb and fast
Delivery semanticsAt-least-once + idempotency keysExactly-onceAt-least-once โ€” providers retry; dedup on our side
Email sendingImmediate per notificationBatched (digest every 15 min for low)Batched for low โ€” cost and reputation win
SMS providerSingle (Twilio)Primary + failoverFailover โ€” provider outages are a matter of time

Failure modes

Provider outage (Twilio down)Circuit breaker โ†’ failover provider; queue depth absorbs hours of outage
Rate-limit breach โ†’ carrier banToken buckets at 80% of carrier limits, per sending number; alerts at 70% utilization
Retry storm after partial outageExponential backoff + jitter; cap retries; DLQ after N attempts
Template bug spams usersCanary: new templates go to 1% first; global kill switch per template
Priority inversion: incident floods critical queueWeighted draining (v2) so low-priority isn't starved forever; separate on-call paging path

What I'd actually build

Opinionated. Steal this for the interview.

Stack: Go API, Kafka with 4 priority topics, Redis token buckets (per channel, per sending number) + idempotency keys + prefs cache. One worker pool per channel with circuit breakers; Twilio primary with a backup SMS provider; FCM + APNs for push; SES for email with low-priority batching. Quiet hours enforced at dispatch, not at enqueue (a queued promo hitting quiet hours waits). Dashboards: delivery rate per channel, bucket utilization, queue age per priority, DLQ depth. Pager fires on bucket saturation, not just errors.

Interview tips

๐Ÿ”ฌ Interactive: notification pipeline

Compose notifications with priorities, then fire a burst and watch the pipeline triage: critical jumps the queue, low-priority waits for tokens.

๐Ÿ”ด critical 0

๐ŸŸ  high 0

๐Ÿ”ต normal 0

โšช low 0

โ–ผ dispatcher: strict priority โ–ผ

๐Ÿ’ฌ SMS 3/s ยท pool of numbers

๐Ÿ”” Push 20/s ยท FCM/APNs

โœ‰๏ธ Email 8/s ยท SES

โ–ผ workers โ–ผ

๐Ÿ’ฌ SMS worker โ€” recently sent

๐Ÿ”” Push worker โ€” recently sent

โœ‰๏ธ Email worker โ€” recently sent

0sent
0throttle waits
0waiting in queues

Watch: the ๐Ÿ”ด critical items clear instantly while โšช low items pile up โ€” and when the SMS bucket runs dry, even high-priority SMS waits. That's throttling working as designed, not a bug.

v2026.10.03-01