System Design #22: Design a Stock Broking System

SYSTEM DESIGN #22 Β· INTERVIEW GUIDE

Design a Stock Broking System

Stock markets are the fastest, most demanding distributed systems on the planet: the NYSE processes 5 million orders per second, each with sub-microsecond matching latency, perfect ordering guarantees, and zero tolerance for data loss. This Stock Broking System covers the complete trade lifecycle: a FIX Protocol gateway for institutional order entry, a real-time risk engine that checks margin, position limits, and circuit breakers in under 500 microseconds, a Red-Black tree order book that maintains price-time priority, and a single-threaded matching engine that eliminates lock contention entirely. Every matched trade is written to an immutable Kafka log before acknowledgement, and settlement flows through a T+1 DTCC integration. After this guide you will know how to design the kind of ultra-low-latency financial system that moves trillions of dollars every trading day.

FIX ProtocolRed-Black TreeKafkaLMAX DisruptorDTCC

πŸ’‘

The Gist β€” What Problem Are We Solving?

Where software directly controls real money β€” correctness is non-negotiable

When you tap ‘Buy 100 shares of AAPL’, a chain of correctness-critical operations must execute flawlessly: risk checks, order book entry, price-time priority matching, trade reporting, and settlement. The matching engine that pairs buyers and sellers is single-threaded for determinism β€” two identical order sequences must always produce identical results. Any correctness bug can cause financial losses, regulatory fines, and destroyed trust.

πŸ’¬

Think of it as an ultra-high-stakes scoreboard where every entry is permanent, legally binding, and audited by regulators for 6 years β€” and where microseconds determine fairness.

βœ…

Functional Requirements

These are the capabilities the system must deliver β€” what users and operators can actually do with it.

πŸ“‹
Order Management

πŸ“‹Submit, cancel, modify orders (market, limit, stop-limit, trailing stop)

🏦
Order Book

🏦Price-time priority matching; best bid/ask display; depth-of-market

βš™οΈ
Matching Engine

βš™οΈExecute trades at crossing prices; FIFO within same price level

πŸ›‘οΈ
Risk Engine

πŸ›‘οΈPre-trade: margin check, position limits, pattern day trader rule

πŸ“Š
Market Data

πŸ“ŠReal-time price, volume, order book; OHLCV candles; historical data

⚑

Non-Functional Requirements

These define how well the system must perform β€” the quality attributes that separate a toy from a production system.

⚑ Matching Latency

⚑<10 microseconds order-to-match

πŸ”’ Correctness

πŸ”’Deterministic β€” identical order sequence always produces identical trades

πŸ’Ύ Durability

πŸ’ΎKafka immutable event log; zero trade loss

πŸ›‘οΈ Risk

πŸ›‘οΈPre-trade risk checks in <1ms β€” in-process, no DB lookup

πŸ“‹ Compliance

πŸ“‹6-year audit log (FINRA); surveillance for spoofing/wash trading

πŸ“Š

Key Metrics β€” The Numbers That Define This System

The headline numbers to know cold β€” and be ready to explain how each one is achieved.

<10ΞΌs
matching latency
Single-threaded
determinism
Kafka
source of truth
T+1
settlement
6yr
audit retention
πŸ—οΈ

System Architecture Diagram

Full data flow from source to serving. Each layer scales independently.

Ingestion
Stock Broking Flow
Client Order

β†’
FIX Gateway

β†’
Risk Engine
in-process, <1ms

β†’
Order Book
in-memory Red-Black tree

↓
Processing
Matching Engine
single-threaded

β†’
Kafka
immutable trade log

β†’
Market Data Feed + Settlement Service + Position Store
Postgres

πŸ—ΊοΈ

End-to-End User Journey

Trace a single request end-to-end β€” the story interviewers want you to tell fluently.

1
Investor submits limit order

β€” Client app: Buy 100 AAPL @ $180. FIX protocol message to order gateway.

2
Risk checks

β€” In-process risk engine: margin available? Position limit OK? PDT rule? Wash trade? All pass in <500ΞΌs.

3
Order enters book

β€” Order added to AAPL order book: bid side, sorted by price DESC then time ASC. Position: 3rd in queue at $180.

4
Matching event

β€” Seller submits: Sell 100 AAPL @ $179. Best bid ($180) β‰₯ ask ($179) β†’ match. FIFO: first buyer at $180 matched first.

5
Trade executed

β€” Clearing price = $179.50 (second-price). Trade event published to Kafka. Both parties’ positions updated.

6
Market data published

β€” Trade price and volume broadcast to all subscribers via market data feed (sub-millisecond). Order book depth updated.

πŸ”­

High-Level Design β€” Component Breakdown

Core components β€” each with a single, well-defined responsibility. The key architectural insight: each layer scales independently, and failure in one component is isolated from the rest.

1
FIX Gateway
β†’
2
Risk Engine
β†’
3
Order Book
β†’
4
Matching
β†’
5
Kafka
β†’
6
Mkt Data Feed
1 β€” FIX Gateway

Terminates WebSocket connections. Maintains connection registry in Redis. Routes incoming messages to target connection via gRPC. Scales horizontally with sticky routing via consistent hash on user_id.

2 β€” Risk Engine

Validates every order against position limits, margin requirements, and fat-finger rules in <500Β΅s using pre-loaded in-memory tables. Branch-free comparison logic β€” no I/O on the critical path.

3 β€” Order Book

Red-Black tree per side (bid/ask) keyed by price level. FIFO deque per price level for time priority. Best bid/ask: O(1). Limit insert: O(log N). Pre-allocated object pool β€” zero malloc during trading.

4 β€” Matching

Scores and ranks candidates by ETA, rating, and acceptance rate. Dispatches offers via WebSocket with 15-second acceptance window. Falls back to next candidate on decline. P95 match time <30 seconds.

5 β€” Kafka

Distributed event bus with RF=3 for durability. Partitioned by user_id_hash for per-user ordering. LZ4 compression reduces storage cost by 60%. Exactly-once semantics via idempotent producers and transactional consumers.

6 β€” Mkt Data Feed

Handles responsibilities for the Mkt Data Feed layer. Designed for independent horizontal scaling β€” additional instances added without architectural changes. Communicates asynchronously with adjacent components to maximise throughput and fault isolation.

πŸ”¬

Low-Level Design β€” Deep Dives

Deep dives worth explaining in detail in any senior engineering interview. For each: know the data structure, the algorithm, the why, and the trade-off you made.

1 β€” FIX Protocol Gateway
Order Entry Β· <500Β΅s

FIX 4.4/5.0 gateway accepts institutional orders via persistent TCP sessions. FIX message parsing uses a pre-allocated field table (no dynamic allocation). Session management: heartbeat every 30 seconds, sequence number resync on gap. Order validated: symbol exists, quantity > 0, price in tick-size increments. Valid orders published to disruptor ring buffer (LMAX pattern) for lock-free inter-thread communication with risk engine. Gateway latency: <50Β΅s for order acknowledgement.

// FIX tag parsing (zero-allocation)
struct FixMessage {
char symbol[8];
double price;
int64_t quantity;
char side; // ‘1’=buy, ‘2’=sell
char order_type; // ‘1’=market, ‘2’=limit
};
void parse_fix(const char* raw, FixMessage* msg) {
// Direct field extraction by FIX tag offsets
}

2 β€” Risk Engine
Pre-computed Margin Β· <500Β΅s

Risk engine runs on the critical path between gateway and order book. Checks: (1) Position limit: current_position + order_qty ≀ max_position (memory-mapped position table, O(1)). (2) Margin: (current_exposure + order_value) ≀ account_balance Γ— margin_factor (pre-computed per instrument at market open). (3) Fat-finger: order_qty ≀ 10Γ— ADV (average daily volume). All checks are branch-free comparisons β€” no I/O. Risk rejection is O(1). SIMD instructions for portfolio-level checks.

bool check_risk(const Order& o, const Account& acct) {
// All data pre-loaded to CPU cache
auto pos = position_table[o.symbol_idx];
auto margin = margin_table[o.symbol_idx];
return (pos + o.qty <= MAX_POSITION[o.symbol_idx]) && (acct.exposure + o.value <= acct.balance * margin); }

3 β€” Red-Black Tree Order Book
O(log N) insert Β· O(1) best

Order book: two std::map (Red-Black tree) for bids (descending) and asks (ascending). Each price level: std::deque of orders (FIFO within price). Best bid: rbegin(), best ask: begin() β€” O(1). Limit order insert: O(log N) tree insert. Market order match: O(K) where K=orders filled. Order cancellation: O(log N) tree lookup + O(1) deque remove via order_id β†’ iterator map. Memory: pre-allocated pool (object pool pattern) β€” no malloc during trading hours.

struct PriceLevel {
double price;
std::deque orders; // FIFO
int64_t total_qty;
};
std::map> bids; // descending
std::map asks; // ascending

Order& best_bid() { return bids.begin()->second.orders.front(); }

4 β€” Matching Engine
Single-thread Β· Price-Time Priority

Single-threaded matching loop processes orders from disruptor ring buffer in strict FIFO order. For each incoming order: iterate opposite side of book from best price outward, fill against resting orders until order is fully filled or no more matching orders. Each fill generates an Execution Report written to Kafka (immutable audit log) via async publisher. Matched trade published to settlement service. Throughput: 500K orders/sec on a single core (no lock contention).

void match_order(Order& incoming) {
auto& opposite = (incoming.side == BUY) ? asks : bids;
for (auto& [price, level] : opposite) {
if (!price_matches(incoming, price)) break;
while (!level.orders.empty() && incoming.qty > 0) {
fill(incoming, level.orders.front());
if (level.orders.front().qty == 0)
level.orders.pop_front();
}
}
}

βš–οΈ

Trade-offs & Decision Log

Every senior interview comes down to these decisions. Know the exact trade-off, the reasoning, and the specific numbers that justify each choice.

βš–οΈ Single-Threaded vs Multi-Threaded Matching Engine

βœ“
Single-Threaded βœ… Chosen
  • Zero lock contention β€” no mutex overhead
  • Deterministic order: all orders processed in strict arrival sequence
  • Cache-friendly β€” entire order book fits in L3 cache (instrument-partitioned)
  • Cannot parallelise β€” throughput capped by single core speed (~500K orders/sec)
β†’
Multi-Threaded (lock-based)
  • Multiple cores available for higher theoretical throughput
  • Lock contention at high load degrades performance below single-thread
  • Non-deterministic order creates audit and regulatory issues
  • Deadlock risk between bid/ask queues

πŸ’‘

Decision: One matching thread per instrument partition; 500 instruments Γ— 1M orders/sec/thread = 500M orders/sec total throughput

βš–οΈ Write-Ahead Log vs Event Sourcing for Order State

βœ“
Kafka Immutable Event Log βœ… Chosen
  • Immutable audit trail required by SEC/FCA regulations
  • Replay entire order book from log for disaster recovery
  • Any downstream system (risk, settlement, reporting) subscribes independently
  • Log compaction reduces storage while preserving final state
β†’
Postgres WAL (traditional DB)
  • ACID transactions with row-level locking
  • WAL not designed for consumption by downstream systems
  • Replication lag causes inconsistency between risk and matching engine
  • Point-in-time recovery possible but slow

πŸ’‘

Decision: Kafka as system of record for all order events; Postgres as materialised view of current order state; settlement reads from Kafka directly

🎯

Interview Questions β€” Answered
The exact questions interviewers ask β€” with production-grade answers

Q1
How does the risk engine check margin in under 500 microseconds?

Ultra-low latency achieved by three techniques: (1) In-process memory: position data, margin requirements, and account balances are pre-loaded into a memory-mapped file on the risk engine’s CPU cache β€” no network round-trip. (2) Branch-free logic: risk checks use SIMD instructions for portfolio calculations β€” vectorised over all positions simultaneously. (3) Pre-computed margin: margin requirements are pre-calculated for each instrument at market open and updated every 60 seconds (not per-order). Per-order check is: current_position Γ— pre_computed_margin ≀ available_margin. This reduces the per-order check to a single comparison. Orders failing risk are rejected before ever reaching the order book β€” typical risk check latency: 15-50 microseconds.

Q2
How is the order book implemented for maximum throughput?

Red-Black tree per side (bid/ask) keyed by price level, each price level containing a doubly-linked list of orders (FIFO priority within price): (1) Best bid/ask retrieval: O(1) β€” rbegin()/begin() of tree. (2) Limit order insertion: O(log N) where N = number of price levels. (3) Market order matching: O(K) where K = number of orders matched. (4) Order cancellation: O(log N) to find price level + O(1) with order pointer stored in order lookup map. Memory layout: orders allocated from a pre-allocated pool (no malloc during trading) β€” allocation is O(1). Entire order book for a single instrument fits in L3 cache (~8MB on modern server). This gives 500K orders/sec on a single core.

Q3
How is T+1 settlement handled for matched trades?

Settlement pipeline: (1) Matched trade event published to Kafka immediately with trade_id, buyer, seller, quantity, price, timestamp. (2) Settlement service consumes trade events and creates settlement obligations in a Postgres ledger: {trade_id, counterparty, net_position, settle_date = trade_date + 1}. (3) Netting: all trades for the same instrument between the same counterparties on the same date are netted to a single position β€” reduces settlement volume by ~80%. (4) T+1: at 16:00 on settlement date, net positions are submitted to DTCC (via FIX) for central clearing. (5) Fails management: if counterparty fails to deliver, automatic buy-in process is triggered. (6) Reconciliation: all positions reconciled against DTCC confirmation file within 30 minutes of market close.

System Design Series Β· Every Tuesday & Thursday

Level up your system design interviews

Each post covers Gist, Functional & Non-Functional Requirements, Key Metrics, System Diagram, User Journey, HLD, LLD, and Trade-offs & FAQs.

Subscribe to never miss a post β†’


Categories: System Design

Tags: , , , , , ,

Leave a Reply

Discover more from Cloud Wizard Inc.

Subscribe now to keep reading and get access to the full archive.

Continue reading