Problem
Mission-critical asynchronous payment processing and telemetry webhook dispatches require at-least-once delivery guarantees even in the face of sudden node restarts, network partitions, and noisy neighbor spikes.
Role
Designed and implemented the core message broker, consensus coordinator, and backpressure scheduling pipeline.
Architecture & Stack
- Raft Consensus Group: 5-node consensus protocol managing partition elections and dynamic leader assignment with sub-100ms failover.
- Predictive Backpressure Scheduler: Rate-limits external third-party webhook sinks using adaptive sliding token buckets.
- Write-Ahead Logging (WAL): Disk-backed durable WAL in Go paired with memory-mapped queues for microsecond dispatcher throughput.
Outcome & Metrics
- Benchmarked at 100,000 messages/sec dispatch rate on modest 4-core worker nodes with p99 latency under 1ms.
- 0 message loss during chaos injection tests simulating partition failures and killed leaders.