Context
Sector: Commercial ticketing. Role: Software Engineer, DevOps. Environment: Mission-critical ticketing services subjected to massive concurrency bursts during on-sale events and multi-year migration backfills.
Challenge
- Stabilize streaming and storage layers under burst-traffic load that overwhelmed the prior architecture.
- Reduce scale-time and infrastructure cost without sacrificing reliability during on-sale events.
- Diagnose and remediate cross-system bottlenecks across Kafka and downstream storage.
Architecture
Compute & Scaling
- Migrated mission-critical ticketing services to lower-cost spot-instance compute with automated scale-out.
- Reduced node scale-time from 15 minutes to approximately 2 minutes (largely consumer group rebalance).
Streaming & Storage
- Diagnosed downstream DynamoDB write bottleneck causing Kafka Streams backpressure during on-sale bursts.
- Prototyped Cassandra and ScyllaDB alternative storage backends; designs subsequently adopted by the team.
Outcomes
- Sustained 42M messages per minute at peak during on-sale events and multi-year migration backfills.
- Reduced node scale-time from 15 minutes to approximately 2 minutes.
- Reduced infrastructure cost by 40% through migration to lower-cost spot-instance compute.
- Eliminated downstream storage bottleneck through prototyped alternative storage architecture.