Kafka vs RabbitMQ for 100K events/sec with a 5-person team?
Performance ↔ Simplicity
Context
We need to process 100K events/sec.
Context
We're replacing a cron-driven batch pipeline with an event-driven one. Producers are ~30 microservices emitting order, inventory and shipment events. Peak load is around 100K events/sec, average closer to 15K.
Consumers include a fraud scorer, a search indexer and an analytics sink. The analytics team wants to replay the last 7 days when they ship a new model.
Constraints
- 5 backend engineers, nobody has run Kafka in production
- ~$2K/month infrastructure budget for messaging
- Everything is on AWS already
- 99.99% availability target for the order path
Options
Option A — Kafka (probably MSK): partitioned log, consumer groups, retention-based replay.
Option B — RabbitMQ (Amazon MQ or self-hosted): mature routing, simpler mental model, per-message acks.
What I want to understand
Is replay a strong enough requirement to justify Kafka's operational overhead for a team this size? Or would RabbitMQ plus an S3 archive get us 90% of the way?
Constraints
- Team
- 5 engineers
- Budget
- $2K/month
- Cloud
- AWS
- SLA
- 99.99%
KafkavsRabbitMQ
Community verdict
10 engineers · 6 opinions
With these constraints, what would you choose?
One choice per engineer. You can change it any time.
Simplicity →
Trade-offs
Dimensions
- Throughput
- 5Kafka scores 5 of 53RabbitMQ scores 3 of 5
- Simplicity
- 2Kafka scores 2 of 54RabbitMQ scores 4 of 5
- Ordering
- 5Kafka scores 5 of 5
Community discussion
6 comments
Have you made this decision in production? Share your reasoning.
Sign in to commentWhichever you pick, put a thin internal publishing library in front of it. We switched brokers once and the library made it a 3-week project instead of a 3-month one.
Strict per-entity ordering is the hidden requirement here. Order events for the same order ID must be processed in sequence; Kafka gives you that with a partition key. With RabbitMQ you either accept a single consumer per queue or build consistent-hash exchanges and hope nobody re-shards.
For 5 engineers I'd start with RabbitMQ (quorum queues) and treat the analytics replay as a separate concern: tee events to Kinesis Firehose → S3. Your order path stays simple and the analytics team gets Parquet they can query directly.
Kafka is the right answer at 50 engineers. At 5 it's a tax on every feature.
Replay is the deciding factor. With RabbitMQ, once a message is acked it's gone — you end up building a second system (S3 archive + re-publisher) to fake retention, and that's more operational surface, not less.
MSK Serverless removes most of the broker babysitting. Your real learning curve is partition keys and consumer lag, not ZooKeeper.
MSK Serverless pricing gets painful at sustained 100K/sec though. Did you model partition-hours plus throughput? We were at ~3x the estimate.
Fair — at that volume provisioned MSK with 3 brokers is cheaper. Still fits a $2K budget if events are under ~1KB.