arosplatforms™AI consultancy
ar
← All articles

Match Throughput and Team Size: Kafka vs Kinesis for Engineers

Match Throughput and Team Size: Kafka vs Kinesis for Engineers

Kafka and Kinesis throughput title card

Choose Amazon Kinesis when you run a small team or sustained ingest under roughly 50 megabytes per second and want AWS-native integration with minimal operations work. Choose Apache Kafka, self-hosted or through Amazon MSK, when you need high throughput, multi-region replication, sub-100ms p99 latency, long retention, or exactly-once processing across streams. Amazon MSK sits between the two: Kafka’s primitives with less broker babysitting. The real trade-off is control and cost at scale against managed simplicity and AWS integration, with throughput and retention limits varying by platform and configuration.


TL;DR:

  • Kinesis is ideal for small teams handling under 50 MB/sec of sustained ingest and seeking minimal operational effort with AWS integration.
  • Kafka and MSK are better suited for high throughput, multi-region replication, and long retention, with Kafka offering more retention control and MSK providing Kafka compatibility managed by AWS.
  • Throughput limits are fixed at 1 MB/sec or 1,000 records/sec per Kinesis shard, while Kafka partitions can handle higher, hardware-dependent throughput.
  • Kinesis retention caps at 365 days, whereas Kafka retention is configurable and can extend indefinitely, influencing replay and archival strategies.
  • Benchmarking with representative traffic and key distribution is critical before choosing, as workload characteristics mainly determine the most cost-effective and reliable platform.

Arosplatforms
Build AI Around Your Operations
Arosplatforms embeds customized AI operating systems into business workflows, helping teams automate manual processes and manage systems without vendor lock-in.

Table of Contents

Apache Kafka, Amazon MSK, and Amazon Kinesis at a glance

Three platforms cover almost every real-time streaming decision an engineering team faces today. Apache Kafka is the open-source distributed event streaming platform that set the standard for high-throughput, replayable event logs. Amazon MSK runs that same Kafka engine as a managed AWS service, trimming broker operations without changing the underlying semantics. Amazon Kinesis Data Streams is AWS’s own fully managed streaming service, built around shards instead of partitions and wired directly into AWS analytics and serverless tooling.

Dimension Apache Kafka Amazon MSK Amazon Kinesis Data Streams
Best for High-throughput, low-latency workloads needing full control Teams wanting Kafka semantics with less broker overhead AWS-native teams prioritizing fast time-to-prod at moderate scale
Deployment model Self-hosted brokers, any cloud or on-premises AWS-managed brokers, Kafka-API compatible Fully serverless, AWS-managed shards
Throughput per unit Varies by partition and hardware, typically higher than a Kinesis shard Same as Kafka, bound by broker sizing 1 MB/sec write per shard, 1,000 records/sec, per AWS’s documented limits
Scaling model Add partitions and brokers manually or via automation Add brokers and partitions, AWS handles provisioning Provisioned resharding or on-demand autoscaling
Retention Configurable, effectively unlimited on disk Same as Kafka 1 to 365 days
Ordering scope Per-partition Per-partition Per-shard
Exactly-once Producer-side, idempotent producers and transactions Same as Kafka Consumer-side, via checkpointing and idempotent sinks
Ops complexity High, full broker and cluster ownership Moderate, AWS manages brokers Low, fully managed
Ecosystem Kafka Streams, ksqlDB, Confluent Schema Registry, broad connector base Same Kafka ecosystem Kinesis Data Analytics, tight AWS service integration

A few things stand out once you look past the row-by-row comparison:

  • Kafka’s biggest edge is retention flexibility: you decide how long data lives, not a service quota.
  • Kinesis wins on time-to-production for teams already living inside AWS, since there are no brokers to size or patch.
  • MSK is the pick when you need Kafka’s exactly-once and ordering guarantees but do not want to run the cluster yourselves.
  • None of the three is universally cheaper. The right answer depends on sustained throughput and how much engineering time you can spend on operations.

How partitions, shards, and brokers shape scaling

Kafka’s unit of parallelism is the partition, and each partition lives on a broker with a configurable replication factor for durability. Since Kafka 3.3, brokers coordinate through KRaft instead of ZooKeeper, cutting a layer of operational dependency. Amazon MSK runs this same broker and partition model, just with AWS handling patching and provisioning.

Kinesis replaces partitions with shards. Each stream is a set of shards, and records route to a shard based on a partition key’s hash. You choose provisioned mode, where you set shard count directly, or on-demand mode, where AWS scales shard count automatically as throughput grows.

Resharding is where the two models diverge sharply. Splitting or merging Kinesis shards changes the hash key ranges those shards own, which can invalidate sequence-number continuity for affected ranges and force consumer recovery logic to catch up, as comparative analysis of Kafka and Kinesis internals explains. Changing Kafka’s partition count is simpler mechanically but carries its own trap: it can silently change which partition a key maps to, breaking strict ordering if you have not planned for it.

Practical consequences worth planning around:

  • A hot key sends disproportionate traffic to one partition or shard, capping effective throughput regardless of total capacity.
  • Kinesis on-demand mode reacts to sustained load changes but is not instantaneous, so bursty spikes still need headroom.
  • Kafka autoscaling is not native. It is something you build with monitoring and automation, or get from MSK’s provisioning tools.

What the numbers actually say about throughput and latency

A Kinesis shard is capped at 1 MB per second or 1,000 records per second for writes, whichever limit hits first, according to AWS’s own service quotas. Kafka partitions do not carry a fixed cap of that kind. Throughput per partition depends on hardware, batching, and compression, and in practice runs well above a single Kinesis shard under comparable conditions, per benchmark comparisons of the two platforms.

A single Kinesis shard tops out at 1 MB/sec or 1,000 records/sec for writes. (AWS documentation) That ceiling is why sustained high-throughput topics often need dozens of shards, each adding cost and consumer complexity.

Latency profiles differ too. Standard Kinesis consumers poll and share throughput across a stream’s read capacity, while Enhanced Fan-Out gives each consumer a dedicated 2 MB per second per shard pipe with lower, more predictable latency. Kafka’s p99 latency advantage tends to matter most for workloads that cannot tolerate polling delay: fraud detection, trading systems, or real-time bidding.

For capacity math, a workable rule of thumb from Kafka and Kinesis architecture comparisons is to divide your sustained write throughput in megabytes per second by 1 to estimate shard count for Kinesis, then add real headroom for skew and bursts. For Kafka, size partitions and brokers iteratively rather than by formula, since hardware and replication factor both move the ceiling.

Before trusting any benchmark, check a few things:

  • Message size: small messages amplify per-record overhead differently on each platform.
  • Key skew: uneven partition keys distort throughput results in ways aggregate numbers hide.
  • Consumer concurrency: a benchmark run with one consumer thread understates real-world contention.

Retention, replay, and how long your data actually lives

Kafka’s retention is a configuration choice, not a platform limit. Topic-level settings control how long segments stay on disk before deletion, and teams commonly hold data for weeks or keep it indefinitely, per Kafka’s own topic configuration reference. Kinesis retention is bounded between 1 and 365 days, a hard platform limit rather than a tunable default, as AWS’s documented service limits confirm.

Replay mechanics reflect that same split. Kafka consumers reset their offset to any point still on disk and reread history freely. Kinesis consumers replay using TRIM_HORIZON to start from the oldest available record or AT_TIMESTAMP to start from a specific point in time, but resharding can complicate replay across changed hash ranges.

For backfills and long-term storage, a pattern that works on both platforms is to materialize the stream into object storage like Amazon S3. That gives you cheap, unlimited archival and decouples large-scale replay from either platform’s retention window.

  • Kafka retention is a dial you set, Kinesis retention is a ceiling you cannot exceed.
  • TRIM_HORIZON and AT_TIMESTAMP cover most Kinesis replay needs but interact awkwardly with resharding history.
  • Archiving to S3 removes retention pressure from either platform entirely.

Ordering guarantees and where exactly-once really lives

Ordering in Kafka is guaranteed within a partition, never across the whole topic. Kinesis guarantees ordering within a shard, never across the stream. Both systems inherit the same underlying constraint: changing the number of partitions or shards can change which partition or shard a key lands on, and that can break ordering guarantees you were relying on.

Ordered records within partitions and shards

Exactly-once semantics are implemented on opposite ends of the pipeline. Kafka’s model is producer-side: idempotent producers combined with transactions let a write succeed exactly once even after retries, and consumers reading in read-committed mode see only fully committed data. Kinesis has no equivalent producer-side transaction model, so exactly-once there depends on consumer-side work: the Kinesis Client Library checkpoints progress in DynamoDB, and your application logic has to be idempotent to survive reprocessing after a failure, as comparative analysis of exactly-once mechanics lays out.

Failure modes follow from that split. A Kafka consumer group rebalance can briefly pause processing across partitions. A Kinesis Client Library lease handoff between workers can produce duplicate reads if checkpointing lags.

  • Salt hot keys to spread load before it concentrates on one partition or shard.
  • Make downstream sinks idempotent so replays and duplicate reads never corrupt state.
  • Checkpoint frequently in Kinesis consumers to shrink the reprocessing window after a failure.

Pro Tip: Treat “exactly-once” as “effectively-once through idempotency,” and design your sinks that way regardless of which platform you pick.

Who’s on call, and what actually breaks at 2am

Self-hosted Kafka means you own broker patching, disk capacity planning, partition rebalancing, and ZooKeeper or KRaft cluster health. That is a real staffing commitment, not a one-time setup cost. Amazon MSK removes broker patching and much of the provisioning burden while keeping the same Kafka failure modes: a consumer group rebalance still pauses processing, a broker still needs monitoring for disk pressure. Kinesis removes almost all of that, since there are no brokers to manage at all, per operational guidance from Last Week in AWS’s platform comparison.

The incidents that actually page an on-call engineer differ by platform. Kafka’s common ones are consumer group rebalance storms, under-replicated partitions after a broker failure, and disk pressure from misconfigured retention. Kinesis’s common ones are throttling from undersized shard counts and Kinesis Client Library lease handoff delays that produce duplicate processing. MSK inherits Kafka’s incident types but not its patching burden.

For teams without a dedicated platform group, Last Week in AWS’s analysis is direct about this: Kinesis reduces operational guesswork, and Kafka’s operational tax is often the deciding factor for organizations that lack the staff to absorb it. That is where managed Kafka through MSK earns its place as a pragmatic middle ground: you keep Kafka’s guarantees and ecosystem while handing broker operations to AWS.

  • Self-hosted Kafka needs a team comfortable with JVM tuning, disk I/O, and cluster upgrades.
  • MSK cuts patching and provisioning work but keeps Kafka’s consumer-side failure modes.
  • Kinesis’s main on-call burden is shard sizing and Kinesis Client Library lease management, not broker health.

Planning capacity and infrastructure work at this level is exactly what a platform’s plumbing decisions come down to: the platform you choose determines the staffing shape you need behind it.

Modeling the real cost of each platform

Kinesis pricing is built from shard-hours, PutRecords or PUT payload unit charges, and an optional Enhanced Fan-Out fee per consumer per shard. Self-hosted Kafka’s costs are less itemized but no smaller: EC2 or bare-metal compute, EBS or local disk, cross-AZ network transfer, and the engineering time to run it all. MSK sits in between, billing for broker instances and storage while removing most of the labor line.

A single Kinesis shard is priced per shard-hour plus per-payload-unit charges, while Kafka’s costs are dominated by compute, storage, and staff time rather than a per-unit meter. (Comparative TCO analysis)

At low sustained throughput, a small team without a platform engineer almost always comes out ahead on Kinesis: no cluster to size, no staff hours to budget against broker maintenance. As sustained throughput climbs and shard counts multiply, the shard-hour and payload-unit charges start compounding, and self-hosted Kafka’s flatter compute costs, or MSK’s broker pricing, begin to look more competitive against a growing Kinesis bill and the fixed labor cost of running Kafka in-house.

To build your own model:

  1. Measure sustained and peak throughput in megabytes per second and records per second.
  2. Estimate shard or partition counts needed at both figures, including headroom for skew.
  3. Price out the managed option’s per-unit charges against self-hosted compute, storage, and a realistic fraction of an engineer’s time.
  4. Recalculate at your expected throughput twelve months out, not just today’s baseline.

Staff cost is the variable most teams underweight. A managed platform’s higher per-unit price often loses to Kafka’s lower unit cost once you account for the SRE hours a self-hosted cluster demands.

Schema registries, connectors, and stream processing tools

Kafka’s ecosystem runs on Confluent Schema Registry for most production deployments, giving you strict schema evolution rules and broad client library support. Kinesis pairs more naturally with AWS Glue Schema Registry, which integrates directly with other AWS services but has a smaller footprint outside that ecosystem.

Connector coverage favors Kafka by breadth: Kafka Connect has a mature library of sink and source connectors for databases, Amazon S3, and data warehouses. Kinesis leans on AWS-native integrations, most commonly Kinesis Data Firehose for landing records into S3 or Amazon Redshift, which is simpler to configure but narrower in scope.

For stream processing, Kafka’s ecosystem includes Kafka Streams and ksqlDB for in-cluster processing, alongside broad support from Apache Flink. Kinesis Data Analytics offers native processing without a separate cluster, but with a smaller feature surface for complex event-time logic compared to Flink.

  • Confluent Schema Registry has wider adoption outside AWS, AWS Glue Schema Registry wins for teams fully inside AWS.
  • Kafka Connect’s connector library covers more sinks and sources than Kinesis’s Firehose-centered integrations.
  • Flink is the deeper stream processing option, Kinesis Data Analytics is faster to set up for simpler jobs.

Planning a migration or a hybrid streaming setup

Moving between platforms is rarely instant. A realistic migration involves dual-writing to both systems during cutover, translating consumer offsets or checkpoints between the two models, and rebuilding any connectors tied to the platform you are leaving. None of that is exotic engineering, but all of it takes calendar time you should budget for explicitly.

A hybrid pattern that works well for bursty workloads: run Kafka as your steady-state baseline for latency-sensitive or exactly-once traffic, and route overflow above a defined threshold to Kinesis, which absorbs bursts without you provisioning Kafka capacity for peaks you rarely hit. This only works if routing logic and idempotent sinks are in place on both paths, since a record processed twice across two systems is worse than a record processed twice within one.

  1. Confirm ordering and exactly-once requirements survive the target platform’s guarantees.
  2. Build offset or checkpoint translation before cutover, not during it.
  3. Run both systems in parallel long enough to validate throughput and latency under real load.
  4. Set explicit gating criteria (error rate, latency, cost) before decommissioning the old path.

Pro Tip: Never treat a hybrid Kafka-plus-Kinesis setup as temporary. Budget for it as a permanent pattern if burst traffic is a recurring part of your workload, not a one-off event.

Running a benchmark you can actually trust

Numbers in a vendor blog post are a starting point, not a decision. Run your own test against representative traffic before committing capacity budget to either platform.

Collect these metrics at minimum: sustained throughput, p50 and p99 latency, error and throttle rates, CPU utilization, and disk I/O on any self-hosted component. Standardize your test conditions across runs: fixed message size, a defined key distribution (including deliberate skew), a set consumer concurrency level, and a documented retry policy.

Watch for these red flags in published benchmarks, including independent throughput comparisons worth reading critically: missing harness details, no disclosed message size or key distribution, and cost figures quoted without the throughput level they assume.

  • A benchmark without message size and key skew disclosed is not reproducible.
  • Cost comparisons that omit staff time are measuring infrastructure spend only, not total cost.
  • Any number without a stated test topology should be treated as directional, not decisive.

What most teams get wrong about this decision

Most teams pick Kafka or Kinesis based on which one they already know, not which one fits their actual sustained throughput. That is backward. The workload should dictate the platform, and the workload rarely needs what the team assumes it needs on day one.

The decision flow that holds up in practice: measure your real sustained throughput and key skew first, prototype a small MSK cluster or Kinesis stream alongside one Kafka topic to validate cost and latency assumptions against real traffic, then decide between managed and self-hosted based on team size and the SLA you are actually contractually bound to, not the one you aspire to. Bringing in outside engineering help to validate that measurement step before committing to a platform is often the cheapest insurance available.

— arosplatforms team

Getting expert help choosing and running your streaming backbone

Picking between Kafka, MSK, and Kinesis is a measurement problem before it is a platform problem, and most teams do not have the spare engineering hours to run that measurement properly while also shipping product work. Some AI consultancies build custom systems for logistics, healthcare, and other operationally intensive industries, and that work often starts with infrastructure decisions: what should ingest events, what should process them, and what a team can realistically operate long-term.

Relevant services for teams facing this decision include:

  • A readiness assessment or scoping engagement to map your actual throughput and skew before committing to a platform.
  • A proof-of-concept build to validate cost and latency assumptions against your real traffic, not a vendor’s sample dataset.
  • Production deployment support once the platform choice is validated.
  • Ongoing managed AI services for teams that want the operational burden handled after launch.

If you want a second set of eyes on your benchmark before you commit budget or headcount to either path, reach out through our services page to scope an assessment or a pilot.

Sources

FAQ

Is Kafka cheaper than Kinesis?

It depends on sustained throughput and staffing. At low volume, Kinesis’s managed model usually costs less once you count engineering time, while at high sustained throughput, Kafka’s flatter compute costs can undercut Kinesis’s per-shard and per-payload charges, according to comparative TCO analysis.

Do companies like Netflix use Kafka?

Large-scale streaming users across the industry commonly run Apache Kafka for high-throughput event pipelines, since it is the open-source standard for this workload. Specific production architectures at any given company are not detailed in the sources this article draws on.

Is Kafka similar to Kubernetes?

No, they solve different problems. Kafka is a distributed event streaming platform for moving and storing data in real time, while Kubernetes is a container orchestration system for deploying and scaling applications, and the two are often used together rather than compared directly.

What is the AWS equivalent of Kafka?

Amazon Kinesis Data Streams is AWS’s native streaming service, while Amazon MSK runs actual Apache Kafka as a managed AWS offering. Teams that need Kafka’s exact semantics on AWS typically choose MSK, while teams prioritizing a fully serverless model choose Kinesis.

How do I decide between Kafka, MSK, and Kinesis for my team?

Match the platform to sustained throughput and team size: Kinesis suits small teams or ingest under roughly 50 megabytes per second with AWS-native needs, while Kafka or MSK suits high throughput, long retention, or strict exactly-once requirements. Running a small benchmark against your own traffic, as described in this article’s methodology section, settles most remaining doubt.

Related: custom AI development.

Match Throughput and Team Size: Kafka vs Kinesis for Engineers