For nearly a decade, Apache Kafka operated under a strict constraint: the broker’s local disk was the source of truth, the performance bottleneck, and the primary scaling factor. As data volumes surged into the petabyte range, this tight coupling between compute and storage forced engineering teams into expensive, over-provisioned clusters just to manage retention requirements. The introduction of tiered storage fundamentally breaks this dependency, allowing Kafka to behave more like a distributed log with an infinite shelf life.
This article provides the technical blueprint for implementing tiered storage in production environments, moving beyond documentation to address the real-world trade-offs in latency, cost, and operational complexity. Whether you are managing multi-tenant clusters or long-term analytical streams, we will examine how to transition from local-only storage to a hybrid architecture that leverages cloud-native object stores.
The Evolution of Kafka Storage and Data Placement
Historically, where does Kafka store data? The answer was simple: the local filesystem. Kafka storage relies on the log.dirs configuration, where each partition is represented as a directory containing append-only log segments. This design prioritized sequential I/O performance, which is excellent for throughput but creates a rigid relationship between the number of brokers and the total retention capacity.
Engineering Callout: The shift toward tiered storage represents the most significant architectural change to the Kafka log since its inception. By decoupling local log segments from long-term archival, you effectively separate the performance-sensitive ‘hot’ data from the cost-sensitive ‘cold’ data.
When you rely solely on local disks, your cluster size is dictated by the total volume of data you need to retain, not by the amount of traffic you process. This leads to ‘storage-bound’ clusters where you pay for high-performance NVMe drives that are mostly sitting idle, simply to satisfy retention policies. Understanding this limitation is the first step toward justifying the move to a tiered storage architecture.
Core Mechanics: How Kafka Tiered Storage Works
Kafka tiered storage functions by delegating the lifecycle management of log segments to a RemoteLogManager. Once a segment reaches a specific age or size threshold, the manager asynchronously uploads the segment to remote object storage, such as Amazon S3 or Google Cloud Storage, while keeping a reference in the local metadata index.
The process follows a specific lifecycle:
- Active Segment: Data is written to the local disk as usual.
- Offloading: The
RemoteLogManagertriggers an upload of closed segments. - Metadata Sync: The broker updates the remote log index, ensuring consumers can still fetch data transparently.
- Local Deletion: Once successfully uploaded, the local copy is purged to reclaim disk space.
[Broker] <---> [Local Log Segments (Hot)] <---> [RemoteLogManager] ---> [S3 / GCS (Cold)]
From the consumer’s perspective, the transition is seamless. When a consumer requests an offset that no longer exists on the local broker, the broker automatically fetches the corresponding data block from the remote object store, caches it locally if necessary, and serves the request.
Optimizing Kafka Data Storage Throughput and Costs
Choosing between traditional retention and tiered storage requires a clear understanding of your workload’s access patterns. Traditional storage is ideal for low-latency, high-frequency read scenarios, while tiered storage is optimized for long-term retention and cost efficiency.
| Metric | Traditional Storage | Tiered Storage |
|---|---|---|
| Latency | Sub-millisecond | 10ms – 100ms (Remote Fetch) |
| Cost per GB | High (NVMe/SSD) | Low (Object Storage) |
| Rebalance Time | Slow (Data Movement) | Fast (Metadata Only) |
| Retention | Limited by Disk | Virtually Unlimited |
By moving Kafka data storage to object stores, you reduce the cost-per-gigabyte by orders of magnitude. However, you must accept a slight latency penalty for reading ‘cold’ data, as the broker must perform an I/O operation against the object store API.
Production Deployment and Configuration Checklist
Enabling Kafka tiered storage in version 3.6+ requires careful coordination between your broker configuration and your object storage provider. Follow this checklist to ensure a stable deployment.
- Verify Provider Support: Ensure your storage backend implements the
RemoteStorageManagerinterface. - Configure Remote Log Manager: Enable the feature in your
server.propertiesfile. - Define Retention Policies: Set your local retention separately from the total retention.
- Monitor Offload Lag: Track the
remote-log-manager-tasksmetrics to prevent backlogs.
Example configuration for server.properties:
remote.log.storage.system.enable=true
remote.log.storage.manager.class.name=org.apache.kafka.server.log.remote.storage.S3RemoteStorageManager
remote.log.metadata.manager.class.name=org.apache.kafka.server.log.remote.metadata.storage.TopicBasedRemoteLogMetadataManager
remote.log.manager.task.interval.ms=30000
Always perform a load test after enabling this feature. The RemoteLogManager consumes CPU and network bandwidth during the upload process; failing to account for this can result in increased consumer latency or partition rebalance failures.
Frequently Asked Questions
Where does Kafka store data by default?
By default, Kafka stores data on the local file system of the broker in log directories defined by the log.dirs configuration. Each partition maps to a directory, and segments are written as active append-only log files on the local disk storage of the broker node.
How does Kafka tiered storage change cluster performance?
Tiered storage offloads older log segments to remote object storage like S3. This reduces the burden on local disks, speeds up cluster rebalancing, and allows for massive retention periods without needing to scale the number of broker nodes proportionally to the total data size.
Is Kafka data storage reliability maintained with tiered storage?
Yes, reliability is maintained by using remote index files and metadata management. When a segment is moved to remote storage, Kafka maintains a remote log index, ensuring that consumers can still read historical data seamlessly while the broker manages the lifecycle of both local and remote segments.
Tiered storage is no longer an experimental feature; it is the standard for modern, high-scale Kafka deployments. By moving away from the ‘disk-per-broker’ constraint, you gain the flexibility to retain data for months or years without forcing your hardware footprint to expand linearly with your data volume.
As you implement these changes, prioritize monitoring your remote log manager throughput and ensure that your object storage IAM policies are configured with the least-privilege access required. The transition to tiered storage is a significant architectural pivot, but it is the most effective lever for controlling long-term cluster costs and operational complexity.