Cloudflare K2 enters public beta as serverless event streaming on R2

Cloudflare has launched K2 in public beta. It is a durable event streaming primitive on its Developer Platform. You send events to a K2 stream, which stores them as an ordered log, and consumers read them at their own pace. Cloudflare describes it as fully serverless, able to scale to vast quantities of data, and built for long-term retention, so long stretches of consumer downtime do not lose data.
The problem Cloudflare starts from is the coupling in traditional RPC architectures, where producers and consumers have to match each other in scale and in time. If producers send more than consumers can handle, or consumers and downstream services go down, events are dropped. Multiple independent consumers make it worse. The post's example is an ecommerce backend that emits an event when a transaction completes, which both an analytics system and a fraud detection service need to read. The fix is a service in the middle that absorbs writes while readers consume independently.
K2 began as an internal need. Cloudflare wanted a durable buffer on the edge, first as the ingestion layer for Basin Pipelines. Pipelines uses a pull-based stream processing engine, so something must store events before they are read, transformed and written to R2, and Cloudflare commits to never dropping events once the Pipelines Stream accepts them. Most companies would reach for Apache Kafka here. But Pipelines runs on the Cloudflare edge, which spans over 335 cities, and Cloudflare says its architecture often means it cannot run traditional distributed systems software like Kafka. Machines come in relatively small slices, are relatively ephemeral, and are often connected over the public Internet; the upside is proximity to users and the ability to scale horizontally.
So K2 implements a partitioned, durable log on top of R2 object storage. The post says object stores like R2 pair extremely durable storage (11 9s) with strongly consistent APIs. Leaving replication and consensus to the storage layer makes the K2 application layer simpler, cheaper and faster, and separating compute from storage lets each scale independently, which keeps large amounts of history cheap to hold.
The awkward part is that R2, like other object stores, does not support appends, the standard log operation. K2 instead accumulates writes in memory on an edge service, waits briefly for data to arrive, then writes all the events as one complete segment file. Ordering and strictly incrementing offsets come from R2's atomic operations, with no separate coordination service. The stated downside is higher produce latency: writing to object storage is slower than a local disk and the batch has to fill first. In the initial release this adds up to about 1 second of produce latency at the 99th percentile. A technical deep dive on the design is promised.
Cloudflare also explains where K2 sits among its existing products. Queues track individual items of expensive or time-consuming work, such as an image-processing request, and support per-item retries, delays and dead-letter queues. K2 is meant for high-scale data movement, long-term retention and fan-out consumption. Messages are produced and consumed in batches, which makes processing efficient but rules out message-level retries, and the batching also means higher producer latency than Queues. Basin Pipelines is a serverless ingestion service that takes JSON events, transforms them and writes them to R2 or a Basin Catalog. Cloudflare recommends Pipelines when the end result is object storage or Iceberg tables, and K2 for custom processing or other destinations.
The walkthrough uses product analytics. You create a stream with the cf command line tool (cf k2 streams create --name app_events --http-enabled); streams can also be created with Wrangler, the dashboard or the API, and an account can hold many streams. The example output shows a retention_seconds value of 604800 and an endpoint for the stream. Producing is done through an HTTP API or a Worker binding; the sample Worker sends a page_view event as bytes and returns a 503 or 500 depending on whether the error is retryable. K2 stores bytes, so any format or encoding works.
Consumers read through subscriptions, which divide work between consumers to give read parallelism. A subscription is created through the HTTP API with a name and a start position (the example uses earliest). Each consumer then polls the consume endpoint with a worker ID and a maximum record count, and receives a batch with a batch ID, a lease expiry and the records. The lease on a batch lasts 5 minutes. The client can ack the batch, so it is not redelivered; nack it, so it is redelivered; or extend the lease if it needs more time. Splitting one subscription across several consumers shares the data between them. Giving each consumer its own subscription gives the pub/sub pattern, where each sees every message. The two can be mixed using several independent consumer pools.
K2 is available now in public beta for accounts with Workers Paid subscriptions, with beta limits of 10GB of storage used and 30 MB/s produce per stream. Higher limits can be requested through Discord or a limit increase form. Usage is not billed during the beta; the post introduces anticipated pricing for when billing starts, but the source text gives no figures. The roadmap for the coming months lists higher write parallelism up to multi-GB/s streams, message keys with key-based ordering guarantees, push-based worker consumers, an Express tier with lower produce and end-to-end latencies, and drop-in support for Apache Kafka clients.
Key facts
- K2 is a serverless durable event streaming primitive in Cloudflare's Developer Platform, in public beta for Workers Paid accounts; it stores events as an ordered, partitioned log on top of R2 object storage.
- Because R2 has no appends, K2 buffers writes in memory on an edge service and writes whole segment files, using R2's atomic operations for ordering and strictly incrementing offsets with no separate coordination service.
- The stated cost is latency: about 1 second of produce latency at the 99th percentile in the initial release, and batch-level rather than message-level retries.
- Consumers read through subscriptions and get a 5-minute lease on each batch, which they can ack, nack or extend; separate subscriptions give pub/sub fan-out.
- Beta limits are 10GB of storage and 30 MB/s produce per stream, with no billing during the beta; planned work includes Kafka client support, key-based ordering and an Express low-latency tier.
Why it matters
Cloudflare is filling a gap in its own platform. Its edge is made of small, ephemeral machines linked over the public Internet, which, Cloudflare says, often rules out running Kafka-style software. K2 gets durability and scale by leaning on R2 for replication and consensus, so the application layer stays simple, and storage and compute scale separately. That also makes long retention cheap. K2 started as the ingestion buffer for Basin Pipelines and is now offered to developers as its own product, next to Queues and Pipelines.
Who it affects
Developers already building on Cloudflare Workers who need to decouple producers from consumers, for example to feed analytics, fraud detection or other independent readers from one event source. Teams with custom processing, or with destinations other than object storage or Iceberg tables, are the audience Cloudflare names for K2. Teams currently using Queues for per-item work, or Pipelines for writing to R2 or Iceberg, are told those remain the better fit for their cases. Teams that rely on Kafka clients will have to wait: drop-in support is only on the roadmap.
How to use it
Create a stream with the cf tool, Wrangler, the dashboard or the API. Produce events through an HTTP API or a Worker binding, as bytes in whatever format suits you. Create a subscription, with a start position such as earliest, and have each consumer poll it for batches. Each batch is leased for 5 minutes; ack it when done, nack it to get it redelivered, or extend the lease. Use one shared subscription to split work across consumers, or one subscription per consumer so each sees every message. K2 is in public beta for Workers Paid accounts, capped at 10GB of storage and 30 MB/s produce per stream, and is not billed during the beta. Higher limits can be requested through Discord or a form.
How solid is it
This is Cloudflare's own launch announcement, so the design claims and the latency figure come from the vendor. The post gives no benchmarks or comparisons of throughput or cost against Kafka or other products, and no customer names or outside reactions. The one hard performance number is about 1 second of produce latency at p99 in the initial release; no median, consume or end-to-end latency is stated. A fuller technical deep dive is promised but not yet published. The product is a public beta, and the roadmap items carry no timescale beyond the coming months.
Risks and caveats
Higher produce latency is the admitted trade-off, and the batching model gives up message-level retries, which Queues provide. Beta limits are tight: 10GB of storage and 30 MB/s per stream. The multi-GB/s write targets, key-based ordering, push-based consumers, the Express tier and Kafka client support are all planned, not available today. Post-beta pricing is not given in the source text, so costs once billing starts cannot be judged. The example output shows retention_seconds of 604800, but the post does not state that as a default retention period.
“Offloading replication and consensus to the storage layer allows us to make the application layer (K2 in this case) radically simpler, cheaper, and higher performance.”
— Cloudflare, announcement post