turbopuffer v3 makes vector ANN a secondary index

turbopuffer, a search engine that began as a serverless vector database, has announced a new storage architecture, v3. The company says v3 changes how documents and indexes are laid out, written, compacted and queried. The stated goal is to make search faster in every respect, including text, regex and vector search, and to lay the foundation for moving many more SQL queries onto turbopuffer and making them fast.

The post walks through the history that led here. turbopuffer launched as a serverless vector database (v1), built to serve extremely cheap and reasonably fast vector searches. Object storage as the source of truth gave the economics, and tiered NVMe SSD and memory caches gave the performance. The earliest customers, including Cursor and Notion, validated those tradeoffs. v2 added very strong text and regex search, and the product is now also used for non-search work, such as Linear's syncing engine. Through all of this the query engine evolved, but the storage architecture stayed largely the same: the ANN vector index was, and still is, the primary index around which every other index and query plan revolves. That design has constrained several query plans, such as GROUP BY and aggregations.

In v1, a document was just an ID and a vector. The team chose a hierarchical clustering index over the then-prevailing graph-based indexes because it plays better with object storage. They started with SPANN and migrated to SPFresh to support incremental indexing. Vectors are grouped into clusters, the cluster centroids are clustered in turn, and the process repeats up to a single root. Everything is stored in a sorted key-value layout keyed by a ClusterId plus a dense LocalId inside the cluster. Together these form the ANN address, which is what it means to say the ANN index is the primary index.

Attribute filtering and full-text search marked the informal move from v1 to v2. Filtering was modelled as an inverted index mapping an attribute value to the ANN addresses of the documents that contain it. Full-text search works the same way: it finds the documents containing the query term (the postings) and stores term count and document length for BM25 scoring. Document attributes were also stored alongside the ID and vector for projections. Aggregations, regex search, fuzzy matching, sparse vector search and attribute ordering were later built on the same vector-primary layout.

The ANN primary index survived this long because it works very well for ANN search on object storage. The company says it has pushed vector search to single indexes of 100B+ vectors, serving 200 ms p99 reads at 1k+ QPS, and that any significant change risks regressions in ANN performance. Those figures describe the current architecture, not v3.

The post names three ways the layout holds back non-vector queries. First, storage amplification: the full contents of each document sit under its ANN address, so a document with multiple vectors (document nesting, late interaction) must have its contents duplicated for each vector. The post says this is the reason for some of the product's more unfortunate limits. Second, write amplification: whenever a document is inserted, updated or deleted, SPFresh may rebalance vectors to keep them well clustered. Because everything is keyed by the ANN address, rebalancing moves the full document contents and every attribute and full-text index that references them, so updating one vector can move hundreds of attributes and their indexes. The post says tuning indexing throughput has started to hit diminishing returns. Third, limited vectorization: modern engines run tight loops over blocks of values. DuckDB works in batches of 2,048 rows, ClickHouse up to about 65k, Lucene's posting blocks hold 256 docs, and turbopuffer's ANN index works best with clusters of around 100 to 200 documents. Every query plan has an optimal block size, but today all are constrained by the ANN primary index, so a plan that wants thousands of documents per block is stuck at 100 to 200.

The company points to its own earlier evidence. The first full-text search version partitioned posting lists along ANN cluster boundaries, and the median block held just about 1.5 postings. FTS v2 moved to fixed blocks of about 256, and the index became 10x smaller while queries got up to 20x faster. Posting lists could change because they are stored separately and point at documents. Aggregations and other scans read the documents themselves, stored one block per cluster, so their block size stays tied to cluster size as long as the ANN address is the primary key.

The fix, in the post's words, is simple: do not key on the ANN address. That is the change v3 makes, and the authors say it is not trivial. They are moving to a new primary index and making ANN just another secondary index. Earlier this month, v3 reached a milestone: 100% of CI passes. So far the work has focused on correctness; the post says the next phase is making it correct and fast. Benchmarks will be shared publicly over the coming weeks, as the team works toward and beyond performance parity before rolling v3 out to production. This is the first of a series of updates, and it does not yet describe the new primary index. The post closes with the company's scale: 1T+ documents hosted, 10M+ writes per second, and 25k+ queries per second.

Key facts

  • turbopuffer is building v3, a storage architecture that stops keying documents and indexes on the ANN address and makes the ANN vector index just another secondary index.
  • The company names three limits of the old layout: storage amplification (document contents duplicated per vector), write amplification (updating one vector can move hundreds of attributes and their indexes), and limited vectorization (blocks stuck at about 100 to 200 documents).
  • v3 passes 100% of CI, which is a correctness milestone; it is not yet in production, and benchmarks are promised over the coming weeks.
  • The earlier full-text search rework to fixed blocks of about 256 made the index 10x smaller and queries up to 20x faster, which the post uses as evidence of what block size can do.
  • The current architecture serves single indexes of 100B+ vectors at 200 ms p99 reads and 1k+ QPS; turbopuffer overall hosts 1T+ documents.

Why it matters

This is a rare public account of a vector database outgrowing its own founding design. turbopuffer started as a specialised vector store, then added filtering, full-text, regex, aggregations and more, all keyed on the vector index. The post argues that this has reached its limit and that the answer is to demote the vector index to a secondary one. The stated aim is to make more SQL-style queries, including GROUP BY and aggregations, fast on the same system, alongside text, regex and vector search.

Who it affects

Current and prospective turbopuffer customers are the most direct audience, especially those whose workloads combine vector search with filters, full-text search, multi-vector documents or aggregations. The post names Cursor and Notion as the earliest customers and Linear as a non-search user. Engineers designing search or retrieval systems on object storage can also read it as a worked example of how storage layout choices limit later query plans.

How to use it

Nothing can be used yet. v3 is not in production, and the post does not say any customer is running it. The post does not give a date for rollout. Readers who want to follow along can watch for the benchmarks the company promises to publish over the coming weeks. For anyone designing a similar system, the practical lesson in the post is to avoid tying the primary key of stored documents to the vector index address.

How solid is it

The post is written by turbopuffer itself, so every claim is the company's own. The design history, the amplification problems and the earlier FTS v2 result (10x smaller index, up to 20x faster queries, a maximum rather than a typical figure) are described in concrete terms. The claim that v3 will unlock significant performance improvement on all query plans is a forecast. The only v3 result is that 100% of CI passes, which says nothing about speed. No v3 benchmarks are given.

Risks and caveats

The company itself notes that any significant change to the ANN index risks regressions in ANN performance, and its stated target is performance parity before the production rollout. The post does not describe the design of the new primary index, and it gives no date or timescale for the rollout. The 100B+ vector, 200 ms p99 and 1k+ QPS figures belong to the current architecture and should not be read as v3 results. The post does not say vector search is being removed.

“We've pushed the vector-primary architecture as far as we can, and it's time to move on.”

— turbopuffer, v3 announcement post