turbopuffer v3: Vector Index Becomes Secondary

Original: RIP, vector database

Why This Matters

Signals that specialized vector databases are evolving into general-purpose search and query engines to stay competitive.

turbopuffer announced a major architectural overhaul called v3, replacing its ANN vector index as the primary storage index with a general-purpose primary index. The shift aims to accelerate text, regex, and vector search while enabling SQL-style queries like GROUP BY and aggregations previously blocked by the vector-first design.

turbopuffer launched as a serverless vector database (v1) built around a hierarchical clustering ANN index — specifically SPANN, later SPFresh — where all storage was keyed by cluster and local IDs ('ANN address'). Object storage held the source of truth; tiered NVMe SSD and memory caches handled performance. Early customers Cursor and Notion validated the model.

V2 added strong text and regex search, and the platform expanded into non-search workloads — Linear uses it as a syncing engine. But the underlying storage layer stayed vector-first throughout, which created hard constraints. GROUP BY, aggregations, and broader SQL-style queries were limited or blocked entirely by the assumption that everything revolves around an ANN index.

With v3, turbopuffer is demoting ANN to 'just another' secondary index and building a new primary index underneath. The goal is to make every query type faster — not just vector search — and to support a much wider range of SQL queries. The team is documenting progress publicly as the migration unfolds. Engineer Dan Harrison framed it plainly: 'We've pushed the vector-primary architecture as far as we can, and it's time to move on.'

Source

turbopuffer.com — Read original →