The headline buries the actual change.
Turbopuffer's post says version 3 will stop keying stored records by an approximate-nearest-neighbor address. Its current layout started with an identifier and vector, then gained attribute filters, full-text search, sparse vectors, regular-expression search, ordering, and aggregations. The vendor says the vector-first layout now causes storage amplification, write amplification, and block sizes that fit some query plans poorly.[1]
The vector database is not dead. One product is moving its vector index out of the primary storage position. That is an architecture change, not a funeral for nearest-neighbor search.
Choose a stable record layout first. Let each retrieval method earn its own index.
Agent memory rarely asks one kind of question.
A coding assistant may need semantic similarity for a vague behavior question, exact text for a symbol or error, a permission filter for repository scope, recency ordering for a session list, and aggregation for an inventory. Turbopuffer's current query API exposes those as separate plans. It lists approximate and exact nearest-neighbor search, BM25 text search, sparse vectors, attribute ordering, lookups, aggregations, and multi-query hybrid search.[2]
That list is useful because it breaks the product label into work. If half the traffic is exact identifiers and filtered recency, a vector benchmark cannot choose the system. If updates are frequent, index maintenance belongs beside read latency. If permissions are part of every request, filtered recall is a release metric.
More indexes are not free.
PostgreSQL's index guide states the trade plainly. Indexes can retrieve rows faster, but they add system overhead. PostgreSQL can combine several indexes for one query, and its planner may skip an available index when extra scans cost more than they save.[3]
The same discipline applies to an agent context store. Do not add dense vectors, sparse vectors, full-text postings, metadata indexes, and recency sort paths because a diagram looks complete. Keep a census of real query shapes. Measure freshness after writes. Check the misses that matter to the task.
Write the workload contract before the migration.
- Name each query shape and its share of traffic.
- Record the filter, ordering, freshness, and recall requirement for each shape.
- Replay inserts, updates, and deletes at the expected rate.
- Compare exact, lexical, semantic, and hybrid results on the same held-out tasks.
- Keep one fallback that works when an index is stale, missing, or too expensive.
A green benchmark for one index proves one lane. The storage decision needs the whole route.