Vector & Hybrid Search
VeltrixDB serves nearest-neighbour vector search, BM25 full-text search and hybrid (vector + text) search from the same database that holds the records they describe. This page covers how it works, how to use it, how much RAM it needs, and what it measured — every number states where it was measured, and none of them is a comparison with another database run on the same machine.
New in v1.11 namespaces, filters, VDEL, int8 quantization, BM25 and hybrid search, cluster-wide search New in v1.12 product quantization, GRAPH disk, the startup-rebuild guard (search_ready). Plain HNSW VSET/VSEARCH on the default namespace has existed since v1.1.0.
One id, one record
A vector, a text document and a KV record that share an id describe one thing. Filters read the KV record; the vector and the text are searched.
PUT doc-42 {"lang":"en","tier":"gold"} # record (filters read this)
VSET doc-42 NS docs 0.12 -0.03 ... # vector in namespace "docs"
TSET doc-42 NS docs TEXT how to rotate keys ... # text document in "docs"
HSEARCH 10 NS docs FILTER lang = en VEC 0.1 ... QUERY rotate keys
Everything is stored as ordinary keys, so it goes through the same WAL, VLog, replication and Raft paths as any other write:
| Reserved key | Holds |
|---|---|
@vec/<ns>/<id> | The vector, float32 little-endian, L2-normalized (similarity is cosine). |
@txt/<ns>/<id> | The text document (UTF-8, at most 1 MiB). |
@vecns/<ns> | Namespace settings (JSON: dimension, quantization, PQ subspaces, graph placement). |
@idxdef/<name> / @idx/<name>/<value>/<id> | Secondary-index definition / entries, used by = filters. |
The search indexes in RAM — one HNSW graph and one BM25 inverted index per namespace — are derived from those keys by hooks that run after each write commits, and are rebuilt from them at startup.
Quickstart session
A real session against a v1.12 server (--data on a local disk, text protocol over nc). The scores are what the server printed for these toy 4-dimensional vectors.
$ nc localhost 9000
VCREATE docs 4 QUANT int8
OK
PUT doc-1 {"lang":"en","tier":"gold"}
OK
PUT doc-2 {"lang":"de","tier":"free"}
OK
PUT doc-3 {"lang":"en","tier":"free"}
OK
VSET doc-1 NS docs 0.9 0.1 0.0 0.2
OK
VSET doc-2 NS docs 0.8 0.3 0.1 0.1
OK
VSET doc-3 NS docs 0.1 0.9 0.4 0.0
OK
TSET doc-1 NS docs TEXT how to rotate api keys safely
OK
TSET doc-2 NS docs TEXT key rotation schedule for the billing service
OK
TSET doc-3 NS docs TEXT onboarding checklist for new engineers
OK
VSEARCH 2 NS docs 0.85 0.2 0.05 0.15
doc-1 0.9903372
doc-2 0.988912
END
VSEARCH 2 NS docs EF 128 FILTER lang = en 0.85 0.2 0.05 0.15
doc-1 0.9903372
doc-3 0.3244192
END
IDXCREATE by_lang lang
OK
IDXQUERY by_lang en
doc-1
doc-3
END
TSEARCH 3 NS docs QUERY rotate keys
doc-1 1.9616585
END
HSEARCH 3 NS docs ALPHA 0.7 CAND 50 VEC 0.85 0.2 0.05 0.15 QUERY rotate keys
doc-1 0.016393442
doc-2 0.011290323
doc-3 0.011111111
END
HSEARCH 3 NS docs FILTER tier = free QUERY rotation
doc-2 0.008196721
END
VDEL doc-2 NS docs
OK
TDEL doc-2 NS docs
OK
VSEARCH 0 NS docs 0.85 0.2 0.05 0.15
doc-1 0.9903372
doc-3 0.3244192
END
VSET doc-9 NS docs 0 0 0 0
ERR vector must be non-zero and finite
VSET doc-9 NS docs 0.1 0.2
ERR vector namespace "docs" already exists with dim 4 (got 2)
VCREATE big 768 QUANT pq PQM 96 PQTRAIN 10000 GRAPH disk
OK
Things the session shows: TSEARCH ... QUERY rotate keys does not match doc-2 (“key rotation”) because there is no stemming; the HSEARCH scores are reciprocal-rank-fusion scores, not cosine or BM25 (doc-1 is first in both lists: 0.7/61 + 0.3/61 = 0.016393); the filtered HSEARCH without VEC runs only its text half; and VSEARCH 0 returns every match (an exact scan).
Commands
Text protocol, exactly as cmd/server/main.go parses it. Search commands answer id score lines then END, or one ERR ... line; writes answer OK.
| Command | Does |
|---|---|
VCREATE ns dim [QUANT none|int8|pq] [PQM m] [PQTRAIN n] [GRAPH memory|disk] v1.11 | Create or reconfigure a namespace. Optional: the first VSET creates a float32 one. PQM, PQTRAIN and GRAPH are v1.12. |
VSET id [NS ns] f1 f2 ... | Upsert a vector. NS is v1.11; without it the namespace is default. |
VDEL id [NS ns] v1.11 | Delete a vector. |
VSEARCH k [NS ns] [EF n] [FILTER field op value] f1 f2 ... | Top-k by cosine; k = 0 returns every match (exact scan). NS/EF/FILTER are v1.11. |
TSET id [NS ns] TEXT free text / TDEL id [NS ns] v1.11 | Upsert / delete a text document (text = rest of the line, at most 1 MiB). |
TSEARCH k [NS ns] [FILTER field op value] QUERY free text v1.11 | Top-k by BM25; k = 0 returns every match. |
HSEARCH k [NS ns] [EF n] [ALPHA a] [CAND n] [FILTER field op value] [VEC f1 ...] [QUERY text] v1.11 | Hybrid, fused by reciprocal rank; k > 0; needs VEC, QUERY or both. |
IDXCREATE name field / IDXDROP name / IDXQUERY name value [LIMIT n] | Secondary index on a record field (used by = filters); IDXQUERY answers one id per line, then END. |
A namespace name must be non-empty and contain no /; the dimension must be 1–4096 and is fixed by VCREATE or the first VSET. Vectors must be non-zero and finite. The option clauses may come in any order, each at most once; the floats (VSEARCH) or the QUERY text come last. The text protocol is whitespace-split and line-based, so filter values cannot contain spaces and TSET text cannot contain newlines — the binary protocol carries both (opcodes in Wire Protocol).
Go client
The Go client (client.BinaryConn, and client.TCPConn for the text protocol) is the only official client with the search API. From other languages, send the text commands above over a TCP socket.
| Method | Command |
|---|---|
VCreateWithOptions(ns, dim, VectorNamespaceOptions{Quantization, PQSubspaces, PQTrainAt, Graph}), VCreate(ns, dim, quant) | VCREATE |
VSetNS(ns, id, vec), VSet(key, vec) (namespace default) | VSET |
VDel(ns, id) | VDEL |
VSearchWithOptions(k, query, VectorSearchOptions{NS, Ef, FilterField, FilterOp, FilterValue}), VSearch(k, query) | VSEARCH |
TSet(ns, id, text), TDel(ns, id), TSearch(k, query, TextSearchOptions{...}) | TSET / TDEL / TSEARCH |
HSearch(k, vec, query, HybridSearchOptions{NS, Alpha, Ef, Candidates, Filter...}) | HSEARCH |
IdxCreate(name, field), IdxDrop(name), IdxQuery(name, value, limit) | IDXCREATE / IDXDROP / IDXQUERY |
Results are []VectorResult{ID, Score}.
Memory layouts
The full float32 vectors are always persisted in the VLog on NVMe. What RAM holds is a per-namespace choice, set with VCREATE:
| Layout | RAM holds | Search |
|---|---|---|
| float32 (default) | 4 × dim bytes + graph | Exact scores. |
QUANT int8 v1.11 | dim + 4 bytes (codes + scale) + graph | Graph walk on int8 (beam ≥ 64); top 4 × k re-ranked against the float32 vectors. |
QUANT pq [PQM m] v1.12 | m bytes (default m = dim/8, range 1–dim) + graph, plus one codebook per namespace | ADC lookup tables; max(4 × k, ef) candidates — the whole beam — re-ranked against the float32 vectors. |
... GRAPH disk v1.12 | As above, minus layer-0 edges | Layer-0 edges (32 per node, ~97 % of edge memory) in 132-byte records in a memory-mapped scratch file the kernel can page to NVMe. |
The re-rank reads the float32 vectors back with one parallel MultiGet (LIRS-cached); a candidate deleted meanwhile is dropped. The disk graph's file is created under <first data dir>/vector-scratch and unlinked at once (it is never read after a restart). If it cannot be created, or on a platform without mmap, the setting is accepted and the graph stays on the heap.
Measured RAM — 768-dim, 10,000 synthetic clustered vectors, Go heap per vector (TestVectorMemoryTable, macOS, 2026-09-30):
| Layout | Heap / vector | Mapped file / vector | vs float32 |
|---|---|---|---|
| float32 | 3,383 B | — | 1× |
| int8 | 1,078 B | — | 3.1× less |
| pq, m = 96 | 493 B | — | 6.9× less |
| pq, m = 96 + disk graph | 378 B | 216 B | 8.9× less heap |
A PQ namespace stores float32 until it holds PQTRAIN vectors (default 10,000, minimum 256), then trains its codebook (k-means, 256 centroids per subspace, 12 Lloyd iterations, at most 10,000 training vectors) in the background and re-encodes; searches keep working throughout. Running VCREATE on an existing namespace with a different quantization, PQM or graph placement re-encodes it from the persisted vectors (changing only PQTRAIN does not; changing the dimension is refused).
How filters work
FILTER field op value (ops = != > < >= <= contains) is evaluated on the KV record whose key is the vector / document id — a JSON object or k=v pairs separated by ,, ; or space. A record without the field never matches. The ordering ops compare numerically when both sides parse as numbers, else lexicographically.
=on an indexed field (IDXCREATE) is served from the index: up to 2,048 matching ids are all scored without a graph walk (exact for float32; int8 / pq score the codes and re-rank as above); larger sets restrict the graph walk to members. Every candidate is still checked against its live record.- Other filters are checked during the walk, one read per candidate that would enter the results, with the walk capped at 10,000 + 20 × ef visited nodes so a filter matching almost nothing cannot scan the whole graph.
- BM25 filters are checked best-first on the scored documents until k match.
BM25 and hybrid search
- Tokenizer: lowercase runs of Unicode letters, digits and combining marks (so Devanagari, Tamil and other Indic words stay whole), no stemming, no stop words. Scripts without spaces (CJK, Thai) index whole runs.
- Scoring: Okapi BM25, k1 = 1.2, b = 0.75, idf = ln(1 + (N − df + 0.5) / (df + 0.5)). A query with no tokens is refused (
query has no searchable terms). - Hybrid: the vector and text searches each return
CANDcandidates (default max(50, 4k), capped at 10,000); the lists are fused by weighted Reciprocal Rank Fusion,score = α/(60 + rank_vec) + (1 − α)/(60 + rank_text)with 1-based ranks and α =ALPHAin [0, 1] (default 0.5). Ranks, not raw scores, are fused, so cosine and BM25 never need a common scale. Either half may be omitted (or missing: no vector namespace, no text tokens); it is an error only when neither can run.
Durability and restarts
A vector or document write is acknowledged after the same WAL + VLog fdatasync as any other write. The HNSW graph and the inverted index are not saved: at startup they are rebuilt from the persisted keys on all cores (GOMAXPROCS workers, namespace settings first).
New in v1.12 Until that rebuild finishes, vector / text / hybrid searches are refused — they would otherwise miss vectors — including when the rebuilding node is a peer in a fan-out:
ERR search indexes are still rebuilding after restart (loaded x of y; retry, or start the server with --search-allow-partial)
INFO shows search_ready=0 search_loaded=x/y during the rebuild and search_ready=1 once done. --search-allow-partial answers anyway. KV traffic, QUERY and IDXQUERY are not affected.
Measured: 3 SIGKILL-during-writes cycles (8.5K, 8.4K and 11.9K acknowledged writes) — every acknowledged vector, document and record was present after the rebuild, and no acknowledged delete came back (TestSearchCrashRecovery, macOS; the nightly workflow runs 10 cycles on Linux). Writes in flight at the kill are not checked. A 60-second soak on macOS (TestSearchSoak, 4 writers + 4 searchers): 201,573 writes and 292,079 searches, 288,025 self-queries with 0 misses, RSS flat at ~1.0 GB, exact oracle match before and after a clean restart.
Clusters New in v1.11
- Writes:
VSET,TSET,VCREATE,IDXCREATE, … go through the normal raft / replicated write path. - Routing keys co-locate: the ring places
@vec/<ns>/<id>,@txt/<ns>/<id>and@idx/<rule>/<value>/<id>on the node that owns<id>, so after a rebalance each node holds whole records and filters locally.@vecns/and@idxdef/keys are copied to every node. - Fan-out: every vector / text / hybrid search,
QUERYandIDXQUERYruns on all non-failed nodes and is merged. BM25 first sums every node's corpus statistics so all nodes score with the same idf. - Fail-closed: a peer that does not answer within
--search-timeout-ms(default 2000, per search phase; BM25 has two) fails the request and is named in the error (search incomplete: n of m peers did not answer: ...).--search-allow-partialreturns the rest instead (logged). Nodes the failure detector marked failed are not queried.--search-fanout=falsekeeps searches local. - Signing: node-to-node search traffic uses the transfer listener (
POST /internal/search). Sign it with--cluster-secret-fileor envVELTRIXDB_CLUSTER_SECRET(same secret on every node, at least 16 bytes; HMAC-SHA256 over time, method, path and body; 5-minute clock-skew window), or use mTLS (--cluster-tls-cert/--cluster-tls-key/--cluster-tls-cawith--cluster-mtls). Without either, the server logs that the listener is unauthenticated.
Measured performance
Real embeddings — GloVe-100
First 100,000 words of GloVe-100 (angular), 100 dimensions, 500 queries, exact cosine ground truth. In-process, one query at a time, macOS (Apple silicon, GOMAXPROCS = 18), 2026-09-30 (TestRealEmbeddings). Build = 18 parallel durable PutVector writers. Not a Linux NVMe server, and not a comparison with any other database.
| Layout | Build | ef | recall@10 | p50 | p99 | QPS (1 thread) |
|---|---|---|---|---|---|---|
| float32 | 57 s | 64 | 0.842 | 0.25 ms | 0.41 ms | 3,933 |
| 128 | 0.906 | 0.42 ms | 0.63 ms | 2,351 | ||
| 256 | 0.953 | 0.77 ms | 1.12 ms | 1,307 | ||
| 512 | 0.983 | 1.43 ms | 1.95 ms | 712 | ||
| int8 | 50 s | 64 | 0.839 | 0.61 ms | 0.81 ms | 1,655 |
| 128 | 0.905 | 0.55 ms | 0.80 ms | 1,905 | ||
| 256 | 0.952 | 0.72 ms | 1.08 ms | 1,356 | ||
| 512 | 0.983 | 1.18 ms | 1.84 ms | 834 | ||
| pq (m = 25) | 42 s | 64 | 0.797 | 0.56 ms | 0.90 ms | 1,771 |
| 128 | 0.870 | 0.72 ms | 1.06 ms | 1,367 | ||
| 256 | 0.927 | 1.17 ms | 2.91 ms | 780 | ||
| 512 | 0.969 | 1.56 ms | 3.44 ms | 604 | ||
| pq + disk graph | 45 s | 64 | 0.788 | 0.55 ms | 0.79 ms | 1,811 |
| 128 | 0.875 | 0.66 ms | 0.99 ms | 1,506 | ||
| 256 | 0.931 | 1.06 ms | 2.18 ms | 907 | ||
| 512 | 0.966 | 1.60 ms | 2.97 ms | 613 |
The quantized layouts are slower per query than float32 at this size because they re-rank from the VLog (a parallel MultiGet — before that change int8 at ef = 64 was p50 1.9 ms / p99 13.8 ms). Their point is RAM: at 100-dim the vectors are small and graph edges dominate; at 768-dim the vector part is 3–9× smaller (see Memory layouts).
Over the network — vecbench
GloVe-100, first 20,000 words, 300 queries, int8, 8 concurrent clients, server and client on one macOS machine (bench/compare/cmd/vecbench):
| ef | recall@10 | QPS | p50 | p99 |
|---|---|---|---|---|
| 64 | 0.894 | 24,637 | 0.28 ms | 0.68 ms |
| 256 | 0.985 | 15,629 | 0.48 ms | 0.91 ms |
Load: 6.4 s (3,146 durable vector writes/s).
CI quality gates (synthetic, every PR)
5,000 × 64-dim clustered vectors, 100 queries, recall@10 (TestSearchQualityGate, job node-7-search). A PR fails if a value drops below its gate.
| Case | Measured | Gate |
|---|---|---|
| float32 | 0.927 | 0.90 |
| float32, ef = 256 | 0.998 | 0.97 |
| int8 | 0.923 | 0.89 |
| pq | 0.882 | 0.85 |
| pq + disk graph | 0.881 | 0.85 |
| Indexed filter on the graph path | 0.994 | 0.96 |
| After deleting 40 % | 0.96 | 0.92 |
Nightly (New in v1.12): recall gates on GloVe-100, a 20-minute search soak against an oracle, and 10 SIGKILL / restart cycles. Full methodology: docs/vector-search.md in the main repo.
Tuning
EFis the recall / latency knob. On GloVe-100, ef = 128 gives ~0.90 recall@10 and ef = 256 ~0.95.- int8 keeps float32 recall at ~⅓ of the RAM; pq trades 1.4–4.5 points of recall@10 on GloVe-100 (ef 512 → 64) for ~⅐ of it at 768-dim. Raise
PQM(more bytes per vector) for recall. GRAPH diskhelps once edges dominate RAM (many vectors, small codes).- Filters: create an
IDXCREATEindex on fields used with=. - Clusters: a search costs one round trip to every node; with
--search-fanout=falseit is local only — correct only if every node holds every record, i.e. no rebalance has moved keys.
Limitations
DEL does not delete a record's vector or text Deleting the KV record removes only the record: the vector and document stay searchable (an unfiltered search still returns the id; a filter on that record no longer matches). Use VDEL / TDEL as well.- The graph and inverted index are rebuilt at every start (not persisted); searches are refused until the rebuild ends. Budget for it on large namespaces.
- Graph nodes, ids and the upper HNSW layers stay on the Go heap even with
GRAPH disk; the constant ~280–400 B per vector does not shrink with quantization. - Tested up to 100K vectors (real) and 10K at 768-dim; nothing at 10M+ yet.
- Dimension 1–4096. Text search has no stemming, phrase queries or per-field weighting; one document is at most 1 MiB.
- No same-hardware comparison with other vector databases.
bench/compare/cmd/vecbenchcontains only a VeltrixDB driver: Aerospike Vector Search and ScyllaDB vector search use client APIs that were not available to test, and VectorDBBench has no client for either. - Only the Go client has the search API; other SDKs use the text commands over a socket.