raggio: a small ray of RAG

Table of Contents
Sooner or later every RAG system sends you the same bill, and it’s written in gigabytes of RAM.
The model wants memory for its weights and its KV cache. The vector database wants memory for every embedding you ever stored, usually as uncompressed floats, usually all loaded at once, whether anyone is asking about them or not. On a big cloud VM you pay and move on. On a machine that also has to run the LLM, or in a department that needs search over a few million documents and owns exactly one modest VM, that bill decides whether the project happens at all.
raggio is my attempt at a smaller bill. The name is Italian for ray, and yes, it starts with RAG, I couldn’t resist. It’s one container with a REST API that stores, indexes and searches chunks for retrieval-augmented generation, and it’s built to stay small: on 2.55 million real arXiv abstracts it serves search from 1.8 GB of RAM, while Weaviate with its default settings used 23.4 GB for the same vectors. And the vector results are the same an exact float32 search would give you: recall@10 of 1.000.
I had three reasons to build it, and they turned out to be the same reason. At work I run an enterprise RAG platform, and the vector database is the piece I always wanted to shrink the most. On a local LLM box, like the DGX Spark I benchmarked on, every gigabyte the vector store keeps for itself is a gigabyte the model doesn’t get. And every time a small team asked me how to put search over their documents, the options were a hosted service they didn’t want to pay for, or a weekend of hand-rolled FAISS that nobody wanted to maintain afterwards. Every time, the problem was memory.
This isn’t a feature tour, the docs already do that. It’s about the three ideas that keep the footprint small, and what each of them costs (spoiler: none of them is free).
What it is, in one screen
raggio is a Python service (FastAPI) on top of turbovec, a Rust vector index that implements Google Research’s TurboQuant. As the README puts it: raggio is built on TurboVec, which is built on TurboQuant. The turbo is under the hood now.
Data lives in collections. Each one is a separate folder on the volume, with its own vector index, its own SQLite database for text, metadata and the job queue, and (optionally) its own API key. Give every team a key and they can share one deployment without seeing each other’s data, and without anyone having to set up an auth server. You can ingest your own pre-computed vectors, or point raggio at any OpenAI-compatible /embeddings endpoint (OpenAI, Azure AI Foundry, a local vLLM) and just send plain text. Search has three modes: vector, text (BM25) or hybrid.
What keeps the RAM flat is that collections aren’t all loaded at once. A collection loads on its first request, and at most four stay resident (that’s the default). When a fifth one is needed, the least recently used idle one goes back to disk, and anything nobody touched for 15 minutes gets offloaded anyway. So the total data can be much bigger than the container’s memory: only the collections people are actually querying pay rent.
The storyboard follows the real eviction rule: least recently used first, but never a collection that still has ingest jobs pending. When it’s over, tap any collection on the volume to search it yourself. The limits are the defaults for MAX_RESIDENT_COLLECTIONS and COLLECTION_IDLE_TTL.
Getting it running looks like this:
podman build -t raggio .
podman run -p 8000:8000 -v raggio-data:/data \
-e ROOT_API_KEY=change-me \
-e EMBEDDING_BASE_URL=https://api.openai.com/v1 \
-e EMBEDDING_API_KEY=sk-... \
-e EMBEDDING_MODEL=text-embedding-3-small \
raggio
# a collection with its own key, then a hybrid search on it
curl :8000/collections -H "x-api-key: change-me" \
--json '{"name": "sales", "collection_key": "sales-secret"}'
curl :8000/collections/sales/search -H "x-api-key: sales-secret" \
--json '{"query": {"text": "acme pricing terms"}, "mode": "hybrid", "k": 5}'
Now, the three ideas.
Idea 1: four bits, then check the answer
An embedding model gives you floats. Qwen3-Embedding-0.6B, which I used for the arXiv run, gives you 1,024 of them per chunk: 4,096 bytes per vector as float32. For 2.55 million abstracts that’s 10.4 GB before any index structure, and a graph index like HNSW wants all of it in RAM, plus the graph on top.
TurboQuant goes after those 4,096 bytes. It applies a random rotation to each vector, which makes every coordinate follow a known and concentrated distribution, and once you know the distribution in advance, a simple per-coordinate quantizer is close to optimal. turbovec stores 4 bits per coordinate: 512 bytes per vector, 8x smaller, 1.3 GB for the whole arXiv corpus. Search then scans those codes with SIMD instructions, front to back, which is exactly the access pattern memory likes best.
The rotation sounded like black magic to me the first time I read about it, so here it is in slow motion:
A 256-dimension toy vector with three loud dimensions, on purpose: plenty of embedding models have a few. The rotation is a randomized Hadamard transform (random sign flips, then eight butterfly stages), a fast stand-in for a truly random rotation. The grey band is a 16-level codebook designed for the bell curve a rotated coordinate follows. Every number in the figure is computed from the toy.
What you pay for 4 bits is ranking error. Two candidates with nearly the same similarity can swap places after rounding, and a true neighbor can end up just outside the top k. On the arXiv corpus the plain 4-bit scan had a recall@10 of 0.949: on average, one result out of twenty wasn’t one of the true top 10.
The fix is old and boring, and I like boring. raggio keeps an fp16 copy of every vector in SQLite, on disk only (it needs those copies anyway to rebuild indexes, more on that in idea 3). At query time the 4-bit scan over-fetches: asked for k results, it keeps max(50, 2k) candidates, capped at 2,000. Then it reads the fp16 originals of just those candidates and re-ranks them by exact cosine similarity. Quantization only decides which candidates get considered, never how the returned hits are ordered, and the scores you get back are exact.
Part 1 is to scale, for one 1,024-dimension vector. Part 2 is a toy with k = 5 and hand-picked scores (bars start at 0.80, so small differences are visible). Change how many candidates the scan keeps to see what a too-shallow over-fetch loses. raggio’s real floor is 50 candidates, not 10.
On 2.55 million arXiv vectors, rescoring took recall@10 from 0.949 to 1.000, for about 0.1 to 0.3 ms per query. I didn’t invent any of this: Qdrant (rescore with oversampling), Weaviate (rescoreLimit), Milvus (refine_k) and ScaNN all ship the same pattern. What matters is where the work happens. The expensive part, touching every vector, runs on the small representation, and the precise part runs on 50 rows.
Two footnotes. First, 1.000 means “quantization stopped costing recall on this corpus, with this model”, it doesn’t mean it beats HNSW everywhere. A corpus where some of the true top 10 fall outside the candidates the scan keeps would lose recall, and quietly, and so far the over-fetch depth has been validated on one dataset only. Second, turbovec can calibrate its quantizer to your data (TQ+). raggio does that exactly once, automatically, when a collection reaches 10,000 vectors. Calibrating later means re-quantizing codes that are already quantized, which measurably loses recall, so the policy is: calibrate early or never.
Idea 2: BM25 without WAND
Vector search finds meaning, and it’s also weirdly bad at the exact strings people actually care about: a contract number, an error code, a surname. BM25, the classic keyword ranking, has the opposite problem. So raggio’s hybrid mode runs both legs at the same time and merges them with Reciprocal Rank Fusion: every document scores 1 / (60 + rank) in each list where it appears, and the scores add up. Only positions count, so there’s no need to reconcile a cosine similarity with a BM25 score (good luck with that), and a document found by both legs beats one that only a single leg loved.
Tap a document in either leg, then move it with the buttons or the arrow keys. raggio fuses 100 candidates per leg, the same depth Weaviate uses, and six are shown here. The constant 60 comes from Cormack, Clarke and Büttcher (2009).
The text leg lives in SQLite’s FTS5, inside each collection’s database. FTS5 is excellent, with one gap that matters here: it has no WAND. WAND (short for “weak AND”, from Broder et al. (2003)) is how dedicated search engines avoid scoring documents that can’t make the top k. Every query term has a ceiling, the most it can add to any document’s score, and the engine keeps track of the score it takes to enter the current top k. A document whose terms can’t reach that score even at their ceilings gets skipped without being scored. BM25 gives near-universal words almost no weight, so their ceilings are tiny, and documents that match only those words are skipped almost for free. Ding and Suel’s BlockMax WAND sharpens the ceilings by keeping one per block of each term’s posting list, instead of one per term.
FTS5 instead scores every row that matches any query term. Put one near-universal word in a query and it ranks most of the corpus: 540 to 670 ms per query on an earlier 553k-vector run, seconds on 2.5 million rows.
My first fix was pruning the query: keep tokens rarest-first until their document frequencies add up to 2% of the collection, and drop the rest. It worked beautifully on the first corpus, then on real arXiv titles it quietly fell apart. The median title has 10 tokens and pruning kept 3, so searching for a paper by its exact title ranked papers by one or two rare words, and the paper itself sank. The hybrid “text-hit” rate (how often the paper whose title you typed makes the top 10) was 0.732, against Weaviate’s 0.978. Ouch.
The second fix keeps the pruning but demotes it. Stage 1 only generates candidates, at a bounded cost: a ranked OR over the kept tokens (top 500), plus an unranked AND over every token (up to 1,000), which FTS5 answers by intersecting posting lists, so its cost follows the rarest word. Stage 2 re-scores those candidates in Python with the full query, using FTS5’s own formula and constants (k1 = 1.2, b = 0.75, same IDF), plus a small bonus when query words appear next to each other and in order, a light version of Metzler and Croft’s sequential dependence model.
Scenes 1, 2 and 4 use made-up document frequencies and titles to show the mechanism. Scene 5 is measured: 120 real titles on the 2.55M-row collection, counting how often the target lands in the text leg’s top 5, which is what the fused top 10 comes down to.
One rule I learned the hard way: never ask FTS5 to rank a query restricted to a list of row ids. Its bm25() function rebuilds its IDF cache for every row it checks, so ranking 200 candidates that way took 1,155 ms instead of 2.7. Filtering by row id is fine, ranking has to happen in Python.
The result is a text-hit rate of 0.984 in the full benchmark, up from 0.732. Before believing that number I wrote an audit of it, and it’s worth a summary, because it’s less flattering than the headline:
- About 80% of the gain comes from full-query rescoring, which is a real fix for a real bug
- The AND branch is at its best case in this benchmark, because the queries are exact titles: a single typo and it comes back empty
- The bigram bonus uses a weight (0.2) higher than the canonical model’s, validated on one corpus
- 0.984 against 0.978 is three queries out of 500. That’s a tie, and I’m calling it a tie
Idea 3: no index is the default index
The default answer to “how do I search millions of vectors fast” is an approximate index, usually HNSW: every vector becomes a node in a layered graph, and a search hops from node to closer node. It’s fast and accurate, but the graph and the (usually uncompressed) vectors need to live in RAM, and every hop is a random memory read that depends on the previous one, so throwing more bandwidth or more cores at it doesn’t help much.
A flat scan is the opposite: it compares the query with every stored code, in order. That’s linear in the size of the collection, sure, but it reads memory sequentially, and raggio batches concurrent queries into a single pass over the codes. On an earlier 553k-vector run the 4-bit flat scan answered in 9.0 ms at the median against HNSW’s 21.8 ms, inside a 1 GiB container against Weaviate’s 8 GiB.
So of course I tried adding an IVF index anyway, ScaNN-style: k-means splits the collection into cells, each cell becomes a small quantized index (raggio calls them shards), and a query scans only the nprobe cells with the closest centroids. At 553k vectors the best configuration that kept recall was 1.63x faster, below the 2x bar I had set for myself, and most configurations were slower than the flat scan. Every probed shard has a fixed cost of about 0.4 ms, and an indexed search can’t batch concurrent queries like the flat scan does.
But linear is linear, and eventually it catches up with you. On a 2.2M-vector probe the flat scan took 18.1 ms per query, and IVF with nprobe=16 took 5.1 ms at a recall of 0.980. So IVF became an optional object: you attach it to a collection with one POST /collections/{name}/index, remove it with a DELETE, and both run online through the job queue while searches keep being served. That’s only possible thanks to the fp16 copies from idea 1: turbovec can’t turn codes back into vectors, so every rebuild, in either direction, starts from the originals on disk.
The toy splits 1,200 points into 64 cells. Probing 4, 2 or 1 of them scans the same fraction of the data as nprobe 16, 8 or 4 of the real index’s 256 cells. The toy’s recall is harsher than the measured one (two dimensions are not a thousand), so look at its shape, not its numbers. The measured column comes from the 2.2M-vector probe in the indexing docs.
On the full 2.55M arXiv run, attaching the index cut median search latency from 23.0 ms to 11.2 ms, with recall@10 going from 1.000 to 0.994. It isn’t free, though: concurrent throughput went slightly down (129 to 116 queries per second), filtered searches got much slower (29.3 to 72.3 ms, because large filters have to be intersected shard by shard), and building the index took 5.6 minutes. My rule of thumb: below about a million vectors, don’t index. Above about two million, if single-query latency is what you care about, attach one. In between, try both on your own data.
ACID 202
None of this matters if the data isn’t there after a crash. Ingest in raggio is asynchronous: you POST documents and get back 202 Accepted with a job id. The contract is that the payload is written to the SQLite journal before that 202 goes out, and the job is marked done only after the vector index has been synced to disk. If the container dies in between, the job is replayed on boot, and since every record is upserted by id, replaying it twice still leaves exactly one copy.
Press “Pull the plug” to kill the container halfway through a job and watch the replay. The lanes match the code: the API journals into meta.db, and a per-collection worker does the rest.
Getting that contract right took two bugs I’d rather not have found under load.
The first one: a SQLite write connection loosely shared between the event loop and the worker threads let statements join each other’s transactions, and a full-corpus ingest lost one journaled job out of 2,211. The client got its 202, and the row was gone. Not great. The fix is strict single-writer discipline: every write transaction runs entirely under one lock, in a worker thread, never on the event loop.
The second one: PRAGMA auto_vacuum is silently ignored if you set it after journal_mode=WAL, so the database kept every page the ingest backlog ever touched and grew to 7.3 GB, 93% of it empty. And as a bonus, Python’s sqlite3 execute() steps a statement only once, so incremental_vacuum freed exactly one page per call until I added a .fetchall(). After both fixes the same database is 990 MB.
Durability made ingest slower, and I took that trade on purpose.
The scorecard, including the losses
Everything below comes from one run on an NVIDIA DGX Spark: 2,549,119 arXiv abstracts embedded with Qwen3-Embedding-0.6B at 1,024 dimensions, raggio capped at 4 GiB, Weaviate 1.38.11 capped at 32 GiB with its defaults (HNSW, uncompressed float32). Both were queried over REST, with the same vectors and the same 500 held-out queries.
Each bar sits inside an outline that shows its container’s memory cap. The caps were fixed before the run and enforced by podman, so an engine that outgrew its cap would have failed, not quietly used more.
raggio wins on memory (13x less under load), recall (1.000 against 0.995), tail latency for vector search (p95 28.7 ms against 50.3, p99 33.9 against 82.5) and, barely, disk (16.5 GB against 18.4). The hybrid text-hit rate is a tie. Most of the rest goes to Weaviate, and not by a little. The only row raggio takes back needs the optional index:
| Metric (2.55M vectors) | raggio | raggio + IVF | Weaviate |
|---|---|---|---|
| Vector search, median | 23.0 ms | 11.2 ms | 15.3 ms |
| Queries per second, 8 in flight | 129 | 116 | 1,050 |
| Filtered search, median | 29.3 ms | 72.3 ms | 12.8 ms |
| Hybrid search, median | 104 ms | 91.4 ms | 35.8 ms |
| Ingest | 1,357 vec/s | n/a | 2,614 vec/s |
| Cold start to first query | 30.4 s | 28.3 s | 12.8 s |
None of this is mysterious, it’s architecture. raggio is a single Python asyncio process, while Weaviate is Go and uses every core in parallel, which is why its throughput under concurrency is about 8x higher. The hybrid gap is the price of FTS5 not having WAND, paid in stage 2. Ingest journals every batch and syncs the index after every job: batching those syncs is a known lever I haven’t pulled, because durability was the point. And a cold start has to load and reconcile the index before it can answer the first query.
Now the caveats, because I’d rather say them myself before someone else does. It’s one machine, and the DGX Spark’s unified memory is unusually kind to bandwidth-bound scans. It’s one embedding model and one kind of document. And it compares defaults with defaults: Weaviate can compress vectors too, and a tuned Weaviate would use far less than 23 GB. The full table, the method and the steps to reproduce it are on the arXiv benchmark page.
How it was built
A short note on process, because it shaped the result more than any single idea. Every performance decision in raggio has an architecture decision record with the measurement that justified it, and the probe scripts that produced those measurements are committed next to the code. Rejected ideas get recorded too, numbers included. Coding agents did a lot of the typing and some of the adversarial reviewing, while my job was mostly deciding what to measure, and not believing a number until it had been reproduced.
If there’s one habit I’d recommend: when you win a benchmark, audit the win. ADR 0003 ends with an “anti-benchmaxxing” section that takes its own text-search gain apart, and it’s the part of the repo I trust the most.
Try it
raggio is open source under Apache-2.0, on GitHub, with documentation covering the API, search modes, storage, sizing and both benchmarks. It runs anywhere podman or Docker does. I’d especially like to hear from anyone running the benchmark on hardware that isn’t a DGX Spark, or on a corpus where the rescoring floor turns out to be too shallow. Issues and pull requests are welcome.
Big vector databases are built for a world where RAM is cheap and the index is always hot. Most of the places where I want retrieval to happen don’t live in that world, so the whole trick, if there is one, is this: keep the approximate thing small, keep the exact thing on disk, and pay for precision only where someone is actually looking.