Case study
Local-first document intelligence — hybrid retrieval, reranking, and answers that cite the exact page they came from.
cited Q&A1 / 2
This approach grew out of a small RAG system I already run in production for a 5-person firm — a few hundred real documents (leases, receipts, applications) in Russian. This write-up is a separate, harder validation exercise: a larger English-language corpus, stress-tested the way I'd want a skeptical client to check it, before trusting the same method on the smaller real deployment. The numbers and demo below are from that stress test, not the production system.
Most "I built a RAG app" portfolio pieces show a screenshot of one good answer. That proves the happy path exists — it doesn't prove the system holds up once you stop being gentle with it. This project was checked the way a skeptical client would check it: deliberately trying to break it at scale, on content it had never seen, and — twice — going back to ask whether the test itself was honest before trusting its own numbers.
generation: local Qwen by default — DeepSeek cloud is opt-in, gated behind an explicit consent countdown before any document text leaves the server
The decline in the last two numbers is reported, not hidden — a single GPU runs reranking by design, and ranking quality on a narrow "quote the exact document ID" scenario does degrade as the corpus grows. That same 57K-document scale test is also where an eval-metric bug was found: the original "Recall@5 = 100%" headline for it was a miscounted metric, corrected to the honest 96%→73% above. The 100% on the left is a separate, smaller-corpus result the same bug never had room to reach.
Read the full validation report → · cross-domain eval methodology →
A "Recall@5 = 100%" headline turned out to be a miscounted metric — the correct document was actually landing in a 6-document window, not 5. Every number was recomputed and re-verified live against the same 57,000-document corpus before being republished as "Recall@6."
4.6% of pages across a 597-document arXiv sample carried a literal NUL byte in their extracted text — a font-encoding artifact invisible in any PDF viewer. One bad page aborted its whole document's ingestion: 36% of that sample failed outright. Fixed at the extraction layer, with a second independent guard at chunking and six new regression tests.
The two largest documents in a 762-document cross-domain corpus failed ingestion — a single upsert request exceeded Qdrant's 32 MB limit by sending every chunk in one call. Fixed by batching on actual serialized payload size instead of document count.
A code path meant to guarantee an exact-cited figure/table survives reranking only fired when ordinary search had already found it weakly — the opposite of when the guarantee was needed. Fixing it raised evidence-chunk recall from 19.9% to 95.5%.
Crash-safe by construction, not by hope. Every mutation across Postgres, Qdrant, and the filesystem goes through the same shared lock a full backup takes exclusively — proven under real concurrent load (a backup measurably blocks a competing writer, not just in theory), with deliberate crash-and-restart drills (killing the DB mid-operation, stopping the search engine mid-delete) in the test suite itself.
A backup you've actually restored, on a schedule. verify_restore.sh runs a full automated restore drill into a disposable Postgres + Qdrant, and a partial/failed restore is verified to surface as a clear failure — never a false "restored successfully."
Local by default, cloud by explicit consent. Qwen via Ollama never leaves the server. DeepSeek is opt-in per request, gated behind an administrator flag and a client-side consent countdown — and streams token-by-token from either backend, not just the local one.