Picture an archive of tens of thousands of scanned pages — typed records, handwriting, faded carbon copies, press clippings. No person could read it end to end, and the answer to the question that matters is usually hiding in three pages that sit thousands of pages apart.
This is a research engine built to interrogate collections like that: ask a question, get an answer with the exact page it came from — or an honest "not found in the archive".
How it works:
- Read without inventing. Scans become located, typed text: every word knows its page and its region on that page, and anything low-confidence goes to a human-review queue instead of being guessed.
- Understand the structure. A document map, small-to-big chunking, entity resolution and a knowledge graph turn pages into people, events, places and claims.
- Assemble before reading. Instead of pouring the whole archive into a model, retrieval lanes (dense, keyword and graph) are fused and reranked to narrow millions of words down to a small, cited candidate set — and only then does a model read.
- Answer with receipts. Every answer cites real pages, and citations are checked against what was actually retrieved.
- Find what nobody asked for. A discrepancy engine surfaces the same fact told two different ways, far apart in the archive, and a council of AI agents investigates it and reports back.
The rule it lives by: never invent text. Provenance and confidence travel with everything, and the cardinal failure — quoting the wrong document with total confidence — is measured at every step.
Under the hood (for the engineers): OCR with located regions and confidence gating, small-to-big chunking and enrichment, an entity registry and knowledge graph, hybrid retrieval with reciprocal rank fusion and reranking, an NLI screen for contradictions, multi-agent analysis, tool access for agents over MCP, a single Postgres + pgvector backbone, and an evaluation harness with committed scorecards.
Why it matters (for everyone else): it turns "somewhere in this archive" into "page 12,408, second paragraph — and here's where page 3,117 says the opposite."
Skills: Retrieval-Augmented Generation (RAG) · Knowledge Graphs · Document AI & OCR · Multi-agent Systems · LLM Evaluation