Project · Data & LLMs

Archive Research Engine

Jul 2026 – Present
Archive Research Engine

Picture an archive of tens of thousands of scanned pages — typed records, handwriting, faded carbon copies, press clippings. No person could read it end to end, and the answer to the question that matters is usually hiding in three pages that sit thousands of pages apart.

This is a research engine built to interrogate collections like that: ask a question, get an answer with the exact page it came from — or an honest "not found in the archive".

How it works:

The rule it lives by: never invent text. Provenance and confidence travel with everything, and the cardinal failure — quoting the wrong document with total confidence — is measured at every step.

Under the hood (for the engineers): OCR with located regions and confidence gating, small-to-big chunking and enrichment, an entity registry and knowledge graph, hybrid retrieval with reciprocal rank fusion and reranking, an NLI screen for contradictions, multi-agent analysis, tool access for agents over MCP, a single Postgres + pgvector backbone, and an evaluation harness with committed scorecards.

Why it matters (for everyone else): it turns "somewhere in this archive" into "page 12,408, second paragraph — and here's where page 3,117 says the opposite."

Skills: Retrieval-Augmented Generation (RAG) · Knowledge Graphs · Document AI & OCR · Multi-agent Systems · LLM Evaluation