Abstract

Every published RAG tutorial begins at the same moment: the corpus is already in the vector store. The documents arrived, the chunks exist, the embeddings are computed. The retrieval begins there. What no tutorial covers is the bridge — the layer that takes a live, growing, real-world data source (an email archive, an export, a mailbox) and moves it into a vector store without loss, without duplication, and without breaking when the machine it runs on dies mid-run.

This paper documents that bridge: the msgvault adapter, which moves a SQLite email archive into a local Chroma vector store with RFC822 message-id normalization, watermark-native checkpointing, and crash-survival engineering. A five-pass prior-art search (GitHub; Reddit/Hacker News/PyPI; an unweighted problem-space search; an adversarial pass using an independent AI collaborator; and an API-driven sweep of Hacker News and PyPI, October 2026) found the components scattered across repositories and tutorials — normalization in one PR, deterministic IDs in another — but no published, resilience-engineered bridge between a SQLite mail archive and a local vector store. This is that bridge, published.

I. The Gap in the Literature

The retrieval-augmented generation literature concerns itself with what happens after ingestion: chunking strategies, embedding models, retrieval quality, reranking, generation. The ingestion layer is treated as a solved problem — a for-loop over documents.

For a personal archive, ingestion is not a solved problem. It is the problem. The archive is live: emails arrive daily, the source database grows, the embedding model runs on the same 8GB machine that also runs the operator’s other work. The pipeline will be interrupted — by crash, by curfew, by the operator needing the machine for something else. Every interruption is an opportunity for duplication, corruption, or silent loss.

The nearest published neighbors confirm the gap. RAGmail [1] ingests massive mailboxes into a local semantic index but targets mbox, not a SQLite archive, with no curfew or per-message crash recovery. Email-RAG [2] has the strongest deduplication model — Message-ID identity preventing duplicates across formats — but runs PostgreSQL/pgvector, not ChromaDB, with no time-budget shutdown. LlamaIndex’s IngestionPipeline [3] offers docstore deduplication, but its default transformation cache is in-memory and resets on restart, re-embedding identical documents. WPipe [4] delivers durable SQLite WAL checkpointing for generic ETL, with no email, embedding, or vector-store awareness. async-doc-ingestion [5] demonstrates crash-safe idempotent resume as a reference architecture, not installable infrastructure. Each solves a piece; none combines a SQLite mail source, ChromaDB, crash-safety, curfew-awareness, and guaranteed deduplication.

Prior-art search was conducted across five passes between October 4 and October 10, 2026: GitHub, Reddit/Hacker News/PyPI, an unweighted problem-space search, an adversarial pass using an independent AI collaborator (Leo, Per Scholas AI stack), and an API-driven sweep of Hacker News and PyPI. The adversarial pass, conducted blind with repo names withheld, surveyed the five repositories above and found the combination absent: “the pieces exist, but the combination — SQLite email source → ChromaDB, crash-safe, curfew-aware, dedup-guaranteed — is not a published tool” [6]. The same pass found the author’s own repositories did not surface in blind keyword searches, attributing the failure to distribution surface area rather than merit — the discovery failure this paper’s publication addresses. No published combination of these components was found at the time of writing.

II. The Architecture

Watermark-native checkpointing

The source schema carries an embed_gen column — a watermark built into msgvault itself. NULL means “needs embedding”; stamped after successful upsert. The checkpoint lives in the data, not in a sidecar file. Crash, restart, curfew, resume — the watermark holds.

The companion doctrine

Every load-bearing mechanism gets a companion. The DB stamp is paired with a JSONL journal written in the same instant. If the DB stamp silently fails, startup adoption recovers the stamped IDs from the journal. No re-embedding. No loss.

Bounded embed requests

The embed request is bounded to five chunks per call after root-causing a crash: per-message requests reaching 450K characters against a 2048-token compute buffer. The bound is in the code, with the root cause documented.

Deterministic identity

Every chunk’s ID derives from the normalized RFC822 message-id (hashed). Same message, same ID, idempotent upserts. Re-ingest is structurally impossible, not merely avoided.

Curfew

The ingester stops cleanly at a configured time — stamp, exit, resume on relaunch — because unattended jobs must respect the machine’s daytime obligations. The schedule is the operator’s week, expressed as flags.

III. Results

Smoke run: 200 messages → 23,500 chunks, ~25 minutes, on an 8GB Intel MacBook Air. Run 2 started at row 201 — rows 1–200 never re-ingested. Progress lives in the data itself.

IV. Conclusion

The hardest part of sovereign RAG is not retrieval. It is getting the corpus in, whole, without loss, on hardware that serves two masters. The bridge is published: AGPL-3.0, with a commercial lane.

And the bridge now runs in both directions. Since this paper’s architecture was built, the same practice has produced an orchestrator that dispatches each question to four workers — three vector collections and the source SQLite archive itself — and a reconstruction tool that returns a document from the store the way it went in: whole, in order, seam-deduplicated. A 4.6-million-character conversation, ingested in 5,163 chunks, has been reassembled and delivered complete. Ingestion and delivery are the same door, opened from either side.

The next paper documents that layer: the overseer pattern, the worker fleet, and the archive that answers with its owner’s own words — whole.

References

AI Collaboration Disclosure

Developed in collaboration with AI. The author directed the architecture, build, and argument; AI assisted with drafting and organization. All claims and conclusions are the author’s own.