30 translations
SQLite + JSON cache
Thirty public-domain translations pulled, normalised, tagged, correlated and exported into a single navigable graph.

Hebrew, Greek, Aramaic and Latin sit alongside the English derivatives, so agreement can be measured across the whole transmission chain rather than within one branch of it. Tap any to inspect.
The first source, getbible.net, died on an expired SSL certificate. Swapped to the scrollmapper/bible_databases repo — complete JSON per translation, no auth. Fully resumable: cached files don't re-download, ingested translations don't re-import.
Twenty themes across the shipped corpus. Stage 8 retagged to thirty and reached 253,164 assignments before the run was interrupted. Pick one to see where a theme concentrates.
A Bayesian score over cross-translation agreement. Corpus mean lands at 42.5%. It measures how strongly a verse is attested across independent transmission lines — nothing more, and the vault says so in plain language.
The shipped model uses Jaccard similarity. Stage 8 moves it to TF-IDF cosine — Jaccard lets short verses score artificially high on shared function words.
Every node is a markdown file with wikilinks, so Obsidian's graph view renders the whole corpus as a navigable structure rather than a search box over a database.
Pipeline lesson learned the hard way: never block on long sleep chains over the SMB mount, and never write logs to the share. Launch with nohup, log to /tmp, check in separately.