Using GPT-6 and Opus 5.5 to trace alchemical knowledge and decode 17th century letters
44 comments
I've been doing family history on my aboriginal Australian side, there were a bunch of Lutheran missions, and ever so kindly they digitized 800 pages for me, but it was all written in German. I transcribed and translated all of it.
https://drive.google.com/file/d/1en9gDgZRP7EmSHU72k6g2Iz3RRf...
I sent it back and they didn't acknowledge it, probably for various reasons, and I doubt they would like me linking it above.
Beyond being fascinating in general, I also found that from a different state in Australia there was an aboriginal who became a man of letters came to my ancestral state and was the first to write the language of the tribe, I think (not something easy to prove) it's the first ever written version of the tribal language circa ~1870 of Kuku Yalanji.
Because if you study this stuff long enough, you'll find the LLMs hallucinate in this space far more than they do in a discourse or programming project.
It can be close enough to seem useful, but make critical errors that can cause rippling problems.
I've done a lot of the same sort of thing with Latin and chancery and while it can really help in understanding it on a personal level, it takes combing over what it has said and zooming in, comparing it with other information, and understanding the historical context. Otherwise it can completely mangle clauses, contexts, etc.
It is remarkable how well they can start to read things like chancery script, but again, many errors. Not reliable, especially in such large volumes. It would require extensive review. And they are probably getting flooded wth things like that now. And not all of them will be benevolent. History is laden with intentional corruptions and the flat, context-less reading that LLMs give can bury that even further.
Don't inundate researchers, churches, and records offices wth this stuff. But it is absolutely a space where it can help, with proper organization and engagement.
There is a farm cited as an ancestor origin by multiple deceased genealogist. Its a common POI that people with this shared ancestor try to find. Finding it could solidify the established line theory and finally debunk a small ancestor fraction's alternate theory.
Different LLM searches keep returning the same colony with invalid/made up references that don't even cite a partial name match. Manual searches throughout that area returns nothing as well. But since the LLM said so, a century+ of various research and work by professional genealogist gets severed from a crowd sourced public tree because "AI" returned a colony name that ended up helping the alt line.
I did eventually find what I believe is the origin farm, and it strengthens the history written by previous genealogist -- I tried to send the information to the tree maintainers and was outright ignored. LLMs fabricating locations apparently beats a listing in the National Heritage List for England of the exact place name, buildings from that time, and in one of the areas the larger english family is known to have been present.
I think this post is missing a lot of context and background. I haven't the foggiest what an "established line theory" is
Make sure you document why/how you determined that a previous research conclusion was a mistake -- generations after you may discover additional information and should be able to weigh the same evidence rather than simply accept (or reject) your novel conclusions.
So how do you use the models more specifically? I have so far used them to write a - in some ways - better genealogy program, with some features I've sorely missed especially with respect to DNA genealogy. I've also tried to use them to help with the tedium of transcribing horrible handwriting, but the results have been so poor and with such a high level of hallucinations that I've not tried that for a while.
- ask it to generate Lean 4 proofs based off centimorgans
- extract subject, object, subjects from all source texts and create a graph database (maybe someone mentioned a red dog in their oral history, and someone else mentioned a sick dog in a eulogy, ask ai to explore weakly correlated connections, sometimes pays off)
To date, it is the largest collection of agent-accessible translations on the web. The mcp pulls texts and illustrations — and the API provides access to the embeddings. It’s free.
We are based at the Embassy of the Free Mind in Amsterdam, a UNESCO-recognized library of alchemy, magic and mysticism (among other related topics)
And, it’s worth saying, I’m very much inspired by Res Obscura’s line of curiosity-driven humanist inquiry…
Dive into some historical mysteries, there are so many!
One question I would have, why not put the extracted English txt contents onto github or some other repository? The API and MCP are definitely awesome features, but providing an LLM ready repository seems like it would be very useful as well. Your service then still serves the purpose for original scans/images/illustrations but the bulk of the content is then immediately searchable by an agent locally using the standard cli tools (grep/rg, etc).
Read the full thread on Hacker News →
Related stories
- A 17th-century font in a 21st-century thesislinyangchen.comLobsters · 50 points · about 3 years ago
- Hacker News · 3 points · 3 days ago
- Stripe's Knowledge AI Platformstripe.devHacker News · 172 points · 8 days ago
- Hacker News · 1 points · 8 days ago
- Hacker News · 12 points · 12 days ago
- Hacker News · 3 points · 6 days ago