Can We Trust AI to Read Old Norwegian Handwriting?
Perhaps I am becoming the party pooper of AI genealogy. My interest in AI transcription of old Norwegian handwriting is not because I am opposed to using artificial intelligence in genealogy. Quite the opposite. I use AI myself, and I am quite open about it.
I use AI for sourcing and research, and when I need help organizing information. I also use it when writing for my blog. Since English is not my first language, I have no problem using AI to edit my English, smooth out awkward passages, and help me express what I am trying to say clearly. The research, ideas, and conclusions remain mine. AI is simply another tool in the process.
I become much more cautious when I ask AI to read historical handwriting.
That is a different matter. An AI can produce a transcription that looks convincing, but the important question is whether it’s actually correct. If I cannot read the original well enough to check it, how can I know?
That question drove a small experiment I ran with several AI tools.
Two documents, two very different tests
I used two Norwegian historical documents for the experiment, and it is important not to treat them as simply two examples of the same thing.
The first was a page from the Solør og Østerdalen sorenskriveri, Tingbok 1702, pages 1b–2a. This is the difficult one. The handwriting is hard for me to decipher. I can pick out individual words and recognize some letters, but I cannot confidently read the entire passage.
There is an important distinction here. I am Norwegian, and I am familiar with the older Danish-Norwegian written tradition found in Norwegian historical sources. The spelling, vocabulary and grammatical forms are not the main obstacle for me. The problem here is reading the handwriting itself.
The second document was a page from the Veøy church book, 1765–1799, pages 238–239. This is a very different proposition. I can read this handwriting with relative ease. I can follow the text, recognise the names and understand what the entries are saying.
That difference is central to the experiment.
The Solør document tests whether AI can help when I cannot confidently read the handwriting myself. The Veøy document allows me to do something much more useful: check the AI against a document that I can actually read.
Those are not equivalent tests. In the first case, the AI can potentially give me access to information that I am struggling to extract myself. In the second, I can assess whether the machine has actually read what is on the page.
That also exposes a problem with many demonstrations of AI handwriting recognition. If the person demonstrating the system cannot read the original document, they can show that the AI has produced a plausible transcription, but they cannot necessarily demonstrate that it is an accurate one.
The Solør court book: when I cannot read the whole thing
The Solør court book is the more interesting document from the point of view of practical genealogy.
I am comfortable with the language used in the document. I understand the older Danish-Norwegian written tradition, including its spelling, vocabulary and grammatical forms. If the words were clearly written, I would have no particular difficulty understanding them.
The difficulty is the handwriting.
I can read enough to recognise individual words and to get some sense of whether an AI transcription is plausible. That is useful. If an AI produces a name or phrase that I can also identify on the page, I have at least some evidence that it is reading the manuscript rather than simply generating convincing historical language.
But I cannot check every word.
That leaves me in an uncomfortable position. An AI may produce a transcription that looks excellent, and I may be able to confirm several parts of it, while still having no way of knowing whether the difficult sections are accurate.
That is precisely where AI transcription of old Norwegian handwriting becomes both useful and potentially dangerous.
It may give me a way into a document that I would otherwise struggle to understand. But the parts I most need help with are also the parts I am least able to verify.
The Veøy church book: when I can check the machine
The Veøy church book gives me the opposite experience.

I can read this document myself. I can follow the entries and recognize the names, dates and descriptions. That means that when an AI produces a transcription, I am not dependent on it to tell me what the page says. I would never used AI on this document in my real workflow.
I can test it.
This makes the Veøy document a much better test of accuracy.
If Claude gives me a particular name, I can look at the handwriting and decide whether the letters support that reading. If Copilot produces a strange word, I can check whether it is actually on the page. If Transkribus gives me something that looks plausible but does not correspond to the handwriting, I can see the problem immediately.
The difference may sound subtle, but it isn’t.
With the Solør document, I am asking:
Can AI help me read this?
With the Veøy document, I am asking:
Did AI actually read this correctly?
The first is a test of usefulness. The second is much closer to a test of accuracy.
For genealogy, those are two very different things.
What the AI tools produced
Claude produced the strongest continuous transcription in my test. Its attempt at the Veøy church-book entry was surprisingly detailed. It identified material concerning a church service, Tresfjord Church and Væsnes Church, as well as marriages, betrothals and deaths. It also produced several recognizable names, including Marit Larsdatter, Christopher Andersen, and Ole Olsen.
That makes the result genuinely useful. It gives a researcher something to work with and provides considerably more information than simply saying that the handwriting is too difficult to read.
At the same time, I would not treat the transcription as a finished piece of evidence. There are uncertain readings, and some words appear to be interpretations rather than readings that can be accepted without checking the manuscript.
ChatGPT took a more cautious approach to the difficult material. Instead of producing a complete transcription of the harder document, it acknowledged that it could only read parts of the passage reliably and avoided guessing at names and places.
Copilot recognized some words and fragments from the church-book material, but its longer transcription became increasingly unreliable.
Transkribus was particularly interesting because it is designed specifically for historical handwriting. In this test, however, its first output contained substantial amounts of text that were clearly incorrect, while the second produced a mixture of recognizable names and badly distorted words.
I would not take this experiment as evidence that Transkribus cannot handle Norwegian historical handwriting. The model being used matters considerably. What the test does show is that a specialist handwriting-recognition system does not automatically produce a transcription that can be trusted.
The real problem is the plausible mistake
The easiest AI transcription errors are not necessarily the most worrying ones. If a system produces obvious nonsense, I can reject the result.
The more serious problem is a transcription that is mostly correct but contains a few plausible mistakes. A surname may be wrong. A farm name may be misread. A place name may be changed into another place that looks reasonable. A relationship may be misunderstood.
The resulting text can still look perfectly convincing.
That is particularly problematic when the researcher cannot read the handwriting independently. The AI may be supplying the only version of the text the researcher has available. There is then no reliable way to distinguish what the machine actually read from what it guessed.
This is the problem I am most concerned about.
An obviously bad transcription is unlikely to become part of someone’s family history. A transcription that is ninety-five per cent correct is a different matter. The five per cent that is wrong may be the name, place or relationship that determines where the researcher looks next.
That is why I keep coming back to the same question: how do we know when the AI is wrong?
When I look at AI transcription of old Norwegian handwriting, I am therefore less interested in how impressive the finished paragraph looks than in whether I can compare it with the original manuscript.
The problem is not simply knowing the language
It is also worth making clear that this is not primarily a problem of understanding old Norwegian.
I understand the Danish-Norwegian written tradition used in these records. I am familiar with the spelling, vocabulary and grammatical forms. If I can see the words clearly enough, I can understand them.
But language knowledge does not automatically make difficult handwriting readable.
The Solør document demonstrates that quite neatly. I can understand the kind of language I am looking at, but there are parts of the handwriting that I cannot confidently decipher.
The Veøy document is different because the handwriting itself is accessible to me.
That is why anyone using AI transcription of old Norwegian handwriting needs to be particularly careful about the difference between understanding a transcription and verifying a transcription.
You may understand every word in the AI’s output and still have no way of knowing whether those were the words written in the manuscript.
Perhaps I am the party pooper
There is considerable enthusiasm for using AI in genealogy, and I understand it.
AI can translate difficult passages. It can explain unfamiliar terminology, provide historical context. and suggest research strategies and help organize a complicated investigation. Those are all uses I have written about before, and they are uses I will continue to explore.
In my earlier article, Using AI in Norwegian Genealogy Research, I looked at practical ways AI can assist with genealogical research. I also wrote about Using AI to Build a Norwegian Genealogy Research Strategy, where I explored how AI can help structure a research strategy and identify possible avenues to investigate.
Those articles are not arguments for accepting everything AI produces. They are about using AI as a research tool.
That distinction is important here.
There is a danger of confusing the ability to produce an answer with the ability to verify an answer.
AI has made it much easier to obtain a proposed transcription of difficult handwriting. It has not necessarily made it easier to determine whether that transcription is accurate.
And perhaps that is where I become the party pooper.
While someone else is saying, “Look how well AI can read this!”, I am inclined to ask, “Can you read it well enough to know whether AI has read it correctly?”
It is not a particularly glamorous question, but genealogy has enough wrong ancestors without us manufacturing a few more.
I am still going to use AI
None of this means that I am abandoning AI transcription. I certainly will not.
If I am struggling with a document, I will ask an AI to have a go. It may identify a word I have missed or suggest a reading that helps me understand the passage. It may give me enough information to locate another source that can confirm or reject the reading.
That is useful.
If I can read the document myself, I can use AI as a second opinion. If I can read only part of it, I can use the AI output as a starting point and concentrate on checking the uncertain sections. If I cannot read the handwriting at all, I can still use the transcription as a lead, but I need to be much more cautious about treating it as fact.
The important thing is to keep the roles clear. The original document is the source. The AI transcription is an interpretation of that source.
That is also how I approach the other ways I use AI in genealogy. I am happy to use it to help me understand a source, decide where to search next, explain unfamiliar terms, organize information and improve my writing. What I am not prepared to do is let the AI quietly become the source itself.
The genealogist still has a job
Perhaps this is the part that makes me the party pooper.
When I look at AI transcription of old Norwegian handwriting, I am not particularly interested in whether a machine can produce a page of text that looks convincing. What interests me is whether a genealogist can determine whether that text actually corresponds with the manuscript.
The Solør court book and the Veøy church book have shown me why that distinction matters. With Solør, I can use AI to help me approach handwriting that I struggle to read. With Veøy, I can use the same technology and then check its work against my own reading.
Those are two very different relationships with the machine.
AI has made it much easier to obtain a proposed transcription of difficult historical handwriting. It has not necessarily made it easier to determine whether that transcription is accurate.
For me, that means I will continue to use AI for genealogy, just as I use it for sourcing, research, editing and improving my English. I simply draw a line when the machine’s interpretation of an original source cannot be checked.
The best approach, as far as I can see, is to use the machine, question the result, and go back to the original whenever possible.
The real test is not whether AI can produce a transcription. It clearly can. The real test is whether we can tell when the transcription is wrong.
For that, at least for now, I still want a human being who can read the document. Preferably someone who understands both the handwriting and the language.
That human being might be me. It might be another genealogist.
But I am not quite ready to let the machine have the last word.
I don’t normally write “Like and share” at the bottom of my blog posts. In this case, though, I am inclined to ask you to share this one.
I have already seen some grim results from the misuse of AI transcriptions. If we use these tools without checking their work, we risk filling the internet with yet more erroneous genealogical information. Once those errors are copied from one website, database or family tree to another, they can become surprisingly difficult to get rid of.
So please share this article, especially with genealogists who are beginning to use AI for historical documents. A little scepticism now may save a lot of corrections later.



What is really needed is a way to train the AI to recognize the handwriting of a particular scribe. I once asked ChatGPT whether it could learn a scribe’s handwriting if I provided both the original handwritten document and a verified transcription of the same passage. It assured me that it could, but when I tried it, the resulting transcriptions were not accurate enough to be useful.