A look under the hood of my genealogy database (what the stats quietly reveal)
Every now and then I like to stop doing “new research” and simply look at what I’ve already built. In this post, I’ll be taking a look under the hood of my genealogy database to share what I’ve learned.
Not the exciting stories. Not the shiny discoveries.
Just the bones of the tree: the numbers, the patterns, the odd little bulges where one century dominates… and the occasional statistic that makes me squint at the screen and say, “Well, that can’t be right.”
Legacy has a Family File Statistics report that does exactly that. It’s basically a peek under the hood—what’s in the database, where it’s strong, and where it’s still held together with placeholders and good intentions.
Before I dive into the numbers, I should admit something: I’m one of those people who keeps coming back to the Genealogical Proof Standard—not because I enjoy rules for the sake of rules, but because I’ve learned (the hard way) that a family tree is only as strong as the evidence holding it up.
And yes… I am whining a little about sources.
Because in genealogy, “I saw it on someone’s tree” is not a source. A family story is not a source (even when it’s a wonderful story). And a screenshot you can’t re-find later is just a tragedy waiting to happen. If I can’t point to where something came from—properly, clearly, so someone else can retrace the steps—then it doesn’t really belong in the tree yet. At best, it’s a clue. At worst, it’s a future error that will quietly multiply.
So how does that ideal look in my own database?
Honestly: it’s a mix of discipline and messy reality. Parts of my tree are rock-solid, anchored in parish registers, censuses, probate records, and the kind of boring-but-beautiful documentation that holds up under scrutiny. But the “under the hood” statistics also reveal the places where I’ve been… optimistic. People with incomplete dates. Families that look suspiciously small. A few absurd outliers that practically scream “wrong person” or “typo.” Those aren’t just database quirks—they’re reminders that proof is a practice, not a personality trait.
In other words: I believe in the GPS, I aim for good sourcing, and my own statistics politely show me where I’ve still got work to do.
Now for the a look under the hood of my genealogy database.
The headline: a big tree… with an honest gap
My file currently holds 12,740 individuals.
That number looks impressive on paper, but here’s the more revealing detail: 2,232 people in the file have insufficient date information.
And I actually like that the report says it so plainly, because it confirms what we all know: genealogy databases often grow faster than our documentation does.
It also explains why the “Births by Era” chart totals 10,508—it’s basically showing the dated portion of the tree (12,740 minus 2,232). In other words, the era chart isn’t a perfect picture of my whole database; it’s a picture of the part that has enough dates to behave nicely on a timeline.
Where my tree really “lives”: the 1700s and 1800s
Once you look at the dated portion, the center of gravity jumps out:
-
1700–1799: 4,369
-
1800–1899: 3,559
-
1600–1699: 1,439
-
1900–1999: 983
-
Before 1600: just a handful
So if you want the plain truth: my database is a very Norwegian kind of tree—thickest in the centuries where the records are most consistently useful. It also reflects that I have been working extensively on a cluster genealogy for my childhood municipality Vestnes. I started, and am still adding records from the 1700s.
The 1700s and 1800s are where the paper trail tends to hold steady: parish registers, confirmations, marriages, burials, and a growing ecosystem of supporting sources. The 1600s exist, but it’s more “careful work” and less “steady coverage.” And the 1900s… well, that’s often where living memory and privacy considerations complicate things.
The stats that made me laugh (and then open the file)
Now for the part I always enjoy: the “this can’t be right” numbers.
My report includes:
-
a 123-year lifespan
-
and a 156-year marriage
If you’ve kept a genealogy database for any length of time, you already know what those numbers are.
They aren’t miracles.
They’re breadcrumbs leading to errors.
Usually it’s something very human:
-
a death date attached to the wrong person (same-name confusion)
-
a typo that shifts a date by a century
-
an estimated date entered as a fact
-
a duplicate marriage event or mislinked spouse
And here’s the thing: these outliers aren’t just embarrassing—they’re useful. Because the stats page hands me a practical task list: fix the biggest outliers, and the whole database becomes more trustworthy.
Marriage patterns: the tree behaves like history
My file currently has 3,565 marriages where both spouses are recorded.
But I also have 1,380 marriages with an unknown date.
That’s a reminder that even in a big database, a lot of the structure is still “floating” until we anchor it with dates.
Still, the age patterns for the marriages that are dated look familiar:
-
husbands most commonly in their late 20s and early 30s
-
wives most commonly in their early-to-late 20s
That’s not just a statistic—it’s a small window into how family formation often worked in the periods my research covers. These statistics confirms information from other sources about typical ages for marriage.
Children per family: a quiet clue about what I haven’t finished yet
This one is sneaky, and it’s the part that always feels most personal.
The report shows 9,105 children, and a surprising number of families that appear to have:
-
one child
-
or even zero children
Now, historically, small families exist. Infant mortality exists. Remarriages exist. Childless couples exist.
But when “one-child families” become the largest category in a database, it’s usually not a demographic truth—it’s a workflow truth.
It often means:
-
I’ve found the couple and the baptism of my direct ancestor, but haven’t swept the parish for siblings yet
-
some children are still sitting in my notes, not linked in the file
-
a branch is “good enough for now” but not completed
So the stats page isn’t just reporting history—it’s quietly pointing at the parts of my process that are unfinished. This number will improve a bit as I progress with my afore mentioned cluster genealogy.
Names and places: this database has a strong geographic heartbeat
If you’ve followed my blog for a while, this won’t surprise you: the place list confirms that my tree has a very clear home region.
The most frequent places include:
-
Vestnes kyrkje / Veøy prestegjeld
-
Tresfjord kyrkje
and a set of farm names that show up again and again.
Even the “common surnames” list reads like a map: Vike, Helland, Vestnes, Sætre, Aas, Daugstad.
Which is exactly how Norwegian genealogy works. In many periods, what looks like a surname is often just a place name someone carried for a while—and my database statistics reflect that beautifully.
What I take away from this (and maybe you will too)
Looking at this report, I’m reminded of three things:
-
A good tree has a shape. Mine is thick in the 1700s and 1800s because that’s where the records give me the best footing.
-
Outliers are not annoyances—they’re leads. The “123-year” people are basically neon arrows saying: check me.
-
Database stats reveal your habits. The “one-child families” don’t just describe the past; they describe where I still need to do the unglamorous sweep work.
So yes—this is a look under the hood. Not just at my ancestors, but at my own research workflow.
And if you’ve never run a statistics report on your own tree, I recommend it. It’s one of the few times your genealogy software stops being a storage box and becomes something better:
A mirror.
Conclusion
So that’s the view under the hood: not just a pile of names, but a living research project with strong beams, a few loose boards, and some parts still waiting to be finished properly.
If anything, the statistics report reminds me why the Genealogical Proof Standard matters in everyday practice. It’s easy to feel “done” once a person is added. But the real work is making sure each connection is supported by sources someone else could follow—and that the dates, places, and family structures actually make historical sense.
For me, the next steps are simple and old-fashioned: hunt down missing dates, clean up the outliers that can’t be true, and go back to those suspiciously small families to see what the records really say. Because the goal isn’t a bigger tree. It’s a sturdier one.
Maybe your genealogy software has a similar function and this article can be a inspiration to take a look at the strengths and possible weaknesses of your database.
My book Scenes from a Life Left Behind The World Our Norwegian Ancestors Knew are now available on Amazon https://amzn.to/3UlHM0N


