How to Track DNA Matches in a Spreadsheet (So the List Becomes Evidence)

Your results came back. There are several thousand matches. The top dozen are people you recognise, and then there is a long tail of names you have never heard of, each with a number next to it.

Six months later the list looks exactly the same. You have scrolled it eleven times. Nothing has moved.

This is the standard failure mode, and it is not a knowledge problem — it is a tracking problem. A match list sorted by shared centimorgans is sorted by size, and size is not the same as what to do next. The biggest match is usually the one you already know. The most valuable match is the biggest one you cannot place, and no default view surfaces that.

Here is the column layout that fixes it, and the one status column that does most of the work.

The Tracker Layout

Column Example What it does
Match Name R. Kavanagh As displayed. Initials are fine — this file may be shared.
Platform Site A Same person can appear on several; this prevents double-counting.
Shared cM 892 The raw number. Everything else is derived from or tested against it.
Segments 31 Context for the cM figure.
Estimated Band 1st cousin Read from an editable lookup — see below.
Tree Available? Yes — 40 people Decides whether contact is worth the message.
Assigned Ancestor (blank) The column that matters.
Side Unknown Paternal / maternal / unknown. Narrows the search enormously.
Shared Matches 3 — see notes Who else matches both of you.
Next Action ASSIGN ANCESTOR Derived, not typed.
Contacted No
Notes Reasoning, dead ends, what they replied.

The Column That Does the Work: Assigned Ancestor

Everything above is context. Assigned Ancestor is the research.

A DNA match is not evidence of anything until it is attached to a named person in your tree. Until then you have a stranger and a number — an interesting fact about biology that contributes nothing to your genealogy. The moment you can say “this person and I both descend from Mary Hale, born about 1870”, the match becomes usable: it confirms a line, or it contradicts one, and either is progress.

So the Next Action column should be derived from that cell rather than typed:

The effect of deriving it is that a large unassigned match cannot quietly disappear down the list. It keeps raising its hand. In a working file, the count of close matches attached to nobody is one of the two or three numbers genuinely worth putting on a dashboard — it is a direct measure of research owed.

Set the threshold yourself. There is no universally correct number, and it depends on how much of your tree is already built and how much time you have.

The Estimated Band, and Why It Should Be a Lookup You Control

Every platform converts shared cM into a suggested relationship. Those suggestions are useful and they are routinely over-read.

Two things are true about them:

The ranges overlap heavily. The amount of DNA two people share in the same relationship varies considerably from pair to pair, because of how recombination works. The practical consequence is that a single mid-range figure is simultaneously consistent with several different relationships. A band is a set of hypotheses ranked by likelihood, not an identification.

Endogamy inflates shared cM substantially. If your ancestors come from a community that intermarried across many generations — and a great many communities did — you and a match may share DNA through several distant paths at once. The figure comes out looking like a much closer relationship than the paper trail supports. If this applies to your lines, treat every band as an upper bound on closeness and lean harder on records.

The design conclusion: keep the band in an editable lookup table in your own file, not baked into a formula. You should be able to widen the bands, shift them for an endogamous line, or add your own notes about which ranges you have learned not to trust. A hard-coded band is a borrowed opinion you cannot argue with.

Working a Match: The Four-Step Loop

A repeatable process beats inspiration, especially for the matches you have already stared at.

1. Establish the side. Do you share this match with known paternal relatives or known maternal ones? Shared matches answer this faster than anything else, and it halves the search space in one move. Record it even when the answer is “unknown” — that is a status, not a blank.

2. Read their tree for surnames and places, not for people. You are unlikely to recognise a name. You are much more likely to recognise a surname in a parish you are already working in. Note the overlap in the notes column even when you cannot use it yet, because the same surname turning up on three separate matches is a pattern, and patterns are what a spreadsheet is for.

3. Form a hypothesis and write it down. “Possibly via the Cork Hales — would make them a 2C1R.” Written in the moment, while the context is in your head. An unwritten hypothesis has to be rebuilt from scratch on every visit, which is why unassigned matches stay unassigned.

4. Test it against records, then assign — or disprove and say so. The match becomes evidence when documents connect you both to the same named ancestor. If the hypothesis fails, record the failure in the notes. A disproven theory kept is a theory you do not get briefly excited about again next spring.

A Worked Example

A hypothetical row, to show the shape:

R. Kavanagh — 892 cM, 31 segments, Site A. Band reads 1st cousin. Assigned Ancestor: blank. Side: unknown. Tree available: 40 people.

892 cM is a close-family-range match. Whoever this is, they are not a distant curiosity. And they are attached to nobody in the tree.

That single row outranks every other research task in the file — not because the number is large, but because a close match with no assignment means there is a branch of the tree you are wrong about, or one you do not know exists. The records can wait. This cannot, in the sense that it is the fastest route to finding out something substantial.

(Assumption: a fictional match, used to show how the status columns interact. Real close matches deserve more care than a worked example implies.)

Before You Share Any of This

Two cautions that belong in the same breath as the method.

DNA results reveal things families did not know. Misattributed parentage and previously unknown close relatives turn up far more often than most people expect. This is worth thinking about before testing, and worth thinking about again before messaging a close unassigned match, because your research question may be someone else’s private life.

Living people need protecting. A match tracker naturally accumulates living relatives’ names, locations and approximate ages, and a family tree file holds birth dates, maiden names and birthplaces — which is, almost exactly, the standard set of security questions. Keep the file offline if you can, use initials for matches where a full name adds nothing, and think carefully before uploading a tree that includes living people.


Where the match tracker fits in the rest of the research file — the individuals table it assigns ancestors from, the sources that confirm the paper trail, and the research log that records the theories you disproved: how to build a family tree spreadsheet that still calculates at five generations.


Featured on ReadySheetGo

Family Tree, Genealogy Research & Ancestry Record Organizer — $13.99

A dedicated DNA Matches tab holding shared cM, segments and platform, with a relationship band read from a lookup table you can edit yourself — then the part every other sheet skips: a Next Action column that keeps asking ASSIGN ANCESTOR, CONTACT, Follow up or Confirmed, and never lets a close-family match sit unattached at the bottom of the list. The count of unassigned close matches is raised on the Dashboard.

It connects to the rest of the file rather than sitting alone: assign a match to a Person ID and it joins the same tree the Pedigree builds from, backed by the Sources tab’s Proven / Probable / Possible / Disproven rating and the Research Log that records the theories you tested and disproved.

Plus Individuals with a completeness score and six row checks, a five-generation pedigree with 31 ahnentafel-numbered slots, Census Grid flagging where an enumerator’s recorded age disagrees with the birth year you hold, Timeline, Interviews ranked by age, Heirlooms, Research Costs, and a Dashboard running 24 checks against your own entries. 16 tabs, 15,700 formulas, a fictional sample family loaded.

Works in Excel and Google Sheets. No macros, no add-ons, no subscription — and it works offline, which matters more than usual for a file holding living relatives’ details.

The relationship estimated from shared centimorgans is a probability range, not an answer. Bands overlap heavily and endogamous populations inflate shared cM substantially. Every band in this file is a hypothesis to test against records.

Get the Family Tree & Genealogy Research Organizer →

Frequently Asked Questions

How accurate is the relationship a DNA site estimates from shared cM?

It is a probability range, not an answer, and the ranges overlap heavily. A single mid-range shared-cM figure can be consistent with several quite different relationships at once — a second cousin, a first cousin twice removed and others — because the amount of DNA shared varies substantially between pairs of people in the same relationship. Endogamous populations, where communities intermarried over many generations, inflate shared cM further. Treat every estimated band as a hypothesis to test against records, never as a conclusion.

What should you actually do with a DNA match you do not recognise?

Give it a row and a status rather than leaving it in the site's list. The useful statuses are ASSIGN ANCESTOR (large enough to matter, not yet attached to anyone in your tree), CONTACT (worth messaging), Follow up (messaged, waiting) and Confirmed (attached to a named ancestor with supporting records). A match is not evidence until it is attached to a specific person in your tree — until then it is a stranger with a number next to their name.

Why do close DNA matches get ignored?

Because match lists are usually sorted by shared cM and then read from the top until attention runs out, which means the ones you cannot immediately place get scrolled past repeatedly. A tracker sorted by status rather than size fixes this: a match with no ancestor assigned stays visible regardless of how many times you have already looked at it. The largest unassigned match in a list is almost always the highest-value research lead available.

Can DNA results reveal things a family did not know?

Yes, and more often than most people expect — misattributed parentage and previously unknown close relatives both turn up regularly. It is worth thinking about before testing, and worth real care before sharing a tree or a match list that includes living people. Birth dates, maiden names and birthplaces are exactly the information used to answer security questions, so a file full of living relatives deserves more caution than a file of people born in 1870.

25 of 31 Ancestor Slots Filled. 11 People With No Source Behind Them.

The Family Tree, Genealogy Research & Ancestry Record Organizer — 16 linked tabs and 15,700 working formulas, with a fictional sample family already loaded — 30 people, 26 sources, 12 DNA matches, 23 census rows and a tree reaching 1836 — so you can see the model running before you type anything. An **Individuals** tab is the single master list every other tab reads from: one row per person, entered once, carrying status, age, source count, a completeness score out of eight and a row check that catches a death dated before a birth, a parent ID that is not in the sheet, a lifespan over 110, a person listed as their own father, and the circular parent link that would otherwise fill a pedigree with the same two people forever. A **Pedigree** tab builds five generations and 31 ahnentafel-numbered ancestor slots entirely from the Father ID and Mother ID columns — change the root person and the whole chart redraws — reporting completeness per generation and showing empty slots in red, because an empty slot is not an error but a person who existed and whose name has not been found yet. A **Sources** tab rates every record Proven, Probable, Possible or Disproven, which is the mechanism that stops a family story quietly becoming a fact, and flags anything still being leaned on after you disproved it. A **Research Log** holds the open question, the record to search next and the target date — and records what you searched and found NOTHING in, so the same dead end is not paid for twice. A **DNA Matches** tab reads a relationship band from an editable lookup and then keeps asking the question every other sheet skips: ASSIGN ANCESTOR, CONTACT, Follow up or Confirmed, so a close match never sits unattached at the bottom of a list. A **Census Grid** records each entry exactly as the enumerator wrote it, misspellings included, derives the birth year the recorded age implies, compares it against the year you hold and flags the gap at a tolerance you set. Plus **Timeline** with age-at-event checks, **Interviews** ranking living relatives by age and raising anyone past your urgency threshold, **Heirlooms** flagging valued objects with no photograph, **Research Costs** with renewal warnings and an annual run rate, and a **Dashboard** running 24 data checks against your own entries. Every historical date is stored twice — as written ("abt 1870", "bef 1901") and as a plain year in a number column — so the arithmetic works in the 1700s exactly as it works for last year. Works in Excel and Google Sheets. No macros, no add-ons, no subscription, works offline. It is a research organiser, not genealogical proof: a clear check means your entries agree with each other, not that they are correct.

View on Etsy — $13.99