Sitemap

How I Inferred the Family Relationships of Nearly One Million People from Wikidata

8 min readSep 9, 2026

--

Scaling Symbolic Reasoning with Knowledge Graphs, Horn Rules and SPARQL

Press enter or click to view image in full size
Primos terceros de Enrique VIII de Inglaterra

What if you could start with only a few basic family relationships and reconstruct an entire network of kinship?

That was the idea behind this project.

I took a subset of Wikidata containing 944,334 people and used a small number of explicitly stated relationships — mainly parent-child relationships — to infer millions of additional family relationships.

The final knowledge graph contains more than 40 million triples, including over 14 million newly inferred family relationships.

The interesting part, however, is not the number of people.

It is how the inference was performed.

The project became an experiment in large-scale symbolic reasoning over a Knowledge Graph.

The Problem: Wikidata Knows More Than It Says

Wikidata contains explicit family relationships.

For example:

  • A is the child of B.
  • B is the child of C.
  • D is the spouse of A.
  • E is the sibling of B.

But the graph does not necessarily contain all the relationships that can logically be derived from these facts.

It may know who someone’s parents are without explicitly stating:

  • who their grandparents are,
  • who their uncles and aunts are,
  • who their cousins are,
  • who their nephews are,
  • or which people are related through marriage.

The information is there.

It is simply distributed across the graph.

This creates an interesting Knowledge Graph problem:

Can we derive the missing relationships automatically?

And more importantly:

Can we do it at the scale of almost one million people?

From Facts to Rules

Once the primitive relationships were available, the next step was to formalize the reasoning.

I represented family relationships using logical rules based on Horn clauses.

Consider the relationship between an aunt and her niece or nephew.

Conceptually:

If A is a woman, A is the sibling of B, and B is the mother of C, then A is the aunt of C.

In logical notation:

Woman(A) ∧ Sibling(A,B) ∧ Mother(B,C)
→ Aunt(A,C)

The important point is that the graph does not need to explicitly state that A is C’s aunt.

The relationship can be inferred.

This changes the role of the Knowledge Graph.

It is no longer just storing information.

It is executing a logical model over that information.

The First Attempt: Jena Rules

My first implementation used the rule-based inference capabilities of Apache Jena.

The approach was conceptually elegant.

I could define rules such as:

(?a rdf:type fami:Woman)
(?a fami:isSiblingOf ?b)
(?b fami:isMotherOf ?c)

→

(?a fami:isAuntOf ?c)

The inference engine would derive the new relationships automatically.

For experimentation and small datasets, this worked extremely well.

It was transparent.

It was easy to modify.

And it made the logic behind each deduction explicit.

But there was a problem.

Everything happened in memory.

And my graph contained nearly one million seed people.

The Combinatorial Explosion

This is where the project became interesting.

Inferring a parent or grandparent relationship is relatively cheap.

But as we move further through a family network, the number of possible relationships grows rapidly.

The target was not simply:

Who are this person’s parents?

I wanted to reach relationships such as:

  • grandparents,
  • great-grandparents,
  • uncles and aunts,
  • nephews,
  • cousins,
  • second cousins,
  • third cousins.

At that point, naive inference becomes extremely expensive.

A graph containing one million people does not contain one million possible relationships.

The number of potential connections is enormous.

The challenge therefore became:

How can symbolic inference be scaled without keeping the entire reasoning process in RAM?

Moving the Reasoning into the Database

The solution was to change the architecture.

Instead of allowing the inference engine to dynamically calculate everything in memory, I converted the logical rules into SPARQL materialization operations.

In other words:

Rule → SPARQL INSERT → persistent knowledge

For example, the aunt relationship became a constructive SPARQL query:

INSERT {
?a fami:esTiaDe ?c .
?a rdf:type fami:Tía .
}
WHERE {
?a rdf:type fami:Mujer .
?a fami:esHermanaoDe ?b .
?b fami:esMadreDe ?c .
}

The difference is fundamental.

The inference is no longer temporary.

The newly discovered relationship becomes part of the graph itself.

The graph learns the result and stores it.

Cascading Inference

The next problem was determining the order in which the rules should run.

The solution was a layered inference pipeline.

Level 1 — Primitive relationships

Parents and children.

Level 2 — Derived relationships

Siblings and spouses.

Level 3 — Extended family

Uncles, aunts, grandparents and grandchildren.

Level 4 — Cousin relationships

Cousins, second cousins and eventually third cousins.

Each stage uses knowledge materialized by previous stages.

This turns the inference process into a form of batch reasoning.

Instead of trying to calculate everything simultaneously, the graph is progressively enriched.

Raw Wikidata
↓
Primitive relationships
↓
First-order inference
↓
Second-order inference
↓
Extended family
↓
Cousins
↓
Third cousins

This approach allowed the system to persist each intermediate result and avoid the memory limitations of the original implementation.

From One Million Facts to More Than 14 Million New Relationships

The results were considerably larger than the original dataset.

The final graph contains:

  • 944,334 people
  • 40,593,509 total triples
  • 1,986,827 new class assertions
  • 14,647,450 inferred property assertions

The system reached relationships as distant as third cousins.

The most striking statistic is perhaps this:

Starting from roughly one million explicit parent-child relationships, the system generated more than 14 million additional family relationships.

That means that for every explicit relationship, the reasoning process produced many additional pieces of structured knowledge.

This is the real power of symbolic inference.

But How Do You Know the Inferences Are Correct?

Press enter or click to view image in full size

Generating millions of triples is easy compared with proving that they make sense.

So I built a validation mechanism around historical categories.

For example, I used a category containing Roman emperors of the 1st century and asked the graph to search for family relationships between its members.

The result provided an interesting sanity check.

The system correctly separated members belonging to different dynasties when there was no direct genealogical connection between them.

The category therefore became a kind of semantic test set for the inference engine.

This suggests an interesting approach to Knowledge Graph validation:

Use the semantic structure of the graph itself to test the consistency of inferred knowledge.

The Graph Became More Than a Genealogy

Press enter or click to view image in full size

Once the family relationships had been materialized, I could use them as a starting point for other dimensions of exploration.

The resulting application does not simply display a traditional family tree.

It allows the user to explore a person through four dimensions:

Family

Who is related to whom?

Time

When were those people born?

Space

Where were the members of the family born?

Cultural context

What historical or cultural categories are associated with them?

This transforms genealogy from a static tree into a multidimensional Knowledge Graph.

From Genealogy to Historical Network Analysis

Press enter or click to view image in full size

Once family relationships are explicit, many other questions become possible.

For example:

  • Which historical families were geographically concentrated?
  • How were political dynasties connected?
  • Which historical figures share ancestors?
  • How did family networks spread geographically?
  • Which people from a cultural category are related?
  • Where do family networks overlap with historical periods?

I also experimented with identifying couples who were already related — for example, cousins or more distant relatives.

This opens the door to studying historical endogamy and dynastic relationships using graph reasoning.

An Unexpected Discovery Layer

Press enter or click to view image in full size

Another interesting consequence was the ability to navigate the graph through DBpedia categories.

A category such as:

Women of Ancient Greece

can become a starting point for exploration.

From there, users can select an individual, expand their family network, inspect related people, move through historical periods, and explore geographic distributions.

The graph therefore becomes both:

a reasoning engine and a discovery engine.

The same infrastructure that calculates relationships can also drive the user interface.

Graph Engineering

This project reinforced several lessons.

1. Minimal models can be surprisingly powerful

A small ontology containing the right primitive relationships can generate a large amount of derived knowledge.

2. Inference does not always need to be dynamic

For large graphs, materializing inferred knowledge can sometimes be more practical than calculating it on every query.

3. Rules are architecture

The logical rules aren’t merely implementation details.

They define how the graph understands its domain.

4. Scale changes the engineering problem

A reasoning strategy that works perfectly on 10,000 entities may become completely impractical at one million.

5. Graphs can generate their own analytical dimensions

Once relationships are explicit, time, geography, categories and social networks can be explored together.

Beyond Genealogy

Although genealogy was the domain of this experiment, the underlying architecture is much more general.

The same approach can be applied wherever a small set of explicit relationships can generate richer implicit knowledge.

Potential applications include:

  • historical network analysis,
  • legal reasoning,
  • organizational knowledge graphs,
  • fraud detection,
  • scientific knowledge discovery,
  • biomedical relationship analysis,
  • recommendation systems,
  • semantic search,
  • and symbolic AI.

The domain changes.

The principle remains the same:

Start with explicit facts. Define the rules. Materialize the consequences.

Final Thoughts

This project started with a simple question:

What can we infer if we take the family relationships already present in Wikidata and systematically reason over them?

The answer was much larger than expected.

Almost one million people became the starting point for a graph containing more than 40 million triples and over 14 million newly inferred relationships.

But the most important result was not the size of the graph.

It was the change in perspective.

A Knowledge Graph doesn’t have to be a passive collection of facts.

With an ontology, a set of formal rules and a scalable materialization strategy, it can become a system that derives new structured knowledge from what it already knows.

That is one of the reasons I find Knowledge Graphs so interesting.

They provide a bridge between data and reasoning.

And perhaps one of the most interesting directions for AI is not choosing between symbolic reasoning and statistical methods, but combining them.

Let embeddings find patterns.
Let graphs provide context.
Let rules provide reasoning.

Demo [https://javiermurcia.tech/lab/genealogias-deducidas/index.html]

--

--

Javier Murcia Zomeño
Javier Murcia Zomeño

Written by Javier Murcia Zomeño

I am currently developing projects and experiments around knowledge graphs, semantic discovery and data integration.