dataarchitect.studio

Field Notes

Ontology vs Data Model: Meaning You Can Reuse vs Structure You Can Query

Here is the distinction, before the philosophy: a data model defines how data is structured so a particular system can store and query it. An ontology defines what concepts mean and how they relate, independently of any system that stores them. The data model is committed to an implementation — these tables, these keys, this warehouse. The ontology is deliberately not: it’s meant to be reused by many applications, and to be formal enough that a machine can reason over it rather than merely read it.

That distinction has been an academic footnote for twenty years. It stopped being one the moment we started pointing language models at enterprise schemas and asking them to interpret meaning they were never given.

The comparison, honestly

  Data model Ontology
Answers How is this stored and queried? What does this mean, and how does it relate?
Scope One system or application A domain, across systems
Optimized for Storage, integrity, query performance Shared meaning, reuse, inference
Typical artifacts Tables, columns, keys, constraints Classes, properties, relationships, axioms
Expressed in DDL, ERDs, dbt models OWL/RDFS, SHACL, or a formal spec
Change driver New query patterns, new sources New understanding of the domain
Reused across apps Rarely — it’s built for one By design — that’s the point
Machine can infer from it No (it can only execute queries) Yes (that’s what the axioms are for)
Validates data by Constraints, in the database A separate shape language (SHACL)
Fails by Being wrong for new questions Being too abstract to implement

The same problem, two altitudes

Take the case every data architect has lived: three systems disagree about what a customer is. Billing has accounts, the CRM has leads that sometimes become accounts, the product database has users. The data model solution is structural — reconcile them into one dimension with a surrogate key, keep each source’s natural key as an attribute, and you have something queryable:

-- Data model: a structure optimized for this warehouse's queries
CREATE TABLE dim_customer (
  customer_key    BIGINT PRIMARY KEY,   -- surrogate, warehouse-owned
  billing_acct_id TEXT,                 -- natural keys retained
  crm_lead_id     TEXT,
  product_user_id TEXT,
  segment         TEXT,
  valid_from      DATE,  valid_to  DATE -- SCD Type 2 history
);

The ontology solution is semantic — it doesn’t tell you where to put the rows, it tells you what the concepts are and how they relate, so any system can implement it consistently:

Ontology (informally rendered, OWL-flavoured):

Class:    Party
Class:    CompletedOrder ⊑ Order
Property: placedOrder  Party → Order
Property: hasAccount   Party → Account

-- Defined classes: membership is DERIVED, not asserted
Customer ≡ Party ⊓ ∃ placedOrder.CompletedOrder
Prospect ≡ Party ⊓ ¬∃ placedOrder.CompletedOrder
-- note: disjointness now follows from the definitions.
-- It doesn't need declaring, and declaring it separately is how
-- people accidentally make their ontology inconsistent.

Those two lines are the tell. They aren’t columns, and they aren’t constraints on a table — they’re definitions, and a reasoner given a Party with one completed order will classify it as a Customer on its own. Nobody writes that logic again in the next system. A data model can enforce a rule inside one database; an ontology states what a thing is, once, for everywhere.

The subtlety that trips up most first ontologies is worth stating plainly, because it’s the difference between a model that works and one that quietly does nothing. OWL derives; it does not validate. It is monotonic — new facts can only add conclusions, never retract them — so if you assert Customer disjointWith Prospect and then hand the reasoner a Prospect who has placed an order, you don’t get a helpful reclassification. You get a logical inconsistency, and the reasoner stops being useful. And because OWL is open-world, an axiom like “every Customer has at least one Account” does not flag a Customer whose Account is missing from your data; it cheerfully infers that one exists somewhere. If what you wanted was a check that fails, that’s a different tool:

Shape (SHACL — validation, closed over your data):

CustomerShape
  targetClass:  Customer
  property hasAccount  { minCount: 1 }   -- FAILS if the data omits it
  property hasContact  { minCount: 1; class: Person }

Same vocabulary, opposite jobs: OWL says what follows, SHACL says what must be present. Reach for the first when you want a machine to conclude something, the second when you want it to complain — and the OWL 2 primer and SHACL specification are the authoritative references when a vendor’s “semantic layer” blurs the two.

One ontology above many data models A single ontology at the top defines Customer as a kind of Party, disjoint from Prospect, with a hasAccount property. Below it sit three separate implementations of that same meaning: a warehouse star schema with dim_customer, a CRM object model with Lead and Account, and an application's third-normal-form users table. Ontology — meaning Customer ⊑ Party · hasAccount · disjointWith Prospect one definition, application-independent, machine-reasonable warehouse model dim_customer, star schema CRM model Lead, Account objects app model users table, 3NF Data models — structure: three implementations of one meaning the ontology is what they should all agree on; the models are how each system honours it get the meaning wrong once and all three inherit the mistake
One ontology, many data models: meaning above, structure below.

What your semantic layer already is

Here’s the part most ontology writing misses, and most data teams need to hear: if you have a well-built semantic layer, you already own a lightweight ontology. Defining revenue once, in one governed place, so every dashboard and notebook returns the same number, is an act of ontology in everything but formalism — conformed dimensions are the same instinct applied to dimensions, and a data contract is that agreement made enforceable.

What a true ontology adds beyond a semantic layer is threefold: formal semantics (axioms a reasoner can act on), relationship modelling as a first-class citizen (not just measures sliced by dimensions), and application independence (meaning that outlives the warehouse it was first implemented in). Each of those has a cost, which is why the honest answer for most teams is strengthen the semantic layer first, and reach for formal ontology only when one of those three is genuinely the binding constraint.

Why this stopped being academic

Ontologies had a long winter for a good reason: they were expensive to build, hard to keep in sync with reality, and mostly consumed by other ontologists. What changed is the reader. When the consumer of your schema was an analyst, tribal knowledge filled the gaps — everyone knew that cust_flg_2 meant enterprise tier. When the consumer is a language model generating SQL, there is no tribal knowledge; there is only what the model can read. This is measurable, not rhetorical: the BIRD text-to-SQL benchmark was built specifically to test queries over messy real-world schemas, and its headline finding is that supplying external knowledge — the human explanation of what a column actually means — moves execution accuracy substantially, while accuracy still degrades sharply on questions spanning several entities and relationships. That is exactly the territory where a data model is silent and an ontology speaks.

That’s the real thesis, and it’s older than the technology: the discipline of declaring meaning explicitly, before anyone queries anything, is the thing that makes data trustworthy. Kimball called it declaring the grain and conforming dimensions. The semantic layer calls it defining a metric once. Ontologists call it formalising a domain. Same instinct, three vocabularies, and the machines just made it non-optional.

When to reach for which

Build a data model — always. It’s not optional; something has to store the rows, and structural discipline (declared grain, star schemas, keys) is what keeps analytics honest.

Build a semantic layer when more than one tool queries the same data and you need one definition to win.

Build a formal ontology when meaning must be shared across many systems or organisations, when the relationships between entities are the analytical product (fraud rings, supply chains, drug interactions), when regulation requires machine-checkable definitions, or when you’re grounding AI systems that must reason over connections rather than scan columns.

And beware the failure mode that gave ontologies their reputation: a beautiful, formally correct model of a domain that no system implements and no query touches. A data model can be wrong for tomorrow’s questions; an ontology can be too abstract to ever be right about today’s. The strongest architectures use both, at their proper altitudes — meaning declared once above, structure optimised deliberately below.

Common questions

What is the difference between an ontology and a data model?

A data model defines how data is structured for a particular system — tables, columns, keys, constraints, optimized for storing and querying. An ontology formally defines what concepts mean and how they relate, independent of any system that stores them. The data model answers 'how do I store and query this here'; the ontology answers 'what is a customer, in this business, everywhere.'

Is an ontology just a conceptual data model?

It's the closest classical analogue, but two things differ. An ontology is deliberately application-independent — it's meant to be reused across systems rather than implemented once — and it carries formal semantics: machine-readable axioms about class hierarchy and properties from which a reasoner can derive new facts. A conceptual model usually stops at a diagram; an ontology keeps going until a machine can act on it. Note that deriving facts (OWL) and validating data (SHACL) are separate jobs — conflating them is the most common first mistake.

Do I need an ontology if I already have a semantic layer?

Often no. A semantic layer already does the essential job — defining business concepts once, in one governed place, so every consumer gets the same answer. An ontology is the heavier, more formal version, and it earns its cost when meaning must be shared across many systems and organisations, when relationships matter as much as measures, or when machines need to reason over the model rather than just query it.

Why are ontologies suddenly relevant again in the AI era?

Because LLMs need grounding. A model given raw table names guesses at meaning; a model given explicit, machine-readable definitions of concepts and relationships doesn't have to. That's the same argument dimensional modelers have made about declared grain and conformed dimensions for thirty years — the audience changed from analysts to machines, and the requirement got stricter.