▣ dataarchitect.studio

Field Notes

HOOK vs Data Vault: What the Hook Pattern Actually Is

If you search for the “hook pattern in Data Vault”, you will find very little, and the reason is that the question contains a category error. Data Vault has no construct called a hook. It has hubs, links and satellites, and that is the whole vocabulary.

HOOK is a different data warehousing approach, created by Andrew Foad, explicitly as an alternative to Data Vault rather than an extension of it. In his own words, “HOOK is my attempt to solve problems that I encountered with data vault.” The two are competitors.

That confusion is worth clearing up properly, because the underlying question people are actually asking — is there a lighter way to do what Data Vault does? — is a good one.

Why HOOK exists

Foad is unusually candid about the motivation, and about the originality of the idea. From the cookbook:

There is nothing inherently wrong with the current data warehousing approaches; Data Vault and Dimensional modelling work well if done right. Unfortunately, they are rarely done right.

His argument is not that Data Vault is wrong. It is that Data Vault demands a level of rigour that real projects abandon the moment delivery pressure arrives, and a method that only works when followed perfectly is a method that mostly fails. He also says plainly that HOOK “is nothing new, not really… a re-imagining of ideas put forward by” Inmon, Linstedt and Kimball.

That is a more honest framing than most methodology pitches manage, and it sets up the actual design goal: keep the integration benefits, drop the up-front restructuring.

The four components

HOOK has four parts, and genuinely only four.

The business glossary is the foundation, and it is a hard gate: you may only align data to a business concept once that concept has been formally defined. If the organisation cannot say what it means by Customer, HOOK’s position is that you have no business integrating customer data yet. This is the same instinct as a semantic layer, applied earlier.

A hook is a structure recording the business keys that identify instances of a concept. A Customer hook records all the business keys for customers. Hooks can also express relationships — a Sales hook associating Customer and Product. The detail that surprises people: hooks can be physical tables but do not have to be. Their existence may simply be implied by the keys projected onto the data.

A bag is a virtual construct — normally a view — that aligns raw, source-shaped data to glossary concepts without changing it. The source table stays exactly as it landed; the bag wraps it and exposes qualified business keys alongside the original columns.

A note on vocabulary, because the sources disagree: the cookbook (2022) calls this construct a bag. Foad’s more recent talks call it a frame. Same thing, newer name. If you read both and think you have missed a fifth component, you haven’t.

A keyset qualifies business keys across sources. Customer #123 in System X is very probably not Customer #123 in System Y, so each keyset has an identifier that prefixes the keys originating from that system. This is the same problem that surrogate keys solve in a dimensional model, handled by qualification rather than by substitution.

The actual difference from Data Vault

Data Vault restructures; HOOK annotates Two approaches shown one above the other from the same source tables. Data Vault, first, transforms source data into hubs, links and satellites, so the physical model is rebuilt before loading and modelling must happen up front. HOOK, below, leaves the source tables unchanged and defines a bag view over each one that adds qualified business keys, so the structure is untouched and modelling is deferred to the consumption layer. Data Vault — restructure source tables transform hubs links satellites the physical model IS the deliverable modelling happens before loading HOOK — annotate source tables wrap in a view bag (a view over the source) same columns + qualified business keys hook: customer hook: product source structure untouched modelling deferred to consumption The trade Data Vault pays up front and delivers a queryable integrated model. HOOK pays later and delivers annotated source data. Data Vault restructures; HOOK annotates Two approaches from the same source tables, side by side. Data Vault, on the left, transforms source data into hubs, links and satellites, so the physical model is rebuilt before loading and modelling must happen up front. HOOK, on the right, leaves the source tables unchanged and defines a bag view over each one that adds qualified business keys, so the structure is untouched and modelling is deferred to the consumption layer. Data Vault — restructure source tables transform hubs links satellites the physical model IS the deliverable modelling happens before loading HOOK — annotate source tables wrap in a view bag (a view over the source) same columns + qualified business keys hook: customer hook: product source structure untouched modelling deferred to consumption Data Vault pays up front and delivers a queryable integrated model. HOOK pays later and delivers annotated source data.
The same integration idea, on opposite sides of the load. One restructures; the other annotates.

Data Vault restructures. Source data is decomposed into hubs (business keys), links (relationships) and satellites (descriptive attributes over time). The physical model is the deliverable, and it has to be designed before you can load anything into it.

HOOK annotates. The source table lands and stays as it is. A bag view sits on top and adds qualified business keys as extra columns. Foad’s summary is ELO rather than ELT — extract, load, organise, with no transformation step. Loading and modelling become independent: the model is a layer of definitions over loaded data, not a structure maintained alongside it.

  Data Vault HOOK
Core objects Hubs, links, satellites Glossary, hooks, bags, keysets
Source structure Decomposed Preserved
When you model Before loading Deferred to consumption
Physical footprint New tables Usually just views
Changing your mind Refactor the model Add or drop a column
What you hand a consumer An integrated model Annotated source data
Maturity 20 years, certification, tooling One author, small community

The last row is the one that decides most real decisions, and I’ll come back to it.

Where the argument is strongest

The genuinely good idea here is decoupling integration from restructuring.

Data Vault’s integration value comes from agreeing business keys and recording relationships. Almost none of that value strictly requires physically splitting tables into hubs and satellites — that is an implementation choice which happens to be expensive, and which has to be right before the first load. HOOK’s claim is that you can get the subject-orientation and integration by qualifying and projecting keys, leaving the bytes alone.

That is a real insight, and it maps onto something the industry already learned elsewhere: open table formats also won by adding a metadata layer over untouched files rather than by rewriting data into a proprietary shape.

It also makes being wrong cheap. Getting a business key wrong in Data Vault means refactoring hubs and reloading satellites. In HOOK it means editing a view.

Where I would be careful

The lakehouse claim is contested, including by his own readers. Foad argues that data lake storage already satisfies Inmon’s non-volatile and time-variant criteria, so layering HOOK on top supplies the remaining two and yields a lakehouse. A commenter on that very chapter pushes back that Inmon’s four criteria describe one integrated layer rather than a checklist to tick off separately, and that storing files immutably does not make them a warehouse. I find the objection reasonable. It does not damage HOOK as a method; it just means the equation is rhetoric rather than a definition.

Deferring modelling is not the same as avoiding it. The structural work moves to the consumption layer under the banner of “just enough modelling”. For a small number of well-understood consumers that is a genuine saving. For many consumers with conflicting requirements, it may just relocate the cost — and relocate it to the layer where it is least likely to be governed. Data Vault’s rigour is expensive precisely because it forces that conversation early.

Adoption is the real constraint. Data Vault has two decades of books, certification, consultancies and automation tooling, and you can hire for it. HOOK is essentially one person’s body of work. Adopting it means no hiring pool, no certification path, no mature tooling, and a bus factor of one for the methodology itself. That is a serious cost for a warehouse expected to outlive several teams.

So which should you use

For most organisations, Data Vault if you need Data Vault’s guarantees and can actually sustain its discipline; dimensional modelling if you cannot. That second clause is doing a lot of work, and it is the same conclusion as Data Vault vs dimensional modelling: Data Vault earns its keep on auditability and multi-source integration, and is overkill almost everywhere else.

HOOK is worth reading rather than adopting wholesale, and it is most worth reading if Data Vault’s rigour is currently failing on your project — because that is the precise failure mode it was designed around. The keyset idea in particular transfers cleanly into other methods. You can qualify business keys across source systems without adopting anything else.

What you should not do is go looking for hooks in your Data Vault implementation. They are not in there.

Common questions

Is HOOK part of Data Vault?

No. This is the most common misunderstanding, and the phrasing 'hook pattern in Data Vault' contains a category error. HOOK is a separate data warehousing approach created by Andrew Foad, explicitly as an alternative to Data Vault rather than an extension of it. Data Vault has hubs, links and satellites; it has no construct called a hook. Foad describes HOOK as his attempt to solve problems he encountered on Data Vault projects, which makes the two competitors rather than components.

What is a hook in the HOOK methodology?

A hook is a structure that records the business keys identifying instances of a business concept. If you have a Customer concept, a Customer hook records all the business keys for customers. Hooks can also represent relationships between concepts, so a Sales hook might associate Customer and Product. Importantly, hooks can be physical tables but do not have to be: their existence may simply be implied by the keys projected onto the data.

What are bags and keysets in HOOK?

A bag is a virtual construct, usually a view, that aligns raw source-aligned data to concepts in the business glossary without changing the underlying data. A keyset qualifies business keys so that Customer #123 from System X is never confused with Customer #123 from System Y; each keyset has its own identifier used to prefix keys from that source. Together with the business glossary and hooks, they are the four components of the method. Note that Foad's later talks call the bag a frame; it is the same construct under a newer name.

How is HOOK different from Data Vault in practice?

Data Vault restructures incoming data into hubs, links and satellites, so modelling happens before loading and the physical model is the deliverable. HOOK leaves the source structure alone and adds qualified business keys as extra columns, usually through a view, so loading and modelling stay independent. Foad summarises it as ELO rather than ELT: extract, load, organise, with no transformation step. The practical trade is that HOOK defers structural work to the consumption layer instead of doing it up front.

Should I use HOOK instead of Data Vault?

For most teams, probably not yet, and the reason is adoption rather than design. Data Vault has two decades of books, certification, consultancies and automation tooling behind it, and hiring for it is possible. HOOK is largely one person's body of work with a small community, so you would be adopting it without a hiring pool, a certification path or mature tooling. The ideas are worth reading regardless, particularly if Data Vault's rigour is failing on your project, which is the exact problem HOOK was built to address.

Does HOOK turn a data lake into a lakehouse?

Foad argues it does, on the reasoning that data lake storage already satisfies Inmon's non-volatile and time-variant criteria, and that HOOK supplies the subject-oriented and integrated criteria without moving any data. That claim is contested in the comments on his own cookbook, on the grounds that the four criteria describe one integrated layer rather than a checklist to be satisfied piecemeal. Treat it as an interesting argument rather than a settled definition, particularly since 'lakehouse' is defined loosely by most vendors anyway.

Who created HOOK and where is it documented?

Andrew Foad, who publishes it as The Hook Cookbook, a chapter-by-chapter Substack, alongside a series of HOOK vs Data Vault articles. It is a single-author body of work rather than a standards-backed methodology, which is both its strength, since it is coherent and readable, and its weakness, since there is no independent body maintaining it.