dataarchitect.studio

Topic — 15 essays

Data Architecture

Warehouses, lakes, lakehouses, medallion layers, and the modern data stack — where data should live, how it should be layered, and which structures earn their complexity.

01

Field Notes

Knowledge Graph vs Data Warehouse: Aggregation vs Connection

A warehouse answers how much; a knowledge graph answers how things connect. They're different query shapes, not rivals — and most serious estates end up with both.

Jul 24, 20267 min
02

Field Notes

Parquet vs ORC vs Avro: Which Data File Format, and When

Parquet and ORC are columnar formats for analytics; Avro is row-based for streaming and schema evolution. A clear comparison and a decision rule that holds.

Jul 23, 20266 min
03

Field Notes

Apache Iceberg vs Parquet: They're Not the Same Layer

Iceberg and Parquet aren't rivals — Parquet is a file format, Iceberg is a table format that manages Parquet files. Here's the difference, and why it matters.

Jul 23, 20267 min
04

Field Notes

The Inmon Methodology: The Corporate Information Factory, with an Example

Inmon's method builds one normalized enterprise warehouse first, then serves dimensional marts from it. How the CIF works, a worked example, and where it fits today.

Jul 17, 20265 min
05

Field Notes

How to Choose an Iceberg Catalog: Unity vs Polaris vs Glue vs Nessie

The table format war is settled; the catalog decides governance and lock-in now. How Iceberg catalogs work, compared honestly — and a decision rule that holds.

Jul 08, 20266 min
06

Field Notes

Iceberg vs Delta Lake: How to Actually Choose in 2026

Iceberg and Delta Lake have converged on capabilities — ACID, time travel, deletion vectors. The real decision is your platform and catalog, not the format.

Jul 06, 20266 min
07

Field Notes

What Is an Open Table Format? Iceberg, Delta, and Hudi Explained

An open table format is a metadata spec that turns raw files in object storage into real tables — with ACID transactions, schema evolution, and time travel.

Jul 05, 20267 min
08

Field Notes

Data Lake vs Lakehouse: What Changed and Which to Use

A data lake stores raw files cheaply and guarantees nothing. A lakehouse adds a table layer over the same files for ACID, schema, and trust. When each fits.

Jun 20, 20267 min
09

Field Notes

Kimball vs Inmon: Two Ways to Build a Data Warehouse

Kimball and Inmon differ on one decision: build dimensional marts bottom-up, or a normalized enterprise warehouse top-down. The trade-off, and the modern blend.

Jun 14, 20268 min
10

Reconsidered

Is the Modern Data Stack Dead?

The modern data stack isn't dead, but the era of bolting together a dozen best-of-breed SaaS tools is ending. What's actually happening, and what to keep.

Jun 09, 20264 min
11

Field Notes

Data Warehouse vs Data Lake vs Lakehouse: A Clear Comparison

A data warehouse stores structured, modeled data. A data lake stores raw data cheaply. A lakehouse tries to be both. A side-by-side, and how to choose.

May 30, 20267 min
12

Essay

What Is a Semantic Layer, and Why Does Your Data Stack Need One?

A semantic layer is the single governed place where business metrics are defined once, independent of any dashboard. What it is, and what it actually fixes.

May 28, 20265 min
13

Reconsidered

The Medallion Architecture, Reconsidered

Bronze, silver, gold is a useful default and a dangerous dogma. A second look at what the layers get right, and where they quietly fall apart.

May 27, 20264 min
14

Field Notes

OLTP vs OLAP: Why You Shouldn't Run Analytics on Your App Database

OLTP handles many small transactions fast; OLAP scans huge volumes for analysis. Opposite optimizations, which is why analytics on production is a trap.

May 23, 20265 min
15

Manifesto

The Shape of Data

Every dataset has a shape. The only question is whether you chose it, or whether it happened to you.

Apr 22, 20264 min