13 Principle Three — Every Artifact Has an Address
In an auditable pipeline, the atomic unit is not the file. It is the artifact: any immutable product of work — a dataset, a figure, a table, a report, a log, a manifest. The principle is simple to state: every artifact has an identity, a lineage, and a home.
Identity means content addressing. The artifact’s name is derived from the hash of its content, the way a fingerprint is derived from a person. Same content, same address, everywhere, forever. Different content, different address — automatically, without a committee naming things.
Lineage means each artifact records its parents: the input data snapshot, the specification version, the code version, the environment, the parent artifacts it consumed. The collection of these edges forms a directed graph — the pipeline’s true history, written by the pipeline itself.
A home means artifacts live in a store that is addressable and durable, so that “the output of run 42” is not a folder someone maintains but a query the system answers.
13.1 What addressing buys
The properties that fall out of this single principle are disproportionate to its cost.
Integrity is structural. Two artifacts claiming the same address must be byte-identical; if they are not, one of them is lying, and the hash says so without argument. Corruption, truncation, and silent partial writes — the plague of shared drives — become detectable by construction rather than by spot check.
Deduplication is free. Re-rendering with unchanged inputs produces unchanged addresses, which means the system recognizes “we already computed this” instead of recomputing it. Cache correctness, normally a hard distributed-systems problem, becomes a tautology.
Provenance is a traversal. “Where did this number come from?” is a walk up the lineage graph, and “what did this snapshot affect?” is a walk down it. Impact analysis, the nightmare of every data correction, becomes a query with a definitive answer.
Replay is mechanical. Because the graph records exactly which parents produced which child, reproduction is not a reconstruction from documentation; it is re-execution along recorded edges. This is the foundation the trust layer’s replay duty stands on.
13.2 The discipline it imposes
Content addressing is honest in a way humans often find uncomfortable. It cannot be negotiated with. An artifact either is what it claims or is not, and there is no version of “close enough.” Teams adopting it must therefore accept two rules without exception:
- No unaddressed side doors. If a result can reach a consumer by any path that bypasses the store — an email attachment, a folder copy, a “quick export” — the graph silently stops describing reality, and everything built on it inherits the lie. The address system must be the only road out.
- Timestamps are context, not identity. When a number was produced matters, but it must never affect its identity. Content is identity; time is metadata. Systems that blur this invent a new artifact every midnight and destroy their own deduplication.
A useful mental model: the artifact graph is the pipeline’s double-entry bookkeeping. Every debit (consumed parent) and credit (created child) is recorded at the moment it happens, by the actor itself. Books that balance themselves cannot be cooked after the fact — and, unlike auditors’ books, they are never written from memory.
The test. Pick any figure in your latest report and ask: “Is there a string I can quote that names exactly this artifact’s content, and can the system walk from it to every input that produced it?” If the best available identifier is a filename with a date in it, your pipeline does not have memory. It has souvenirs.