22  Case Study: A Validation Gateway

The second case occupies the empty cell directly: a validation gateway — the evidence layer of the reference architecture — run as a production service by a two-to-three-person data function, and documented across the companion paper (Human-Governed Validation of R Packages for Regulated Statistical Computing) and a PHUSE paper on two-tier validation strategy. It is the deepest published example this book knows of evidence generation as a system, and its numbers are instructive precisely because they are modest.

22.1 The organizational question first

The gateway exists because of a question the guidance literature leaves open: a fully outsourced biotech — all GxP production at vendors — produces no validated environment of its own, yet sponsor accountability does not outsource. Who validates what, and when is buying validation cheaper than building the competence? The answer adopted: a two-tier strategy. Tier 1, internal and permanent: riskmetric-based screening wired into package-intake SOPs, engineered to cost near zero or it will silently stop happening. Tier 2, triggered: full function-level validation, purchased or machine-assisted, fired only by a role change — the moment the company takes on a validated QC role — not by any property of a package.

22.2 The gateway itself

Tier 1 runs as automated collectors producing every report number as an evidence-ID-tagged fact, under one hard drafting rule — no statement without an evidence ID — with sections lacking evidence suppressed rather than padded, and report bundles shipping with risk JSON, the renv lock, and the container image digest, making each image a self-evidencing environment reference.

Tier 2 is the specification-driven engine of Chapter 6, in production: versioned requirements with acceptance criteria; an isolated execution sandbox combining known-answer, metamorphic, property-based, fuzz, and mutation testing; characterization baselines machine-captured, never LLM-imagined; records written to a tamper-evident hash chain; an optional local LLM drafting specification structure only, with authority boundaries enforced in code (fifteen of fifteen boundary checks held under replay, including a prompt-injection red-team set); and every verdict approved by a named human. A daily canary self-qualifies the framework itself — injected defects must be caught, clean packages must stay clean — and the canary has, in fact, caught an engine bug, which is the system working.

The honest ledger from its pilot: a pristine synthetic package passed eleven of eleven cases across four techniques while four injected defect variants each failed their targeted requirement; production replays on real CRAN packages surfaced a documented serialization default that did not meet the declared intended use; mutation scores are reported with process failures classified as undetected-by-design rather than detections; and surviving mutants are labeled equivalence-undetermined. The claims are bounded on purpose: test-oracle construction and fail-closed governance, demonstrated — not package quality, not regulatory compliance, not LLM superiority.

22.3 What this case adds to the book

It completes the reference architecture’s evidence layer with a running implementation of its three hardest elements: independent recomputation (the metamorphic engine answers “would your tests notice?”), locked references (hash-chained, replayable), and governed AI (fail-closed by construction). Combined with the first case — the same team operating layers 2–5 in production — the distance remaining to the full seven-layer product is integration, not invention. That distance, and the business of crossing it, is Part V.

The test. Ask any validation tooling the canary question: “When did the tool last catch itself — and where is that record?” Tools that have never failed in public are tools whose failure mode is untested; the chain should contain the tool’s own name.