26  Gaps and Open Questions

A review owes the field its bill of what is missing. These are the gaps this book’s literature survey (Part II) surfaced, each stated as a research-and-building agenda item rather than a lament — and each one a place where a small team can still matter.

26.1 1. No published trust-layer product architecture

The parts exist: comparison engines, lockfile discipline, audit chains, pipeline tools, catalogs. What no one has published — as of this book’s survey — is the binding: an operating environment where every execution emits accuracy, reproducibility, and traceability evidence as a side effect, shipped as one product. The field publishes components and papers; nobody publishes the factory. The reference architecture of Chapter 18 is this book’s proposal; its open questions are the hard ones — cross-engine comparison semantics, replay across language boundaries, and evidence formats that survive years.

26.2 2. The organizational layer is under-documented

The PHUSE open-source guidance itself leaves its cost chapter and its business-model chapter as open questions — who should validate what, and when is internal assessment enough. The two-tier strategy of Chapter 20 is one answer for one company shape. Missing: equivalent decision frameworks for CROs, for mid-size sponsors with hybrid sourcing, and for academic cooperative groups, each of which faces the same wall with different resources.

26.3 3. LLM governance lacks shared benchmarks

Fail-closed designs exist (and this book’s companion work is among them), but the field has no shared red-team benchmark for “did the model exceed its authority?” — the equivalent of a SAT solver’s unsatisfiable instances, runnable by any reviewer. Until one exists, every governance claim is a vendor’s own homework, graded by the same vendor. Building and maintaining that benchmark, neutrally, is high-leverage service work.

26.4 4. Standards are converging but not converged

ARD, Dataset-JSON, and pipeline-evidence formats are all in motion (Chapter 8). The risk is not divergence — it is that early adopters bake private formats into working systems, and the eventual standard arrives as a migration tax. The mitigation is boring and effective: keep every private format behind a versioned contract so the standard, when it lands, is one more consumer.

26.5 5. The small-sponsor SCE is barely studied

The literature centers on large pharma and CRO internal environments. The fastest-growing population — small, fully outsourced sponsors running internal R tooling for review and QC — has almost no published architecture guidance, and what exists (the cases of Part IV) stops self-consciously at the GxP line. The question “what is the minimum credible SCE for a 3-person biotech” has no canonical answer. It should.

26.6 The pattern inside the gaps

Notice what all five share: none requires new statistics, new languages, or new regulation. They are integration and honesty problems — the kind that yield to a small team with a taste for evidence and no interest in rebuilding commodity parts. That is either reassuring or motivating, depending on how much of your calendar is free.

The test. Any roadmap you write should be checkable against gaps like these: which field-level gap does this item close, and could a reviewer tell? Work that closes no gap — however technically pleasant — is the subsidy-to-commodity error of Principle Eight, wearing a Gantt chart.