8 The Literature II: From Screening to Intended-Use Evidence
The newest stratum of the literature — still thin, moving fast — attacks the gap the screening framework leaves open: demonstrating that a package, or any pipeline component, is fit for your intended use, with evidence rather than scores. This chapter reviews that movement, because it is where the field is heading and where this book’s own research sits.
8.1 The three questions that define the gap
Recent specification-driven work (including the companion paper to this book, Human-Governed Validation of R Packages for Regulated Statistical Computing) sharpens the gap into three questions a validation dossier must answer for anything in the execution path:
- What did you require? A versioned statement of requirements with acceptance criteria, specific to how you use the component — not the package’s generic documentation.
- How do you know it meets them? Executed evidence linking each requirement to concrete test cases and results.
- How do you know your tests would catch a relevant fault? Evidence about the test suite itself — mutation-style checks that inject known defects and confirm the suite notices.
Question three is the one that separates validation from verification theater, and it is almost never asked in practice.
8.2 The test-oracle problem, and the metamorphic answer
The obstacle that kept question two expensive is the oracle problem: for arbitrary inputs you cannot know the correct output without re-implementing the software. The research lineage that unlocks it is metamorphic testing — from Chen’s foundational work through the software-testing literature — which replaces “is this output correct?” with “does this pair of outputs satisfy a known relation?” Round-trip a serialization and you should get the original back; permute row order and a correct statistic should not change; scale inputs and a location statistic should shift accordingly.
Applied to R packages through a fixed, auditable relation library driven by deterministic seeded generators, this yields test designs whose expected results are constructed rather than guessed. The companion paper’s public replays are illustrative: a round-trip relation surfaced that a serialization default’s documented precision semantics did not meet a declared intended use; a synthetic time-to-event capsule reproduced prespecified Kaplan–Meier and Cox outputs and passed row-order invariance across fifty checks; four pre-injected defect variants in a pristine package each failed exactly the targeted requirement, while the clean package passed all cases — which is question three answered empirically.
8.3 Records that resist revision
The same movement pays attention to where the evidence lives. Executed records enter append-only, hash-chained audit stores (Merkle-style), so that tampering is detectable rather than merely forbidden — a design that connects directly to Part 11’s audit-trail instincts and to this book’s artifact-addressing principle. Related work by the same author (CAVE-Onc, PLOS One 2026) applies identical chain-of-custody machinery to submission data validation itself, writing every validation trace to a tamper-evident store.
8.4 LLMs enter, governed
The most active frontier is where large language models meet validation — and the literature is already warning about its own failure mode: prototypes that put the LLM in the loop from the start, drafting assessments and verdicts, with human review as an afterthought. The counter-design is fail-closed, human-governed: the model drafts structure only; every expected value is either a machine-captured characterization baseline or independently derived; authority-boundary rules are enforced by code, not by prompting; and approval remains a named human’s act. In the companion paper’s public contract replay, fifteen of fifteen authority-boundary checks held — including a prompt-injection red-team set — precisely because the boundaries were mechanical. The same governance pattern runs through the author’s adjacent work: LLM-assisted independent code generation that eliminates QC programming duplication (PharmaSUG 2026, AI-201) and an AI-assisted methodology for clinical trial data processing and statistical programming (ClinAgent, 2026).
8.5 Where this leaves the field
Assemble the strata and the picture is: a mature screening layer (Chapter 5), an emerging intended-use evidence layer (this chapter), a rich component ecosystem that disclaims both — and, crucially, no published architecture that binds these into an operating environment. That binding is what Part III’s principles and Part IV’s cases attempt.
The test. For any “validated” component, ask for one artifact: the requirement-to-test traceability row and the fault-injection result for the requirement most likely to hurt you. If either is missing, what you have is screening with better manners.