7  The Literature I: Risk-Based Validation of Open Source

The first thing the field solved was how to think about open-source software in regulated settings — and the answer, refined over a decade, is risk-based assessment. This chapter reviews that line of work, because its concepts are the vocabulary everything else assumes.

7.1 The white paper that set the frame

The canonical statement is the R Validation Hub’s white paper, A Risk-based Approach for Assessing R package Accuracy within a Validated Infrastructure (Nicholls, Bargo, Sims, and the R Validation Hub, 2020). Its architecture of ideas has aged remarkably well:

  • Separate the infrastructure from the software. Environment validation (servers, OS, runtime) follows standard practice like GAMP change control; the paper’s subject is package accuracy within a validated infrastructure.
  • Classify by relationship to the user, not by dependency distance. Packages actually loaded by users — Intended for Use — get risk assessment; transitive Imports need only be managed for reproducibility. This one distinction collapses an intractable dependency tree into a tractable review surface.
  • Assess four criteria: purpose (statistical packages are riskier, because their bugs hide in math), maintenance good practice, community usage, and testing. Metrics feed a subjective assessment by a qualified assessor and a qualified reviewer — deliberately, because an opaque aggregate score hides the judgment an auditor needs to see.
  • Respond proportionately: low-risk packages need no remediation; high-risk ones get requirements-linked tests; and over time, authors and collections can earn “trusted resource” status (the R Foundation and the tidyverse team being the named examples).

The Hub then built the tooling this implies — the {riskmetric} package for metric collection, the {riskassessment} Shiny app for organizational review, the newer {riskscore} — and, more recently, has been moving toward a public regulatory repository: standardized quality measures, transparently assessed and publicly available, as a central forum for regulatory software-quality expectations. Real adoption is documented in case studies; a good example is SCHARP’s study-specific risk framework, presented at PHUSE US Connect 2026, built around study purpose, the software development lifecycle, community usage, and testing, and mandated first for high-profile trials because resources are finite.

7.2 What the ecosystem itself says

It is worth pausing on a quietly remarkable fact: the pharmaverse project — the curated network of ~50 clinical R packages from admiral through xportr — answers its own FAQ question “Is this validated/GxP/regulatory assured?” with a plain “No.” The component ecosystem is explicit that it ships capability, and that assessment for intended use is the adopter’s job. This is not a failure of the ecosystem; it is a correct division of labor. But it means the most important sentence in the open-source clinical stack is that “No,” and everything downstream of it is unbuilt.

7.3 The limits of screening

Risk-based screening answers: which packages deserve attention, and how much? It does not answer the questions that follow: what did you require of the package for your use, how do you know it meets those requirements, and would your tests detect a relevant fault if one were present? Those are questions about evidence of intended use, and the screening literature is candid about stopping at the border. The white paper’s own remediation advice — write tests linked to requirements — gestures across that border without crossing it.

Crossing it is the subject of the next chapter, and of a growing body of work (including this book’s companion research) on specification-driven, multi-technique validation with auditable records.

The test. Ask an organization how it adopts an open-source package for regulated work. If the answer ends at “we scored it with riskmetric and filed the report,” they have a triage system, not a validation system — the score told them where to look, and nobody looked.