In most industries, competitors do not ship software together. In pharma, some of the most critical clinical-trial code in the world — packages that build the datasets behind submissions — is maintained by employees of companies that compete trial-for-trial, molecule-for-molecule. A statistician at one large pharma reviews a pull request from a rival’s programmer before it reaches the CRAN release. This is not idealism. It is engineering economics, and understanding it changes how you should adopt every package you use.
This part maps the ecosystem: the packages, the companies, the governance, and the practical question every reader eventually asks — how do I get in?
TL;DR — The pharmaverse is a governed alliance, not a package pile: pharma companies co-fund foundational libraries (admiral for ADaM data, tern for displays, teal for exploration) because validation economics make shared infrastructure cheaper than 20 internal forks. The ecosystem has an intake path — templates, working groups, open issues — and knowing who maintains what is itself a validation argument. This part gives you the map, the governance, and the entry ladder.
The fundamentals
Why competitors share code
The argument runs in four steps, and each step is a fact you can verify in any large pharma’s engineering budget:
- The standards are shared. CDISC is the same for everyone. An ADaM derivation that satisfies traceability at Roche satisfies it at GSK. The differences are in therapeutic strategy, not in how a treatment date is imputed.
- Validation dominates cost. Qualifying an internal package for GxP use costs far more than writing it. Qualifying a shared package spreads that cost across the alliance.
- Forks decay. Every company that kept a private ADaM macro library spent its life re-applying regulatory updates. A shared upstream gets fixes once, for everyone.
- Talent is mobile. Programmers move between companies and bring expectations. An industry-standard stack is cheaper to hire for than a proprietary one.
The result: a stack where foundational layers are co-owned and company-specific layers sit on top as thin, private skins. You see the same pattern in every part of this series — nobody competes on a date-derivation function; everybody competes on the molecule.
The map, layer by layer
| Layer | Packages | What it solves | Built around |
|---|---|---|---|
| Data in | {admiral}, {pharmaversesdtm}, {sdtm.oak} | SDTM→ADaM derivation bricks | Roche-led, multi-company |
| Metadata | {metacore}, {metatools}, {xportr} | Spec as code; labels, lengths, transport | GSK, Roche, Atorus |
| Analysis | {tern}, {mmrm}, {ggsurvfit} | Standard displays and models | Roche, openstatsware |
| Exploration | {teal}, NEST applications | Interactive, reproducible apps | Roche/NEST community |
| Results | {cards}, {gtsummary} | Analysis results as data | Sjoberg (MSK), DFCI |
| Validation | {riskmetric}, riskassessment app | Package qualification | R Validation Hub |
Two things to note on the map. First, the layers are loosely coupled by design: metacore does not require admiral, cards does not require teal. You can adopt one layer and keep proprietary code elsewhere — most companies do. Second, several packages you might file under “general R” (gt, pointblank) have pharma maintainers or pharma-funded roadmaps; the ecosystem bleeds outward more than it looks.
Governance: who actually decides
The pharmaverse is not a legal entity. It is a brand over a cluster of commitments: packages follow a shared development standard (testing coverage, release cadence, vignettes, versioning policy), participate in shared working groups (under the R Consortium’s pharma axis and the R Validation Hub), and agree on interop conventions like the admiral template flow. Individual packages keep their own governance — most are company-backed with external contributors, a few are individual-maintained with corporate sponsors. When you evaluate a package, you are really evaluating three things: maintainer, backing, and bus factor — and the ecosystem’s public health metrics (part 9’s risk tooling) exist precisely to make that evaluation mechanical.
The modern workflow
Installing a study-ready stack
The ecosystem’s own recommendation is template-first: do not assemble packages by hand, start from the admiral template and grow. But knowing what the template gives you is the map in executable form:
# The data-chain spine (parts 4-5)
install.packages(c("admiral", "admiraldev", "metacore", "metatools", "xportr"))
# Analysis displays (part 6-7 foundations)
install.packages(c("tern", "cards", "gtsummary", "ggsurvfit"))
# Validation tooling (part 9)
install.packages("riskmetric")
# The riskassessment Shiny app is GitHub-only, not on CRAN:
remotes::install_github("pharmaR/riskassessment")
Every one of those lines installs a package with a public validation story — maintenance metrics, test coverage, a risk assessment you can pull yourself. That is the difference between this stack and random CRAN.
The template flow
The admiral template — admiral::use_ad_template() — generates a single ADaM program script following the canonical structure the ecosystem has converged on:
library(admiral)
# Create an ADaM program template (one script per dataset)
use_ad_template("adsl") # writes adsl.R: load packages → read SDTM → derivations → save output
The generated script encodes the ecosystem’s opinion of a pipeline: load packages and source datasets (conforming SDTM), apply derivation bricks, then save the ADaM output — one script per ADaM dataset. Companies layer their templates on top of this; if your shop’s structure looks different, it is almost certainly a variation of this shape with more history attached.
Reading a package’s health in five minutes
Before adopting any pharmaverse package, run this due-diligence loop — it is the same evidence your QA will eventually ask for:
library(riskmetric)
pkg_assessment <- pkg_ref("admiral") |>
as_tibble() |>
pkg_assess()
pkg_score(pkg_assessment)
The score aggregates maintenance, community, and testing signals. It will not make your adoption decision — risk-based validation (part 9) never outsources judgment — but it turns “we heard it is good” into a dated, reproducible artifact.
The entry ladder
The ecosystem’s intake path is unusually honest, and it is the same for a company and for a person:
| Rung | What it looks like | What it earns you |
|---|---|---|
| 1. Use | Adopt packages via templates | Working software with a validation story |
| 2. Report | File issues with reproducible examples | Fixes; your name in NEWS files |
| 3. Contribute | PRs: tests, vignettes, small features | Review relationships; internal credibility |
| 4. Co-own | Company assigns you hours to a package | Voice in roadmaps; the real network |
| 5. Seed | Propose a new package through a working group | The ecosystem’s rarest achievement |
Most readers of this series will live on rungs 1–3, and that is fine: the ecosystem runs on them. What distinguishes pharma from general open source is that rung 4 is a budgeted day job at a growing list of companies — the strongest signal that the alliance is real.
The agentic way
The ecosystem map is exactly the kind of knowledge LLMs hold shallowly: they will name the packages, misattribute the maintainers, and invent governance bodies with plausible names. The map is small, human-scale, and changes monthly — the worst combination for a static model. Yet agents are excellent map builders: given a package’s repository, an agent can draft a maintainer timeline, summarize its test coverage trend, or diff its validation story against a sibling package’s — the mechanical half of due diligence.
The agentic way — Use agents to assemble ecosystem evidence (repos, release cadence, issue latency), never to recall it from memory. Attribution errors in a validation memo are findings; the same errors in a chat window are just noise.
Rule: every ecosystem fact in a governed document carries a link a human clicked.
Volatile layer — last verified 2026-10-12. Re-verify before relying on tool specifics.
Key takeaways
- The pharmaverse is an engineering alliance, not a library: shared standards plus validation economics make co-ownership cheaper than private forks.
- Six layers, loosely coupled: adopt by layer; the admiral template is the canonical assembly.
- A package’s governance (maintainer, backing, bus factor) is a first-class adoption criterion, and the ecosystem ships tooling to measure it.
- The entry ladder — use, report, contribute, co-own, seed — is the same for companies and individuals, and rung 4 is increasingly a funded role.
- Ecosystem knowledge is your validation argument’s context: “who stands behind this package” is half of “may we use it.”
FAQ
Is the pharmaverse only for big pharma? No. Small sponsors and CROs are arguably the biggest beneficiaries — they inherit qualification-grade infrastructure they could never fund alone — and the entry ladder starts at zero cost. The smallest shops in the ecosystem run on the template flow alone.
What is the difference between pharmaverse and the R Validation Hub? The pharmaverse is the package alliance; the Validation Hub is the working group that builds the qualification tooling (riskmetric, the risk assessment app, white papers). They intersect — packages want to be assessable, and the Hub makes assessment mechanical — but they answer different questions: “what to build” versus “how to trust what is built.”
Should my company build or adopt? Adopt the foundations, skin the specifics. If you find yourself maintaining a private date-derivation library, you are paying the alliance’s shared costs alone. If you find yourself unable to name what your internal layer adds, that is the finding.
Do I need to know the whole map to use one package? No — but you need the layer around it. Using admiral without metacore knowledge works; using admiral without understanding what xportr does to your outputs eventually surprises someone at a review meeting. Learn layers, not the whole map.
Next in the series: the ledger nobody publishes — what SAS-to-R migrations actually cost, in the words of the companies that ran them.