Git for SAS Programmers: Version Control in a GxP World
Why filename-versioned SAS programs are an audit-trail liability, the minimum Git vocabulary a clinical programmer needs, and how branch-per-output maps to QC.
Writing / Research & practice
Notes on LLM agents, clinical trial programming automation, and CDISC data standards — extended companions to my published research.
Browse by section, or search titles, descriptions and tags.
No results. Try a broader keyword, or clear the search to browse all sections.
Why filename-versioned SAS programs are an audit-trail liability, the minimum Git vocabulary a clinical programmer needs, and how branch-per-output maps to QC.
How to turn the SDTM-to-ADaM-to-TLF chain into pipeline as code: explicit dependencies, pinned environments, hash-verified outputs, and stage-level failure isolation.
What the job is, who gets hired, and how to train for it: a realistic 2026 guide to becoming a clinical statistical programmer, from SAS base to AI-assisted practice.
▶ video
The 2026 CDISC AI Innovation Challenge winners present today in Denver. None published code — except one adjacent R package. We read synadam's source and compare it to our own synthetic ADaM pipeline.
One pipeline, five first-hand sources: 3 deals, 0 FDA approvals and 12 upcoming readouts, collected and charted without an LLM.
▶ video
A 15-part series on R in pharma: pharmaverse ADaM, ARD tables, risk-based validation, targets pipelines, GxP Shiny, LLM agents — collected into a free bilingual book.
▶ video
A complete ADaM derivation built from composable bricks: the derive_* mental model, ADSL and BDS patterns, and why company templates sit on top of admiral instead of replacing it.
▶ video
ARD turns tables from pixels into data — machine-reviewable, traceable, reusable. How the cards standard works, how gtsummary speaks it natively, and why reviewers will eventually ask for it.
▶ video
From chat to autonomous workflow: how the Model Context Protocol standardizes what tools agents may touch, what an agent-native clinical pipeline looks like, and the guardrails that keep autonomy auditable.
▶ video
One map of every dataset, standard, and hand-off between the clinic and the regulator — SDTM, ADaM, define.xml, TLFs, and the submission package — and where R now sits at each stage.
▶ video
The series compressed into one argument: standards absorbing tooling, tooling absorbing validation, validation absorbing AI — where the field goes, and the book this series becomes.
▶ video
One adverse-event table built three ways — the structural differences, the validation story of each framework, and a decision tree for which framework earns which role.
▶ video
The complete path for a clinical Shiny app: rhino structure, shinytest2 suites, secure backends, async patterns, and the validation evidence file that lets QA sign.
▶ video
AI pair programmers for ADaM, LLM-assisted tables, QC drafting — the production case ledger from 2024-2026, and the exact wall each case hit when the rubber met validation.
▶ video
Spec as code: how metacore, metatools, and xportr turn your define metadata into checks, labels, and transport files — killing inconsistency at the source instead of in QC.
▶ video
NL-to-CDISC generation, LLM-built analysis apps, auto slide decks — how close is the fully automatic submission, and which bottleneck is regulatory rather than technical?
▶ video
admiral, tern, teal, metacore, xportr — who builds which package, how they interlock, who governs them, and how a company or a solo programmer enters the ecosystem.
▶ video
Parameterized CSR automation: branded documents, multi-study templating, the SAS engine for legacy code, and where auto-generated TLFs flow into review-ready reports.
▶ video
The R Validation Hub method, hands-on: risk dimensions, the public risk-assessment app, package qualification records, and what an inspector actually asks for when your stack is open source.
▶ video
What large pharma migrations really cost — the training trap, the dual-run tax, the retraining ledger — and the ROI logic that survives an audit committee.
▶ video
Incremental compute, crew parallelism, frozen environments — a pipeline that re-runs in minutes, replays exactly a year later, and answers an inspector by existing.
▶ video
How I designed the governance and validation layer for ClinVista, an internal Shiny data-review platform — and why it survived the decision to replace the platform itself.
▶ video
Build clinical TFL mock shells, review variable bindings, and export documents and SAS/R program scaffolds with a free browser tool and a practical walkthrough.
▶ video
One pipeline, five first-hand sources: 32 deals, 25 FDA approvals and 12 upcoming readouts, collected and charted without an LLM.
▶ video
Sep 8–9: 17 deal events, 0 FDA approvals, 12 queued readouts; Novartis's $12B Avidity bet hits a Phase 3 miss and Monte Rosa signs a $2.1B license.
▶ video
Six months after closing the $72/share Avidity buy: del-zota sits under FDA priority review, del-brax won Phase 1/2, and del-desiran failed HARBOR — NVS fell 13.9%.
▶ video
Sep 7–8: $14.6B in disclosed totals, zero FDA approvals, and Novartis's $12B Avidity buyout's lead asset fails Phase 3. Monte Rosa licenses VAV1 degraders to Novartis.
▶ video
HARMONi-2's OS win over pembrolizumab is real but detail-light: SMMT +17.3% on 2.7x volume, Akeso +15.3% in two sessions — and the HR itself lands at WCLC this month.
ADTTE step by step: event and censoring definitions from the SAP, the censoring date cascade, CNSR semantics, partial dates at the event, and QC listings per subject.
▶ video
Sep 6–7: zero deals, zero FDA approvals, and 12 Phase 2/3 readouts due within 90 days. The calendar, not the tape, is the story.
AgomAb priced $200M at the midpoint in February, broke 8.4% on debut, fell to a −39.9% close by June, then rallied 85% on no data — now −1.4% vs offer as of Sep 4.
▶ video
Aktis priced 2026's first biotech IPO at the top of range with Eli Lilly buying $100M in the deal; the stock is +50% by September. The anchor playbook the class repeated.
Alamar upsized 20% to $219.9M with the shoe, popped +29.4% on debut, and trades +81.2% vs offer — a proteomics tools story priced on revenue, not trial risk.
Apnimed upsized 20% to $192M at the top of range, popped +56.3% on debut, and trades +71.6% vs offer — with an FDA decision on its sleep-apnea pill due Feb 28, 2027.
Attovia priced an upsized $289M IPO at the $17 top of range on August 4, popped +28.8%, and never closed below offer — Phase 1b itch data, runway into 2030.
▶ video
Avalyn upsized 41% to $300M at the top of range, popped +63.8%, and doubled by September on enrollment execution — with inhaled pirfenidone data due H2 2027.
BlossomHill upsized to $150M at $16.00 on August 6, 2026, traded flat on debut, then ground to +37.6% by the September 4 close on Phase 1 EGFR data and a dated path.
▶ video
Braveheart priced $382.5M above range for one Hengrui-in-licensed HCM drug, popped +65.6%, and holds +48.6% a month in — where Kailera round-tripped. The NewCo anatomy.
▶ video
Eikon upsized to $381M at the top of range and broke −16.7% on debut, with no negative catalyst since. The class's clearest platform discount — and its H2 2026 binary.
▶ video
Generate priced $400M at range midpoint, fell −20.9% on debut — 2026's worst first day — and fought back above water by September on Phase 3 execution.
▶ video
Hemab upsized 42% to $347M, popped +88.9%, then climbed on ISTH data and an FDA-endorsed Phase 3 path in Glanzmann thrombasthenia. A +147% tape, anatomized.
▶ video
Kailera priced a then-record $625M at the top of range, popped +62.5%, and round-tripped to +1.4% by September. The Hengrui in-license anatomy and what reprices it.
Kardigan priced 25M shares at the $16 top of range, popped +37.5% on debut, and holds +30.3% vs offer as of Sep 4 — with all three cardiac readouts guided to H1 2027.
Latigo priced $345.6M at the top of range on Aug 6, 2026 — $397.4M after a full greenshoe — behind a Phase 3-ready oral NaV1.8 pain drug; +28.0% vs offer as of Sep 4.
Odyssey priced an upsized $279M IPO at the top of range on May 7, broke 8.8% on debut, then re-rated +59.8% by the Sep 4 close on two earnings prints.
▶ video
Parabilis priced a record $670M biotech IPO above range after two upsizes; twelve weeks later the tape pays for one desmoid-tumor asset, not the Helicon platform.
Scribe priced an upsized IPO at the $15 top on Jul 23, closed at $148M with a full shoe plus a $7.5M Sanofi placement, and trades +91.1% vs offer.
Seaport upsized 20% to $254.9M at the top of range, closed up 10.2% on debut, and trades +38.1% vs offer as of Sep 4 — on prodrug chemistry, not platform adjectives.
▶ video
SpyGlass priced $172.5M at midpoint with no upsize, popped +65%, and holds +81% by September on a 505(b)(2) glaucoma implant with two Phase 3s enrolling.
▶ video
Veradermics priced $295M at $17, popped +122%, then two pivotal readouts and a $442M follow-on at $100 took it to +479% of offer. The class's best tape.
Vogenx raised $93.4M all-in at $13 — the smallest 2026 biotech IPO — then ran +144.1% by the September 4 close with zero post-IPO disclosures. Float, not data.
▶ video
Sep 4–5: Menarini/Gan & Lee $771M GLP-1 deal, NeuShen $80M financing, Novartis Lp(a) miss — China-originated assets took 100% of disclosed dollars.
▶ video
21 US biotech IPOs raised ~$6.5B in 2026 through Aug 11 — median check $295M, 18 of 21 above offer as of Aug 28. What the class tells practitioners.
▶ video
Sep 3–4: GSK-Hutchmed $110M upfront license, Medicus $1B ADC deal, ivonescimab Phase 3 OS win — China-originated assets dominated both deal flow and the tape.
As of Sep 4, 2026, the 20 largest pharma companies total $5.12T; Lilly alone is 20% of the table, Novo sits at #9, and WuXi AppTec leads YTD movers at +90%.
▶ video ✎ interactive lesson
ADAE from the OCCDS side: one row per event, the AE-to-ADSL merge, treatment-emergent flags driven by TRT01SDT, serious flags, and the QC defects reviewers catch.
▶ video ✎ interactive lesson
The BDS skeleton behind ADaM analysis datasets: PARAM/PARAMCD/AVAL, baseline flags, change from baseline, and how ADVS and ADLB are built visit by visit.
A worked SDTM AE domain mapping example: MedDRA coding, serious flags, partial ISO dates, AESEQ derivation, SUPPAE, and CORE validation triage from a mock raw extract.
How to write an SDTM mapping specification: column anatomy, a row-by-row VS domain walk, hygiene rules, and how specs become define-XML and machine-usable code.
SDTM domains explained: domain classes, the topic-timing-qualifier variable pattern, USUBJID, controlled terminology, SUPPQUAL, and a day-one reading order.
▶ video ✎ interactive lesson
How ADSL is built: deriving treatment dates and population flags from DM/EX/DS/SV, the one-row-per-subject rule, and the QC checks that catch real discrepancies.
▶ video ✎ interactive lesson
A systematic path into clinical statistical programming: CDISC fundamentals, modern SCE workflow, and where AI agents fit — every part with runnable takeaways.
▶ video ✎ interactive lesson
How clinical SAS interviews actually work: four rounds from SAS mechanics to GxP habits, real-format questions, and the signals strong answers contain.
▶ video ✎ interactive lesson
What a statistical computing environment is for in clinical trials, why desktop SAS and shared drives failed audits, and how cloud SCE workflow changes day one.
▶ video ✎ interactive lesson
How clinical TLF outputs ship: reading mock shells, PROC REPORT tables and listings, RTF conventions, and the four-pass QC order that catches defects cheapest first.
Where ChatGPT, Claude, and Copilot genuinely save time in SAS and R clinical programming, where they fail, and how to use them inside auditable GxP workflows.
#ai-coding-assistants#sas-programming#gxp#clinical-programming#llm-tools
CDISC CORE is a free, open-source rule engine for SDTM and ADaM validation. How its YAML rules work, how to run it, and where rule-based checking stops.
How to tell demo-ware from deployable agentic AI in clinical trials: determinism, audit trails, 21 CFR Part 11, constrained architectures, and vendor questions.
#llm-agents#clinical-trials#gxp#audit-trail#agent-architecture
A tiered survey of LLM evidence in clinical trial statistical programming: benchmarked results, promising single-team studies, vendor hype, and open gaps for 2026–2027.
#llm#clinical-trials#statistical-programming#survey#benchmarks#gxp
69.7% of oncology CORE rules port cleanly to a graph constraint language. The 30% that don't are exactly where the dangerous errors live.
Setting temperature to zero feels like determinism. It isn't — and in GxP work the difference will find you during an audit, not during development.
Why pharma is moving from SAS to R, what breaks in practice (procedures, XPT files, QC), and how to migrate a validated macro library without a rewrite.
#sas-to-r#pharmaverse#admiral#clinical-programming#xpt#validation
Rule-based SDTM validators structurally miss cross-domain contradictions. SHACL-SPARQL graph constraints plus a deterministic agent layer catch all 20 archetypes.
Independent QC re-programming costs 30–50% of clinical programming effort. An AI framework matched 97.1–100% of variables while keeping independence intact.
Five LLM generation methods benchmarked across 1,999 bootstrap experiments on ICH E3-conformant TLF templates: hybrid RAG with reranking beats direct prompting.
#rag#llm-benchmarking#tlf-templates#clinical-trials#r-code-generation
Free-form agent loops break down in GxP clinical programming. Structuring the work as a typed process DAG makes LLM agents reliable, replayable, and auditable.
A GRADE-rated review of 2020–2025 automation evidence in statistical programming: real gains, mostly Low to Very Low quality evidence.
#statistical-programming#automation#evidence-quality#clinical-trials
A bridge map, typed IR, and orchestrator wrap a legacy SAS TFL library unchanged — AI-ready JSON on day one, 80%+ cell-level parity, optional 92% code cut.
#sas#legacy-modernization#clinical-trials#llm-integration#statistical-programming
ClinAgent splits clinical programming capability into five layers, keeping MCP tools stateless and packaging domain expertise as testable skills.
#llm-agents#mcp#clinical-trials#statistical-programming#agent-architecture
Schema-only synthetic ADaM generation plateaus at 0.45 overall quality; enriching schemas from protocol/SAP/CRF knowledge graphs plus templates reaches 0.70.
#synthetic-data#adam#knowledge-graphs#llm#statistical-programming
Base Llama 3.1 8B scores 0.36 on admiral code generation. LoRA fine-tuning plus knowledge-graph validation gets it to 0.82 — without sending data to an API.
Interactive explainer
How a messy EDC extract becomes a submission package — eight animated scenes covering SDTM mapping, SUPPQUAL, ADSL, BDS baseline and windowing, TLF production and QC, and Define-XML. Play it like a lecture, one beat at a time.
Interactive explainer
Why domain-scoped SDTM validators structurally miss cross-domain RECIST contradictions — seven scenes from one impossible response to the graph, the SHACL shapes, the deterministic agent, and the 20/20 scoreboard with its fine print.
Interactive explainer
ClinAgent's five-layer stack — thin stateless MCP tools, thick testable skills, compliance infrastructure — traced through one tool call, the STUDY-A validation bench, and the study-specific failure mode. Seven scenes.
Latest
Start here
The pieces that best represent what this blog is about.
Intro · ▶ video
A 15-part series on R in pharma: pharmaverse ADaM, ARD tables, risk-based validation, targets pipelines, GxP Shiny, LLM agents — collected into a free bilingual book.
Intro · ▶ video ✎ interactive lesson
A systematic path into clinical statistical programming: CDISC fundamentals, modern SCE workflow, and where AI agents fit — every part with runnable takeaways.
Rule-based SDTM validators structurally miss cross-domain contradictions. SHACL-SPARQL graph constraints plus a deterministic agent layer catch all 20 archetypes.
Independent QC re-programming costs 30–50% of clinical programming effort. An AI framework matched 97.1–100% of variables while keeping independence intact.
Research Deep-Dives
Extended companions to the published papers — methods, results, and honest limitations.
▶ video
The 2026 CDISC AI Innovation Challenge winners present today in Denver. None published code — except one adjacent R package. We read synadam's source and compare it to our own synthetic ADaM pipeline.
▶ video
How I designed the governance and validation layer for ClinVista, an internal Shiny data-review platform — and why it survived the decision to replace the platform itself.
Rule-based SDTM validators structurally miss cross-domain contradictions. SHACL-SPARQL graph constraints plus a deterministic agent layer catch all 20 archetypes.
Independent QC re-programming costs 30–50% of clinical programming effort. An AI framework matched 97.1–100% of variables while keeping independence intact.
Five LLM generation methods benchmarked across 1,999 bootstrap experiments on ICH E3-conformant TLF templates: hybrid RAG with reranking beats direct prompting.
#rag#llm-benchmarking#tlf-templates#clinical-trials#r-code-generation
Free-form agent loops break down in GxP clinical programming. Structuring the work as a typed process DAG makes LLM agents reliable, replayable, and auditable.
A GRADE-rated review of 2020–2025 automation evidence in statistical programming: real gains, mostly Low to Very Low quality evidence.
#statistical-programming#automation#evidence-quality#clinical-trials
A bridge map, typed IR, and orchestrator wrap a legacy SAS TFL library unchanged — AI-ready JSON on day one, 80%+ cell-level parity, optional 92% code cut.
#sas#legacy-modernization#clinical-trials#llm-integration#statistical-programming
ClinAgent splits clinical programming capability into five layers, keeping MCP tools stateless and packaging domain expertise as testable skills.
#llm-agents#mcp#clinical-trials#statistical-programming#agent-architecture
Schema-only synthetic ADaM generation plateaus at 0.45 overall quality; enriching schemas from protocol/SAP/CRF knowledge graphs plus templates reaches 0.70.
#synthetic-data#adam#knowledge-graphs#llm#statistical-programming
Base Llama 3.1 8B scores 0.36 on admiral code generation. LoRA fine-tuning plus knowledge-graph validation gets it to 0.82 — without sending data to an API.
Clinical SP Bootcamp
A systematic training series: fundamentals, modern SCE workflow, and the agentic way — from raw data to TLF delivery. Key parts ship a downloadable Claude skill. Every part pairs with an interactive lesson (✎ badge). Series DOI: 10.5281/zenodo.22233175.
Intro · ▶ video ✎ interactive lesson
A systematic path into clinical statistical programming: CDISC fundamentals, modern SCE workflow, and where AI agents fit — every part with runnable takeaways.
Part 1 · ▶ video ✎ interactive lesson
What a statistical computing environment is for in clinical trials, why desktop SAS and shared drives failed audits, and how cloud SCE workflow changes day one.
Part 2 · ▶ video ✎ interactive lesson
How ADSL is built: deriving treatment dates and population flags from DM/EX/DS/SV, the one-row-per-subject rule, and the QC checks that catch real discrepancies.
Part 3 · ▶ video ✎ interactive lesson
How clinical TLF outputs ship: reading mock shells, PROC REPORT tables and listings, RTF conventions, and the four-pass QC order that catches defects cheapest first.
Part 4 · ✎ interactive lesson
SDTM domains explained: domain classes, the topic-timing-qualifier variable pattern, USUBJID, controlled terminology, SUPPQUAL, and a day-one reading order.
Part 5 · ✎ interactive lesson
A worked SDTM AE domain mapping example: MedDRA coding, serious flags, partial ISO dates, AESEQ derivation, SUPPAE, and CORE validation triage from a mock raw extract.
Part 6 · ✎ interactive lesson
How to write an SDTM mapping specification: column anatomy, a row-by-row VS domain walk, hygiene rules, and how specs become define-XML and machine-usable code.
Part 7 · ▶ video ✎ interactive lesson
The BDS skeleton behind ADaM analysis datasets: PARAM/PARAMCD/AVAL, baseline flags, change from baseline, and how ADVS and ADLB are built visit by visit.
Part 8 · ▶ video ✎ interactive lesson
ADAE from the OCCDS side: one row per event, the AE-to-ADSL merge, treatment-emergent flags driven by TRT01SDT, serious flags, and the QC defects reviewers catch.
Part 9 · ✎ interactive lesson
ADTTE step by step: event and censoring definitions from the SAP, the censoring date cascade, CNSR semantics, partial dates at the event, and QC listings per subject.
Part 10 · ▶ video ✎ interactive lesson
How clinical SAS interviews actually work: four rounds from SAS mechanics to GxP habits, real-format questions, and the signals strong answers contain.
Part 11 · ✎ interactive lesson
What the job is, who gets hired, and how to train for it: a realistic 2026 guide to becoming a clinical statistical programmer, from SAS base to AI-assisted practice.
Part 12 · ✎ interactive lesson
Why filename-versioned SAS programs are an audit-trail liability, the minimum Git vocabulary a clinical programmer needs, and how branch-per-output maps to QC.
Part 13 · ✎ interactive lesson
How to turn the SDTM-to-ADaM-to-TLF chain into pipeline as code: explicit dependencies, pinned environments, hash-verified outputs, and stage-level failure isolation.
Clinical R in Practice
The open-source stack behind regulated drug development: the pharmaverse ADaM toolchain, regulatory-grade tables and ARD, risk-based validation, reproducible pipelines, GxP Shiny, and where LLM agents honestly fit. The complete series is also a free companion book.
Intro · ▶ video
A 15-part series on R in pharma: pharmaverse ADaM, ARD tables, risk-based validation, targets pipelines, GxP Shiny, LLM agents — collected into a free bilingual book.
Part 1 · ▶ video
One map of every dataset, standard, and hand-off between the clinic and the regulator — SDTM, ADaM, define.xml, TLFs, and the submission package — and where R now sits at each stage.
Part 2 · ▶ video
admiral, tern, teal, metacore, xportr — who builds which package, how they interlock, who governs them, and how a company or a solo programmer enters the ecosystem.
Part 3 · ▶ video
What large pharma migrations really cost — the training trap, the dual-run tax, the retraining ledger — and the ROI logic that survives an audit committee.
Part 4 · ▶ video
A complete ADaM derivation built from composable bricks: the derive_* mental model, ADSL and BDS patterns, and why company templates sit on top of admiral instead of replacing it.
Part 5 · ▶ video
Spec as code: how metacore, metatools, and xportr turn your define metadata into checks, labels, and transport files — killing inconsistency at the source instead of in QC.
Part 6 · ▶ video
One adverse-event table built three ways — the structural differences, the validation story of each framework, and a decision tree for which framework earns which role.
Part 7 · ▶ video
ARD turns tables from pixels into data — machine-reviewable, traceable, reusable. How the cards standard works, how gtsummary speaks it natively, and why reviewers will eventually ask for it.
Part 8 · ▶ video
Parameterized CSR automation: branded documents, multi-study templating, the SAS engine for legacy code, and where auto-generated TLFs flow into review-ready reports.
Part 9 · ▶ video
The R Validation Hub method, hands-on: risk dimensions, the public risk-assessment app, package qualification records, and what an inspector actually asks for when your stack is open source.
Part 10 · ▶ video
Incremental compute, crew parallelism, frozen environments — a pipeline that re-runs in minutes, replays exactly a year later, and answers an inspector by existing.
Part 11 · ▶ video
The complete path for a clinical Shiny app: rhino structure, shinytest2 suites, secure backends, async patterns, and the validation evidence file that lets QA sign.
Part 12 · ▶ video
AI pair programmers for ADaM, LLM-assisted tables, QC drafting — the production case ledger from 2024-2026, and the exact wall each case hit when the rubber met validation.
Part 13 · ▶ video
From chat to autonomous workflow: how the Model Context Protocol standardizes what tools agents may touch, what an agent-native clinical pipeline looks like, and the guardrails that keep autonomy auditable.
Part 14 · ▶ video
NL-to-CDISC generation, LLM-built analysis apps, auto slide decks — how close is the fully automatic submission, and which bottleneck is regulatory rather than technical?
Part 15 · ▶ video
The series compressed into one argument: standards absorbing tooling, tooling absorbing validation, validation absorbing AI — where the field goes, and the book this series becomes.
Learning Center
The bootcamp as interactive lessons — scenes, checkpoint quizzes, hands-on practice, progress saved in your browser.
Open the Learning Center → 8 lines · 21 lessons All 21 lessons live. Each pairs with its bootcamp article, adds checkpoint quizzes and hands-on practice, and saves your progress locally.Field Guides & Hot Topics
Practical explainers on the tools and debates shaping clinical programming today.
Where ChatGPT, Claude, and Copilot genuinely save time in SAS and R clinical programming, where they fail, and how to use them inside auditable GxP workflows.
#ai-coding-assistants#sas-programming#gxp#clinical-programming#llm-tools
CDISC CORE is a free, open-source rule engine for SDTM and ADaM validation. How its YAML rules work, how to run it, and where rule-based checking stops.
How to tell demo-ware from deployable agentic AI in clinical trials: determinism, audit trails, 21 CFR Part 11, constrained architectures, and vendor questions.
#llm-agents#clinical-trials#gxp#audit-trail#agent-architecture
Why pharma is moving from SAS to R, what breaks in practice (procedures, XPT files, QC), and how to migrate a validated macro library without a rewrite.
#sas-to-r#pharmaverse#admiral#clinical-programming#xpt#validation
Interactive Explainers
Self-contained animated walkthroughs of the series' flagship topics — play them like a lecture, one beat at a time. All explainers →
2026-09-01 · 8 scenes · interactive
How a messy EDC extract becomes a submission package — eight animated scenes covering SDTM mapping, SUPPQUAL, ADSL, BDS baseline and windowing, TLF production and QC, and Define-XML. Play it like a lecture, one beat at a time.
2026-09-02 · 7 scenes · interactive
Why domain-scoped SDTM validators structurally miss cross-domain RECIST contradictions — seven scenes from one impossible response to the graph, the SHACL shapes, the deterministic agent, and the 20/20 scoreboard with its fine print.
2026-09-02 · 7 scenes · interactive
ClinAgent's five-layer stack — thin stateless MCP tools, thick testable skills, compliance infrastructure — traced through one tool call, the STUDY-A validation bench, and the study-specific failure mode. Seven scenes.
State of the Field
Quarterly surveys of the evidence — what's proven, what's hype, with full references.
A tiered survey of LLM evidence in clinical trial statistical programming: benchmarked results, promising single-team studies, vendor hype, and open gaps for 2026–2027.
#llm#clinical-trials#statistical-programming#survey#benchmarks#gxp
Notes
Short, frequent observations — tools, papers, and field notes. Longer arguments live in the sections above.
▶ video
Build clinical TFL mock shells, review variable bindings, and export documents and SAS/R program scaffolds with a free browser tool and a practical walkthrough.
69.7% of oncology CORE rules port cleanly to a graph constraint language. The 30% that don't are exactly where the dangerous errors live.
Setting temperature to zero feels like determinism. It isn't — and in GxP work the difference will find you during an audit, not during development.
Biotech IPO & Markets
IPO debuts and pharma market tables — pricing, first-day pops, and what the tape says afterwards. Slower-moving companions to the daily brief.
▶ video
Six months after closing the $72/share Avidity buy: del-zota sits under FDA priority review, del-brax won Phase 1/2, and del-desiran failed HARBOR — NVS fell 13.9%.
▶ video
HARMONi-2's OS win over pembrolizumab is real but detail-light: SMMT +17.3% on 2.7x volume, Akeso +15.3% in two sessions — and the HR itself lands at WCLC this month.
AgomAb priced $200M at the midpoint in February, broke 8.4% on debut, fell to a −39.9% close by June, then rallied 85% on no data — now −1.4% vs offer as of Sep 4.
▶ video
Aktis priced 2026's first biotech IPO at the top of range with Eli Lilly buying $100M in the deal; the stock is +50% by September. The anchor playbook the class repeated.
Alamar upsized 20% to $219.9M with the shoe, popped +29.4% on debut, and trades +81.2% vs offer — a proteomics tools story priced on revenue, not trial risk.
Apnimed upsized 20% to $192M at the top of range, popped +56.3% on debut, and trades +71.6% vs offer — with an FDA decision on its sleep-apnea pill due Feb 28, 2027.
Attovia priced an upsized $289M IPO at the $17 top of range on August 4, popped +28.8%, and never closed below offer — Phase 1b itch data, runway into 2030.
▶ video
Avalyn upsized 41% to $300M at the top of range, popped +63.8%, and doubled by September on enrollment execution — with inhaled pirfenidone data due H2 2027.
BlossomHill upsized to $150M at $16.00 on August 6, 2026, traded flat on debut, then ground to +37.6% by the September 4 close on Phase 1 EGFR data and a dated path.
▶ video
Braveheart priced $382.5M above range for one Hengrui-in-licensed HCM drug, popped +65.6%, and holds +48.6% a month in — where Kailera round-tripped. The NewCo anatomy.
▶ video
Eikon upsized to $381M at the top of range and broke −16.7% on debut, with no negative catalyst since. The class's clearest platform discount — and its H2 2026 binary.
▶ video
Generate priced $400M at range midpoint, fell −20.9% on debut — 2026's worst first day — and fought back above water by September on Phase 3 execution.
▶ video
Hemab upsized 42% to $347M, popped +88.9%, then climbed on ISTH data and an FDA-endorsed Phase 3 path in Glanzmann thrombasthenia. A +147% tape, anatomized.
▶ video
Kailera priced a then-record $625M at the top of range, popped +62.5%, and round-tripped to +1.4% by September. The Hengrui in-license anatomy and what reprices it.
Kardigan priced 25M shares at the $16 top of range, popped +37.5% on debut, and holds +30.3% vs offer as of Sep 4 — with all three cardiac readouts guided to H1 2027.
Latigo priced $345.6M at the top of range on Aug 6, 2026 — $397.4M after a full greenshoe — behind a Phase 3-ready oral NaV1.8 pain drug; +28.0% vs offer as of Sep 4.
Odyssey priced an upsized $279M IPO at the top of range on May 7, broke 8.8% on debut, then re-rated +59.8% by the Sep 4 close on two earnings prints.
▶ video
Parabilis priced a record $670M biotech IPO above range after two upsizes; twelve weeks later the tape pays for one desmoid-tumor asset, not the Helicon platform.
Scribe priced an upsized IPO at the $15 top on Jul 23, closed at $148M with a full shoe plus a $7.5M Sanofi placement, and trades +91.1% vs offer.
Seaport upsized 20% to $254.9M at the top of range, closed up 10.2% on debut, and trades +38.1% vs offer as of Sep 4 — on prodrug chemistry, not platform adjectives.
▶ video
SpyGlass priced $172.5M at midpoint with no upsize, popped +65%, and holds +81% by September on a 505(b)(2) glaucoma implant with two Phase 3s enrolling.
▶ video
Veradermics priced $295M at $17, popped +122%, then two pivotal readouts and a $442M follow-on at $100 took it to +479% of offer. The class's best tape.
Vogenx raised $93.4M all-in at $13 — the smallest 2026 biotech IPO — then ran +144.1% by the September 4 close with zero post-IPO disclosures. Float, not data.
▶ video
21 US biotech IPOs raised ~$6.5B in 2026 through Aug 11 — median check $295M, 18 of 21 above offer as of Aug 28. What the class tells practitioners.
As of Sep 4, 2026, the 20 largest pharma companies total $5.12T; Lilly alone is 20% of the table, Novo sits at #9, and WuXi AppTec leads YTD movers at +90%.
Pharma Daily
A daily pharma/biotech finance brief built from first-hand sources only — SEC filings, FDA and trial registries, wire releases. Latest issue below; all daily briefs →
One pipeline, five first-hand sources: 3 deals, 0 FDA approvals and 12 upcoming readouts, collected and charted without an LLM.