Start Here All lessons Scene 1 / 12

Start Here · Interactive Lesson

Clinical SP Bootcamp Roadmap

12 scenes· ~22 min· pairs with the article

Step through the scenes, pass the checkpoint quizzes, and try the hands-on exercises. Progress saves locally in this browser — no account, no tracking.

Scene index · 12 scenes
  1. ConceptDay One at the CRO
  2. ConceptA Path That Replaces Ambush
  3. ConceptThree Layers, Three Shelf Lives
  4. ConceptThe Skeleton You Get in Every Part
  5. ConceptTwo Honest Syllabi
  6. CheckpointCheckpoint: Know Your Layers
  7. ConceptOne Real Derivation: ADSL Treatment Variables
  8. Hands-onHands On: Map armcd to Treatment
  9. ConceptWhat Changes in a Validated Environment
  10. ConceptThe Agentic Way: Draft, Review, Verify
  11. CheckpointFinal Check: Raw Data to TLF Path
  12. ConceptStart Where Your Syllabus Says
Concept1 / 12

Day One at the CRO

Day One at the CRO

No syllabus — five weeks to database lock.

• Day one: mapping spec + shared drive

• SDTM arrives with an AE domain on your desk

• Validation: an auditor asks who signed Program v3

• Macros at 2 a.m. — last week’s run eats the file

Is there a deliberate order

that replaces ambush-learning with a path?

Lesson part: opening (Start Here). Grounds learners in the concrete situation a junior clinical statistical programmer faces: no syllabus, tight deadlines, and knowledge acquired by ambush.

Speaker notes

On day one at a Contract Research Organization, or CRO, you do not get a syllabus. You get a mapping specification, a shared drive, and a database lock in five weeks. Then a Study Data Tabulation Model, or SDTM, adverse event domain lands on your desk, and you learn by doing. Validation shows up when an auditor asks who signed off on program version three. Macros get learned at two in the morning, when last week's run eats the deliverable. So is there a deliberate order from raw data to Tables, Listings, and Figures, or TLF, that replaces ambush learning with a real path?

Concept2 / 12

A Path That Replaces Ambush

A Path That Replaces Ambush

Every part runs the same three layers: timeless CDISC fundamentals · modern SCE workflow · honest AI fit

The roadmap groups parts into lines, not publication order — start where your gaps are

Roadmap lineWhat it covers
SDTM / TLFProduction spine: raw data → standard domains → analysis datasets → deliverable tables
ADaMAnalysis layer that makes output-building reliable
SCEEnvironment and reproducibility: consistent, traceable runs
CareerGetting in, and moving up, as a clinical programmer

Lesson part: concept arc (Start Here orientation). Introduces the bootcamp's structural answer: self-contained parts organized into roadmap lines, each part running the same three layers.

Speaker notes

The bootcamp roadmap replaces ambush with a clear path. Every part runs the same three layers: timeless Clinical Data Interchange Standards Consortium, or CDISC, fundamentals; the modern Statistical Computing Environment, or SCE, workflow; and where Artificial Intelligence, or AI, assistance honestly fits. The roadmap groups parts into lines, not publication order. The main line is Study Data Tabulation Model and Tables, Listings, and Figures — the production spine from raw data to standard domains to analysis datasets to deliverable tables. Then the Analysis Data Model, or ADaM, handles the analysis layer, while the SCE line covers environment and reproducibility for consistent, traceable runs. The career line is about getting in and moving up, and remember, publication order is not learning order — each part opens with enough context to stand alone, so start where your gaps are.

Concept3 / 12

Three Layers, Three Shelf Lives

Three Layers, Three Shelf Lives

Layer & shelf lifeStudy / treat layer asInterview probe
L1 · Fundamentals — ≈10 yrLearn deeply — CDISC structures · derivation logic · GxP reasoning“Walk me through building the AE domain from raw data.”
L2 · Modern workflow — 3–5 yrHabits transfer; menus do not — cloud SCE · Git · reproducible runs · SAS/R/Python mix“What changes about your code in a validated environment?”
L3 · Agentic practice — ≈6 moVolatile · AI drafting · review · debugging — treat anything >6 months as unverified“How would you use an AI assistant here?”

One undifferentiated pile → years of ambush-training.

Lesson part: concept arc (L1/L2/L3 model). Teaches the layer stack and its very different shelf lives, which is the core mental model for deciding what to study deeply versus loosely.

Speaker notes

This slide shows three layers of clinical statistical programming, each with a different shelf life. Layer one, fundamentals, covers Clinical Data Interchange Standards Consortium (CDISC) structures, derivation logic, and Good Practice (GxP) reasoning. It holds for ten years, so learn it deeply. Layer two, modern workflow, includes cloud Statistical Computing Environment (SCE), Git, reproducible runs, and Statistical Analysis System (SAS), R, and Python. This is good for three to five years because habits transfer, but menus do not. Layer three, agentic practice, involves Artificial Intelligence (AI) assistants drafting, reviewing, and debugging. It is volatile, so treat anything older than six months as unverified. Treating all three as one pile is why ambush-training takes years. Interviewers probe with building the adverse event (AE) domain from raw data, what changes in a validated environment, and how you would use an AI assistant.

Concept4 / 12

The Skeleton You Get in Every Part

PART ARCHITECTURE · READING CONTRACT

The Skeleton You Get in Every Part

L1

## The fundamentals

Stable — still reads fine in ten years

L2

## The modern workflow

Evolving — expect a rewrite by 2030

L3

## The agentic way

Volatile — Bootcamp agentic layer: last verified 2026-08-30

If you read later: L1 stays true · L3 = hypothesis to re-check

Tool specifics: Bootcamp agentic layer verified 2026-08-30 — re-verify before relying

Lesson part: concept arc (reading contract). Explains that every part follows one skeleton so learners always know which layer they are in and how to trust or distrust it.

Speaker notes

Every part in this roadmap follows the same three-layer skeleton, so you always know what kind of claim you are reading. Layer one, the fundamentals, is the stable layer. It is written to still read fine in ten years, so treat it as true even if you arrive late. Layer two, the modern workflow, is the evolving layer, and you should expect it to be rewritten by 2030. Layer three, the agentic way, is the volatile layer, and it carries an era callout naming the date its claims were last verified against real tools — in this bootcamp, the thirtieth of August, 2026. If you read a part months after that date, treat layer three as a hypothesis to re-check before you rely on any tool specifics, while layer one stays true.

Concept5 / 12

Two Honest Syllabi

Two Honest Syllabi

6-WEEK SPRINT

12-MONTH SWITCH

Offer in hand

Career switch

1. Part 4 — SDTM basics

2. Part 2 — ADSL

3. Part 3 — TLF

4. Part 1 — environment

5. Part 10 — interview drill

1. SDTM line — pilot data

2. ADaM line

3. Env + Git + pipeline

4. Part 11 — entry plan

New to CDISC? Read Part 4 before Part 2.

6 weeks — useful on day one, not finished.

12 months — competitive, not guaranteed.

Real study data compresses every horizon.

Lesson part: concept arc (planning). Gives the learner concrete entry paths from the roadmap: a six-week sprint or a twelve-month career switch, plus the correct reading order for CDISC newcomers.

Speaker notes

Two honest syllabi: an offer in hand with six weeks, or a career switch over about twelve months. In the six-week sprint, start with Part 4 for Study Data Tabulation Model, or SDTM, basics, then Part 2 for the Subject-Level Analysis Dataset, or ADSL, then Part 3 to produce a Tables, Listings, and Figures, or TLF, deliverable, then Part 1 for environment, and Part 10 as the interview drill. For the twelve-month switch, work the SDTM line slowly on public pilot data, then the Analysis Data Model, or ADaM, line, then environment, Git, and pipeline, and use Part 11 to plan entry. If Clinical Data Interchange Standards Consortium, or CDISC, is new, read Part 4 before Part 2; six weeks makes you useful on day one, not finished, and twelve months makes you competitive, not guaranteed, because real study data compresses every horizon.

Checkpoint6 / 12

Checkpoint: Know Your Layers

1 In the layer model used for clinical reporting, L1 is the tabulation/source layer (for example, the Study Data Tabulation Model, SDTM), L2 is the analysis-ready layer (for example, the Analysis Data Model, ADaM), and L3 is the reporting/publication layer (for example, clinical-study tables and journal displays). Which statement correctly describes their shelf lives?

2 When you are learning to apply the layer model to a clinical analysis, which two study strategies match the advice to study L2 deeply and treat L3 more loosely? Select all that apply. (select all that apply, then Check)

3 A clinical publication often shows results tables and figures before the underlying data. Which action follows the rule that publication order is not learning order?

Speaker notes

This is the Checkpoint: Know Your Layers quiz, covering the shelf lives of L1, L2, and L3, which layer to study deeply versus loosely, and why publication order is not learning order. In the layer model, L1 is the tabulation or source layer, for example the Study Data Tabulation Model, SDTM; L2 is the analysis-ready layer, for example the Analysis Data Model, ADaM; and L3 is the reporting or publication layer. The correct answer is C: L3 has the shortest shelf life because displays are tied to a specific analysis snapshot. When learning to apply the layer model, which two strategies match studying L2 deeply and treating L3 more loosely? The correct answers are A and C: trace how safety (SAF) and intention-to-treat (ITT) population flags are derived in an ADaM dataset before using them to interpret a table, and begin with a mock shell to follow the displayed result back to L2 ADaM variables and, when needed, to L1 SDTM source data. A clinical publication often shows results tables and figures before the underlying data. The correct answer is C: use a published result as the entry point, then trace that result from L3 outputs back through L2 analysis data to L1 source data when needed, because publication order is not learning order.

Concept7 / 12

One Real Derivation: ADSL Treatment Variables

One Real Derivation: ADSL Treatment Variables

data adsl; set dm; length trt01p $40 trt01pn 8;

if armcd = 'XYZ001' then do; trt01p = 'Study Drug 50 mg'; trt01pn = 1; end;

else if armcd = 'PLC' then do; trt01p = 'Placebo'; trt01pn = 0; end;

else trt01p = 'Screen Failure'; run;

Branch conditionTRT01P valueTRT01PN value
XYZ001Study Drug 50 mg1
PLCPlacebo0
else / any other armcdScreen Failuremissing

Behind the code: spec decision + controlled terminology + review question.

Interviews grade reasoning, not syntax recall.

Lesson part: real-data walkthrough (L1). Steps through the ten-line ADSL treatment derivation exactly as supplied in the roadmap, revealing the spec decision, controlled-terminology choice, and review question behind the syntax.

Speaker notes

Here is a real ADSL derivation — that's the Subject-Level Analysis Dataset — for the treatment variables TRT01P and TRT01PN, the planned treatment and its numeric code. In ten lines, we set the dataset from the demographics domain, DM, declare TRT01P as character with length forty and TRT01PN as numeric, then branch on the planned arm code, ARMCD. If ARMCD is XYZ001, TRT01P becomes Study Drug 50 mg and TRT01PN is 1. If it is PLC, we get Placebo and 0. Any other value falls through to Screen Failure, leaving TRT01PN missing. But behind those ten lines sits a specification decision, a controlled-terminology choice, and a review question — and in interviews, reasoning beats syntax recall.

Hands-on8 / 12

Hands On: Map armcd to Treatment

Hands-on interactive — if it does not load, open the paired article and try the exercise there.

Speaker notes

For this part of the lesson you will not just watch, you will work in the browser, because the hands-on exercise runs on the website at jaimeyan.com/learn. There you practice a spec-driven derivation, completing the if and else-if branches so that Arm Code (ARMCD) maps to the right Planned Treatment for Period 01 (TRT01P) and Planned Treatment for Period 01 Number (TRT01PN) inside the Subject-Level Analysis Dataset (ADSL) snippet: 'XYZ001' becomes 'Study Drug 50 mg' with a value of 1, 'PLC' becomes 'Placebo' with a value of 0, and any other value falls through to 'Screen Failure', using only the values shown in the snippet. So once the video ends, head over to the site and try it yourself, and let the snippet, not your imagination, drive every branch you write.

Concept9 / 12

What Changes in a Validated Environment

What Changes in a Validated Environment

SCE = Statistical Computing Environment — the engine changes; CDISC/GxP logic stays.

Hiring test: “What changes about your code in a validated environment?” — habits, not brands.

Engine reality: mostly SAS today; R/pharmaverse momentum is real (R Consortium pilots).

Engine-agnostic core: CDISC logic & GxP reasoning transfer unchanged — L1 stays.

Entry point: start with the engine your target jobs list, then add the second.

Free practice: SAS Studio in the browser — SAS OnDemand for Academics.

Shelf life: the SCE itself is L2 — replaced within a contract or two; train habits, not menus.

Lesson part: concept arc (L2, modern workflow). Positions the SCE layer: environment skills are real but shorter-lived, and the underlying CDISC/GxP logic is engine-agnostic.

Speaker notes

In a Statistical Computing Environment, or SCE, the real hiring test is one question: what changes about your code in a validated environment? That question checks habits, not brand names. Validated environments still run mostly SAS, and the R-based pharmaverse stack has genuine regulatory momentum, including R Consortium submission pilots. But Clinical Data Interchange Standards Consortium, or CDISC, logic and Good Practice, or GxP, reasoning are engine-agnostic, so Layer 1 transfers unchanged; start with the engine your target jobs list, then add the second. Free practice is available: SAS Studio runs in a browser through SAS OnDemand for Academics, covering everything the early parts ask you to type. Remember the shelf life: the SCE itself is Layer 2, replaced within a contract or two, so train the habits, not the menus.

Concept10 / 12

The Agentic Way: Draft, Review, Verify

The Agentic Way

Draft · Review · Verify

Draft + debug code → evidence-backed

Autonomous submission generation → no benchmark

Value shift: specify + verify outputs

AI raises L1 stakes: read code before release

Use AI to explain/critique; code until routine

Era callout: volatile layer — verified 2026-08-30

Re-verify before relying on tool-specific capabilities.

Lesson part: concept arc (L3, agentic practice, flagged time-sensitive). States honestly where AI assistance has evidence behind it and where it stops inside regulated programming.

Speaker notes

Let's talk about where Artificial Intelligence (AI) can safely touch a Good Practice (GxP) workflow. Drafting derivation code and debugging have benchmarked results behind them, but autonomous submission generation has no benchmark. Your value shifts from typing programs toward specifying them precisely and verifying their output. You cannot verify code you cannot read, so the AI layer raises the stakes on Level 1 (L1) knowledge. Use assistants to explain and critique, but write the code yourself until each line is routine. Era callout: this volatile layer was last verified on 2026-08-30, so re-verify before relying on tool-specific capabilities.

Checkpoint11 / 12

Final Check: Raw Data to TLF Path

1 Which statement is true about the canonical ARMCD-to-treatment mapping used to produce ADSL?

2 Which statement best describes the three-layer stack and its shelf lives in the raw-data-to-TLF roadmap?

3 In a validated clinical reporting environment, which statements correctly describe where AI assistance can be used and where it cannot? Select all that apply. (select all that apply, then Check)

4 Your runnable ADSL derivation has just completed for a new study, and the next step is to build the first safety TLF. Describe your production-readiness plan before starting that TLF. Include at least four of the following: unique USUBJID records, the ARMCD-to-TRT01P/TRT01PN mapping, TRT01SDT and the SDTM exposure source, Safety/ITT population flags, define.xml/mock shell traceability, and where AI assistance is allowed in a validated environment. Do not write code. (reflect, then reveal)

Reveal analysis
A strong answer should say: confirm ADSL has exactly one row per USUBJID; review a frequency of ARMCD and the TRT01P/TRT01PN mapping for missing or out-of-range values; confirm TRT01SDT comes from the SDTM exposure domain and matches the planned treatment window; confirm Safety and ITT flags agree with the statistical analysis plan; align ADSL variables with define.xml and mock shell columns; and note that AI may assist in drafting or checking but a qualified programmer must validate and approve all code for production.
Speaker notes

This is your final checkpoint quiz, tying the raw-data-to-Tables, Listings, and Figures (TLF) path together. First, on the canonical treatment arm code (ARMCD) to treatment mapping: the correct answer is A. One ARMCD value should resolve to one planned treatment for period 01 (TRT01P) and numeric counterpart (TRT01PN) pair in the Subject-Level Analysis Dataset (ADSL), so every downstream TLF uses that consistent pair. Next, on the three-layer stack and its shelf lives: the correct answer is A. Raw data are locked inputs, Analysis Data Model (ADaM) datasets and runnable derivation programs are the middle layer, and TLF outputs are downstream deliverables that can be regenerated. Third, in a validated clinical reporting environment, where can artificial intelligence (AI) assistance be used? The correct choices are A and B: AI can help draft or explain ADSL mapping code in an exploratory, controlled review before formal validation, and it may be used inside the validated pipeline only after a qualified programmer reviews, tests, and approves the generated solution. Finally, for the short answer: after your runnable ADSL derivation completes, a strong production-readiness plan before the first safety TLF confirms exactly one row per unique subject identifier (USUBJID), reviews the ARMCD-to-TRT01P/TRT01PN mapping for missing or out-of-range values, confirms treatment start date (TRT01SDT) comes from the Study Data Tabulation Model (SDTM) exposure domain and matches the planned treatment window, confirms Safety (SAF) and intention-to-treat (ITT) population flags agree with the statistical analysis plan, aligns ADSL variables with define.xml and mock shell columns, and states where AI assistance is allowed in a validated environment.

Concept12 / 12

Start Where Your Syllabus Says

Start Where Your Syllabus Says

▸ Depth: L1 deep ~10-yr · L2 solid 3-5-yr · L3 light ~6-mo

▸ Parts open standalone — start where your gaps sit

▸ New to CDISC? Start Part 4 · SDTM domain basics

▸ Set up environment first? Start Part 1 · SCE guide

▸ Six-week sprint: useful Day 1 · 12-month full spine: competitive

▸ Verification tracks reading fluency — AI rewards knowing what “correct” looks like

Next: continue into the full bootcamp roadmap and paired parts

Lesson part: summary (Start Here). Pulls together the key takeaways and links back to the bootcamp roadmap article and its paired parts as the next learning step.

Speaker notes

This is the part where you stop planning and pick a starting point. Depth has a shelf life: level one skills stay sharp for roughly ten years, level two for three to five years, and level three for about six months. Because every part opens with enough context to stand alone, you should begin wherever your own gaps sit. If you are new to the Clinical Data Interchange Standards Consortium, start with Part 4 and its Study Data Tabulation Model domain basics. If you would rather build your workspace first, start with Part 1, the statistical computing environment guide, before anything else. Six weeks on the sprint makes you useful on day one, while twelve months on the full spine makes you competitive. And remember that verification skill sits downstream of reading fluency, because the artificial intelligence layer rewards programmers who already know what correct looks like.

✓

Lesson complete

Nice work — every scene seen. Keep the momentum going.

← → Space to navigate · progress is saved locally in your browser

AI Tutor

Ask the tutor