Career All lessons Scene 1 / 12

Career · Interactive Lesson

Clinical SP Career Guide 2026

12 scenes· ~22 min· pairs with the article

Step through the scenes, pass the checkpoint quizzes, and try the hands-on exercises. Progress saves locally in this browser — no account, no tracking.

Scene index · 12 scenes
  1. ConceptYour First Spec: One Week In
  2. ConceptWhat This Job Actually Is
  3. ConceptA Little Code, A Lot of Intent
  4. ConceptThree Doors In
  5. ConceptThe Five-Rung Skill Ladder
  6. ConceptThe Parts Nobody Advertises
  7. CheckpointCheckpoint: The Role and the Ladder
  8. Hands-onBuild the Derivation from the Spec
  9. Concept2026 Workflow: SCE, R, and AI
  10. ConceptThe 90-Day Plan
  11. CheckpointCheckpoint: Roadmap and AI Reality
  12. ConceptKey Takeaways and Where to Go Next
Concept1 / 12

Your First Spec: One Week In

2026 CAREER MAP • CLINICAL STATISTICAL PROGRAMMING

Your First Spec

One Week In — a 2026 Career Map

Today's Map

• What the day contains

• Who gets hired

• The skill ladder

• The 2026 workflow

• The 90-day plan

Anchor — SAP §6.2

• Rule: TRT01SDT = first EX dose date

• None → randomization date + flag

• No invented data

[opening] A junior statistical programmer receives a reviewer's note citing SAP section 6.2 and must implement the ADSL treatment-start derivation. This lesson turns that moment into a realistic 2026 career map: what the job is, who gets hired, how to train in order, and what AI does and does not change.

Speaker notes

Welcome to your first specification, one week into the job. The deliverable here is a regulatory submission, not a model, so we start from that reality. Today's map covers what the day contains, who gets hired, the skill ladder, the 2026 workflow as of August 30, 2026, and the 90-day plan. This matters for the Subject-Level Analysis Dataset (ADSL), because our concrete anchor is Statistical Analysis Plan (SAP) section 6.2. There, treatment start uses TRT01SDT, which equals the first Exposure (EX) dose date. If there is no EX dose date, use the randomization date and flag that imputation, and remember that any number or Statistical Analysis System (SAS) line you see is traceable to the source guide, with no patient data invented.

Concept2 / 12

What This Job Actually Is

What This Job Actually Is

The role: spec-driven data pipelines under GxP — not data science.

1 · Build the data pipeline

Collected data → SDTM domains → ADaM per SAP → tables, listings, figures (TLFs) — all under GxP.

2 · Traceability and review

Every number traceable; every program has a reviewer. Reproducibility outranks exploration.

3 · Auditable deliverable

An auditable artifact, not an insight. Deadlines follow the trial calendar — database lock, not a product sprint.

4 · Fit check

Open-ended experimentation? Wrong job. Work that is inspected and depended on? Right job.

[concept · L1] Clarify the role: clinical statistical programmers sit in the middle of a trial pipeline, building spec-driven data pipelines under GxP — and this is not data science.

Speaker notes

The role here is a spec-driven data pipeline, not data science. You map collected data into Study Data Tabulation Model domains, derive Analysis Data Model datasets exactly as the Statistical Analysis Plan describes them, then program the tables, listings, and figures, the TLFs, that carry the results, all under Good Practice regulations, GxP. Every number is traceable and every program has a reviewer, because reproducibility outranks exploration. What you hand over is an auditable artifact, not an insight, and deadlines follow the trial calendar and database lock, not a product sprint. If you want open-ended experimentation, this is the wrong job. If you want work that gets inspected and depended on, it is the right one.

Concept3 / 12

A Little Code, A Lot of Intent

L1 · DATA WALKTHROUGH

A Little Code, A Lot of Intent

/* SAP section 6.2: treatment start = first EX dose date;

   if none, use randomization date and flag the imputation */

data adsl;

  set adsl_base;

  trt01sdt = coalesce(first_ex_dtc, rfstdtc);

  if missing(first_ex_dtc) then trt01dtf = "Y";

run;

Rule — TRT01SDT = first EX dose date; randomization date is only a flagged fallback.

Reviewer — exact SAP rule; coalesce priority; flag marks the judgment call.

Candidate — deliverable shape: a little code, a lot of defensible intent.

Evidence — no simulated trial rows supplied; rule + snippet are the evidence.

[data walkthrough · L1] Real-code walkthrough of the canonical ADSL derivation from the source guide. No simulated trial rows were supplied for this lesson, so no patient-level values are shown; the rule and the exact snippet are the evidence.

Speaker notes

Let's read this snippet from the Subject-Level Analysis Dataset, or ADSL, the way a reviewer would. The Statistical Analysis Plan (SAP), section 6.2, says treatment start is the first dose date in the Exposure (EX) domain; if there is none, we fall back to the randomization date and flag that imputation. That's why TRT01SDT comes from coalesce of first_ex_dtc and rfstdtc, in that priority order. The flag TRT01DTF equals Y only when the fallback is used, so the judgment call is visible. As a candidate, this is the deliverable shape you want: a little code carrying a lot of intent, all of it defensible. One data discipline note: we teach the rule from the canonical snippet because no trial rows were supplied, so no patient numbers are invented.

Concept4 / 12

Three Doors In

Three Doors In

CRO Junior RolesSponsor / CRO InternshipsAdjacent Switches
Volume entry — sponsors outsource at scaleConverts to offers when placements go wellLateral move from DM · biostatistics · IT / analytics

EVERY DOOR TESTS THE SAME EVIDENCE

Read a spec · write code · show your QC

Certs are keyword filters at screening; a QC'd portfolio on public pilot data speaks louder.

[concept · L1] How people actually enter the field: CRO junior roles, internships, and adjacent-career switches. No door needs a perfect resume; every door tests the same evidence.

Speaker notes

Here are three realistic doors into clinical statistical programming, and they all lead to the same evidence bar. Door one is the volume entry point: a contract research organization, or CRO, junior role, because sponsors outsource at scale, and a CRO junior can touch more studies in a year than most sponsor-side programmers see in three. Door two is a sponsor or CRO internship, which converts to an offer when the placement goes well. Door three is an adjacent switch from data management, biostatistics, or information technology and analytics. Every door tests the same thing: can you read a specification, write defensible code, and show your quality control? Statistical Analysis System, SAS, base and advanced certificates help as keyword filters at screening, but a quality-controlled portfolio on public pilot data speaks louder than any certificate.

Concept5 / 12

The Five-Rung Skill Ladder

The Five-Rung Skill Ladder

RungMastery targetHonest horizon
1SAS base fluency: DATA step, key PROCs, merges, macro literacy2–3 months
2CDISC SDTM (Study Data Tabulation Model) & ADaM (Analysis Data Model) — read a spec; map domains; derive ADSL/BDS/OCCDS6–12 months
3TLF (Tables, Listings, Figures) & QC (Quality Control) — tables from shells; honest independent QC1–2 years
4SCE & Git — absorbed alongside Rungs 2–3Alongside
5AI-assisted practice (as of 2026-08-30) — draft with assistants; verify like a reviewerContinuous; churns fastest

Climb in order — order matters more than speed

Ranges are honest: mentoring/real data compress; gaps stretch

Rung-1 evidence: read a spec, defensible code, show QC

[concept · L1] Present the ordered skill ladder the bootcamp teaches against, with honest horizon ranges for each rung.

Speaker notes

The five-rung skill ladder is climbed in order, because each rung has a different shelf life. Rung one is Statistical Analysis System, or SAS, base fluency: DATA step, key procedures, merges, macro literacy, in two to three months. Rung two is the Clinical Data Interchange Standards Consortium, or CDISC, data standards: read a spec, map domains, derive key datasets, over six to twelve months. Rung three is tables, listings, and figures, or TLF, plus quality control, or QC: tables from shells, honest independent QC, over one to two years. Rung four, statistical computing environment, or SCE, and Git, rides alongside rungs two and three, while rung five, artificial intelligence, or AI, assisted practice as of August thirtieth, twenty twenty-six, is continuous and churns fastest. Entry-level evidence at rung one: read a spec, write defensible code, and show your QC.

Concept6 / 12

The Parts Nobody Advertises

The Parts Nobody Advertises

Audit Pressure

• Inspectors read logs, programs, reviews

• Findings land on named people

• Accountability beats most software jobs

Stress or Clarity

• Stressful to some

• Clarifying to many

• Work matters; the system shows where

Database-Lock Crunch

• Pre-lock weeks compress everything

• The trial doesn't negotiate

• Ask about lock culture in interviews

Reading Load

• Protocols, SAPs, mapping specs

• Data management guidelines, query threads

• Reading takes more hours than coding

[concept · L1] The realistic job features hiring pages omit: audit pressure, database-lock crunch, and the reading load behind the code.

Speaker notes

Now let's talk about the parts nobody advertises, starting with audit pressure. Inspectors really do read your logs, your programs, and your review records, and findings land on named people, so accountability here is heavier than in most software jobs. That same weight reads two ways: some people find it stressful, while many find it clarifying, because the work matters and the system shows exactly where. Then there's the database-lock crunch, where the weeks before lock compress everything, and the trial does not negotiate, so ask about lock culture in interviews. And expect a heavy reading load: protocols, Statistical Analysis Plans, mapping specifications, data management guidelines, and query threads. Reading takes more hours than coding, so treat programming as the visible third of the work.

Checkpoint7 / 12

Checkpoint: The Role and the Ladder

1 An Analysis Data Model (ADaM) subject-level analysis dataset (ADSL) specification tells you how to populate TRT01SDT, the date of first exposure to protocol treatment. The flag TRT01DTF is used when the start date is not an exact observed date. Which reading of the derivation rule is correct?

2 The concept arc separates auditable reporting artifacts from exploratory data science. Which statements match the role/ladder model? Select all that apply. (select all that apply, then Check)

Speaker notes

Time for a checkpoint on the role and the ladder. First, an Analysis Data Model (ADaM) subject-level analysis dataset (ADSL) specification asks how you populate TRT01SDT, the date of first exposure to protocol treatment, and when the flag TRT01DTF applies. The correct reading is B: prefer the observed protocol-treatment dose date from EX, fall back to the randomization date only when a usable EX date is not present, and set TRT01DTF, because without that flag the audit trail would imply that a real dose occurred. Second, the concept arc separates auditable reporting artifacts from exploratory data science, so which statements match the role and ladder model? The correct answers are A, C, and D: a table, listing, or figure (TLF) that supports a regulatory conclusion belongs to an auditable artifact chain, you can enter through source data, Study Data Tabulation Model (SDTM), or ADaM but the evidence bar is set by the final use of the result, and the ladder order is source data, SDTM, ADaM, analysis output, and submission, with traceability retained from each rung to the next.

Hands-on8 / 12

Build the Derivation from the Spec

Hands-on interactive — if it does not load, open the paired article and try the exercise there.

Speaker notes

This one is not a video exercise — it runs hands-on on the website at jaimeyan.com/learn. You will match each clause of the Statistical Analysis Plan (SAP) section 6.2 to the right statement of the canonical code, and assemble the derivation for the Subject-Level Analysis Dataset (ADSL) the way a reviewer would read it. Try it right after the video, and notice how the primary source first_ex_dtc, the fallback rfstdtc, and the TRT01DTF imputation flag each earn their place.

Concept9 / 12

2026 Workflow: SCE, R, and AI

2026 Workflow: SCE, R, and AI

L2 · Where the work runs now

• SCE + Git literacy — expected at hire or shortly after

• SAS remains the anchor in most validated environments

• R + pharmaverse packages read well — start with the engine in target job listings

L3 · AI assistance — as-of 2026-08-30

• Proven: debugging code · drafting independent QC · template generation

• Hype: autonomous submission generation — no public quantitative evaluation

• Shrinks typing-heavy junior work; values spec-reading, QC judgment, validation instincts

Planning note — direction matters more than magnitude: no invented percentages.

[concept · L2/L3] Where this work runs now — validated environments, Git, multi-engine skills — and the evidence-tiered view of AI assistance. AI claims are flagged as time-sensitive.

Speaker notes

By 2026, where the work runs matters as much as what you code, so Statistical Computing Environment, or SCE, and Git literacy are increasingly expected at hire or shortly after. SAS stays the anchor in most validated environments, while R with the pharmaverse packages reads well and signals range — start with the engine your target job listings name. For artificial intelligence, or AI, keep an evidence tier in mind: debugging code, drafting independent quality control, or QC, and template generation are proven, as of the twenty twenty-six framing. Autonomous submission generation is still hype, with no public quantitative evaluation. AI shrinks the typing-heavy share of junior work — first-draft mapping code, boilerplate, log triage — and appreciates spec-reading, QC judgment, and validation instincts. There are no invented percentages here; direction matters more than magnitude for your planning.

Concept10 / 12

The 90-Day Plan

The 90-Day Plan

TIMELINEFOCUSBUILD / SOURCE
Days 1–30SAS base + SDTMParts 4 & 5 — map two raw files into standard domains
Days 31–60ADaM lineParts 2, 7, 8 — derive ADSL, one BDS, one OCCDS on pilot data
Days 61–80TLF craft + QCPart 3 — program five tables from shells, then QC independently
Days 81–90Environment + positioningParts 1, 12, 10, 11 — SCE habits, Git basics, interview drill, portfolio

▪ Pace: halve the pace, double the calendar — the order survives.

▪ Day 90: evidence point, not finish line.

▪ Target: ≈1 hr/day → ~3 months portfolio-ready; junior roles 6–12 months.

[concept · L1/L3] The guide's concrete self-study roadmap, wired to the bootcamp and calibrated for portfolio-ready in about three focused months.

Speaker notes

Here's the ninety-day plan, assuming one hour a day. Days one to thirty: build Statistical Analysis System, or SAS, base and Study Data Tabulation Model, or SDTM, by mapping two raw files into standard domains. Days thirty-one to sixty: Analysis Data Model, or ADaM, line, deriving Subject-Level Analysis Dataset, or ADSL, Basic Data Structure, or BDS, and Occurrence Data Structure, or OCCDS, on pilot data. Days sixty-one to eighty: Tables, Listings, and Figures, or TLF, craft and Quality Control, or QC: program five tables from shells, then QC independently. Days eighty-one to ninety: Statistical Computing Environment, or SCE, habits, Git basics, interview drill, portfolio. Halve pace, double calendar; order survives, day ninety is an evidence point, not the finish line, and at one hour a day, portfolio-ready in three months, while junior roles typically take six to twelve months.

Checkpoint11 / 12

Checkpoint: Roadmap and AI Reality

1 In the 90-day plan, which sequence correctly maps the main deliverables?

2 Which statements accurately describe the evidence-tiered view of AI assistants in statistical programming as of 2026-08-30? Select all that apply. (select all that apply, then Check)

3 Explain the three positioning facts in this lesson: (1) the SCE/Git expectations, (2) the engine question, and (3) why the portfolio from days 61-80 matters. Write 4-6 sentences. (reflect, then reveal)

Reveal analysis
A strong answer should say: SCE/Git expectations mean publishing role-relevant, version-controlled code that reviewers can run and reproduce. The engine question asks whether you present yourself as a SAS-programming candidate, an R/Python candidate, or a hybrid, and whether that engine supports clinical regulatory outputs. The days 61-80 portfolio matters because it demonstrates that you can take ADaM through TLF generation and formal QC, producing submission-ready evidence that a hiring manager can review.
Speaker notes

Let's check your understanding with a checkpoint on the ninety-day roadmap and the artificial intelligence, or AI, reality. The correct sequence is C: Day zero to forty-five builds Study Data Tabulation Model, or SDTM; Day forty-six to sixty builds Analysis Data Model, or ADaM; Day sixty-one to eighty builds Tables, Listings, and Figures, or TLF, plus Quality Control, or QC. That order works because SDTM gives downstream analysis datasets a stable foundation, ADaM is derived in the middle, and the final phase produces TLFs plus QC because those outputs depend on ADaM being complete. For the evidence-tiered view of AI assistants as of August thirtieth, twenty twenty-six, the accurate statements are A and D. Those are correct because evidence supports AI for drafting, summarization, and low-risk exploratory support, but not for replacing final QC or regulatory sign-off, while verification cannot be delegated due to the need for traceable human accountability, clinical context, and a full evidence-chain review. A strong answer says that Statistical Computing Environment, or SCE, and Git expectations mean publishing role-relevant, version-controlled code that reviewers can run and reproduce; the engine question asks whether you present yourself as a Statistical Analysis System, or SAS, programming candidate, an R or Python candidate, or a hybrid, and whether that engine supports clinical regulatory outputs; and the days sixty-one to eighty portfolio matters because it demonstrates that you can take a project from analysis-ready data through to final TLFs and QC.

Concept12 / 12

Key Takeaways and Where to Go Next

Key Takeaways & Where to Go Next

• Map: spec-driven GxP pipelines — SDTM mapping, ADaM derivation, auditable TLFs — not data science.

• Entry: CRO junior roles, internships, or adjacent switches; show spec, clean code, QC.

• Order: SAS Base → CDISC → TLF/QC → SCE/Git → AI-assisted practice.

• Signals: SAS certs pass screens; a QC’d public-pilot portfolio wins interviews.

• AI lens: junior work shifts to spec + verification — the signer stays accountable.

• Next: read Part 11 “Realistic 2026 Guide”; then Part 12 “Git for Clinical Programmers”.

[summary] Consolidate the realistic 2026 map, restate the data-discipline note, and link back to the paired article in the bootcamp series.

Speaker notes

The work is spec-driven pipelines under Good Practice, GxP: Study Data Tabulation Model, SDTM, mapping, Analysis Data Model, ADaM, derivation, and auditable Tables, Listings, and Figures, TLFs — not data science. Enter through Contract Research Organization, CRO, junior roles, internships, or adjacent switches; evidence opens all three doors: read a spec, write defensible code, show your quality control, QC. Climb in order — the Statistical Analysis System, SAS, base, then Clinical Data Interchange Standards Consortium, CDISC, then TLF and QC, then the Statistical Computing Environment, SCE, and Git, then artificial intelligence assisted practice — because order beats speed as of August 2026. Certificates pass screens; a QC'd portfolio on public pilot data wins interviews, and AI shifts junior work toward specification and verification — the signer stays accountable. Next, read Part 11, the Realistic 2026 Guide, then Part 12, Git for Clinical Programmers.

✓

Lesson complete

Nice work — every scene seen. Keep the momentum going.

← → Space to navigate · progress is saved locally in your browser

AI Tutor

Ask the tutor