One journey, four stations
Every clinical study moves its data through the same four stations, each with its own contract.
- Data enters as raw EDC extracts and leaves as TLF deliverables, passing through SDTM tabulation domains and ADaM analysis datasets.
- Every arrow is a governed transformation: validated programs, reviewed specifications, no hand edits.
- Each station has a written contract — mapping spec, ADaM spec, mock shell, define-XML — and code that produces anything the contract does not name is a finding.
- The contracts govern the stations; you can be checked against every row in them.
- Traceability runs both directions: any number in any table can be walked back to the value a site collected.
- This page walks the whole road with one fictional study: Study XYZ, subjects 001–004.
- data station
- contract document
- governed flow
Teaching schematic — not measured. Illustrative structure only; rules per SDTMIG/ADaMIG.
Sources: series roadmap · Part 4, SDTM domain basics.
Raw data as it actually arrives
A mock adverse-event extract: four subjects, eleven raw columns, and every classic mapping problem in one screen.
- The AE extract for Study XYZ arrives on a Tuesday: eleven columns, no standard names, no standard formats.
- Four subjects, four events — and nothing about the table hints at the target shape. The SDTM model does.
- Subject 002's onset is a partial date, month precision only: 2023-02. That is valid, collected data — not a bug.
- Subjects 003 and 004 carry unknown dates (UNK). Unknown is stored as missing and raised as a data query — never invented.
- Severity mixes two scales: CT-assigned Grade 1 (Mild) and Grade 3 (Severe) beside bare MODERATE. Relatedness is free text.
- Subject 004's MedDRA coding is still pending. The row ships with AEDECOD blank, tracked by an open query.
- None of this is unusual. The distance from this table to a submission-grade domain is about forty decisions — and every one belongs in the spec before the code.
- raw column header
- collected row
- mapping decision needed
Teaching schematic — not measured. Table reproduced from the Part 5 mock extract (itself synthetic, fictional Study XYZ).
Source: Part 5, SDTM AE domain mapping.
Mapping into an SDTM domain
Spec rows turn the raw extract into the AE domain: identifiers, verbatim topic, controlled terminology, coded terms, dates, sequence.
- Mapping starts from the raw table: which columns move directly, which need a decision row, which need a data query.
- The mapping specification is the contract between data management, programming, and inspection — target, source, transformation, CT, notes.
- The AE domain takes shape: one row per observation per subject, standard variables, standard names.
- USUBJID is STUDYID + '-' + SUBJID, built in one program and reused verbatim everywhere — two programs building it is how trailing-blank ghosts are born.
- AETERM keeps the verbatim string ("head ache" stays "head ache"); AEDECOD comes only from the version-pinned MedDRA coding deliverable — 004 stays blank with an open query.
- Dates ship as collected ISO 8601 strings: 2023-02 keeps month precision. Completing a partial in SDTM changes clinical meaning.
- Severity collapses to MILD/MODERATE/SEVERE through decision rows; AESER becomes Y/N; the six seriousness criteria stay blank unless checked, never defaulted to N.
- AESEQ numbers records within subject from a deterministic sort — never EDC row numbers, which reshuffle on every re-extract.
- Result: a submission-grade AE domain where every variable traces to a spec row — the code contains no judgment, only transcription.
- raw extract
- mapping spec rows
- SDTM AE
Teaching schematic — not measured. Variables and rules per SDTMIG; fictional Study XYZ data.
Sources: Part 5, AE mapping · Part 6, mapping specification.
SUPPQUAL: the pressure valve
CRF values with no IG variable go vertical: one row per stored value, linked back through the parent's sequence number.
- Real CRFs collect things the Implementation Guide has no variable for — verbatim relatedness, a local toxicity grade.
- Those values go to a supplemental qualifier dataset, vertical: one row per stored value instead of one extra column.
- Each response becomes a SUPPAE row: QNAM = AERELVER, QLABEL = "Causality, Verbatim", QVAL = the collected string.
- The local grade string rides along as AETOXLOC because safety review uses the raw scale — with QORIG = CRF and QEVAL = INVESTIGATOR.
- The link back to the parent record runs through IDVAR = AESEQ and the sequence number in IDVARVAL.
- That link is exactly why AESEQ must be derived deterministically — reshuffled sequences silently orphan every attached qualifier.
- parent domain record
- SUPP-- row (vertical)
- IDVAR/IDVARVAL link
Teaching schematic — not measured. SUPP-- structure per SDTMIG; values from the fictional Study XYZ extract.
Source: Part 4, SDTM domain basics.
ADSL: one row per subject is the whole job
The Subject-Level Analysis Dataset is the spine every other analysis dataset inherits — and merge discipline is what keeps it one row per subject.
- ADSL carries demographics, treatment variables, key dates, and population flags at exactly one record per subject.
- Sources with many rows per subject — EX, SV, DS — are aggregated before the merge, never after.
- TRT01SDT is a cascade: restrict EX to qualifying records, take the earliest date, then the SAP-defined fallback — not a lookup.
- Every population flag ships with three things: the SAP citation it implements, the derivation, and a QC listing of disagreements.
- Merges land on the spine: rows must equal distinct subjects at every step.
- Subject 002 randomized to A but dosed B: TRT01P and TRT01A disagree — that is data for medical review, never a code fix.
- The running check catches fan-out the moment it happens — a mystery N=204 becomes a five-minute fix.
- Get ADSL right and every downstream dataset inherits the right answer on every row; get it wrong and they inherit the wrong answer just as consistently.
- ADSL row (the spine)
- derivation cascade
- invariant check
Teaching schematic — not measured. Treatment scenarios per the Part 2 discrepancy table (illustrative).
Source: Part 2, ADSL derivation walkthrough.
BDS: parameters, baseline, windows
One subject's systolic blood pressure shows how windowing, the baseline flag, and change from baseline decide whether a BDS dataset is right.
- BDS grain: one record per subject per parameter per analysis timepoint, with ADY anchored to first dose and no Day 0.
- Subject 001's collected SYSBP records: screening 124, Day -2 value 128, Week 1 value 122, a Week 12 visit 118, an unscheduled recheck 120, and a late visit 116.
- Windows are day ranges transcribed from the SAP: Baseline -30..-1, Week 1 1..7, Week 12 64..98, Week 24 141..189.
- The Day 130 visit lands in the gap between windows: it stays in the dataset flagged off — dropping it would destroy the query trail.
- Baseline is the last non-missing value on or before first dose — 128 at Day -2, not the first record (124). Picking the wrong record here corrupts every change-from-baseline table at once.
- BASE copies onto every record; CHG = AVAL - BASE, so the Week 12 visit reads -10 mmHg.
- Two records sit inside Week 12: the SAP names the winner — last record, so the recheck carries ANL01FL = Y and the loser stays, flag off.
- collected record (illustrative mmHg)
- flagged: ABLFL / ANL01FL
- SAP window band
Teaching schematic — not measured. Values invented for teaching; window boundaries per the Part 17 worked example.
Sources: Part 7, ADaM BDS · Part 17, windowing & baseline.
TLF: the shell is a contract
From mock shell to shipped RTF — and the four-pass QC order that catches the N=24/N=26 class of defect at its cheapest point.
- The mock shell is a contract: output ID, titles, population statement, column blocks, denominators, footnotes — every element checkable.
- The program builds the table with denominators pulled from ADSL under the population flag — never from the event dataset.
- The output says N=26 where the shell says N=24: both defensible, only one is the deliverable — caught at Pass 2, data definitions.
- QC runs four passes cheapest-first: format, definitions, numbers, consistency — and never recomputes an output that fails an earlier pass.
- Every discrepancy gets a record — observed, expected, disposition, owner — because a silently fixed discrepancy resurfaces at the next data cut.
- The shipped RTF carries titles inside the file, one style template, page x of y, precision exactly per shell — including the n (100) convention.
- Read the shell into decisions in order, population first — skipping step one is the single most common serious discrepancy on a first QC cycle.
- shell / output
- QC catch
- discrepancy record
Teaching schematic — not measured. The N=24/N=26 scene is the Part 3 worked example (illustrative).
Source: Part 3, mock shell to RTF.
Define-XML and the package that ships
The spec becomes machine-readable metadata, the guide explains what machines cannot, and the whole package travels together.
- The mapping spec's columns are the source: target, source, transformation, CT/origin, notes.
- They generate define-XML almost one-to-one: origin column → ItemDef origin attribute, CT column → CodeList reference, transformation → MethodDef text.
- A rule that varies by parameter — TEMP converted from Fahrenheit, weight passed through — becomes ValueListDef plus WhereClauseDef: value-level metadata.
- Define-XML is generated, never hand-typed — generation is what makes data/metadata disagreement structurally impossible.
- The ADRG covers what machines cannot: orientation, special derivations, population filters, and known issues — disclosed rough edges earn trust.
- The package ships as one unit in eCTD module 5: datasets, define.xml, define.pdf, the guide, and the annotated CRF.
- That closes the journey: every value traceable to collection, every rule to a spec row, every claim checkable — the whole point of the road you just walked.
- spec columns
- define-XML element
- submission package
Teaching schematic — not measured. Element-to-column mapping per the Part 16 table (summarized).
Sources: Part 16, define-XML & the reviewer's guide · Part 6, mapping specification.