SDTM All lessons Scene 1 / 12

SDTM · Interactive Lesson

SDTM Domain Basics

12 scenes· ~22 min· pairs with the article

Step through the scenes, pass the checkpoint quizzes, and try the hands-on exercises. Progress saves locally in this browser — no account, no tracking.

Scene index · 12 scenes
  1. ConceptDay One: Raw Export vs. SDTM
  2. ConceptOne Row per Observation, Data as Collected
  3. ConceptDomain Classes Predict the Variables
  4. ConceptThe Topic–Timing–Qualifier Skeleton
  5. ConceptReal Rows: USUBJID and Study Days
  6. CheckpointCheck Point: Reading the Model
  7. ConceptControlled Terminology: Collected vs. Submission
  8. ConceptSUPPQUAL: The Pressure Valve
  9. ConceptSDTMIG, define-XML, and Your Reading Order
  10. Hands-onHands-On: Match the Domain
  11. CheckpointKnowledge Check: The Whole Model
  12. ConceptKey Takeaways and Where to Go Next
Concept1 / 12

Day One: Raw Export vs. SDTM

Day One: Raw Export vs. SDTM

Raw rows do not show the target shape — the SDTMIG model does.

Raw EDC export

received from data-entry screens

SDTMIG model

defines the table grammar

Submission SDTM

readable with no CRF

Agenda: Domain classes  ·  Topic–timing–qualifier skeleton  ·  USUBJID

Controlled terminology  ·  SUPPQUAL  ·  First-day reading order

Course examples use simulated study 043-18101 — identifiers, dates, and codes are real source data.

Opening scene that places the learner in the real first-day situation of a junior clinical programmer: facing a raw EDC export and the SDTMIG, with the job of producing submission datasets a reviewer can read without the CRF.

Speaker notes

Picture your first morning on a study. You receive a raw electronic data capture export from the data-entry screens, plus the Study Data Tabulation Model Implementation Guide, the SDTMIG, hundreds of pages long. Your job is to turn those raw rows into submission datasets. Raw rows give no clue about the target shape; the Study Data Tabulation Model, SDTM, model does, because it is a fixed table grammar that makes every dataset predictable, so a reviewer who has never seen your case report form, your CRF, can still read it. Today we cover domain classes, the topic–timing–qualifier skeleton, the Unique Subject Identifier, USUBJID, controlled terminology, Supplemental Qualifiers, SUPPQUAL, and a first-day reading order. All examples come from a simulated study.

Concept2 / 12

One Row per Observation, Data as Collected

One Row per Observation,

Data as Collected

SDTM (Study Data Tabulation Model) — tables of observations

Domains: AE adverse events  ·  CM medications  ·  VS vital signs

Rule 1  —  one row per observation

• 3 adverse events reported → 3 AE rows

• DM carries exactly one row per subject

Rule 2  —  data as collected

• RFSTDTC = 2019-05-08T09:23

• No imputation, no analysis logic in SDTM

Derived, analysis-ready values live one step downstream in ADaM — not in SDTM

AE

Concept scene establishing the two governing rules of SDTM: one row per observation per subject, and data transmitted exactly as collected, with no analysis logic.

Speaker notes

In the Study Data Tabulation Model, or SDTM, a domain is one table holding one kind of observation. Adverse events, or AE, are rows in the AE domain; medications, or CM, are rows in CM; and vital signs, or VS, are rows in VS. Rule one: one row per observation per subject, so three reported adverse events generate three AE rows, while Demographics, or DM, carries exactly one row per subject. Rule two: data as collected, meaning no imputation and no analysis logic. For example, DM stores the collected Reference Start Date/Time, or RFSTDTC, as 2019-05-08T09:23, exactly as the electronic data capture, or EDC, export supplied it. Derived, analysis-ready values live one step downstream in the Analysis Data Model, or ADaM, not in SDTM.

Concept3 / 12

Domain Classes Predict the Variables

Domain Classes Predict the Variables

SDTMIG (SDTM Implementation Guide) groups domains into five classes.

ClassWhat the row tells youExamples (SDTMIG)
InterventionsWhat was administeredCM, EX
EventsWhat happenedAE, MH, DS
FindingsWhat was measured or observedLB, VS, EG, PE
Special-PurposeDM: one row per subject; SV: actual visitsDM, SV
Trial DesignProtocol as rows, not patient dataTA, TV, TS

• Class predicts the variable list: identify the class, then confirm in SDTMIG.

• DS and EX feel structural — they still belong to the Events and Interventions classes.

• Trial Design rows are protocol data: TA carries ARMCD AIRIS1 and ETCD CYCLE1.

Concept scene on the SDTM domain classes: Interventions, Events, Findings, Special-Purpose, and Trial Design, with example domains drawn from the course vocabulary and the study extract.

Speaker notes

The Study Data Tabulation Model Implementation Guide (SDTMIG) groups domains into five classes. Interventions answer what was administered, like Concomitant Medications (CM) and Exposure (EX); Events answer what happened, like Adverse Events (AE), Medical History (MH), and Disposition (DS). Findings answer what was measured or observed, such as Laboratory (LB), Vital Signs (VS), Electrocardiogram (EG), and Physical Examination (PE). Special-Purpose covers Demographics (DM) and Subject Visits (SV); Trial Design covers Trial Arms (TA), Trial Visits (TV), and Trial Summary (TS). Class predicts the variable list: identify it, then confirm in SDTMIG; DM has one row per subject, SV records actual visits, and DS and EX still belong to Events and Interventions. Trial Design is protocol data: TA carries ARMCD AIRIS1 and ETCD CYCLE1, and the TV row for CYCLE 1 DAY 5 begins '4 days after start of Treatment Epoch.'

Concept4 / 12

The Topic–Timing–Qualifier Skeleton

The Topic–Timing–Qualifier Skeleton

One repeating row skeleton: identifiers → topic → timing → qualifiers → standardized topic

1 · Identifiers2 · Topic variable3 · Timing4 · Qualifiers5 · Standardized topic
STUDYID, USUBJID, --SEQ--TRT / --TERM / --TESTCD--DTC, --DYe.g., --STAT, --REASND--DECOD (optional)

Swap the prefix by class

• --TRT: interventions (CM)

• --TERM: events (MH)

• --TESTCD: findings (VS)

• Same pattern fits CM/MH/VS

Timing: --DTC and --DY

• --DTC: character ISO 8601

• partial dates ship as-is

• --DY: day 1 = reference start

• no day 0; before: −1, −2

--SEQ: the record key

• --SEQ deterministic per subject

• later tables link back to it

• subject 043-18101-74001-001

• SVSTDY = -6 on 2019-05-02

Concept scene teaching the variable pattern repeated in every domain: identifiers, a topic variable, ISO 8601 timing, qualifiers, and standardized topic.

Speaker notes

In the Study Data Tabulation Model, or SDTM, every domain repeats a skeleton: identifiers, a topic variable, timing, qualifiers, and a standardized topic. Identifiers include STUDYID, the study identifier, and USUBJID, the unique subject identifier, plus --SEQ, the sequence number. The prefix changes by class: --TRT for interventions, --TERM for events, --TESTCD for findings, the same skeleton fits Concomitant Medications, Medical History, or Vital Signs. --DTC is a character International Organization for Standardization 8601 string, because partial dates ship as collected. --DY is study day relative to the reference start date: day one is on or after treatment, negative down to minus one, and no day zero. In the Subject Visits extract, the demo subject shows the subject visit study day, SVSTDY, equal to minus six on 2019-05-02, and --SEQ must be deterministic because later tables link back to it.

Concept5 / 12

Real Rows: USUBJID and Study Days

Real Rows: USUBJID and Study Days

USUBJID: build once, reuse verbatim

DM: 1 row per subject; SV: 5 rows, same USUBJID

043-18101 + 74001-001 → 043-18101-74001-001

Derive SV study day from the DM reference date

Subject 043-18101-74001-001: RFSTDTC 2019-05-08

Cycle 1 Day 1 → SVSTDY = 1; Day 5 → 5; Day 8 → 8

Missing and absent data: follow the rule

Subject 043-18101-74001-005: RFENDTC empty

No matching rows: state the rule; do not fabricate records

Real-data walkthrough page stepping through the provided DM and SV extracts: USUBJID composition, reuse of the same identifier across domains, and study-day derivations.

Speaker notes

The Unique Subject Identifier, USUBJID, is built once in Demographics, DM, and reused verbatim. Each USUBJID is the study-site-subject key: STUDYID, the study identifier, a hyphen, then SUBJID, the subject identifier. DM holds one row per subject; Subject Visits, SV, holds five rows for that same demo subject and carries the identical string, so every domain stays consistent. Study day comes from DM's reference start date, RFSTDTC: Cycle 1 Day 1 on 2019-05-08 gives SVSTDY, the subject visit study day, equal to 1; Day 5 gives 5; Day 8 gives 8. A missing value is a data query, not a programming problem: one demo subject has RFSTDTC 2019-12-19T10:25 and an empty RFENDTC, the reference end date. Where no matching rows exist, we state the rule; we never fabricate patient records.

Checkpoint6 / 12

Check Point: Reading the Model

1 In CDISC (Clinical Data Interchange Standards Consortium) SDTM (Study Data Tabulation Model), a domain is a collection of observations related to one clinical topic. Which statement best describes the one-row-per-observation rule?

2 In CDISC SDTM, most datasets are grouped into three domain classes: Interventions, Events, and Findings. An `AE` dataset contains one row per reported adverse event. Which question does the `AE` table answer?

3 The standard SDTM domain skeleton separates variables into identifier, topic, timing, and qualifier roles. Which statements correctly identify these roles? Select all that apply. (select all that apply, then Check)

Speaker notes

This is your checkpoint quiz on reading the Study Data Tabulation Model, or SDTM, and its domain structure. For a domain, the one-row-per-observation rule means the correct answer is B: each row is one recorded observation, such as one event, one finding, or one dose administration, which sets row granularity so a subject can have many rows in a domain. For the adverse event, or AE, dataset, which contains one row per reported adverse event, the correct answer is B: what happened to the subject during the study, because an adverse event is something that happens to a subject and AE belongs to the Events class. Which statements correctly identify identifier, topic, timing, and qualifier roles in the standard SDTM domain skeleton? The correct answers are A, B, and C: Unique Subject Identifier, or USUBJID, is an identifier; --DTC is a timing variable for when the observation occurred; and in an Events domain, --TERM is the topic variable that names the specific event, following the standard role pattern of identifiers, topic variable, timing variables, then qualifier variables.

Concept7 / 12

Controlled Terminology: Collected vs. Submission

Controlled Terminology:

Collected vs. Submission

RAW COLLECTED

SEX field

code: M

label: Male

MAPPING SPEC

decision rows

M/Male → M

no row = no mapping

SDTM SUBMIT

DM.SEX

value: M

from codelist

Example: SEX = M, F · ETHNIC = HISPANIC OR LATINO | NOT HISPANIC OR LATINO | NOT REPORTED

Rules: map every CT variable — no decision row, no mapping · lock one dated/versioned CT package at specification

Concept scene on controlled terminology: CT-bound variables must use CDISC codelist values, raw collected text is never shipped as-is, and the study locks one dated CT package at specification time.

Speaker notes

Controlled Terminology, or CT, is the dated, versioned set of codelists published by CDISC, the Clinical Data Interchange Standards Consortium, and tied to your Implementation Guide version. Your study locks one CT package at the specification stage. Raw collected values are never moved into SDTM, the Study Data Tabulation Model, as-is. In the raw Electronic Data Capture extract, one column pair holds code M with display label Male; your mapping specification needs decision rows that map collected value to submission value. For SEX that means M and Male map to M, and the submitted DM, Demographics, dataset ships M from the codelist. Remember: no decision row, no mapping; SEX codelist values are M and F, and ETHNIC codelist values are HISPANIC OR LATINO, NOT HISPANIC OR LATINO, and NOT REPORTED.

Concept8 / 12

SUPPQUAL: The Pressure Valve

SUPPQUAL: The Pressure Valve

CRF values with no SDTMIG variable go vertical in SUPP-- — one row per stored value.

SUPP-- field / groupRole and routing
RDOMAIN / USUBJIDPoints to the parent domain and subject that own the row
IDVAR / IDVARVALRoutes back to the parent row key — usually --SEQ; derive --SEQ deterministically
QNAM / QLABELQualifier name — QNAM is at most 8 characters — plus its label
QVAL / QORIG / QEVALValue, origin, and evaluator for the stored qualifier
SUPPAE / SUPPCM / SUPPVSExample SUPP-- targets: long rows, one row per QNAM value

SAS: sort by usubjid idvarval qnam; transpose id=qnam var=qval

Merge: wide SUPP-- back to AE by USUBJID

QC: pre/post row counts match; convert IDVARVAL to numeric --SEQ

Concept and code-walkthrough scene covering supplemental qualifier datasets: when a CRF field has no SDTM variable, it goes vertical into SUPP--, and the canonical SAS snippet shows the long-to-wide transpose discipline.

Speaker notes

Some case report form (CRF) fields have no variable in the Study Data Tabulation Model Implementation Guide (SDTMIG). Those values go vertical into a supplemental qualifier dataset, such as SUPPAE, SUPPCM, or SUPPVS — one row per stored value, not an extra column. RDOMAIN and the unique subject identifier, USUBJID, point to the parent domain and subject; IDVAR and IDVARVAL route back to the parent record, usually through --SEQ. QNAM, at most eight characters, and QLABEL name the qualifier; QVAL, QORIG, and QEVAL carry the value, origin, and evaluator. In SAS, sort by usubjid idvarval qnam, transpose with id=qnam and var=qval, then merge the wide result back onto the adverse event (AE) domain by USUBJID. Then compare row counts before and after that merge, and convert IDVARVAL from character to numeric before comparing it with --SEQ.

Concept9 / 12

SDTMIG, define-XML, and Your Reading Order

SDTMIG, define-XML, and Your Reading Order

SDTMIG — human-readable standard

• Per-domain chapters and assumptions

• Required / expected / permissible variables

• Scope: general — what CM should look like

define-XML — machine-readable profile

• Describes your actual study datasets

• Origins, codelists, computational methods

• Scope: particular — what your CM is

First-day reading order

1Protocol synopsis & schedule of activities
2Annotated CRF
3SDTMIG chapter for each domain
4Study mapping specification

Modern flow: mapping spec → define-XML → CDISC CORE checks every build.

AI drafts are risky — map each variable to a spec row; --DTC byte-for-byte.

Concept scene positioning SDTMIG versus define-XML, then giving the first-day reading order and framing how modern spec-driven clinical programming and validation change the workflow.

Speaker notes

SDTMIG, the Study Data Tabulation Model Implementation Guide, is your human-readable rulebook: per-domain chapters, required, expected, and permissible variables, plus assumptions. define-XML is the machine-readable profile of your actual datasets: variables, origins, codelist bindings, and computational methods. So SDTMIG says what the Concomitant Medications domain, or CM, should look like in general; define-XML says what your CM looks like in particular. First-day reading order: protocol synopsis and schedule of activities, annotated CRF, SDTMIG chapter for each domain, then study mapping specification. Modern flow keeps the mapping spec structured and versioned, generates define-XML from that same metadata, and runs CDISC CORE validation on every build; reviewers and validation engines read it alongside the data. Agentic drafting is fast but risky: mapped variables must trace to a spec row, and --DTC strings must match the collected string byte for byte.

Hands-on10 / 12

Hands-On: Match the Domain

Hands-on interactive — if it does not load, open the paired article and try the exercise there.

Speaker notes

This exercise is hands-on on the website at jaimeyan.com/learn, where you classify twelve domain codes into Interventions, Events, Findings, Special-Purpose, and Trial Design buckets, then match variable-role cards to a domain prefix, and finish with a join-key mini-round. That means you practice the standard terminology like --TERM for events, --TRT for interventions, and --TESTCD for findings, and you choose the identifier that connects Demographics (DM) to Subject Visits (SV). The correct join key is the Unique Subject Identifier (USUBJID), the Study Identifier-Subject Identifier composite built once and reused verbatim; wrong answers trigger the domain-class or skeleton rule as immediate feedback, and no patient data appears beyond the provided extract, so try it yourself after the video.

Checkpoint11 / 12

Knowledge Check: The Whole Model

1 A draft SDTMIG domain uses code XX. The domain has variables --TESTCD and --ORRES, but no planned-dose variables such as --DOSE or --ROUTE. Based on the variable pattern, which SDTM domain class should XX be mapped to?

2 Which statements about USUBJID composition and reuse are true? Select all that apply. (select all that apply, then Check)

3 When reviewing SUPPAE, the supplemental qualifier dataset for the AE domain, you see a record for QNAM = 'QSPID' where IDVAR and IDVARVAL are both empty, while most other records in SUPPAE use IDVAR = 'AESEQ' and IDVARVAL equal to an AESEQ value. Explain what an empty IDVAR and empty IDVARVAL mean in SUPPAE, how you should correctly associate the QSPID value with AE records in a listing, and then name two QC checks you would run before merging SUPPAE with AE to ensure no AE records are duplicated or lost. (reflect, then reveal)

Reveal analysis
In SUPPAE, IDVAR and IDVARVAL identify the parent AE record to which a supplemental qualifier belongs. If both are empty, the qualifier is subject-level: it applies to all AE records for that USUBJID in the AE domain. Therefore, for QSPID with IDVAR = '' and IDVARVAL = '', you should associate the value by USUBJID only, or carry it down to each AE record for that USUBJID; you must not join it to AE.AESEQ. QC checks include sorting and checking for duplicate QNAM records within the same key (USUBJID, RDOMAIN, IDVAR, IDVARVAL, QNAM); verifying that every nonblank IDVARVAL has a matching AESEQ value in AE for the same USUBJID; and comparing observation counts before and after the merge to detect lost or duplicated AE rows.
Speaker notes

This is your knowledge check on the whole Study Data Tabulation Model (SDTM) model. For a draft Study Data Tabulation Model Implementation Guide (SDTMIG) domain with code XX, where the only pattern is --TESTCD and --ORRES and there are no planned-dose variables, the correct answer is A: map XX to the Findings class, because --TESTCD names the measurement or test and --ORRES is the original result, which is the classic observation pattern. For the Unique Subject Identifier (USUBJID) composition and reuse question, the correct answers are A, B, and C: USUBJID is derived once at the subject level and reused unchanged in every SDTM domain, it is the key that links the same subject's records across domains, and in a multi-site study it typically combines site information and a subject number so it stays unique across sites. In the Supplemental Qualifiers for the Adverse Events (AE) domain, called SUPPAE, when IDVAR and IDVARVAL are both empty the QSPID qualifier is subject-level, so you associate it by USUBJID only and carry it down to every AE record for that subject; before merging SUPPAE with AE, run a uniqueness check on the SUPPAE merge key and a record-count or unmatched-record check to confirm no AE records are duplicated or lost.

Concept12 / 12

Key Takeaways and Where to Go Next

Key Takeaways

and Where to Go Next

• Fixed grammar: I/E/F + special & trial-design tables

• Skeleton: identifiers, topic, ISO 8601 timing, qualifiers

• USUBJID: built once from STUDYID-SUBJID, reused verbatim

• No SDTMIG home? Use vertical SUPP--; QNAM ≤ 8 chars

• Read order: synopsis → CRF → SDTMIG → mapping spec

Full article — “SDTM Domain Basics: The Model Every Clinical Programmer Must Know”

Next in series — SDTM AE domain mapping worked example

Reference — sdtm-mapping-conventions SKILL.md

Summary slide linking back to the paired SDTM Domain Basics article: the core rules of the model, the reading order, and the next step in the bootcamp series.

Speaker notes

Study Data Tabulation Model, or SDTM, is a fixed grammar: domains are one row per observation, grouped into Interventions, Events, and Findings, plus special-purpose and trial-design tables. Every domain repeats the same skeleton — identifiers, a topic variable, International Organization for Standardization 8601 timing, and qualifiers. Unique Subject Identifier, or USUBJID, is built once from Study Identifier and Subject Identifier composite and reused verbatim. If a value has no home in the Study Data Tabulation Model Implementation Guide, or SDTMIG, send it vertical into Supplemental Qualifiers, with Qualifier Variable Name of eight characters or fewer, linked to Sequence Number. Read study in order: protocol synopsis, annotated Case Report Form, SDTMIG chapter, then mapping specification. Full article is 'SDTM Domain Basics: The Model Every Clinical Programmer Must Know'; next is Adverse Events domain mapping worked example, and sdtm-mapping-conventions SKILL.md is distilled reference.

✓

Lesson complete

Nice work — every scene seen. Keep the momentum going.

← → Space to navigate · progress is saved locally in your browser

AI Tutor

Ask the tutor