Day 3 of QC: 312 Events vs 315
DAY 3 · QC · TREATMENT-EMERGENT ADVERSE EVENTS
312 Events vs 315
Safety table (312) vs QC recount (315) — gap = 3 events
Disputed 3: onset before first dose, yet TRTEMFL = 'Y'.
Root cause: onset was compared with randomization date, not TRT01SDT.
Lesson: invisible in code, visible in counts — run date-comparison listings.
Agenda: OCCDS grain → disciplined AE-to-ADSL merge → TRTEMFL anchored to TRT01SDT → partial-date rules → QC checks that catch each defect
Opening page (L1 hook): start from the failure a junior programmer actually meets — an independent QC recount disagrees with the safety table on treatment-emergent events.
Speaker notes
Welcome to Day 3 of Quality Control (QC). The safety table reports 312 treatment-emergent adverse events, but the independent QC program counts 315. That three-event gap is our story. All three disputed rows have onset before first dose, yet they were flagged treatment-emergent because the flag compared onset against the randomization date instead of the treatment start date variable, TRT01SDT. This defect is invisible in the code but visible in the counts, which is why reviewers run date-comparison listings, not code reviews alone. Our agenda covers the Analysis Data Model (ADaM) Occurrence Data Structure (OCCDS) grain, a disciplined adverse events to Subject-Level Analysis Dataset (ADSL) merge, the treatment-emergent flag, TRTEMFL, anchored to TRT01SDT, partial-date rules, and the QC checks that catch each defect.
OCCDS: One Row per Adverse Event
OCCDS: One Row per Adverse Event
ADAE is an OCCDS dataset — one row per adverse event.
Five events from one subject contribute five ADAE rows.
| ADaM class | Row grain | Typical content | Typical datasets |
|---|---|---|---|
| ADSL | One per subject | Subject-level data | ADSL |
| BDS | One per subject-parameter-timepoint | Repeated measures | ADVS, ADLB, ADRS |
| OCCDS | One per occurrence | Events, meds, history | ADAE, ADCM, ADMH |
Defining OCCDS mistake: collapsing ADAE
One row per subject, or per body system, belongs in summary tables derived from ADAE.
Concept page (L1 fundamentals): establish the occurrence-data grain before any programming is shown.
Speaker notes
Now look at the Occurrence Data Structure, or OCCDS. In the Analysis Data Model, or ADaM, we work with three grains. The Subject-Level Analysis Dataset, or ADSL, holds one row per subject, while the Basic Data Structure, or BDS, holds one row per subject-parameter-timepoint. The Adverse Events Analysis Dataset, or ADAE, is an OCCDS dataset, so it carries exactly one row per adverse event; five events from one subject give five rows. Collapsing ADAE to one row per subject, or per body system, is the defining OCCDS mistake; those shapes belong in summary tables derived from ADAE, never inside ADAE itself.
Building ADAE: The AE-to-ADSL Merge
Building ADAE: The AE-to-ADSL Merge
data adae_1;
merge ae_sdtm(in=ae) adsl(keep=usubjid trt01sdt trt01edt trt01a saffl);
by usubjid;
run;
• Direction: AE-SDTM left, ADSL right, join BY USUBJID.
• ADSL supplies TRT01SDT, TRT01EDT, TRT01A, SAFFL onto every AE row.
• Parse dates after length check — partial date parsed as complete is a silent defect.
In AE but not on ADSL
• Data-management listing — never silently drop
In ADSL but no AE events
• Expected — contributes to denominators only
Concept page (L1 fundamentals): teach merge discipline, join direction, and what a mismatch means before showing the flag derivation.
Speaker notes
Building ADAE, the Adverse Events Analysis Dataset, means merging adverse event data from SDTM, the Study Data Tabulation Model, with ADSL, the Subject-Level Analysis Dataset. The AE domain sits on the left with many rows per subject, and ADSL sits on the right with one row per subject; they join BY the unique subject identifier, or USUBJID. In a Statistical Analysis System, or SAS, merge, ADSL supplies TRT01SDT, TRT01EDT, TRT01A, and SAFFL onto every AE row, so the first-dose date arrives through TRT01SDT. Then parse character dates to numeric with `aestdt = input(aestdtc, yymmdd10.);` but length-check the string first, because a partial date parsed as complete is a silent defect. A subject in AE but not in ADSL is a data-management listing, never rows to drop silently; subjects in ADSL with no events are normal and contribute only to denominators.
TRTEMFL: Anchor to First Dose, Never Randomization
TRTEMFL: Anchor to First Dose, Never Randomization
• TRTEMFL (treatment-emergent flag) asks: did the AE begin on or after first dose?
• SAP convention: ASTDT >= TRT01SDT; grace window only if the SAP defines one.
• Anchor: first dose — never randomization.
• Between randomization and first dose → NOT treatment-emergent (Scene 1 QC failure).
Canonical SAS logic — core flags
if aestdt >= trt01sdt
and (aeendt = . or aestdt <= trt01edt + 30)
then trtemfl = 'Y';
else trtemfl = 'N';
by usubjid descending trtemfl
aestdt aeseq;
if first.usubjid and trtemfl='Y'
then aoccfl='Y';
Concept page (L1 fundamentals): show the treatment-emergent comparison and connect it back to the opening QC failure.
Speaker notes
The treatment-emergent flag, TRTEMFL, answers a treatment question, not a randomization question: did the adverse event, or AE, begin on or after first dose? Under the Statistical Analysis Plan, or SAP, convention we compare ASTDT to TRT01SDT, and a grace window is added only when the SAP defines one explicitly. The anchor is first dose, never randomization, so an AE starting between randomization and first dose is not treatment-emergent, which is exactly the defect from Scene 1. In SAS, the canonical flag block reads: if aestdt >= trt01sdt and (aeendt = . or aestdt <= trt01edt + 30) then trtemfl = 'Y'; else trtemfl = 'N';. The first treatment-emergent event flag and the first preferred term occurrence flag both come from sort order: by usubjid descending trtemfl aestdt aeseq; then if first.usubjid and trtemfl='Y' then aoccfl='Y';.
Partial Dates: Transcribe the SAP, Don't Choose
Partial Dates: Transcribe the SAP,
Don't Choose
1
Impute per SAP. AESTDTC is ISO 8601 text; partial day/month → ASTDT; record in an AESTDTCF-style flag.
2
Compare after imputation. ASTDT vs TRT01SDT — raw partial text vs date = parse defect.
3
Transcribe, don't choose. Earliest-plausible vs complete date — follow what the SAP prescribes.
4
Dictionary terms ride through. AEDECOD/AEBODSYS unchanged from AE; one MedDRA version per data-cut.
Concept page (L1 fundamentals): cover imputation discipline, dictionary pass-through, and why these rules are never selected by the programmer.
Speaker notes
AESTDTC is text in the International Organization for Standardization 8601 format, so a partial day or month will not parse as a date. Before the treatment-emergence comparison, impute an analysis start date into ASTDT exactly as your Statistical Analysis Plan prescribes, and record that decision in an AESTDTCF-style flag. Impute first, then compare; comparing raw partial text to TRT01SDT is a parse defect. Plans legitimately differ: some accept the earliest plausible date, others require a complete date before calling an event treatment-emergent. Transcribe the rule you were given; do not choose one. Finally, the Medical Dictionary for Regulatory Activities terms, AEDECOD, AEBODSYS, and the rest of the hierarchy, ride through from Adverse Events untouched, with one dictionary version per data-cut snapshot.
Concept Check: OCCDS, Merge, and the TE Anchor
1 In ADAE, the adverse-events analysis dataset, the record grain is one row per adverse event occurrence. If one subject has five adverse event records in SDTM.AE, how many ADAE rows will that subject contribute before any additional filtering or merge with ADSL?
2 You build ADAE by merging the Subject-Level Analysis Dataset (ADSL) with an adverse-event input by USUBJID. A subject appears in the adverse-event input but has no matching ADSL row. Assuming ADAE must contain adverse-event records only for subjects represented in ADSL, which SAS `IN=` condition correctly prevents the unmatched adverse-event rows from entering ADAE?
3 Which statements correctly describe treatment-emergent flag derivation and partial-date discipline in ADAE? Select all that apply. (select all that apply, then Check)
Speaker notes
Welcome to this checkpoint on Occurrence Data Structure (OCCDS), merge direction, and the treatment-emergent anchor in Analysis Data Model (ADaM) adverse-events datasets. In the adverse-events analysis dataset (ADAE), the record grain is one row per adverse event occurrence, so if one subject has five adverse event (AE) records in Study Data Tabulation Model (SDTM) AE, how many ADAE rows will that subject contribute before filtering or merging with Subject-Level Analysis Dataset (ADSL)? The correct answer is B: five rows, because ADAE retains one row per adverse event occurrence, and Unique Subject Identifier (USUBJID) is not unique in ADAE. You build ADAE by merging ADSL with an adverse-event input by USUBJID, but a subject appears in the adverse-event input with no matching ADSL row; assuming ADAE must contain AE records only for subjects represented in ADSL, which Statistical Analysis System (SAS) IN= condition prevents the unmatched AE rows from entering ADAE? The correct answer is B: require both ADSL and the AE input with if inADSL and inAE;, because the unmatched AE rows have inADSL=0 and inAE=1, so they are rejected. Which statements correctly describe treatment-emergent flag (TRTEMFL) derivation and partial-date discipline in ADAE? The correct answers are A and D: for TRTEMFL, Analysis Start Date (ASTDT) is compared with Treatment Start Date (TRT01SDT), and an AE is treatment-emergent when the AE analysis start date is on or after the first treatment date, not randomization date, because a subject can be randomized before the first dose; and partial dates require a planned, documented imputation rule established by the study statistician or lead statistical programmer and recorded in the ADaM specification or metadata. Ad hoc per-record choices are not reproducible, and an AE beginning before TRT01SDT is not treatment-emergent.
Real-Data Walkthrough: Five Events, One Subject
Real-Data Walkthrough | USUBJID 043-18101-74001-001
Five AE Events, One Subject
| AESEQ | AE Term | AESTDTC | AEENDTC | ASTDY | TRTEMFL |
|---|---|---|---|---|---|
| 1 | ALT Increase | 2019-05-20 | — | 13 | Y |
| 2 | AST Increase | 2019-05-20 | — | 13 | Y |
| 3 | Blood Bilirubin Increase | 2019-05-17 | — | 10 | Y |
| 4 | Confusion | 2019-05-17 | — | 10 | Y |
| 5 | Delirium | 2019-05-17 | — | 10 | Y |
Result: TRTEMFL = Y for all five AE rows
AESEQ 1-2: start 2019-05-20, after TRTEDT 2019-05-19 — AEENDTC missing keeps the treatment window open.
Trace is one-to-one: 5 SDTM.AE rows → 5 ADAE rows.
Real-data walkthrough page: step through the actual extract rows for subject 043-18101-74001-001 and apply the canonical flag rule. All numbers are taken verbatim from the CSV extracts.
Speaker notes
Let's walk through the demo subject, who has five adverse event (AE) rows in the extract. The first two, AESEQ 1 and 2, start on 2019-05-20; the last three, AESEQ 3 through 5, start on 2019-05-17. AESEQ 3 through 5 began during treatment with ASTDY 10, while AESEQ 1 and 2 began in safety follow-up with ASTDY 13. All five have no AEENDTC, so the treatment emergent flag, TRTEMFL, is Y for every row, even though AESEQ 1 and 2 start after the treatment end date, TRTEDT, 2019-05-19. The trace is one-to-one: five Study Data Tabulation Model (SDTM) adverse event rows become five Adverse Events Analysis Dataset (ADAE) rows.
Hands-On: Repair the Flag, Not the Count
Hands-on interactive — if it does not load, open the paired article and try the exercise there.
Speaker notes
This scene is a hands-on exercise, so you will work through it on the website at jaimeyan.com/learn rather than just watching me do it. You will practice completing the Subject-Level Analysis Dataset (ADSL) KEEP list so the MERGE carries the right treatment dates and safety flag, checking that the BY variable is usubjid in both datasets, and then you will build the treatment emergent condition and apply it to the adverse event rows in the extract. Please try it right after this video, because repairing the flag yourself is what turns a count you hope is right into one you can defend.
Where the Build Lives: SCE, Metadata, Double Programming
Where the Build Lives
Statistical Computing Environment (SCE) · Metadata · Double Programming
Spec-Driven Build
Dictionary Pipeline
QC & Engines
SCE: spec owns flag names, emergence, analysis populations.
Program transcribes; define.xml from the same metadata.
MedDRA: versioned artifact; coding run recorded.
ADAE stamps one version per data-cut snapshot.
QC: independent double programming.
Defect listings and structural checks per build.
admiral / admiralpython — one spec, same logic.
SAS QC habit: compare any-AE vs treatment-emergent-AE counts
proc freq data=adae_1; tables trtemfl / missing; run;
L2 modern-workflow page: locate the flag build inside a modern statistical-computing environment without changing the derivation rules.
Speaker notes
In the Statistical Computing Environment (SCE), the Adverse Events Analysis Dataset (ADAE) build is spec-driven: the specification owns the treatment-emergent flag (TRTEMFL) name, the emergence convention, and the analysis populations, and define.xml comes from that same metadata. The program transcribes; it does not invent flag logic. MedDRA, the Medical Dictionary for Regulatory Activities, is a versioned artifact: the coding run is recorded, and the ADAE build stamps one version per data-cut snapshot. For quality control (QC), run independent double programming, attach the recurring defect listings to the review, and check structural conformance on every build. A Statistical Analysis System (SAS) habit: run proc freq data=adae_1; tables trtemfl / missing; run; then compare any-AE and treatment-emergent-AE subject counts. In R or Python shops, admiral and admiralpython expose the same occurrence logic — one spec stays the source of truth.
The Agentic Way: Agents Draft, the SAP Decides
The Agentic Way:
Agents Draft, the SAP Decides
Volatile L3 guidance
As of 2026-08-30
Clean one-pass AE-to-ADSL merge — the agent is good here.
Do not outsource: emergence convention, partial-date imputation.
Partial-date month-only to day-01 and a 1-day grace window look sane.
A clean log is not enough — only diff vs the actual SAP sentence exposes the invention.
Accountability: QC signer owns all flag logic, agent-drafted or not.
Verification habit: ask the agent to state the emergence rule it implemented and its source before anyone reads the code.
Hand-check three borderline subjects: pre-dose onset, month-only onset, onset dated the same day as first dose.
Volatile L3 page: show where agent assistance helps and where it invents rules; flag the as-of date for tool-specific claims.
Speaker notes
An agent can draft a clean one-pass adverse event to subject-level analysis dataset merge, and that part it does well. What you must not outsource is the emergence convention and the partial-date imputation rule: a one-day grace window, or imputing a month-only date to day-01, both look sane and leave a clean log. Only a diff against the actual sentence in the Statistical Analysis Plan exposes the invention, so ask the agent to state the emergence rule it implemented and its source before anyone reads the code. Then hand-check three borderline subjects: one pre-dose onset, one month-only onset, and one event dated the same day as first dose. The accountability boundary stays put, whoever signs the quality control document owns the flag logic. This level three guidance was last verified 2026-08-30, so re-verify before relying on agent-specific claims.
Final Check: Flag Logic, Defects, and QC Listings
1 Subject 043-18101-74001-001 in the real extract has five adverse-event records in the source data. Under the ADaM (Analysis Data Model) occurrence data structure (OCCDS) grain used to build ADAE (Adverse Events Analysis Dataset), how many ADAE rows must these five events produce for that subject?
2 During QC of the ADAE build on the real extract, which statements are correct? Select all that apply. (select all that apply, then Check)
3 When ADaM code has been drafted by an AI agent, who owns the treatment-emergence and partial-date decisions? What must the person who signs the QC document verify before the agent-drafted output is accepted? (reflect, then reveal)
Reveal analysis
Speaker notes
This checkpoint quizzes you on the Analysis Data Model (ADaM) occurrence data structure (OCCDS) grain, on quality control (QC) catches in the Adverse Events Analysis Dataset (ADAE), and on ownership of treatment-emergence decisions. For the demo subject with five adverse events, the ADaM OCCDS grain gives answer B: one row per adverse event, so five rows in the ADAE. In QC of the ADAE build, the correct choices are A, B, C, and D. A Confusion record with analysis start date (ASTDT) equal to 2019-05-17 and treatment start date (TRT01SDT) equal to 2019-05-08 should have treatment-emergent flag (TRTEMFL) set to 'Y'; a Y flag with ASTDT before TRT01SDT belongs in the date-comparison listing; a missing ASTDT when adverse event start date/time character (AESTDTC) exists belongs in the parse-failure listing; and a subject split between ADAE and Subject-Level Analysis Dataset (ADSL) belongs in an anti-join check. The treatment-emergence and partial-date decisions are owned by the approved Statistical Analysis Plan (SAP), not by artificial intelligence (AI) agent-drafted code, and the person who signs the QC document must verify that the SAP decision rules were faithfully transcribed into the ADAE derivation before accepting the output.
Takeaways and the Reviewer's Checklist
Takeaways & the Reviewer's Checklist
ADaM OCCDS · RULES TO KEEP
• One row per occurrence — ADAE row count tracks the events served.
• Emergence: compare ASTDT vs TRT01SDT per SAP — never randomization.
• Partial dates: impute per SAP, then flag — never a programmer choice.
• Pin one MedDRA version per snapshot; QC via independent double programming.
• 312-vs-315: unsupported flags are caught by date listings.
RECURRING ADAE DEFECTS · REVIEWER CHECKS
| Defect | Check |
|---|---|
| ADAE subjects not in ADSL | Anti-join |
| TRTEMFL=Y with ASTDT before TRT01SDT | Date comparison |
| ASTDT missing where AESTDTC exists | Parse failure |
| MedDRA version drift | Version trace |
| Duplicate rows after the merge | Key uniqueness |
• Run structural conformance checks on every build.
• Next: Part 9 ADTTE — time-to-event programming.
Summary page: tie the lesson back to the paired ADaM OCCDS article and the bootcamp arc.
Speaker notes
Takeaways: the Adverse Events Analysis Dataset, or ADAE, in the Analysis Data Model Occurrence Data Structure, or OCCDS, keeps one row per occurrence. Treatment emergence compares Analysis Start Date, or ASTDT, against Treatment Start Date, or TRT01SDT, per the Statistical Analysis Plan, or SAP — never randomization; partial dates are imputed per SAP and flagged. Reviewer checks: anti-join ADAE not in Subject-Level Analysis Dataset, or ADSL; date comparison Treatment Emergent Flag, or TRTEMFL, equals Y with ASTDT before TRT01SDT; parse failure missing ASTDT where start date exists; version trace Medical Dictionary for Regulatory Activities, or MedDRA, drift; key uniqueness duplicate rows post-MERGE. Pin one MedDRA version per snapshot; run Quality Control, or QC, as independent double programming. Remember the 312-vs-315 gap: unsupported flags are caught by date listings. Next is Part 9, the Time-to-Event Analysis Dataset, or ADTTE.