# Scheduled observation matrix: contract and adaptation note

Author: Jaime Yan. Original implementation and teaching text use the project personal noncommercial attribution license. External Matplotlib/ggplot2 rights remain independent; no upstream implementation was copied.

## Clinical question and descriptive target

Which scheduled assessments are observed, and are gaps intermittent or at the end of recorded follow-up? This is a descriptive census of the selected fictional roster, not a treatment effect or a missing-data model. It supplements the completion curve by identifying individual observation histories.

The source is the existing independently synthetic 36-subject teaching generator (`scripts/generate_data.py`, seed 20260921); the actual existing input files and hashes are recorded by the execution manifests. No new endpoint or patient data were introduced. Consult the original data contract for deterministic missingness generation. The reused API accepts subjects, domains and tumor frames through the shared validated input contract; domain outcomes are validated but not used in the matrix.

## Fixed specification

- Include every roster subject, assigned Reference/Investigational arm, and scheduled weeks 0, 4, 8, 12. Baseline is week 0. All/F/M strata are independently calculated and overlap.
- Require a unique, complete subject-by-scheduled-visit grid, known arms/sex, positive finite nonmissing diameter in mm, and explicit missing values. An absent scheduled row is an input error, not an inferred missing observation. Boolean measurements are rejected.
- Observed means a nonmissing diameter; Missing means an explicit missing value. No imputation, withdrawal assignment, censoring adjustment, variable follow-up or eligibility logic is performed.
- For each selected arm/visit, n is observed subjects, expected is the full roster, missing is expected minus n. These counts repeat across subject rows and must be deduplicated before aggregation. Empty strata return zero rows and a native plot stating no participants, not invented records or estimates.
- Complete: every visit observed. All missing: no visit observed. Trailing missing: a nonempty observed prefix followed only by missing visits. Intermittent or early missing: all other incomplete histories, including early missingness and histories with both gaps and a trailing run.
- Subjects sort by identifier within arm, without outcome-based clustering. O/X plus fill distinguishes status. Subject and visit axes do not imply pairing across treatment arms.
- No intervals, hypothesis tests or multiplicity claims apply. Counts and classifications match exactly across R/Python; no rounding is applied. Displayed numerical percentages belong to the separate completion template.

## Adaptation from missing-data matrices

Missingness matrices from general data analysis expose co-occurring gaps. In this clinical adaptation, columns are fixed scheduled visits and rows are roster subjects within assigned arms. The temporal order and complete schedule must be preserved; clustering or dropping unobserved rows would change the interpretation. Naniar's missingness visualization paper and inspected upstream project informed the design only; Naniar is not installed/integrated or required. See catalog and paper-provenance records for original sources and license evidence.

A matrix cannot distinguish administrative noneligibility, death, withdrawal and missed assessment using this input. For staggered enrollment or event-dependent visits, add separately defined eligibility states and revisit the denominator before reuse. The current fixture is not appropriate unchanged for those designs. MCAR/MAR/MNAR cannot be diagnosed from its visual pattern.

## Read and reproduce

S01 has a Week-8 gap with a later observation: intermittent. S36 has only baseline observed: trailing, without evidence of the reason. In All participants, each arm has 15/18 observed at Week 8. Reconcile this by counting O marks, then by deduplicating the arm/visit counts in the table, then against the completion curve.

Run `python scripts/render_visit_matrix.py --rscript <existing-Rscript-path>`. It executes separate R and Python raw-data calculations and exports SVG/PDF/PNG/CSV/JSON for each stratum, plus independent known-answer fixtures and `qc-visit-matrix.json`. The functions `prepare_matrix` / `draw_matrix` and `prepare_visit_matrix` / `draw_visit_matrix` return result frames and native editable plot objects.

For authorized private data, map variables and units locally and revise the contract. Publishing subject-level histories needs a separate privacy decision; this demonstration only contains already-public fictional identifiers. Cite Jaime Yan and the actual upstream versions; design-reference attribution does not claim upstream integration.
