# Extension analysis contracts

Original author: Jaime Yan. These templates are descriptive teaching tools, not clinical efficacy claims. Their common input is the existing, independently generated synthetic dataset in `data/`, seed 20260921. The original generation script is unchanged. Input bytes and generator hashes are captured by `scripts/verify_extensions.py`. No third-party data are added.

## Shared input contract

`subjects`: unique nonmissing subject, arm exactly Reference or Investigational, sex F/M. `domains`: unique subject/domain/week; all five domains Fatigue, Pain, Sleep, Appetite, Mobility at weeks 0 and 12 for every subject; numeric score in [0,100] or missing. `tumor`: unique subject/week; weeks 0,4,8,12 for every subject; numeric diameter in mm >0 or missing. Explicit units argument must be `score_unit="points_0_100"`, `diameter_unit="mm"`. Unknown subjects, duplicate/missing keys, infinities, incomplete schedules and empty rosters fail. No implicit deduplication, missing-row filling or unit conversion.

Reference and Investigational denote the assigned fictional arm. Baseline is scheduled week 0, with no visit-window derivation. Separate outputs for All/F/M are independently computed from the raw roster in each language. Filtering by sex changes the population and all its denominators. Empty strata produce explicit n=0 / null estimates, not zero outcomes. Missing observations are never imputed. Structural absence of a scheduled row is an input error, distinct from a retained row with missing value.

## ECDF: distribution of Week-12 Fatigue change

Target: empirical distribution of week-12 minus baseline Fatigue score among complete pairs in the selected arm/stratum. Include every roster subject with both observations; exclude incomplete pairs and report their count. No adjustment, causal treatment-policy estimand, time-to-event censoring or confidence band. F(x)=count(change≤x)/complete pairs, using `scipy.stats.ecdf` and `stats::ecdf` independently. Ties create one jump. Values range [-100,100] points; negative change means less burden. Empty groups have a single explicit unavailable row. No responder threshold is asserted. Compare shapes and quantiles, not the area as a treatment effect. Selective missingness can reverse the apparent arm comparison. Prefer adjusted effects or missing-data sensitivity analysis when inference is required.

## Completion: scheduled diameter observations

Target: observed measurement count / full selected arm roster at each scheduled week, with missing count and percentage. Every subject is scheduled for all four visits in this synthetic design. No exclusions, imputation, censoring adjustment or confidence interval: this is a census of the example roster. Baseline can be missing for this availability analysis, though other original templates require positive observed baseline. Lines connect scheduled aggregate proportions, not individuals. A missed measurement is not a withdrawal. Do not apply this denominator to staggered enrollment, visits not yet due, death, or changing eligibility without explicit state/eligibility variables. Completeness is not a test of MCAR/MAR and cannot correct missing-data bias.

## Domain intervals: mean within-person change

Target: arm-specific arithmetic mean week-12 minus baseline score for each of five invented domains among complete pairs. Each domain has its own n and missing count; exclude incomplete pairs without imputation. For n≥2, use sample SD (n−1), SE=SD/sqrt(n), two-sided pointwise 95% Student t interval, df=n−1 (`scipy.stats.t` / `stats::qt`). For n=1 show the mean but no interval; n=0 mean/SD/limits are null. Zero variance for n≥2 gives a zero-width interval, with no implication of certainty in a broader population. No interval clipping to scale bounds. Independent subjects and an approximately normal sampling distribution are required for interpreting coverage. Five domains × two arms × three overlapping strata are exploratory; there is no multiplicity adjustment. Overlap/non-overlap is not a between-arm significance test. Use an adjusted treatment contrast and appropriate multiplicity strategy for confirmatory work.

## Precision, display and reproducibility

No intermediate rounding. CSV uses at least 15 significant digits; JSON numeric export uses 15–16 digits. Comparison is abs(a−b)≤1e−9+1e−9|b|, exact integers and categorical keys, exact missingness masks. Website numeric display rounds to three decimal places; image labels show counts exactly and percentages to one decimal. Display rounding never feeds analysis.

All three return editable Matplotlib Figure / ggplot objects. Export 11×7 inch SVG/PDF and 1760×1120 PNG. The browser switches between exact precomputed strata and languages; it never recalculates an estimator. Image, summary, table and current-stratum CSV/JSON use the same selection. Full outputs are separately labeled. No upload, Run button or calculation backend.

## Missingness generation and limits

The existing generator deterministically omits all postbaseline diameter values for S36, week-8 values for subject indices divisible by eight, and selected week-12 domain scores by `(i+3*j) % 19 == 0`. This illustrative, index-dependent mechanism is not an empirically justified MCAR/MAR/MNAR model. Arm means are deliberately generated differently for teaching. These are invented scales, not validated PRO instruments. Synthetic findings must never be reported as clinical results.
