# Figure contracts, version 0.1

All examples are independent synthetic teaching data. Treatment names are fictional. The random seed is 20260921, but rendered data are the versioned CSV bytes, not a promise of cross-language random number equivalence. Both languages independently analyze those CSVs. No inferential clinical conclusion is justified.

## Inputs

- `subjects.csv`: one row per subject; `subject`, `arm` (Reference / Investigational), `sex` (F / M), `age` (years), `followup` (weeks), `ongoing` (0/1), `milestone` (week of an illustrative assessment), `alt_baseline`, `alt_week12` (multiples of ULN). Missing lab values are empty cells.
- `tumor.csv`: one row per subject and scheduled `week` (0,4,8,12), `diameter` (mm). Positive baseline required; missing follow-up is empty. Rows retain scheduled missing visits.
- `domains.csv`: one row per subject / `domain` (Fatigue, Pain, Sleep, Appetite, Mobility) / `week` (0,12), `score` (0-100; higher means worse). These are invented comparable domain scales, not validated instruments.

Invalid duplicate keys, unknown subjects, inconsistent schedules, nonpositive baseline diameter, out-of-range observed domain scores and impossible swimmer dates raise errors. The example analysis requires at least two observations per forest arm/subgroup and fails explicitly otherwise.

## Templates

| Template | Question and estimator | Missingness / denominator | Interpretation |
|---|---|---|---|
| Waterfall | Minimum observed postbaseline percentage diameter change, `100*(value/baseline-1)`, ranked ascending with subject tie-break | Exclude subjects with no observed follow-up; retain subject ID and number of assessments | -30% / +20% are visual guides only; not RECIST classification, confirmed response, target-lesion sum or ORR |
| Forest | Week-12 Fatigue change, Investigational minus Reference, overall and sex/age subgroups | Complete baseline/week-12 pairs; no imputation; Welch t 95% CI with Satterthwaite df | Lower is favorable on this invented score; exploratory overlapping subgroups, no interaction test or multiplicity correction |
| Radar | Week-12 arithmetic mean by arm and domain on a fixed 0-100 scale | Observed domain scores; each spoke has its own n | Higher is worse everywhere; no axis reversal, clipping, normalization to observed extrema or polygon-area interpretation; table is the quantitative comparison |
| Swimmer | Observed follow-up duration per subject, sorted within arm; assessment marker and ongoing arrow | All subjects, no survival estimator | Assessment marker is not response; arrow denotes follow-up ongoing at cutoff, not extrapolation |
| Spider | Individual percent diameter change over scheduled weeks | All scheduled rows retained; missing visits break lines, no interpolation; per-visit n in table | Lines are individual trajectories, not a mean or model; same fictional tumor data as waterfall |
| Shift | Baseline vs week-12 ALT/ULN categories: L <0.5, N 0.5-1 inclusive, H >1 | Complete pairs per arm, nine cells including zero; percentages use paired arm n | Educational thresholds only, not universal lab reference ranges; row is baseline, column is week12 |

All exported figures show the full example population. Switching R/Python changes renderer, not the population. The website's table search affects table rows only and is explicitly labeled; it never changes chart statistics. CSV downloads contain the full computed figure table. R/Python comparisons use unrounded numeric values with absolute and relative tolerance 1e-9. Counts and keys match exactly. Display rounding does not feed analysis.

## Design

Stable Okabe-Ito blue (#0072B2) and vermillion (#D55E00), plus distinct line styles / shapes, warm white background, light grid, restrained typography. Figures use installed sans-serif fonts for portable exports; the website uses its own local Inter / Fraunces. Vector exports can be enlarged independently of the page. The accessible result table and written interpretation remain available without hover or color perception.
