# Matched ECDF comparison

Author: Jaime Yan. Executed 2026-09-21 on Windows, Intel Core Ultra 5 125U (12 cores / 14 logical processors). No new dependencies.

Task: plot the same uncensored complete-pair Fatigue change distribution by arm, right-continuous F(x). The example has 34 complete pairs. The dense case has 34,000 deterministically repeated/shifted teaching observations (not additional independent clinical evidence). Input CSVs and scripts are retained; input creation does not download upstream demo data.

One warmup followed by five in-process repetitions. Python measures computation plus Agg canvas rendering; R measures computation/layout/null PDF drawing. File export and import/startup excluded. Devices differ: cross-language runtime rankings are not justified. Library timing includes its denominator annotations and styling; it is not a bare-engine microbenchmark.

| Engine | Installed version | Example median seconds | Dense median seconds | Numerical task |
|---|---|---|---|---|
| scipy | 1.17.1 | 0.1568 | 0.1565 | pass |
| statsmodels | 0.14.6 | 0.1312 | 0.1504 | pass |
| seaborn | 0.13.2 | 0.1614 | 0.1970 | pass |
| library | local extension / SciPy 1.17.1 | 0.1410 | 0.1677 | pass |
| ggplot2 | 4.0.3 | 0.4100 | 0.5300 | pass |

Raw timings, package metadata and installed-distribution byte counts are in [benchmark-python.json](benchmark-python.json) and [benchmark-r.json](benchmark-r.json). Installed-distribution size excludes shared/transitive dependencies and is not download size. Full dependency closure: **not measured**, because no new environment/install was authorized.

## API, export and access comparison

| Aspect | SciPy + Matplotlib | statsmodels + Matplotlib | Seaborn | ggplot2 | Library adapter |
|---|---|---|---|---|---|
| Core API | stats.ecdf + step | ECDF + step | ecdfplot, one grouped call possible | stat_ecdf | prepare + draw, explicit clinical input grid |
| Editable object | Figure | Figure | Axes / Figure | ggplot | Figure / ggplot |
| SVG/PDF/PNG execution | yes | yes | yes | yes | yes |
| Complete-pair denominator metadata | caller supplies | caller supplies | caller supplies | caller supplies | checked roster + complete pairs |
| Native capabilities beyond this case | right-censoring/CI | step conventions, model ecosystem | concise distributions/facets | scales/facets/grammar | three narrow clinical teaching contracts |
| Screen-reader table / mobile interaction | not supplied in this script | not supplied in this script | not supplied in this script | not supplied in this script | local Astro page, separately tested |

A bare plotting engine is not a website. No numeric accessibility or interaction scores are assigned to these incomparable scopes. Export inspection compares actual figures at the same 11?7 inch size; it does not intentionally create poor baselines. No clipping-reduction percentage is claimed. Dense lines remain readable as cumulative distributions, but cannot reveal individual identity; the browser table is optimized for the small teaching case, not a 34,000-row interactive subject explorer.

## Known disadvantages and decision

Seaborn is more concise for generic grouped ECDFs. ggplot2 has a much broader compositional grammar and robust maintained tests. SciPy supports censored distributions/intervals that this contract intentionally does not exercise. The library adds clinical roster validation, explicit missing-pair accounting, dual-language agreement and learning material, at the cost of more input preparation and a narrower API. Those benefits are demonstrated by contract/fixture tests, not by a visual novelty score.

The library benchmark uses an already complete-pair input and sets the benchmark roster equal to that input for comparability. The actual teaching templates use all 36 roster subjects and report the two missing Fatigue pairs. Timing the full validation/three-stratum pipeline against a single bare ECDF would answer a different question, so no such speed claim is made.

Full initial-package installation size, long-label localization stress, cross-browser certification, formal screen-reader testing, user-task completion and external review remain **not-run**. This is a limited reproducible task comparison, not evidence of best-in-field performance.

## Independent review clarification (2026-09-21)

The recorded Python timings include expected-value sorting and numerical validation inside the timed loop. R performs its numerical validation outside the timed block. These recorded measurements therefore have different validation boundaries in addition to their different rendering devices. They are execution-script timings, not pure computation/render benchmarks; no R/Python speed ranking is warranted. The retained results are not retroactively relabeled as timing-equivalent. A future performance comparison must align validation boundaries before collecting fresh measurements.
