# Evidence standard and selection protocol

Author: Jaime Yan. Original research notes use the repository license.

Review date: 2026-09-21. This is a bounded, purposive review, not an exhaustive survey or systematic review. Seed projects come from the task, official ecosystem directories, and related methods. Search official repositories, reference manuals, runnable examples and original methodological sources. Search summaries are discovery aids only. Record failed retrievals and unknowns. Popularity is not a selection criterion.

## Acceptance gates

| Dimension | Observable evidence required |
|---|---|
| Clinical value | A specific question, decision boundary, alternative and added information |
| Analysis | Population, estimand or descriptive target, time, units, keys and exclusions |
| Uncertainty | Named interval method, denominator, follow-up/censoring rules, multiplicity |
| Legibility | Inspect actual exports at final size; no hidden intervals or clipped labels |
| Interaction | State whether display, precomputed stratum or recalculation; test downloads |
| Access | Text alternative, semantic table, keyboard focus, 390/768/1440 px checks |
| API | Native editable object, explicit contract, actionable invalid-input failures |
| Reproduction | Executed source, versioned input, environment and SHA-256 linkage |
| Numerical QC | Independent R/Python calculations plus known answers, not agreement alone |
| Teaching | Why this plot, reading uncertainty, misuse, adaptation and citation |

Use pass / fail / not-run / not-applicable with evidence, never a synthetic overall score. A failed licensing, data integrity or estimator gate blocks adoption. Browser acceptance and statistical agreement are separate from package-risk assessment. No SOTA or regulatory-validation claim.

## Priority matrix (defined before new implementation)

| Clinical question | Structure / required variables | Candidate | Existing coverage / decision |
|---|---|---|---|
| Is improvement broadly distributed? | Subject, arm, baseline and week-12 symptom score | ECDF of complete-pair change | New first vertical slice; reveals distribution beyond a mean |
| Who contributed to longitudinal summaries? | Full scheduled subject × week grid, arm, observed flag | Completion curve and subject missingness matrix | New; protects interpretation of spider/waterfall |
| Which symptom domains change, with what precision? | Subject × domain × time, arm, score | Aligned mean-change intervals | New; supplements radar with linear axes and pointwise uncertainty |
| Baseline balance | Randomized arm, baseline covariates | Standardized differences / distributions | Existing Lab; defer until population contract reviewed |
| Effect modification | Outcome, treatment, baseline covariates | Interaction model + forest | Existing descriptive forest is not interaction evidence; defer |
| AE risk under unequal exposure | Subject event, exposure, censoring, competing event | Risk difference / incidence rate / CIF | Existing synthetic data cannot support; defer, no invented endpoints |
| Laboratory extremes | Analyte, units, ULN/LLN, time | Reference-range trajectories | Existing ALT shift only; defer until longitudinal lab input exists |
| Event-free time gained | Time, event code, horizon, censoring | RMST difference | No suitable synthetic event input; defer |
| Treatment pathways | State entry/exit, competing transitions | Multistate probabilities | Swimmer follow-up is not a transition process; defer |
| Dose and exposure response | Dose, concentration, covariates, outcome | Adjusted exposure-response intervals | No corresponding input; defer |
| Measurement agreement | Paired methods on same specimen/time, units | Bland-Altman | Baseline vs follow-up is not method agreement; defer |
| Model calibration / stability | Predicted risks, outcomes, censoring, resamples | Calibration and sensitivity plots | No predictions; defer |

## Reuse decision

First use installed mature statistical and drawing APIs. Add a small input/analysis/export adapter where clinical semantics are missing. Do not copy source or redistribute example images/data without file-specific permission. Use independent languages to test the contract; do not make a second copy of one precomputed table. New original formulas are limited to standard t-interval assembly and denominator bookkeeping, where additional frameworks would add dependencies without changing the estimator.

## Cross-domain adaptation notes

**Distribution diagnostics → PRO change.** ECDFs in forecasting/measurement show cumulative mass without bin or kernel selection. Here x is week-12 minus baseline on an invented 0–100 symptom scale; y is the proportion of complete pairs with change ≤ x. It is not response probability for the randomized population, nor Kaplan–Meier survival. Missing pairs cannot be treated as censored change values. Benefits over means/rainclouds: exact ties, no bandwidth, all thresholds visible. Cost: tails are sparse and no treatment-effect confidence band is supplied.

**Quality monitoring / temporal heatmaps → visit completeness.** A scheduled-grid matrix exposes temporal gaps; the companion curve uses the full arm roster as denominator. An absent measurement is not proof of withdrawal or an event. All participants here are scheduled at every visit; staggered recruitment, death and not-yet-due visits require distinct eligibility states and a revised denominator. Unlike clustered genomic heatmaps, time must remain ordered and subjects are not clustered by outcomes. Benefit: makes the changing observed-case population inspectable.

**Scientific small multiples → symptom domains.** Linear aligned estimates replace polygon-area comparisons. Within-participant changes preserve pairing; separate arms are not connected as paired subjects. Intervals describe each observed-pair mean, not between-arm treatment effects. Five invented scales share direction and range but are not validated PRO instruments. Pointwise t intervals assume independent subjects and approximately normal sampling of means; no multiplicity adjustment or causal estimand. Use a model-based adjusted difference for confirmatory inference.

## Coverage update

| Clinical question | Template | Required input | Executed data | Boundary |
|---|---|---|---|---|
| Which subjects have intermittent versus trailing gaps? | visit-matrix | Roster, assigned arm, sex, full scheduled grid, observed/missing diameter in mm | Existing synthetic 36-person input | Missingness pattern is not mechanism, withdrawal or censoring |
| What is the complete-case baseline-adjusted between-arm coefficient? | adjusted-domains | Roster, domain, baseline and Week-12 scores | Existing invented 0-100 domains | Common-slope OLS; pointwise intervals, no missing-data correction or multiplicity control |
