# Baseline-adjusted domain contrasts

Author: Jaime Yan. Original teaching adapter, example and documentation use the project personal noncommercial attribution license. Statsmodels and R retain their own licenses. This is not a causal analysis of a real trial.

## Target and population

For each invented domain and All/F/M stratum, estimate the coefficient of Investigational (Reference=0, Investigational=1) in an ordinary least-squares regression of Week-12 score on intercept, treatment and centered baseline. The target is a fitted common-slope between-arm contrast in the domain-specific complete-case sample. It is not an intention-to-treat estimand or an estimate protected against selective missingness. The existing synthetic generator deliberately creates arm differences; the apparent effects are teaching artifacts.

The existing 36-person fictional subjects/domains/tumor input, seed 20260921, is reused without new endpoints. All roster subjects are eligible; exclude a subject from a particular domain model when baseline or Week 12 is missing. Report roster N, complete-case n, missing n and arm-specific complete-case counts. Baseline=scheduled Week 0, outcome=scheduled Week 12, score unit=points_0_100. No visit windows, switching, censoring, imputation or intercurrent-event strategy is modeled. Validate the shared complete input grids first; Boolean outcomes are invalid.

## Calculation contract

Fit `Y12 = beta0 + betaT * treated + betaB * (baseline - complete_case_mean_baseline) + error`. Centering does not change betaT. Return betaT, classical homoscedastic SE, residual df n-3 and the two-sided pointwise 95% Student t interval. In Python call existing statsmodels 0.14.6 OLS with an explicit intercept and use_t=True; in R independently call stats::lm and stats::confint. No precomputed cross-language model results are shared.

Require at least two complete cases in each arm, n>3, full column rank and spectral condition number <=1e6. Otherwise return null estimates/intervals with an explicit unavailable reason, retaining actual counts. A constant baseline is rejected for this prespecified three-column design rather than silently fitting a different model. No p-values or multiplicity adjustment are supplied. Five domains and overlapping strata are exploratory. No intermediate rounding; CSV/JSON retain analysis precision, UI uses three decimals; count/label agreement exact and numerical tolerance abs/rel 1e-9. Intervals are not clipped to the observed scale.

Independent errors, a correctly specified linear mean with common baseline slope, and homoscedastic approximately normal residuals support the small-sample t interval interpretation. The current example does not diagnose or prove these assumptions. Treatment assignment and selection can invalidate causal interpretation. Use trial-specific design, covariates, interaction and missing-data sensitivity specifications before inference on authorized clinical data.

## Why this complements the domain-means plot

The descriptive plot asks how each observed arm changed. This plot asks for an explicit baseline-adjusted between-arm coefficient. Two separate arm intervals cannot simply be subtracted to obtain the model's interval. A coefficient near zero does not prove equivalence, and a pointwise interval excluding zero does not establish multiplicity-controlled efficacy or clinical importance.

The orthogonal known-answer fixture has eight participants, four per arm; baseline values 47,49,51,53 in each arm; errors 1,-1,-1,1 in each arm; follow-up=20+0.5*baseline-4*treated+error. Orthogonality gives betaT=-4, SSE=8, df=5, variance=8/5, SE=sqrt(0.8). Both languages independently verify it, then test empty and singular populations. The fixture is analytically constructed teaching data, not an observed trial result.

## Run and adapt

Run `python scripts/render_adjusted_domains.py --rscript <existing-Rscript-path>`. Functions `prepare_adjusted` / `draw_adjusted` and `prepare_adjusted_domains` / `draw_adjusted_domains` return result frames and editable native figures. Exports include all strata and full overlapping-strata tables, in SVG/PDF/PNG/CSV/JSON. See qc-adjusted-domains.json for actual execution. The website switches exact precomputed models; it does not hide subjects while retaining an overall estimate.

The adapter reuses mature implementations instead of coding regression or covariance from scratch. Sources: [statsmodels OLS documentation](https://www.statsmodels.org/stable/generated/statsmodels.regression.linear_model.OLS.html), [R lm documentation](https://stat.ethz.ch/R-manual/R-devel/library/stats/html/lm.html), installed statsmodels source/license hashes in the catalog. Online documentation can describe newer versions; the executed version is recorded separately. No upstream source code was copied. Cite Jaime Yan plus actual R/statsmodels versions and preserve upstream notices.

Missing-data sensitivity, robust covariance alternatives, diagnostic plots and confirmatory multiplicity control remain future work. Do not relabel this complete-case teaching analysis as any of those methods.
