A tiered survey of LLM evidence in clinical trial statistical programming: benchmarked results, promising single-team studies, vendor hype, and open gaps for 2026–2027.
Free-form agent loops break down in GxP clinical programming. Structuring the work as a typed process DAG makes LLM agents reliable, replayable, and auditable.
A bridge map, typed IR, and orchestrator wrap a legacy SAS TFL library unchanged — AI-ready JSON on day one, 80%+ cell-level parity, optional 92% code cut.
Schema-only synthetic ADaM generation plateaus at 0.45 overall quality; enriching schemas from protocol/SAP/CRF knowledge graphs plus templates reaches 0.70.