Clinical programming / an inspectable draft

admiralagent.

From one specification to code with evidence for every line.

Turn ADaM derivation descriptions into a reviewable Layer IR, then generate admiral R drafts from deterministic templates. See the reasoning, the validation failures, and the decisions that must be made by a human.

01
SpecificationSource text and derivation intent preserved verbatim
02
Layer IRControlled vocabulary · explicit parameters · human review
03
admiral codeDeterministic draft + CHECK + provenance tracking
14registered IR layers
49real ADSL spec variables
Rulesprecomputed · zero online inference

The working notebook

See the results, and the judgment too.

Download data snapshot ↗

Follow one date specification through rule classification, IR validation, code generation, and comparison against a public oracle.

ORACLE 100% · non-missing dates254 / 254 pairs non-missing on both sides match by date. Across the full 306-record basis the rate is 83.0%; missing-status agreement is 100%. Source: pharmaverseadam::adsl; does not prove time-of-day agreement.
ADSL / TRTSDTM PASS
Click the source text, a layer card, or a code line to trace the same piece of evidence
01 / SOURCE

Specification source

Why this classification

rule: impute DM RFSTDTC directly (pilot ADSL convention)

Source text preserved verbatim. Linking shows provenance; it does not mean CHECK has proven the full semantics of the spec.

02 / INTERMEDIATE REPRESENTATION

Reviewable layer pipeline

Rule confidence 1.00
Rule-match metadata, not a probability of correctness

View full IR JSON
{
  "dataset": "ADSL",
  "variable": "TRTSDTM",
  "steps": [
    {
      "layer": "impute_dtc",
      "args": {
        "target": "TRTSDTM",
        "dtc": "RFSTDTC",
        "output_class": "dtm",
        "highest_imputation": "M",
        "date_imputation": "first"
      }
    }
  ],
  "spec_origin": "Date of first study treatment, imputed from DM RFSTDTC",
  "confidence": 1,
  "needs_human": false,
  "rationale": "rule: impute DM RFSTDTC directly (pilot ADSL convention)"
}
03 / DETERMINISTIC OUTPUT

admiral · R

DISCLAIMER · Draft code — must be reviewed by qualified personnel before use.
Select a CHECK to see its code line number, spec sentence, and owning IR layer.
Post-execution validation gatesPASS imputation_flagPASS not_all_na

Loading interactive snapshot…

Architecture / bounded by design

Rules generate; humans judge.

Redrawn from the DESIGN.md pipeline. The public demo uses rules; the engine's optional LLM backend also only produces IR, constrained by the same validation gates and templates.

01Read the spec

P21 / define.xml
Deterministic parsing, source text preserved

02Classify into Layer IR

rules → closed vocabulary
Inexpressible → needs_human

03validate_ir()

Shape, parameter, and semantic checks
Invalid IR is refused rendering

04Deterministic compilation

admiral R + CHECK
DISCLAIMER + JSON sidecar

05Validation & human revision

run_validation() + Oracle
Reviewer edits IR → regenerate

The 14 layers actually registered in aa_layers()

assignmerge_varlookup_joinimpute_dtcdtm_to_dtdurationdate_shiftcompute_paramsummary_recordextreme_flagcodelist_varobs_numbercategorizecompute_var

Evidence / no hidden failures

Real specs, the complete matrix.

Download raw CSV ↗

Of 49 ADSL variables, 16 have comparable results and 14 reach 100%. REVIEW, ERROR, and non-comparable items are shown as well.

pilot1 define.xml → rules → public SDTM → pilot1 adsl.xpt. Joined by USUBJID; click a header to sort.
Comparison basis / notes
STUDYIDEXECUTED100.0%100.0%254char exact
USUBJIDEXECUTED83.0%100.0%306set overlap (join key)
SUBJIDEXECUTED100.0%100.0%254char exact
SITEIDEXECUTED100.0%100.0%254char exact
SITEGR1REVIEW——0not in generated ADSL
ARMEXECUTED100.0%100.0%254char exact
TRT01PEXECUTED100.0%100.0%254char exact
TRT01PNREVIEW——0not in generated ADSL
TRT01AEXECUTED100.0%100.0%254char exact
TRT01ANREVIEW——0not in generated ADSL
TRTSDTREVIEW——0not in generated ADSL
TRTEDTREVIEW——0not in generated ADSL
TRTDURDREVIEW——0not in generated ADSL
AVGDDREVIEW——0not in generated ADSL
CUMDOSEREVIEW——0not in generated ADSL
AGEEXECUTED100.0%100.0%254num tol 0.5 | mean|diff|=0.000
AGEGR1REVIEW——0not in generated ADSL
AGEGR1NREVIEW——0not in generated ADSL
AGEUEXECUTED100.0%100.0%254char exact
RACEEXECUTED100.0%100.0%254char exact
RACENREVIEW——0not in generated ADSL
SEXEXECUTED100.0%100.0%254char exact
ETHNICEXECUTED100.0%100.0%254char exact
SAFFLREVIEW——0not in generated ADSL
ITTFLREVIEW——0not in generated ADSL
EFFFLREVIEW——0not in generated ADSL
COMP8FLREVIEW——0not in generated ADSL
COMP16FLREVIEW——0not in generated ADSL
COMP24FLREVIEW——0not in generated ADSL
DISCONFLREVIEW——0not in generated ADSL
DSRAEFLREVIEW——0not in generated ADSL
DTHFLEXECUTED1.2%1.2%254char exact
BMIBLREVIEW——0not in generated ADSL
BMIBLGR1ERROR——0not in generated ADSL
HEIGHTBLREVIEW——0not in generated ADSL
WEIGHTBLREVIEW——0not in generated ADSL
EDUCLVLREVIEW——0not in generated ADSL
DISONSDTREVIEW——0not in generated ADSL
DURDISREVIEW——0not in generated ADSL
DURDSGR1REVIEW——0not in generated ADSL
VISIT1DTREVIEW——0not in generated ADSL
RFSTDTCEXECUTED100.0%100.0%254char exact
RFENDTCEXECUTED100.0%100.0%254char exact
VISNUMENREVIEW——0not in generated ADSL
RFENDTEXECUTED100.0%100.0%254num tol 0.5 | mean|diff|=0.000
DCDECODREVIEW——0not in generated ADSL
EOSSTTREVIEW——0not in generated ADSL
DCSREASREVIEW——0not in generated ADSL
MMSETOTREVIEW——0not in generated ADSL

Currently in original CSV order.

Numeric variables use a tolerance of |difference| < 0.5; date variables are compared via as.Date; character values must match exactly. Pairs missing on both sides do not count toward the numeric match rate; missing-status agreement is listed separately. USUBJID is a set-overlap rate. Execution status and match rate are shown separately: values already present in the base data may still be comparable, which does not establish a successful derivation. This matrix is not a CDISC compliance certification.

A few important distinctions

About this draft generator.

Stating the boundaries clearly is part of reviewability.

Is this "AI automatically writing submission code"?

No. It is a draft generator + human-in-the-loop. Every output must be reviewed line by line by qualified personnel, independently QC'd, and validated at the study level. A passing CHECK only proves that the implemented checks passed — it does not mean the spec is complete, clinically correct, or submission-ready.

Does this page call an LLM or execute R?

No. All cases are precomputed once locally in R by the rules backend. The browser only loads static JSON and links the views; it receives no study data and needs no API key.

For a version that actually executes online: Try the platform live submits real jobs to the tfl-platform API (synthetic / CDISC pilot data, 5 anonymous runs per day).

What is the difference between confidence, PASS, and Oracle 100%?

confidence is classifier metadata, not a calibrated probability of correctness. PASS is the result of a specific validation check. The Oracle match rate depends on the sample, missing values, and tolerance basis. TRTSDTM's 100% refers only to the 254 pairs non-missing on both sides; across all 306 records the basis is 83.0%.

Why does BMIBL still need human confirmation after revision?

Changing BASELINE to SCREENING 1 provides HEIGHT and WEIGHT, letting the not-all-missing and uniqueness checks pass. Whether that is an appropriate study baseline still requires confirmation. Its pilot1 Oracle match rate is 85.8%, not 100%.

Where are the public data and source code?

The examples use CDISC pilot public data from pharmaversesdtm / pharmaverseadam, plus local pilot1 public submission specs and oracle. The site publishes only specs, code, IR, and aggregate metrics — no subject records.

GitHub repository (placeholder, link pending public release) · Data and provenance snapshot