Smart agent, empty hands
The agent can reason about clinical programming — it just can't touch the artifacts.
- Ask a coding agent to review a SAS log and it does fine.
- Ask it to read a SAS7BDAT dataset or parse an ADaM spec spreadsheet — it can't open the files.
- And it can't reliably tell a real ERROR from a harmless “Unable to copy SASUSER” warning.
- The gap is in the tools, not the reasoning.
- ClinAgent's answer in one line: thin MCP, thick skills.
- the agent (your choice)
- artifact the agent can't touch directly
- the tool gap
Teaching schematic — not measured. Tool gap summarized from the post's opening; the log lines are illustrative.
Source: Thin MCP, Thick Skills (introduction).
Five layers, each with one job
The stack under any MCP-compatible agent — the design bet is where the thickness sits.
- Layer 1 renders results: dashboards, log tables, RTF viewers — deterministic output for human review.
- Layer 2 routes each tool call to the right skill and validates inputs first.
- Layer 3 is thick: nine skill packages, SK-001 to SK-009, each bundling prompts, examples, and rule engines.
- Layer 4 is thin: stateless I/O only — read a dataset, parse a spec, read a log. It never interprets.
- Layer 5 is compliance: audit logging, PHI masking, access control, context minimization.
- The agent reasons; the stack supplies domain expertise and the compliance trail. Swap the agent — the skills still work.
- the design bet: thick skills, thin tools
- supporting layers
Teaching schematic — not measured. Layer structure follows the post's Figure 1 and the preprint's Figure 1.
Sources: the post (the five layers) · ClinAgent preprint, Fig 1.
Anatomy of a single tool call
“Classify the findings in this SAS log” — deterministic classification first, stochastic interpretation last.
- A call arrives: classify the findings in this SAS log.
- The router picks skill SK-005 and checks the input first.
- A thin MCP tool reads the log bytes. It does not know what an ERROR means.
- A deterministic rule engine tags every line — errors, warnings, and known false positives like “Unable to copy SASUSER”.
- Only then does the LLM write the human-readable interpretation of what the rules already decided.
- The infrastructure layer logs the whole call — timestamped inputs and outputs.
- deterministic decision
- stochastic interpretation
- compliance trail
Teaching schematic — not measured. The rule patterns are quoted from the post's rule_engine example; the flow follows the preprint's Figure 2.
Sources: the post (why thick skills beat fat prompts) · ClinAgent preprint, Fig 2.
Four reasons the rules live in JSON, not in a prompt
The alternative is one fat prompt. It loses on all four axes that matter in a regulated shop.
- Why not just write a bigger prompt? Four reasons.
- Testability: rules in a skill are unit-tested with plain fixtures — no LLM in the loop.
- Evolvability: a new warning pattern is one line in a JSON file. No code change, no redeploy.
- Transparency: a senior statistical programmer can read and sign off a JSON rule file. Under GxP, reviewability is the validation story.
- Determinism where it matters: the rule engine classifies; the LLM only explains.
- Adding a capability is mostly configuration — a new skill JSON; the MCP tools carry over unchanged.
- an axis where skill-side JSON wins
Teaching schematic — not measured. Argument structure from the post; no measured data in this scene.
Source: the post (why thick skills beat fat prompts).
Nine skills, one production study
The validation target: artifacts of one real study — with synthetic data standing in for patients.
- The validation target: artifacts of STUDY-A, a production Phase 2 cardiovascular study — eleven ADaM domains, 93,239 synthetic observations, no patient data.
- SK-005 log analysis: every real error and warning caught in ten logs, zero false alarms across 13,595 clean NOTE lines.
- SK-006 data validation: all 56 ADSL variables matched.
- Prompt-based spec generation scored 72.1% derivation accuracy — above 96% on the simple domains.
- SK-007 generated 12 of 16 TLFs; the four skips were missing macro names in the spec — an input-quality gap, not a skill failure.
- The benchmark measures tool correctness, not LLM reasoning — deliberately.
- skill under test
- measured headline number
Measured results — preprint §5.4 and Tables 8–13 (STUDY-A, SK-005/006/007, spec accuracy); layout is schematic.
Sources: the post (what the validation showed) · ClinAgent preprint, Tables 8–13.
Where it breaks — and why that's the point
Derivation accuracy collapses exactly where a study invents its own variables.
- The simple, standardized domains sit at the top — above 96% derivation accuracy.
- ADSL, with its many study-specific derivations, drops to 54.3%.
- ADBASE — almost entirely custom baseline flags — scores 0.0%.
- The correlation is Spearman ρ = −0.867 (p = 0.003); of 115 missing variables, 58.3% were study-specific derivations.
- That knowledge cannot live in a generic prompt. It belongs in thick, organization-specific skills.
- standardized domain
- study-specific-heavy domain
Measured statistics — preprint §6.8.3, Table 18 and Fig 7 (ρ, p); §6.8.4 / Table 19 (58.3%). Axis positions are schematic; point labels are exact.
Sources: the post (the failure-mode paragraph) · ClinAgent preprint, §6.8.
What this did not prove
The limits are part of the result — and the architectural claim survives them.
- One Phase 2 study. That's the whole evidence base.
- The log benchmark rests on a single real error — the Wilson interval runs from 20.7% to 100%.
- End-to-end productivity was never measured, and generation accuracy assumes well-structured input specs — the four skipped TLFs are the proof.
- Within those limits: rules belong in testable skill JSON, MCP tools stay thin, and the agent is interchangeable.
- Thin MCP, thick skills, swappable agent — the post and the preprint carry the full argument, limitations included.
- limitation (quoted from the post)
- claim that survives the limits
Limitations quoted from the post; the Wilson interval [20.7%, 100.0%] is from the preprint's Table 15.
Sources: the post (honest limitations) · ClinAgent preprint, Table 15 · journal version.