jaimeyan.com / interactive explainer

Thin MCP, Thick Skills: Five Layers for Clinical Programming Agents

AI coding agents reason well but can't read a SAS7BDAT, parse an ADaM spec, or tell a real ERROR from a harmless warning. ClinAgent's answer is a five-layer stack — thin stateless tools, thick testable skills — validated on a production study's artifacts. Seven scenes, failure mode included.

Honest label: teaching schematic — measured numbers are cited to the ClinAgent preprint (medRxiv 2026) by figure and table; single-study results are point estimates, and the walkthrough says so where it uses them.

scene — beat 0/0

Space play/pause  → next beat  ← previous beat  R reset

Static mode: animation disabled (reduced motion or no JavaScript) — every scene is shown in its complete final state.

SCENE 1 / 7

Smart agent, empty hands

The agent can reason about clinical programming — it just can't touch the artifacts.

Any MCP-compatible agent Claude Code · Cursor Cline · Augment Code SAS dataset dataset.sas7bdat binary, proprietary ADaM spec adam_spec.xlsx multi-sheet Excel SAS log adsl.log real vs harmless ERROR: File WORK.ADSL not found. WARNING: Unable to copy SASUSER real harmless ✗ ✗ ✗ the gap is in the tools, not the reasoning — thin MCP, thick skills
  1. Ask a coding agent to review a SAS log and it does fine.
  2. Ask it to read a SAS7BDAT dataset or parse an ADaM spec spreadsheet — it can't open the files.
  3. And it can't reliably tell a real ERROR from a harmless “Unable to copy SASUSER” warning.
  4. The gap is in the tools, not the reasoning.
  5. ClinAgent's answer in one line: thin MCP, thick skills.
  • the agent (your choice)
  • artifact the agent can't touch directly
  • the tool gap

Teaching schematic — not measured. Tool gap summarized from the post's opening; the log lines are illustrative.

Source: Thin MCP, Thick Skills (introduction).

SCENE 2 / 7

Five layers, each with one job

The stack under any MCP-compatible agent — the design bet is where the thickness sits.

any MCP-compatible agent — swappable 1 · A2UI — results rendering dashboards · log tables · RTF viewers — deterministic output for human review 2 · Skill router routes each tool call to the right skill · validates inputs first 3 · Skills — domain expertise SK-001–SK-009: prompt templates · few-shot examples · rule engines · tool bindings THICK 4 · MCP tools — data access stateless I/O only: read SAS7BDAT · parse spec Excel · read logs THIN 5 · Infrastructure — compliance AuditLogger · DataMasker · AccessControl · context minimization swap the agent — the skill layer keeps working
  1. Layer 1 renders results: dashboards, log tables, RTF viewers — deterministic output for human review.
  2. Layer 2 routes each tool call to the right skill and validates inputs first.
  3. Layer 3 is thick: nine skill packages, SK-001 to SK-009, each bundling prompts, examples, and rule engines.
  4. Layer 4 is thin: stateless I/O only — read a dataset, parse a spec, read a log. It never interprets.
  5. Layer 5 is compliance: audit logging, PHI masking, access control, context minimization.
  6. The agent reasons; the stack supplies domain expertise and the compliance trail. Swap the agent — the skills still work.
  • the design bet: thick skills, thin tools
  • supporting layers

Teaching schematic — not measured. Layer structure follows the post's Figure 1 and the preprint's Figure 1.

Sources: the post (the five layers) · ClinAgent preprint, Fig 1.

SCENE 3 / 7

Anatomy of a single tool call

“Classify the findings in this SAS log” — deterministic classification first, stochastic interpretation last.

“classify the ADSL log” router → SK-005 log reviewer inputs validated first MCP tool: read_log() stateless bytes in — no opinion rule_engine.json — deterministic ERROR: File WORK.ADSL not found. WARNING: AGE may be uninitialized. NOTE: Unable to copy SASUSER false positive — ignored LLM writes the explanation of what the rules already decided stochastic — allowed here AuditLogger: timestamp + inputs + outputs recorded
  1. A call arrives: classify the findings in this SAS log.
  2. The router picks skill SK-005 and checks the input first.
  3. A thin MCP tool reads the log bytes. It does not know what an ERROR means.
  4. A deterministic rule engine tags every line — errors, warnings, and known false positives like “Unable to copy SASUSER”.
  5. Only then does the LLM write the human-readable interpretation of what the rules already decided.
  6. The infrastructure layer logs the whole call — timestamped inputs and outputs.
  • deterministic decision
  • stochastic interpretation
  • compliance trail

Teaching schematic — not measured. The rule patterns are quoted from the post's rule_engine example; the flow follows the preprint's Figure 2.

Sources: the post (why thick skills beat fat prompts) · ClinAgent preprint, Fig 2.

SCENE 4 / 7

Four reasons the rules live in JSON, not in a prompt

The alternative is one fat prompt. It loses on all four axes that matter in a regulated shop.

why not just write a bigger prompt? Testability rules in JSON unit-test with plain fixtures — no LLM in the loop Evolvability a new warning pattern = one line in warning_patterns.json — no redeploy Transparency a senior statistical programmer can read and sign off the rule file Determinism the rule engine classifies; the LLM only writes the human-readable reading a new capability is a new skill JSON — MCP tools carry over unchanged
  1. Why not just write a bigger prompt? Four reasons.
  2. Testability: rules in a skill are unit-tested with plain fixtures — no LLM in the loop.
  3. Evolvability: a new warning pattern is one line in a JSON file. No code change, no redeploy.
  4. Transparency: a senior statistical programmer can read and sign off a JSON rule file. Under GxP, reviewability is the validation story.
  5. Determinism where it matters: the rule engine classifies; the LLM only explains.
  6. Adding a capability is mostly configuration — a new skill JSON; the MCP tools carry over unchanged.
  • an axis where skill-side JSON wins

Teaching schematic — not measured. Argument structure from the post; no measured data in this scene.

Source: the post (why thick skills beat fat prompts).

SCENE 5 / 7

Nine skills, one production study

The validation target: artifacts of one real study — with synthetic data standing in for patients.

STUDY-A — production Phase 2 cardiovascular artifacts 11 ADaM domains · 93,239 synthetic observations (Faker) · no patient data SK-005 · log analysis 10 logs: 1 error + 7 warnings caught 100% precision 13,595 clean NOTE lines, 0 false alarms SK-006 · data validation 56 / 56 ADSL variables matched deterministic comparison against the specification spec generation · prompt-based 72.1% derivation accuracy 95% Wilson CI [67.1, 76.7] >96% on ADMH / ADEX / ADCM SK-007 · TLF coder 12 / 16 TLF programs generated 4 skipped: spec lacked macro names — an input-quality gap, not a skill failure the bench measures tool correctness, not LLM reasoning — by design
  1. The validation target: artifacts of STUDY-A, a production Phase 2 cardiovascular study — eleven ADaM domains, 93,239 synthetic observations, no patient data.
  2. SK-005 log analysis: every real error and warning caught in ten logs, zero false alarms across 13,595 clean NOTE lines.
  3. SK-006 data validation: all 56 ADSL variables matched.
  4. Prompt-based spec generation scored 72.1% derivation accuracy — above 96% on the simple domains.
  5. SK-007 generated 12 of 16 TLFs; the four skips were missing macro names in the spec — an input-quality gap, not a skill failure.
  6. The benchmark measures tool correctness, not LLM reasoning — deliberately.
  • skill under test
  • measured headline number

Measured results — preprint §5.4 and Tables 8–13 (STUDY-A, SK-005/006/007, spec accuracy); layout is schematic.

Sources: the post (what the validation showed) · ClinAgent preprint, Tables 8–13.

SCENE 6 / 7

Where it breaks — and why that's the point

Derivation accuracy collapses exactly where a study invents its own variables.

derivation accuracy % study-specific variables simple domains >96% ADSL 54.3% ADBASE 0.0% Spearman ρ = -0.867 p = 0.003 115 missing variables: 58.3% study-specific study-specific knowledge belongs in thick, org-specific skills — not a fatter prompt
  1. The simple, standardized domains sit at the top — above 96% derivation accuracy.
  2. ADSL, with its many study-specific derivations, drops to 54.3%.
  3. ADBASE — almost entirely custom baseline flags — scores 0.0%.
  4. The correlation is Spearman ρ = −0.867 (p = 0.003); of 115 missing variables, 58.3% were study-specific derivations.
  5. That knowledge cannot live in a generic prompt. It belongs in thick, organization-specific skills.
  • standardized domain
  • study-specific-heavy domain

Measured statistics — preprint §6.8.3, Table 18 and Fig 7 (ρ, p); §6.8.4 / Table 19 (58.3%). Axis positions are schematic; point labels are exact.

Sources: the post (the failure-mode paragraph) · ClinAgent preprint, §6.8.

SCENE 7 / 7

What this did not prove

The limits are part of the result — and the architectural claim survives them.

Not proven ✗ one Phase 2 study — the whole evidence base ✗ one real error behind “100%”: Wilson CI [20.7%, 100.0%] ✗ productivity gain unmeasured ✗ needs well-formed input specs (the 4 skipped TLFs prove it) Within those limits ✓ rules live in testable skill JSON ✓ MCP tools stay thin + stateless ✓ invest in the skill layer: the agent is interchangeable thin MCP · thick skills · swappable agent — the post and preprint carry the full argument
  1. One Phase 2 study. That's the whole evidence base.
  2. The log benchmark rests on a single real error — the Wilson interval runs from 20.7% to 100%.
  3. End-to-end productivity was never measured, and generation accuracy assumes well-structured input specs — the four skipped TLFs are the proof.
  4. Within those limits: rules belong in testable skill JSON, MCP tools stay thin, and the agent is interchangeable.
  5. Thin MCP, thick skills, swappable agent — the post and the preprint carry the full argument, limitations included.
  • limitation (quoted from the post)
  • claim that survives the limits

Limitations quoted from the post; the Wilson interval [20.7%, 100.0%] is from the preprint's Table 15.

Sources: the post (honest limitations) · ClinAgent preprint, Table 15 · journal version.