jaimeyan.com / interactive explainer

The Contradictions Your Validator Can't See

A rule engine that reads one dataset at a time structurally cannot catch a contradiction that only exists across datasets. Seven scenes: one impossible response, the graph that exposes it, the shapes that patrol it, and the scoreboard — with the fine print.

Honest label: teaching schematic — records are fictional (Study XYZ, Subject 001) — measured numbers are cited to the CAVE-Onc paper (PLOS One, 2026) by figure and table.

scene — beat 0/0

Space play/pause  → next beat  ← previous beat  R reset

Static mode: animation disabled (reduced motion or no JavaScript) — every scene is shown in its complete final state.

SCENE 1 / 7

A response that cannot be true

One fictional subject, three domain records — and a combination RECIST 1.1 does not allow.

TR — target lesions Subject 001 SLD change: -24% PR-range shrinkage TU — non-target / new Subject 001 non-target: SD new lesions: none RS — overall response Subject 001 RSORRES: CR clinically impossible under RECIST 1.1 CORE: clean ✓ Pinnacle 21: clean ✓
  1. Meet Subject 001 in fictional Study XYZ: target lesions shrank 24% — a PR-level change.
  2. Non-target lesions are stable, and no new lesions appeared.
  3. Yet the overall response record says CR — complete response.
  4. Under RECIST 1.1 this combination is clinically impossible.
  5. Both industry validators reviewed this data — and reported it clean.
  • domain record
  • the impossible combination
  • validator verdict

Teaching schematic — not measured. Fictional records (Study XYZ, Subject 001); the scenario follows the RECIST 1.1 example in the post and the paper's introduction.

Sources: The Contradictions Your Validator Can't See · CAVE-Onc, PLOS One 2026 (introduction).

SCENE 2 / 7

One domain at a time

A domain-scoped rule is a filter over the rows of a single table — so every domain passes, and the join stays unchecked.

rule engine — one domain at a time TR TU RS every rule = a filter over the rows of one table PASS ✓ PASS ✓ PASS ✓ cross-domain joins: no rule looks here an expressiveness boundary — not a coverage gap you patch with more rules
  1. A rule engine like CORE sees one domain as one table — a DataFrame.
  2. Every check is a filter over the rows of that single table.
  3. TR passes its checks. TU passes. RS passes.
  4. The contradiction only exists across the tables — in the join nobody evaluates.
  5. This is an expressiveness boundary, not a coverage gap you can patch with more rules.
  • domain table
  • domain-local rules pass
  • cross-domain join, unchecked

Teaching schematic — not measured. Validator behavior simplified for teaching; the post carries the formal argument.

Source: The Contradictions Your Validator Can't See (section: why domain-scoped rules can't express this).

SCENE 3 / 7

Two-thirds port. The dangerous third doesn't.

Measured on the real CORE oncology rule set: what a SHACL porting pipeline could and could not carry over.

122 oncology CORE rules 85 port to SHACL losslessly (69.7%) 31 join across domains 6 row-set uniqueness the dangerous contradictions live in the unportable third
  1. CDISC CORE ships 122 oncology rules.
  2. 85 of them — 69.7% — port to SHACL shapes with zero loss.
  3. 31 cannot be ported: they join across domains.
  4. 6 more need row-set uniqueness that a single-record shape can't state.
  5. The dangerous contradictions live exactly in that unportable third.
  • ports losslessly (85)
  • cross-domain joins (31)
  • row-set uniqueness (6)

Measured, not schematic. Counts from the paper's SHACL shape-porting pipeline (85/122 = 69.7%; 31 cross-domain joins + 6 row-set uniqueness); bar proportions drawn to scale.

Sources: the post · CAVE-Onc, PLOS One 2026 (porting statistics).

SCENE 4 / 7

Turn the submission into one graph

Nine oncology domains leave their XPT silos; RELREC relationships become real edges, and SUPP-- qualifiers become plain properties.

DM EX AE LB TR TU RS SUPPDM RELREC 9 oncology SDTM domains, as XPT files XPT → RDF SUBJ 001 DM TR SLD -24% TU SD · no new RS CR RELREC keys, now explicit SUPPDM QNAM/QVAL unfold onto the parent record as properties
  1. Nine oncology SDTM domains leave their XPT silos.
  2. Each becomes nodes and edges in a single RDF knowledge graph.
  3. Subject 001's records are now nodes you can walk between.
  4. RELREC foreign keys survive as real edges — the joins validators couldn't see are now explicit.
  5. SUPP-- qualifiers unfold onto their parent record as plain properties.
  • XPT domain
  • record node
  • RELREC edge

Teaching schematic — not measured. Graph reduced to Subject 001's records; the nine-domain count and mapping approach follow the paper's Figure 1.

Sources: the post (the two-layer fix) · CAVE-Onc, Fig 1.

SCENE 5 / 7

Shapes patrol the graph

Layer 1 is a library of 111 SHACL shapes; each one is a graph pattern that either matches cleanly or raises a traced flag.

SUBJ 001 TR TU RS SHACL shape library 111 85 · ported CORE rules 8 · RECIST derivations 18 · archetype SHACL-SPARQL one shape = one graph pattern flag: CR contradicts TR/TU evidence trace a trace b trace c every firing lands in an append-only Merkle-chained audit store
  1. Layer 1 is a library of 111 SHACL shapes.
  2. 85 are the ported CORE rules from Scene 3; 8 encode RECIST derivations; 18 are archetype-specific SHACL-SPARQL.
  3. A shape matches the pattern from Scene 1 across TR, TU and RS — and fires.
  4. The flag says what a validator never could: the response contradicts the lesion evidence.
  5. Every firing writes a trace to an append-only Merkle audit chain — the 21 CFR Part 11 foundation.
  • record node
  • shape match / flag
  • shape library / audit store

Library composition measured (111 = 85 + 8 + 18, paper architecture section); the graph and flag layout are schematic.

Sources: the post (the two-layer fix) · CAVE-Onc, Fig 1.

SCENE 6 / 7

The rule that can't be one rule

Archetype A19 — the RECIST Table 7 matrix — is where a single SPARQL rule collapses, and a deterministic state machine takes over.

SPARQL-only attempt 34 nested IFs max nesting 34 0 testable units SHACL-SPARQL forbids lookup tables, so 34 matrix rows become one expression CaveAgent — deterministic state machine query TR (SLD) query RS (resp.) query TU (new) Table 7 lookup compare → trace flat dictionary, 34 rows 22 unit-testable blocks max nesting 7 $0.000 API cost per subject · fully deterministic · exactly replayable for inspectors
  1. One archetype — A19, the RECIST Table 7 matrix — defeats a single SPARQL rule.
  2. SHACL-SPARQL forbids lookup tables: 34 matrix rows become 34 nested IFs.
  3. Zero independently testable pieces — unmaintainable by construction.
  4. The agent layer instead runs a deterministic state machine: three typed queries against the graph.
  5. Then a flat 34-row dictionary lookup and one comparison — 22 unit-testable blocks, maximum nesting 7, not 34.
  6. No LLM call on this path: $0.000 per subject, exactly replayable for an inspector.
  • SPARQL-only attempt
  • deterministic agent (CaveAgent)

Complexity numbers measured (34 nested IFs vs. 22 testable blocks, max nesting 7; $0.000/subject — paper, Discussion / L3 cost); the diagrams are simplified.

Sources: the post (why the agent layer earns its place) · CAVE-Onc, Discussion.

SCENE 7 / 7

What the evaluation showed — and what it didn't

Twenty expert-reviewed contradiction archetypes, injected into clean data; each cell is one archetype.

20 injected contradiction archetypes — expert-reviewed CAVE (L1+L3) 20/20 A19→L3 CORE v0.15 8/20 P21 (FDA) 6/20 cross-domain RECIST subset: engines 0/10 · CAVE 10/10 (McNemar p = 0.002) clean data: flag sets barely overlap — Jaccard 0.004 → augment, not replace fine print — held-out check (shape library frozen): 3/5 caught construction validation · synthetic corpus · one L3 archetype — claims scoped accordingly
  1. Twenty expert-reviewed contradiction archetypes were injected into otherwise clean data.
  2. CAVE caught all 20 — nineteen by the shape layer, A19 by the deterministic agent.
  3. CORE caught 8 — the ones seeded from its own rule corpus.
  4. Pinnacle 21's FDA engine caught 6; on the ten cross-domain RECIST archetypes, both industry engines caught zero.
  5. On clean data the flag sets barely overlap — Jaccard 0.004: this augments validators, it doesn't replace them.
  6. Honest fine print: with the shape library frozen, 3 of 5 held-out archetypes were caught — construction validation on a synthetic corpus.
  • archetype detected
  • missed
  • needed the agent layer (A19)

Measured results — detection counts: paper Table 3; cross-domain subset and p-value: Table 3 / Results; Jaccard 0.004: Table 2; held-out 3/5: Table 4. Cell grid is a schematic echo of Fig 2, not the figure itself.

Sources: the post (evaluation; warnings) · CAVE-Onc, Tables 2–4, Fig 2.