Concept demo· Synthetic data· Company-neutral re-implementation

Platform architecture / ClinVista concept demo

A governed analytics platform
for clinical trials.

At a clinical-stage biotech, I designed, built, and operate a containerized analytics platform that serves clinical reviewers through a single governed entry point — 21 services covering data review, safety monitoring, and AI assistance under one governance layer. This section re-implements the architecture and five of its modules in company-neutral form with synthetic data, so the design can be inspected rather than taken on faith.

Nothing here is production code, configuration, or data. Employer, compound, and study identifiers are deliberately omitted; every dataset on these pages is synthetic.

21

containerized services behind one reverse-proxied entry

100%

replay fidelity, verified daily across 56,265 records in five SDTM domains

~96%

storage reduction — baseline + change log vs. full daily snapshots

0

external API calls — AI inference runs locally; data never leaves the server

Architecture

One door in. Governance at the platform layer. AI that never phones home.

Redrawn for this site from the production design — same topology, generic names. Users authenticate through a partner SSO portal; every request enters through a single TLS endpoint; services hold no data and no permission copies.

Clinical reviewers & analysts browser · corporate network / VPN Partner SSO portal AD login · issues JWT app launch links PLATFORM SERVER · SINGLE VM · 21 CONTAINERIZED SERVICES Nginx reverse proxy · :443 TLS — the only exposed port all internal service ports bind 127.0.0.1 DATA-REVIEW APPLICATIONS Patient Profile Viewer ADaM Explorer Endpoint Analysis Figure Studio TLF Viewer Clinical Tables Viz Explorer ML Insights TOC Editor PLATFORM SERVICES R APIanalysis endpoints Safety Monitorbaseline + change log AI InterfaceNL query · doc Q&A Zero-egress AI — local LLM inference on the same server embedding model baked into the image · indexes built offline · page-level citations · no external API calls Governance layer — applied to every service, surviving any single app tiered GxP classification · tamper-evident hash-chain audit · dual-signer e-signature · per-request authorization Partner data layer — workspace storage (read-only datasets) · SQL database every dataset request authorized per user, per request, by the partner's data service — the platform holds no data and no permission copies JWT authorized, read-only

Scroll horizontally to inspect the full diagram →

Figure 1. Platform topology, redrawn in company-neutral form. Production deployment adds scripted, version-controlled releases with automatic backups and daily restore-manifest checks.

Why one door in

A single reverse-proxied TLS entry means there is exactly one place where identity is established, authorization is checked, and access is logged. Internal services bind to localhost and are unreachable from the network — a misconfigured app cannot accidentally become a public endpoint, and the audit trail cannot be bypassed by going around the front door, because there is no other door.

Why not Posit Connect or ShinyProxy

Both are sound products; the constraint set pointed elsewhere. Reviewers authenticate through a partner CRO's SSO portal, so per-named-user licensing priced the wrong unit. Dataset authorization had to be enforced server-side, per request, by the data owner's own service — no shadow ACLs copied into the platform. And because GxP validation already required pinned, rebuildable environments, containers were the unit of validation; the orchestrator had to be as rebuildable as the apps it serves. Owning the thin proxy layer turned out to be less total complexity than bending a general-purpose product to these three rules.

Why governance lives in the platform layer

Applications come and go — one was already retired and replaced in production while the governance layer carried over unchanged. So the controls that must survive any single app live below it: an append-only, hash-chained audit log where every record is SHA-256-linked to its predecessor (six tamper classes are detected and localized by an automated verifier); a tiered GxP classification that scales review depth to risk — automated screening for exploratory tools, human validation with independent double programming for GxP-critical ones; and dual-signer e-signatures (author + reviewer, each re-authenticated at signing time) that are byte-bound to the exact gated object, so any later edit invalidates the signature and fails closed.

Why AI must be zero-egress

Clinical review data cannot leave the controlled environment, which rules out hosted model APIs by construction rather than by policy. The LLM runs locally on the same server; the embedding model is baked into the container image at build time; document indexes are built offline and shipped as files. Documents are curated by administrators — there is no end-user upload path — and every answer carries page-level citations back to the exact source page. In a regulated workflow, an answer that cannot show where it came from is not an answer.

Interactive modules

Five modules, re-implemented for the browser.

Each demo is a pure front-end re-implementation running entirely in your browser on embedded synthetic data — no backend, no external calls, nothing to configure. They are concept demos: the interactions are real, the data is not.

Concept demo

Patient Profile Viewer

Single-subject review: visits, exposure, adverse events, and concomitant medications on one timeline; laboratory trends with reference-range shading and out-of-range markers; a filterable AE listing. Switch between three synthetic subjects.

Open the demo →
Concept demo

Safety Data Monitor

Baseline-plus-change-log change detection: pick a snapshot date to see new, modified and deleted records with field-level old/new values, the alerts that day would have raised, and the daily replay verification that reconstructs history from the log.

Open the demo →
Concept demo

Safety Signal Review

The four views a safety reviewer opens first: hepatic safety against Hy's law, Kaplan-Meier time to first moderate/severe event, AE incidence by system organ class with risk differences, and a baseline-to-worst lab shift — every step, interval and count computed in your browser from record-level synthetic data.

Open the demo →
Concept demo

Trial Operations Dashboard

A global trial-operations KPI board in the language of feasibility, site selection and recruitment: enrollment vs. plan, screen-failure rates, site activation and cycle time, retention and dropout reasons, and operational-quality signals — filterable by country and site across three synthetic studies.

Open the demo →
Concept demo

Document Q&A — Zero-Egress AI

A mock of the platform's retrieval-augmented document assistant: ask preset questions of a synthetic protocol and watch answers arrive with page-level citations that highlight the exact source passage — the interaction contract of an AI that never sends data out.

Open the demo →
Live demo

Try the Platform Live

Not a browser re-implementation: submit real jobs to the governed tfl-platform API — run a synthetic study review where an independent QC gate can block the report, build an ADaM dataset, or render a figure from a JSON contract, then poll the queue and download hash-pinned artifacts and the ZIP evidence bundle. Synthetic data only; anonymous quota 5 runs/day.

Run a live job →

Validation evidence

Store once, reconstruct many — then prove it daily.

Replay as a standing control

One baseline snapshot plus an append-only change log reconstructs any historical state. A reconciliation job replays the log every day and diffs the result against the export actually received — 100% match on 56,265 records across AE, MH, CM, VS, and DM, on every date with source data. Replay fidelity is not a one-time claim; it is a daily automated check.

The storage math

Full daily snapshots would have cost roughly 7 GB over the 35-day validation window. Baseline plus deltas came to about 250 MB — a ~96% reduction with complete data fidelity and field-level change history retained. The monitor demo reproduces the mechanism at toy scale.

Evidence by construction

Validation deliverables (IQ/OQ/PQ, traceability, release records) map to scripted health checks and automated suites that run on every release, with daily SHA-256-manifest backups and rehearsed restores. Governance survives audits because it produces its evidence as a side effect of running.

Proof beyond this demo

Every claim has a public counterpart.

The platform itself runs behind its employer's firewall, so its mechanisms cannot be shown directly. But each of them has a public, inspectable counterpart on this site — same author, same engineering standard. If you can verify the counterparts, you know how the platform was built.

This site itself runs on the same standard — executed code, verified artifacts, public benchmarks — so the philosophy described on this page is inspectable end to end.

Related writing

The governance thinking, in long form.

What this section deliberately is not

  • It is not production code, configuration, infrastructure detail, or data. Every page is a from-scratch re-implementation running in your browser on synthetic fixtures.
  • The underlying platform runs at a clinical-stage biotech; employer, compound, and study identifiers are omitted by design.
  • The Trial Operations Dashboard is a concept implementation — it demonstrates how trial-operations analytics requirements (feasibility, site selection, recruitment and retention) get productized. It is not a system in production use.
  • Production metrics quoted above (21 services, 100% replay fidelity over 56,265 records, ~96% storage reduction, 0 external API calls) come from the platform's validation records; the demos use their own smaller synthetic numbers, labeled as such.