A 2020 SAS Program Hits the Cloud
A 2020 SAS Program Hits the Cloud
2020-ERA STUDY PROGRAM · FIRST EXECUTABLE LINE
libname sdtm "E:\STUDY\SDTM";
• libname works only where E:\ is mapped
• Colleague or cloud: fails before first PROC
• One box assumption — convention + luck
Statistical Computing Environment (SCE)
Modern workflow — what changes on day one
(Opening page) Put learners in a realistic day-one situation: opening a legacy clinical SAS program, hitting the hardcoded drive-letter libname, and asking what a statistical computing environment is actually for.
Speaker notes
Let's start with a 2020-era study program and its first executable line. That line assigns a libname for the Study Data Tabulation Model (SDTM) library, pointing at the E drive on one specific machine. It ran because exactly one box had that drive mapped. Hand the same program to a colleague, or lift it into a cloud workspace, and it fails before the first PROC step ever executes. The drive letter is only a small symptom of a large assumption: the study lives in one place, held together by naming conventions and luck. That is why we introduce the Statistical Computing Environment (SCE) and a modern workflow, and what changes on day one.
Three jobs, no more
STATISTICAL COMPUTING ENVIRONMENT (SCE)
Three Jobs, No More
Not speed. Not convenience. Compliance instrument first.
Validated Statistical Computing
• Versions & configuration
• Change control
• Run record, not a shrug
Access Control
• Patient-level data: authorized only
• Per-study scope
• Access logged
Reproducibility
• Same code & data → same output
• Your machine, QC's, years later
(Concept arc, L1) Establish the fundamental definition of an SCE as exactly three duties: validated statistical computing, access control, and reproducibility.
Speaker notes
Strip away vendor decks: a Statistical Computing Environment (SCE) has exactly three jobs, no more. The first is validated statistical computing, where known versions and configuration stay under change control, so the answer to what produced this table is a run record, not a shrug. The second is access control: patient-level data is visible only to authorized people, scoped to the study they work on, and access is logged. The third is reproducibility: same code and same input data produce the same output on your machine, on the quality control (QC) programmer's machine, and years later for an inspection response. Speed and convenience are not on this list; an SCE is a compliance instrument first.
The Old Stack and Its Failure Modes
The Old Stack and Its Failure Modes
Old stack elements
• Desktop SAS on one machine
• Shared drive, folder tree
• Versioning by file name
• Audit trail: timestamps, copies
Failure modes
• QC copy drifts from production
• final_v2 edited post-run
• Inputs are left to memory
• Run order breaks on leave
| Dimension | Old stack: desktop SAS + shared drive | Cloud SCE (statistical computing environment) |
|---|---|---|
| Engines | Desktop SAS versions vary by machine | Controlled engine per run |
| Versioning | File names: _v2, _final, _final_v2_JC | Git commits and branches |
| Audit trail | Timestamps, copies, memory | Run record: code, data, environment |
| Data access | Folder paths, manual copies | Controlled data locations per run |
| Review | Output files shared informally | Pull request on code and outputs |
| Reproducibility | Holds only with discipline | System property: rerun reproduces |
(Concept arc, L1) Break down the desktop-SAS plus shared-drive pattern, show how each element fails, and set up the six-dimension comparison against a cloud SCE.
Speaker notes
Here is the legacy stack: desktop SAS on one machine, a shared network drive organized by folder hierarchy, versioning by file suffix, and an audit trail assembled from timestamps, folder copies, and memory. The table contrasts it with a statistical computing environment, or SCE, across six dimensions: engines, versioning, audit trail, data access, review, and reproducibility. Filenames drift into patterns like _v2, _final, _final_v2, and _final_v2_JC, but re-running t_14_1_1_final_v2.sas proves nothing if nobody can establish which inputs it read. That is where failures compound into provenance. The quality control, or QC, copy differs from production's; final_v2 was edited after the output was generated; and the one person who knew the run order is on leave. Every element of the old stack assumes good discipline, while the SCE replaces that assumption with a system property.
Checkpoint: What an SCE Is For
1 Which list names the three duties that define a Statistical Computing Environment (SCE) in a regulated clinical reporting environment?
2 Why is filename-suffix versioning such as `_final_v2` unable to deliver the reproducibility that an SCE requires? Select all that apply. (select all that apply, then Check)
3 Why does a shared-drive stack where any user with the drive letter mapped can read or edit program files fail the access-control duty of an SCE?
Speaker notes
This checkpoint quiz verifies your understanding of what a Statistical Computing Environment (SCE) is for. Question one: which list names the three duties that define an SCE? The correct answer is C, reproducibility, access control, and audit trail, because those three duties are what define an SCE in a regulated clinical reporting environment. Question two: why can filename-suffix versioning such as _final_v2 not deliver the reproducibility an SCE requires? The correct answers are B, C, and D, because a filename suffix cannot create a run record, cannot record the exact state of inputs and package versions at the time of analysis, and cannot identify which version of the analysis program or supporting library produced a particular output. Question three: why does a shared-drive stack where any user with the drive letter mapped can read or edit files fail the access-control duty of an SCE? The correct answer is C, because access control requires per-user permissions and an audit trail, and a broad drive mapping does not restrict actions or show who changed a file.
Anatomy of a Cloud SCE
Anatomy of a Cloud SCE
| # | Piece | What it does |
|---|---|---|
| 1 | Workbench | Browser IDE; code execution is not tied to a specific physical machine |
| 2 | Data connections | Governed links to the study data store; data is mounted with your permissions, never emailed or copied |
| 3 | Git-backed projects | Project code and configuration stay in Git repositories with history and review |
| 4 | Reproducible runs | Run record stores code commit, input data snapshot, and compute environment; outputs can be replayed |
| 5 | Governance layer | Permissions, approvals, and audit trail across workbench, data, code, and runs |
Lens: score SCE vendors (e.g., Domino Data Lab) on all five pieces, not the demo
(Concept arc, L2) Walk through the five pieces shared by cloud SCE platforms and the governance viewpoint behind what sponsors buy.
Speaker notes
Whatever a vendor calls it, a Statistical Computing Environment, or SCE, has five recognizable pieces. The workbench is a browser-based workspace where you edit and execute code without caring which physical machine runs it. Data connections are governed links to the study data store, so data is mounted into your session with your permissions rather than emailed or copied. Git-backed projects keep project code and configuration in Git repositories with history and review. Reproducible runs stamp every execution with the code commit, the input data snapshot, and the compute environment, so the run record lets you replay outputs and defend them. The governance layer adds permissions, approvals, and an audit trail across workbench, data, code, and runs; evaluate vendors like Domino Data Lab against all five pieces, not just the demo.
Walkthrough: The libname That Replaces Drive Letters
Walkthrough: The libname That Replaces Drive Letters
libname sdtm "E:\STUDY\SDTM";
Legacy hard-code | E: mapped on exactly one machine
%let root = %sysget(STUDY_ROOT); /* set by SCE at launch */
libname sdtm "&root./data/sdtm" access=readonly; /* input, read-only */
libname adam "&root./output/adam"; /* output */
proc contents data=sdtm.dm nods; run;
Confirms run reaches the first PROC step unchanged
Old: location | New: program intent
Same code runs in dev, QC, validation — no edits
(Real-code walkthrough, L2) Step through the legacy and environment-driven libname patterns from the material facts verbatim, annotating what each line encodes.
Speaker notes
Here is the legacy line: libname sdtm "E:\STUDY\SDTM"; it worked because exactly one machine had the E: drive mapped. The Statistical Computing Environment (SCE) replaces that machine-specific path with an environment variable: %let root = %sysget(STUDY_ROOT);, which the SCE sets at launch. Your Study Data Tabulation Model (SDTM) inputs become libname sdtm '&root./data/sdtm' access=readonly;, so they cannot be written to. Your Analysis Data Model (ADaM) outputs become libname adam '&root./output/adam';. Then proc contents data=sdtm.dm nods; run; confirms the run reaches the first PROC step unchanged. The old libname encoded where the study lived; the new libname encodes what the program means, so the same code runs in development, quality control, and validation with no edits.
Hands-on: Refactor the Hardcoded Path
Hands-on interactive — if it does not load, open the paired article and try the exercise there.
Speaker notes
This step is a hands-on exercise you complete on the website at jaimeyan.com/learn, not in the video, so open that page when you are ready. There you practice the cheapest portability fix in legacy SAS: taking a LIBNAME that points at a hardcoded drive letter and rebuilding it as an environment-driven pattern, reading the root with %SYSGET and declaring the Study Data Tabulation Model (SDTM) inputs as read-only. Try it right after this video, because once you can assemble that program so that PROC CONTENTS runs cleanly, you have a skill that carries straight into the cloud Study Computing Environment (SCE).
Multi-engine Is the Point
Multi-engine Is the Point
| 1 · Legacy default | SAS was the only engine installed, so SAS became the default. |
|---|---|
| 2 · Cloud SCE | SAS, R, and Python run as peer engines inside the same project. |
| 3 · Current workload | Submissions still lean on SAS; the R pharmaverse (admiral, metacore, tern) and Python toolchains take a growing share. |
| 4 · Validation | Validation follows the artifact — an R-built ADaM needs the same specification, QC, and review as an SAS-built ADaM. |
(Concept arc, L2) Explain why a cloud SCE runs SAS, R, and Python as peers and where the validation obligation actually attaches.
Speaker notes
Let's talk about why multi-engine support is the point of a cloud statistical computing environment, or SCE. In the old stack, SAS was the only engine installed, so SAS simply became the default by circumstance, not by design. A cloud SCE changes that: SAS, R, and Python run as peer engines inside the same project. That said, submission packages still lean on SAS, while the R pharmaverse stack, including admiral, metacore, and tern, along with Python toolchains, take a growing share of analysis work. Here is the part that trips people up: validation follows the artifact, not the language. An analysis data model, or ADaM, built in R needs the same specification, quality control, and review as one built in SAS, and that expectation does not move just because the engine changed.
Day-one Habits
Day-one Habits
01 Repo first
02 Branch + PR
03 No data in Git
04 Run record
Repo is the study home
Session-only work is lost
Branch per deliverable
PR diff = audit trail
Datasets stay outside Git
Git keeps code and specs
Commit · snapshot · env
Trust the stamp
No permission required — inspection-ready habits
(Concept arc, L2) Four habits that separate actually using an SCE from merely being logged into one.
Speaker notes
On day one in the Statistical Computing Environment (SCE), adopt four habits that make your work auditable. First, repository first: the project repository is the study's home, and work that exists only in your session does not exist. Second, branch for every deliverable, then open a pull request (PR); the PR diff and its review become your audit trail. Third, keep data out of Git: datasets live behind the platform's data connections, while version control holds code and specifications. Fourth, after every execution, read the run record—the commit, data snapshot, and environment the platform stamped—and trust that stamp rather than your memory. None of these habits requires anyone's permission, and each one lines up with what an inspector will probe for.
The Agentic Way in a Governed SCE
The Agentic Way in a Governed SCE
Statistical Computing Environment (SCE) — governed AI action: traceable, contained, human-gated.
| Aspect | Governed-agent behavior |
|---|---|
| SCE conversation | Traceability & containment: project-tied sessions; diffs attributed & reviewed; prompts and outputs in the audit trail. |
| Boundary is unchanged | Assistant drafts. Deterministic checks and independent human review decide. A named person owns the validation record. |
| Main failure mode | Plausible output — reads well but never executed/reviewed/gated. Treat like a first draft: run, diff, gate; human approves merge. |
| Time-sensitive layer | Platform/tool specifics last verified 2026-08-30; re-verify before relying; no permanence promise. |
| Governed access | Study data does not leave the SCE by default; check your organization’s policy before pasting. |
(Concept page, L3 — time-sensitive) Frame how AI coding assistants fit inside the SCE governance boundary and flag the volatile-layer asOf framing.
Speaker notes
An artificial intelligence, or AI, assistant inside a Statistical Computing Environment, or SCE, is part of that conversation. What matters is traceability and containment: sessions stay tied to the project, AI-suggested diffs are attributed and reviewed, and prompts and outputs land in the audit trail. The boundary is unchanged: the assistant drafts, deterministic checks and independent human review decide, and a named person owns the validation record. The main failure mode is plausible output: a diff that reads well but was never executed, reviewed, or gated, so treat it like a first draft and run it, diff it, gate it, then let a human approve the pull request. Platform and tool details were last verified on 2026-08-30, so re-verify before relying; no permanence promise. Governed access means study data does not leave the SCE by default; check policy before pasting.
Final Check: SCE and Modern Workflow
1 What must a reproducible-run record capture to make an output reproducible in a Statistical Computing Environment (SCE)?
2 A SAS VM is not automatically an SCE; without governance it is only a VM with better marketing. Which missing governance pieces would turn that VM into a real SCE? Select all that apply. (select all that apply, then Check)
3 Explain why re-running the SAS program `t_14_1_1_final_v2.sas` today cannot establish provenance for the deliverable it formerly produced. In your answer, state what evidence a run record would add. (reflect, then reveal)
Reveal analysis
Speaker notes
Time for the final checkpoint: three questions that integrate the Statistical Computing Environment (SCE) and modern workflow habits you have learned. Question one asks what a reproducible-run record must capture to make an output reproducible in an SCE. The correct answer is B: the exact code commit, the data snapshot, and the compute environment used for that run, because only that triplet identifies what actually executed. Question two asks which governance pieces turn a SAS virtual machine (VM) into a real SCE rather than a VM with better marketing. The correct options are A, B, and C: automated run records for every execution, a Git repository using branch-per-deliverable and peer-reviewed pull requests, and immutable data snapshot management, because more hardware is performance, not governance. For the short-answer question, the model answer is that a file name such as t_14_1_1_final_v2.sas is not provenance, because running it today executes the current file contents against whatever datasets and compute environment are present and produces a new output, while a run record supplies the missing evidence by stamping the Git commit, data snapshot, and compute environment that created the original output.
Key Takeaways and Next Steps
Key Takeaways and Next Steps
SCE = three duties
• Validated tools
• Controlled data access
• Reproducible runs
Not: where SAS is installed
Shared-drive stack
• Versioning by filename
• Access by drive letter
• Provenance by memory
Fails all three duties
On a cloud SCE
• Code in Git
• Data behind governed connections
• Run record = audit trail
not the file name
Material boundary
Module scope: code-pattern facts only
No patient-level rows shown or invented
Rules taught: role-based, study-scoped, logged access
Next in bootcamp
Git for SAS programmers
Branching, pull requests, review
Try the sce-study-bootstrap skill
Repo layout + run checklist
(Summary page) Consolidate the lesson, state the material boundary explicitly, and link forward to the paired bootcamp material and the sce-study-bootstrap skill.
Speaker notes
A Statistical Computing Environment (SCE) is defined by three duties: validated tools, controlled data access, and reproducible runs. It is not defined by having SAS, the Statistical Analysis System, installed somewhere. The shared-drive stack fails all three duties because it versions by filename, grants access by drive letter, and tracks provenance by memory. On a cloud SCE, the study is code in Git plus data behind governed connections, and the run record, not the file name, is the audit trail. This module had code-pattern material facts only, so no patient-level rows are shown or invented; patient-data rules such as role-based, study-scoped, and logged access are taught as rules. Next in the bootcamp comes Git for SAS programmers, where branching, pull requests, and review move work forward; also try the sce-study-bootstrap skill for repository layout and the run checklist.