The Auditor's Question
The Auditor's Question
SIX WEEKS AFTER SUBMISSION · AUDITOR POINTS AT ONE TABLE
“Which program version produced this,
and what changed since the previous run?”
t_14_1_1.sas
t_14_1_1_final.sas
t_14_1_1_final_v2.sas
t_14_1_1_final_v2_JC.sas
Nobody can say which copy ran — the log nobody kept is the honest answer.
Filename versioning is not an audit trail.
Agenda: Git vocab · branch-per-output for QC · .gitignore
Keep data out of version control.
[Opening · L2 modern workflow] Drops learners into the concrete situation from the top of the article: an auditor asks which program version produced a shipped table. Establishes why filename versioning fails on day one in clinical programming.
Speaker notes
Picture a submission package that shipped six weeks ago, and a quality auditor points at one table asking which program version produced it and what changed since the previous run. The study folder answers with a pile of files: t_14_1_1.sas, t_14_1_1_final.sas, t_14_1_1_final_v2.sas, and t_14_1_1_final_v2_JC.sas. Nobody can honestly say which copy actually ran. The only truthful answer lives in a log file that nobody kept. That is the trap: filename versioning is what many clinical teams treat as an audit trail, and it is not one. So today we cover the minimum Git vocabulary, branch-per-output mapped to quality control, and how to keep data out of version control.
Why the Suffix Pile Is a Liability
Why the Suffix Pile Is a Liability
No diff
Manual compare; changes ship unreviewed.
Silent forking
Parallel edits from one base diverge or overwrite.
Provenance is hearsay
No archived log ties file to output.
“Final” always lies
“_final” means final until Friday.
Root cause: version suffix has no enforced metadata
No author · No timestamp · No link to output
[Concept arc · L1] Names the four failure modes hiding in the example folder and the single root cause: filename suffixes carry no enforced metadata.
Speaker notes
On this slide, we look at why piling up filename suffixes becomes a real liability. First, there is no diff: comparing _v1 to _v2 means opening both files and comparing by eye, so changes often ship unreviewed. Second, silent forking happens when two programmers fix two issues from the same base copy, both save final versions, and one overwrites the other or the files diverge. Third, provenance becomes hearsay, because no archived log reliably ties a file to the output it produced. Fourth, final always lies: every experienced programmer knows _final means final until Friday. The root cause is that the version suffix carries no enforced metadata; it is just a label typed by a tired human at 6 p.m., with no author, no timestamp, and no link to the output.
An Audit Trail Is System-Maintained
An Audit Trail Is System-Maintained
Filename versioning
✗ no author
✗ no timestamp
✗ no output link
✗ later edits silent
Audit trail
✓ author recorded
✓ timestamp recorded
✓ output linked
✓ later edits detected
Filename versioning fails as an audit trail.
Git answers by construction — never by filename.
Next: the minimum Git vocabulary actually needed.
[Concept arc · L1] Contrasts a remembered filename suffix with a system-maintained record, and previews why Git answers the audit question by construction.
Speaker notes
Imagine an auditor asks which program version produced a table. With filename versioning, the suffix is something a person remembers to type. It gives you no author, no timestamp, no link to the output it produced, and later edits stay silent. An audit trail has to be system-maintained, so it records the author, records the timestamp, links the output, and detects later edits. That is the difference: a version suffix has none of those four properties. Git answers those questions by construction — never by the filename. It tells you who changed what, when, and starting from what. So filename versioning fails as an audit trail. Next question: what is the minimum Git vocabulary a clinical programmer actually needs?
The Five-Word Git Vocabulary
The Five-Word Git Vocabulary
| # | Shared-drive habit | Git word | Git’s stricter record |
|---|---|---|---|
| 1 | Copy the study folder to your machine | clone | Working copy + full change history |
| 2 | Save a new version of the file | commit | Timestamped record: author, diff, message |
| 3 | Work on a copy without breaking shared state | branch | Isolated line of work, mergeable later |
| 4 | Email the folder for review | pull request | Proposed merge: diff visible, commentable, approvable |
| 5 | Ask “Who changed this line?” | review / git log / git blame | The answer, for every line, permanently |
• SCE = Statistical Computing Environment · Read row by row: old habit → Git word → stricter record.
• Deliberately small — clone, commit, branch, pull request, review — enough for modern SCE work.
[Concept arc · L1 → L2] Maps the five shared-drive habits learners already know to the five Git words that replace them, using Table 1 from the source.
Speaker notes
You already know these five habits: copying the study folder to your machine is clone — but now you also get the full change history. Saving a new version of the file is commit — a timestamped record with author, diff, and message. Working on a copy without breaking shared state is branch — an isolated line of work you can merge later. Emailing the folder for review is a pull request — a proposed merge where the diff is visible, commentable, and approvable. And asking who changed this line is review, git log, or git blame — the answer, for every line, permanently. That's the whole vocabulary, and it's deliberately small — enough for modern Statistical Computing Environment, or SCE, work.
Match the Habit to the Git Word
Hands-on interactive — if it does not load, open the paired article and try the exercise there.
Speaker notes
This next piece is a hands-on exercise, so you'll do it on the website at jaimeyan.com/learn rather than in the video. In it you practise the five-word vocabulary by matching each shared-drive habit to the Git word that replaces it: copy folder to clone, save a new version to commit, work on a copy to branch, email folder for review to pull request, and ask who changed this line to review, git log, or git blame. Pause the video here and give it a try, because pairing each habit with its Git word is the fastest way to make the vocabulary stick.
Checkpoint: The Audit-Trail Case
1 Two programmers start from the same base file. Programmer A saves a copy as final.sas, and Programmer B also saves a copy as final.sas. What is the clearest failure mode in this filename-versioning scenario?
2 Why does a descriptive filename suffix such as program_final_v2.sas fail as an audit trail, even when the name looks informative? Select all that apply. (select all that apply, then Check)
3 A programmer needs to produce evidence that a program version was saved with its author, timestamp, change message, and diff. Which single Git vocabulary word names that timestamped change record?
Speaker notes
This checkpoint is the audit-trail case, where you apply Git vocabulary to concrete version-control choices. Two programmers start from the same base file and both save final.sas, and the clearest failure mode is B: one save can overwrite or orphan the other programmer's changes, and the identically named files carry no system-maintained record of which version was used. For a descriptive suffix such as program_final_v2.sas, the correct choices are A, C, and D: the suffix is only a user-typed label, it does not show who changed the file or when or what content changed, and it does not link the file to a version-control event for comparison with the previous version. When a programmer must show that a program version was saved with its author, timestamp, change message, and diff, the single Git word is B, commit. A commit is the timestamped change record, while a clone copies a repository, a branch is a movable line of development, and a pull request, or PR, is a review-and-merge proposal built from commits.
The Daily Loop, One Output at a Time
The Daily Loop, One Output at a Time
SETUP — ONE BRANCH PER OUTPUT
git switch -c tlf/t_14_1_1 # one branch per output
$EDITOR programs/tlf/t_14_1_1.sas
git add programs/tlf/t_14_1_1.sas
git commit -m "t_14_1_1: denominator to SAF per spec v3"
git push -u origin tlf/t_14_1_1
gh pr create
/* Study XYZ — t_14_1_1.sas (no suffix; history lives in Git) */
/* Purpose: AE summary by treatment group */
/* Change control: git log -- programs/tlf/t_14_1_1.sas */
%include "setup.sas";
[Workflow walkthrough · L2] Steps through Study XYZ's daily loop from the source: a branch per output, the Git commands in order, and a SAS program header that no longer carries the version story because Git owns the history.
Speaker notes
Now let's walk the daily loop, one output at a time. You begin with one branch per deliverable: git switch -c tlf/t_14_1_1. Then edit programs/tlf/t_14_1_1.sas with $EDITOR, git add that same file, and commit with the message t_14_1_1 colon denominator to SAF, the safety analysis set, per spec v3. Push it with git push -u origin tlf/t_14_1_1, then open a pull request, a PR, with gh pr create or the Statistical Computing Environment, the SCE, interface. The header now carries no suffix, because history lives in Git, and change control is git log -- programs/tlf/t_14_1_1.sas, followed by %include setup.sas. The repository owns the version story, so the _final pile is gone, and every commit carries an author, a timestamp, a diff, and a message.
Branch per Output, Mapped to QC
Hands-on interactive — if it does not load, open the paired article and try the exercise there.
Speaker notes
This one is a hands-on exercise you do on the website at jaimeyan.com/learn, where you click through the branch-per-output flow and watch where each piece of quality control (QC) evidence lands. It practices the skill of mapping the two halves of double programming onto Git roles and artifacts, so you can see the production programmer open a pull request (PR) and the independent QC programmer review that PR with the spec open. Try it right after the video, while the ideas are still fresh.
The Pull Request Is the QC Record
The Pull Request Is the QC Record
| Auditor asks | In the approved PR |
|---|---|
| What changed? | The diff |
| Who reviewed it? | Reviewer approval |
| When? | Timestamps |
| Against which spec version? | PR description |
| With what QC evidence? | Results recorded on the PR |
Shared-drive alternative
Email chain
Tracked-changes file
Plus hope
Reviewer independence
PR approver ≠ branch author
Same rule as double programming
Branch-per-output
Merge conflicts are rare
Two people rarely edit same deliverable at once
[Concept · L2] Shows how an approved pull request mechanically answers the auditor's questions, and why reviewer independence is the part worth protecting.
Speaker notes
An approved pull request, or PR, is your quality control, or QC, record. The diff shows exactly what changed, the approval shows who reviewed it, and the timestamps show when. The PR description ties the work to a specific spec version, and the results recorded on the PR are your QC evidence. Compare that to a shared drive, where you get an email chain, a tracked-changes file, and hope. For reviewer independence, the PR approver must not be the branch author, exactly like double programming. And with branch-per-output, merge conflicts are rare because two people rarely edit the same deliverable at once.
Data Never Enters Git
Data Never Enters Git
SCE = Statistical Computing Environment · repository boundary: source, not data
✓ Commit to Git
code, specs, shells, docs
✗ Never commit data
data — patient-level or derived
Why keep data out?
Datasets are large and access-controlled elsewhere; Git copies quietly defeat that model.
Study XYZ — .gitignore
*.sas7bdat
*.xpt
*.csv
logs/
output/
• Data arrives via SCE governed connections
• Reference datasets by path or snapshot ID
• Run records log the exact snapshots read
• Access is scoped and logged — no drive letter
[Concept · L2] States the repository boundary rule: code, specs, shells, and docs yes; data, patient-level or derived, never. Shows the exact .gitignore and how governed SCE data connections replace the shared drive.
Speaker notes
In this scene, the rule is simple: data never enters Git. The repository boundary is source, not data, so commit code, specs, shells, and documentation. Never commit data, whether patient-level or derived, because datasets are large, access-controlled elsewhere, and Git copies quietly defeat that access model. The example study's .gitignore makes this mechanical with five lines: *.sas7bdat, *.xpt, *.csv, logs/, and output/. Data reaches your session through the Statistical Computing Environment (SCE) and its governed connections, and programs reference datasets by path or snapshot identifier. Run records tie each execution to the exact data snapshot it read, so access is scoped and logged instead of granted by drive letter.
Final Check: Working with Git in GxP
1 In a Git-backed clinical study, an auditor asks which program version generated a shipped output. Which artifact gives the definitive answer?
2 Which controls make Git history acceptable to auditors in a GxP environment? Select all that apply. (select all that apply, then Check)
3 Datasets are analysis inputs, yet they usually stay out of Git in a regulated Statistical Computing Environment (SCE). Explain why repository boundaries exclude datasets, and describe how the audit link to the version of the dataset used for an output is preserved. (reflect, then reveal)
Reveal analysis
Speaker notes
This is the final checkpoint on working with Git in regulated good-practice environments (GxP). For the auditor asking which program version generated a shipped output, the definitive artifact is A, the commit hash that the run record stamps when the output was created. That hash identifies the exact program version used to produce the output, while a review event, a timestamp, or a branch name does not uniquely pin down the run. In a GxP environment, the controls that make Git history acceptable to auditors are A, protected branches that prevent direct pushes, B, recorded approvals on pull requests (PRs) before merge, and C, access controls that restrict who can rewrite history. Together these prevent casual or unauthorized changes to history, which is what makes a Git log trustworthy for an audit. In a regulated Statistical Computing Environment (SCE), datasets stay out of Git because they are often too large, are governed by separate access controls, and may require data-masking or security protections a code repository is not designed to handle, while the audit link is preserved through governed data connections so the dataset version is known via a snapshot or data-management record and the run record stamps the version used for the output.
The Repo Is the Record
The Repo Is the Record
Summary · paired source article · no patient-row walkthrough shown
1 · Filename versioning is not an audit trail
2 · clone → commit → branch → pull request → review
3 · Branch per output; QC programmer reviews PR
4 · Code in Git, governed data outside
SCE (Statistical Computing Environment) connections + .gitignore enforce it
5 · L3 agentic flow is volatile (as of 2026-08-30) — agents can draft the change, commit message, and PR description, but a human approves on QC evidence, not prose. Re-verify tool specifics.
[Summary · L2 + L3 volatile] Consolidates the source's key takeaways, links back to the paired source article, flags the L3 agentic question as time-sensitive, and honestly notes that no patient-row walkthrough appears because the material contains no patient-level extracts.
Speaker notes
Let's pull the lesson together, because the repository is the record and this pairs with the source article. Filename versioning is not an audit trail; there is no diff, no provenance, silent forking, and final is a label a tired human typed. Five Git words cover our clinical workflow: clone, commit, branch, pull request (PR), and review, mapped one-to-one from shared-drive habits. Branch per output, with the independent quality control (QC) programmer as the PR reviewer, turns double programming into an exportable record. Data never enters Git; the Statistical Computing Environment (SCE) connections carry it, and .gitignore makes that rule mechanical. Finally, the L3 agentic way is volatile, as of 2026-08-30: agents can draft the change, the commit message, and the PR description, but a human approves on QC evidence, not prose, so re-verify tool specifics before relying on them.