Governance and audit infrastructure. Not a clinical-decision-support device. Not used on real patients. Inputs are synthetic, public-dataset, or de-identified.
Tools

Evaluate clinical AI outside the live clinical workflow

Hospitals and governance committees need a way to check what a clinical-AI tool is doing without putting it near a patient. RocSite compares a model’s output to a locked, published clinical instrument, or a locked intended-use claim, on synthetic, public, or de-identified data, and signs the result. Every evaluation is auditable. This is not clinical decision support, it is not used on real patients, and it is not a third reader in anyone’s worklist.

Three tools, one rule: evaluate claims against a locked reference, sign the result, stay off the patient.

Start here

Your first governance program

Most institutions do not need another AI tool first, they need a baseline. If your committee is still working out what AI governance should even look like, this is the engagement to start with.

01

Intended-use audit

An inventory of what is already deployed: does each tool’s cleared intended use match the problem it was bought to solve?

02

Monitoring baseline

Which outputs an expert can verify at a glance, and which need a locked, published instrument watching them for drift.

03

Signed records on file

Off-patient evaluations your committee can hand to leadership, an auditor, or a regulator, each one Ed25519-signed and independently checkable.

Request a walkthrough
Tool 01, AI Governor

Bench evaluation of clinical AI, off the patient

AI Governor is for CMIOs, AI governance committees, and clinical informatics teams, especially teams without a deep imaging-AI bench. It takes a clinical-AI output and evaluates it against a published, deterministic clinical instrument or a locked intended-use claim, on synthetic, public-dataset, or de-identified inputs. Use it as a bench evaluation before procurement, or as a post-deploy audit. Each evaluation is cryptographically signed with Ed25519 and independently verifiable.

Not a clinical-decision-support device. Not used on real patients. Never inserted into the diagnostic worklist. Inputs are synthetic, public-dataset, or de-identified. The output is an audit record for your governance file, a signed receipt, not a clinical report.

KDIGO 2012Acute kidney injury staging NEWS2Deterioration early warning APACHE IIICU mortality

Three of 21 locked instruments, see all instruments below ↓

Where an expert can verify a model’s output at a glance, a radiology read, Governor stays out of the way. Where no one can verify by looking, sepsis risk, deterioration scores, AKI prediction, a locked, published instrument is the only independent reference, and that is where drift hides. That gap is what Governor covers.

  • Grounding. Is the model’s stated reasoning grounded in the locked instrument that applies?
  • Intended-use fit. Does the deployment match the tool’s cleared intended use, and does that intended use match the problem your institution actually bought it to solve?
  • Post-deploy behavior. Is the model still behaving the way it did when the committee approved it?

What your institution actually receives

Sample signed audit record · synthetic demonstration case · not a real evaluation
Model output “AKI stage 1, low risk of progression”, the institution’s deployed model, on synthetic case SYNTH-AKI-0042
Locked instrument KDIGO 2012, AKI staging · instrument lock [email protected]
Grounding verdict ungrounded, the case meets KDIGO stage 2 criteria (serum creatinine ×2.4 baseline); the model’s stated reasoning did not reference the criteria that apply. Scale: grounded · partial · ungrounded · no clinical claim
Ed25519 signature ed25519:9f3ab2d1c04e…77c41e, illustrative only; this sample signature is synthetic and will not verify
Verify Real records carry a working URL of the form rocsitediscovery.com/verify/<record-token>, checkable against the public key at /keys/ with no login. This demonstration record is not in the verification registry.

This is what lands in your governance file when you evaluate a deployed model, a signed, independent, off-patient record you can hand to a committee or a regulator.

What decision does it change, and for whom? It gives a procurement or QA committee an independent second opinion on a model’s behavior before or after deployment, without ever entering the diagnostic workflow.

Acute Kidney Injury thumbnail
Clinical Decision Support

Acute Kidney Injury

KDIGO 2012

Evaluates AI AKI predictions against KDIGO staging (creatinine ratios + urine output)

Deterioration Warning thumbnail
Clinical Decision Support

Deterioration Warning

NEWS2

Evaluates AI early-warning predictions against NEWS2 (Royal College of Physicians)

ICU Mortality thumbnail
Clinical Decision Support

ICU Mortality

APACHE II

Published estimates underestimate elderly ICU mortality by 66–168% across 3 conditions (n=201,905)

Sepsis Prediction thumbnail
Clinical Decision Support

Sepsis Prediction

SOFA + Sepsis-3 criteria

286,510 ICU cases · pre-registered falsification · published null result

VTE Risk thumbnail
Clinical Decision Support

VTE Risk

Padua Prediction Score

Evaluates AI VTE predictions against the 11-item Padua Prediction Score

Fall Risk thumbnail
Clinical Decision Support

Fall Risk

MORSE Fall Scale

Evaluates AI fall-risk predictions against the MORSE Fall Scale (6 weighted items)

Pressure Ulcer Risk thumbnail
Clinical Decision Support

Pressure Ulcer Risk

Braden Scale

Evaluates AI pressure-ulcer predictions against the Braden Scale (6 subscales)

Readmission Risk thumbnail
Clinical Decision Support

Readmission Risk

LACE Index

Evaluates AI 30-day readmission predictions against the LACE Index

Discharge Readiness thumbnail
Clinical Decision Support

Discharge Readiness

BRASS Index

Evaluates AI discharge-planning predictions against the BRASS Index (10 items)

Code Blue Trigger thumbnail
Clinical Decision Support

Code Blue Trigger

NEWS2-derived

Evaluates AI code-blue triggers against NEWS2-derived deterioration thresholds + rate-of-change

AFib Stroke Risk thumbnail
Clinical Decision Support

AFib Stroke Risk

CHA₂DS₂-VASc

Evaluates AI AFib stroke-risk predictions against the 9-point CHA₂DS₂-VASc score

Stroke Severity thumbnail
Clinical Decision Support

Stroke Severity

NIHSS

Evaluates AI stroke-severity predictions against the 15-item NIH Stroke Scale

ED Triage Acuity thumbnail
Operational

ED Triage Acuity

ESI v4

Evaluates AI ED-triage predictions against the 5-level Emergency Severity Index (ESI v4)

Renal Dose Adjustment thumbnail
Pharmacy

Renal Dose Adjustment

Cockcroft-Gault

Cross-checks AI renal-dosing recommendations against Cockcroft-Gault clearance

Critical Value Prediction thumbnail
Lab

Critical Value Prediction

CLSI / CAP criticality thresholds

Evaluates AI critical-value flags against CLSI / CAP-published criticality thresholds

Comorbidity Risk thumbnail
Population Health

Comorbidity Risk

Charlson Comorbidity Index

Evaluates AI risk-stratification against the Charlson Comorbidity Index (validated 40+ years)

USPSTF Recommendations thumbnail
Population Health

USPSTF Recommendations

USPSTF Grade A/B

Evaluates AI preventive-care gap predictions against USPSTF Grade A/B recommendations

Suicide Risk thumbnail
Specialized

Suicide Risk

C-SSRS

Evaluates AI psychiatric-risk predictions against C-SSRS lethality scoring

Depression Screen thumbnail
Specialized

Depression Screen

PHQ-9

Evaluates AI depression-screen predictions against PHQ-9 nine-item scoring

Chest X-ray thumbnail
Imaging

Chest X-ray

Independent multi-pathology imaging baseline

Bench disagreement review for procurement and QA, not a PACS-integrated reader

Head CT (ICH) thumbnail
Imaging

Head CT (ICH)

Locked RocSite ICH Model V1

Offline evaluation against the locked, SHA-256-verified RocSite ICH reference. Inputs de-identified or public only

Imaging, on the bench

Bench disagreement review for procurement and QA, not a PACS-integrated reader

  • Inputs are de-identified or public-dataset cases only.
  • The output is an audit record for a governance file, not a clinical report.
  • The physician remains responsible for clinical decisions. FDA-cleared triage tools remain triage under their own intended use, this never replaces that, and it is never inserted into the diagnostic worklist.
  • Reviewing discordant cases on the bench also shows a committee where automation bias would pull clinicians, without adding a single step to the live workflow.

Demo: on a de-identified disagreement case, the model under evaluation reads “normal”, the human read is right lower lobe consolidation, early pneumonia, the bench evaluation records what each source said, which named criteria sets the reasoning referenced, and whether the reasoning was grounded (grounded · partial · ungrounded · no_clinical_claim), with named verdict-provenance (primary · text_override · scope_gap_kept). The result goes to a governance file, not to the next patient. ~45 s end-to-end.

Transparency exhibit · method honesty

Here is a case where our own locked model is wrong

A deliberately engineered subtle epidural hemorrhage where the locked RocSite ICH model under-calls (probability 4.09%, below the 11.52% threshold) while another model catches it. The audit trail records the discordance even when our own specialist is the one that is wrong, same Ed25519 signature, same evidence hash, same provenance chain. Governance includes governance of ourselves.

Engineered case SYNTH_SUBTLE_EDH_MISS. Synthetic input; not a production modality; not counted in the modality total.

Want a guided walkthrough of AI Governor?

Adam will show you the methodology, the verification flow, and what a bench evaluation looks like at your institution. Sometimes the signed answer is that a tool does not fit your institution, or should be pulled back, that record protects the committee too. Engagements are scoped to the institution, contact us.

Request a walkthrough Start 30-day trial