Our Standard

How we rate AI risk — transparently, and defensibly.

The Attestor Generative and Agentic AI Assurance Standard turns a complex AI system into a single, defensible risk rating. Its defining feature is transparency: every rating is reproducible and traceable to specific components — never a black-box number.

The method

A single, defensible rating in three steps.

1

Decompose the system into its components

Each AI system is mapped as a whole and every component is classified by its function — the reasoning model, retrieval, action tools, guardrails, human checkpoints, and controls. Governance follows function, so the same discipline applies whether a component is built in-house or bought from a vendor.

ComponentRole in the score
OrchestratorThe reasoning model — sets the base risk level.
Retrieval / RAGGrounds the model in data — adds complexity and can raise data sensitivity.
Action toolsWhat the system can do in the world — drive autonomy and complexity.
Guardrails & rulesBright-line and model-based limits — reduce risk when present and tested.
Human-in-the-loopA human checkpoint — caps autonomy and earns credited oversight.
Monitoring, kill-switch, evidence vaultDetective and preventive controls — credited when operational.
2

Compute a residual-risk score

Components drive a structured score that captures the system’s inherent risk, its autonomy, and the quality of its controls:

Residual Risk = [ (Base + Data Sensitivity + Complexity) × Autonomy ] − Controls
Base and Autonomy use a high-water-mark rule across components; a floor keeps the score realistic.
InputWhat it captures
BaseThe inherent risk of the system’s reasoning — from internal-only assistance to autonomous, consequential decisions.
Data sensitivityThe sensitivity of the data it can access — from anonymized to customer PII to material non-public information.
ComplexityThe attack surface — how many tools and data sources the system connects to.
AutonomyHow far it can act without a human — read-only, human-gated, or fully autonomous execution.
ControlsCredited reductions for controls that are present and tested — rules, guardrails, human oversight, monitoring, kill-switches, and immutable evidence.
One standard, the full spectrum. A generative assistant scores at the low-autonomy end of this scale; an autonomous agent scores at the high end. The same standard rates both — so an institution governs its assistants and its agents with one consistent method, not two.

Because the score is built this way, it is defensible. An institution — or its examiner — can see exactly why a system scored as it did, and exactly which control changes would move it. There is no arbitrary judgment to explain away.

3

Assign a risk tier that drives governance

The residual score maps to a tier that sets how often the system is validated and how it is reported — so oversight is proportionate to risk.

Risk tierValidation cadenceReporting
LowAnnualAnnual summary
MediumBi-annualSemi-annual status
HighQuarterlyQuarterly dashboard; CRO escalation
CriticalMonthlyBoard Risk Committee; sign-off to operate
UnacceptableNot deployable as designedRedesign required before deployment

Regulatory alignment

Mapped to the frameworks examiners use.

The standard is built to satisfy the intent of SR 26-2 and to align with the NIST AI Risk Management Framework, so the output slots directly into an institution’s existing model risk governance.

NIST AI RMF functionHow the standard addresses it
GOVERNRoles, review cadence, and sign-off thresholds tied to the risk tier.
MAPSystem decomposition and component classification — context and categorization.
MEASUREThe residual-risk score and assembled-system testing — evaluation of trustworthiness.
MANAGEOngoing monitoring, material-change triggers, and tiered escalation and reporting.

See the standard applied to your AI.

A short scoping conversation lets us decompose one of your systems, score it, and show you the evidence the standard produces.

Start a conversation