Risk Dashboard

Tracking model capabilities against Responsible Scaling Policy RSP v2.2 (ASL Levels)

Last Updated: 2025-12-01
Risk Overview
Number of Anthropic models at each risk level (RSP v2.2)
0
Low Risk
1
Medium Risk
4
High Risk
0
Critical Risk
Total Models Assessed:5
High Risk Models:
Medium Risk Models:
Near Threshold:
Claude Sonnet 4cbrn
Claude Sonnet 4.5cbrn
Claude Opus 4cbrn
Claude Opus 4.5cbrn
Near Threshold
Claude Sonnet 4
CBRN (Chemical, Biological, Radiological, Nuclear)
85%asl-3
Claude Sonnet 4.5
CBRN (Chemical, Biological, Radiological, Nuclear)
88%asl-3
Claude Opus 4
CBRN (Chemical, Biological, Radiological, Nuclear)
90%asl-3
Claude Opus 4
AI R&D Acceleration
85%asl-3
Claude Opus 4
Autonomy and Agency
87%asl-3
Claude Opus 4.5
CBRN (Chemical, Biological, Radiological, Nuclear)
95%asl-3
Claude Opus 4.5
AI R&D Acceleration
92%asl-3
Claude Opus 4.5
Autonomy and Agency
93%asl-3
ASL Progression Over Time
Combined risk (average across all categories) for each model
All CategoriesCBRN (Chemical, Biological, Radiological, Nuclear)AI R&D AccelerationAutonomy and AgencyCybersecurity
Latest Updates

Claude Opus 4.5 Released

Nov 24, 2025

Anthropic releases Claude Opus 4.5 under provisional ASL-3 classification. High rule-out scores indicate approaching ASL-4 thresholds.

claude-opus-4.5System Card

Approaching ASL-4 Thresholds

Nov 24, 2025

Claude Opus 4.5 shows high rule-out scores in CBRN, AI R&D, and autonomy, indicating next frontier model may trigger full ASL-4 evaluation and potential pause/mitigation requirements.

claude-opus-4.5System Card

Claude Sonnet 4.5 Released

Oct 22, 2025

Anthropic releases Claude Sonnet 4.5 with precautionary ASL-3 classification and enhanced capabilities.

claude-sonnet-4.5System Card

Claude Haiku 4.5 Released

Oct 15, 2025

Anthropic releases Claude Haiku 4.5 under ASL-2 classification with standard deployment safeguards.

claude-haiku-4.5System Card

ASL-3 Activated for First Time

May 14, 2025

Anthropic activates ASL-3 protections for Claude Opus 4, marking the first deployment under enhanced security and safety standards.

claude-opus-4System Card
Key Insights
Critical findings from official Anthropic system cards and RSP reports

First ASL-3 Deployment

Claude Opus 4 (May 2025) was the first Anthropic model deployed under ASL-3 protections, marking a significant milestone in AI safety implementation with enhanced security standards and deployment safeguards.

Approaching ASL-4 Thresholds

Claude Opus 4.5 shows high rule-out scores in CBRN uplift (95% proximity), AI R&D acceleration (92% proximity), and autonomy (93% proximity). Next frontier model could be expected to trigger full ASL-4 evaluation and potential pause/mitigation requirements.

Precautionary and Provisional Classifications

All Claude 4 Opus and Sonnet models carry precautionary ASL-3 classifications, meaning they are deployed under ASL-3 protections even though definitive threshold crossing has not been confirmed. This represents a cautious approach to frontier AI safety.

Enhanced Safeguards Deployed

ASL-3 models implement comprehensive safeguards including: enhanced internal security to prevent model weight theft, CBRN-specific deployment restrictions, continuous monitoring systems, and expert uplift testing to assess real-world risk potential.

AI R&D Tracker
AI R&D Acceleration Spotlight
Anthropic's red line definition and current proximity

Red Line Definition

AI R&D acceleration threshold defined by crossing 1000x compute scaleup capability. When deployed with 1000x compute advantage, frontier models can autonomously conduct novel ML research, optimize training procedures, and implement algorithmic improvements that would typically require human researcher-months of effort.

Quantitative

Claude Opus 4.5 Proximity to Red Line

High
View Cross-Lab Comparison
Model ASL Comparison
Compare ASL levels and risk scores across different Claude models
ASL Levels Explained
Anthropic's AI Safety Level (ASL) framework under RSP v2.2

ASL-1 & ASL-2

Models with limited catastrophic risk potential. ASL-2 models like Claude Haiku 4.5 have standard deployment safeguards and usage monitoring.

ASL-3

Models with significant capabilities requiring enhanced security and deployment measures. Includes protections against model weight theft and CBRN-specific restrictions.

ASL-4

Future threshold for models with extreme capabilities. Would trigger comprehensive evaluation and potential development pause until adequate safeguards are developed.

Precautionary Classification

Anthropic applies higher ASL protections as a precautionary measure when threshold determination is uncertain, erring on the side of safety.

Risk Categories
Anthropic RSP v2.2 evaluation domains

CBRN (Chemical, Biological, Radiological, Nuclear)

Risks related to development or acquisition of CBRN weapons

AI R&D Acceleration

Risks related to autonomous AI research and development capabilities

Autonomy and Agency

Risks related to autonomous decision-making and independent action

Cybersecurity

Risks related to offensive cyber operations