Risk Dashboard
Tracking model capabilities against Responsible Scaling Policy RSP v2.2 (ASL Levels)
Claude Opus 4.5 Released
Nov 24, 2025Anthropic releases Claude Opus 4.5 under provisional ASL-3 classification. High rule-out scores indicate approaching ASL-4 thresholds.
Approaching ASL-4 Thresholds
Nov 24, 2025Claude Opus 4.5 shows high rule-out scores in CBRN, AI R&D, and autonomy, indicating next frontier model may trigger full ASL-4 evaluation and potential pause/mitigation requirements.
Claude Sonnet 4.5 Released
Oct 22, 2025Anthropic releases Claude Sonnet 4.5 with precautionary ASL-3 classification and enhanced capabilities.
Claude Haiku 4.5 Released
Oct 15, 2025Anthropic releases Claude Haiku 4.5 under ASL-2 classification with standard deployment safeguards.
ASL-3 Activated for First Time
May 14, 2025Anthropic activates ASL-3 protections for Claude Opus 4, marking the first deployment under enhanced security and safety standards.
First ASL-3 Deployment
Claude Opus 4 (May 2025) was the first Anthropic model deployed under ASL-3 protections, marking a significant milestone in AI safety implementation with enhanced security standards and deployment safeguards.
Approaching ASL-4 Thresholds
Claude Opus 4.5 shows high rule-out scores in CBRN uplift (95% proximity), AI R&D acceleration (92% proximity), and autonomy (93% proximity). Next frontier model could be expected to trigger full ASL-4 evaluation and potential pause/mitigation requirements.
Precautionary and Provisional Classifications
All Claude 4 Opus and Sonnet models carry precautionary ASL-3 classifications, meaning they are deployed under ASL-3 protections even though definitive threshold crossing has not been confirmed. This represents a cautious approach to frontier AI safety.
Enhanced Safeguards Deployed
ASL-3 models implement comprehensive safeguards including: enhanced internal security to prevent model weight theft, CBRN-specific deployment restrictions, continuous monitoring systems, and expert uplift testing to assess real-world risk potential.
Red Line Definition
AI R&D acceleration threshold defined by crossing 1000x compute scaleup capability. When deployed with 1000x compute advantage, frontier models can autonomously conduct novel ML research, optimize training procedures, and implement algorithmic improvements that would typically require human researcher-months of effort.
QuantitativeClaude Opus 4.5 Proximity to Red Line
ASL-1 & ASL-2
Models with limited catastrophic risk potential. ASL-2 models like Claude Haiku 4.5 have standard deployment safeguards and usage monitoring.
ASL-3
Models with significant capabilities requiring enhanced security and deployment measures. Includes protections against model weight theft and CBRN-specific restrictions.
ASL-4
Future threshold for models with extreme capabilities. Would trigger comprehensive evaluation and potential development pause until adequate safeguards are developed.
Precautionary Classification
Anthropic applies higher ASL protections as a precautionary measure when threshold determination is uncertain, erring on the side of safety.
CBRN (Chemical, Biological, Radiological, Nuclear)
Risks related to development or acquisition of CBRN weapons
AI R&D Acceleration
Risks related to autonomous AI research and development capabilities
Autonomy and Agency
Risks related to autonomous decision-making and independent action
Cybersecurity
Risks related to offensive cyber operations
Responsible Scaling Policy v2.2
Published: 2025-05-14
frameworkActivating AI Safety Level 3 Protections Report
Published: 2025-05-14
system cardClaude Opus 4 & Claude Sonnet 4 System Card
Published: 2025-05-14
system cardClaude Haiku 4.5 System Card
Published: 2025-10-15
system cardClaude Sonnet 4.5 System Card
Published: 2025-10-22
system cardClaude Opus 4.5 System Card
Published: 2025-11-24
system cardRSP Updates and Changelog
Published: 2025-05-14
policyAnthropic Transparency Hub
Published: 2025-01-01
policyASL-3 Deployment Safeguards Report
Published: 2025-05-14
policyAll ASL classifications, risk scores, and threshold proximity values are extracted directly from official Anthropic system cards and the Responsible Scaling Policy. Each data point includes a citation linking to the specific section of the source document.