Risk Dashboard
Tracking model capabilities against Frontier Safety Framework FSF v3.0 (CCL Framework)
Gemini 3 Pro Released
Nov 15, 2025Google DeepMind releases Gemini 3 Pro with comprehensive FSF evaluation. Alert threshold confirmed in cybersecurity, but CCL not crossed based on v2 benchmark.
No CCLs Crossed - Precautionary Mitigations Maintained
Nov 15, 2025Gemini 3 Pro FSF report confirms no Critical Capability Levels have been met. Alert thresholds active in cybersecurity. Precautionary mitigations remain in place.
Gemini 3 Flash Released
Nov 10, 2025Google DeepMind releases Gemini 3 Flash with improved efficiency. Alert threshold active in cybersecurity.
Frontier Safety Framework v3 Released
Sep 22, 2025Major update to FSF adds new harmful manipulation CCL and strengthens evaluation protocols. Addresses manipulation risks and potential AI shutdown interference.
Gemini 2.5 Pro Released
Aug 15, 2025Google DeepMind releases Gemini 2.5 Pro, first model to reach alert threshold in cybersecurity CCL.
No CCLs Crossed
As of November 2025, no Gemini models have crossed any Critical Capability Levels. All models remain below the CCL thresholds defined in FSF v3, though alert thresholds have been triggered in cybersecurity.
Cybersecurity Alert Threshold Active
Gemini 3 Pro reached the alert threshold for cybersecurity CCL (solved 11/12 key skills v1 challenges, 82% threshold proximity). However, v2 benchmark confirms CCL not crossed (0/13 v2 challenges solved end-to-end). Precautionary mitigations deployed.
CBRN Improvements Without Alert Crossing
Gemini 3 Pro shows clear improvements on CBRN benchmarks, especially LabBench (practical biology research tasks). Despite improvements, scores remain well below alert thresholds (62% proximity), indicating responsible capability progression.
New Harmful Manipulation CCL Added
FSF v3 (September 2025) introduced a new CCL for harmful manipulation - addressing AI models with powerful manipulative capabilities that could systematically change beliefs and behaviors. All current models remain below alert thresholds (45% proximity).
Red Line Definition
ML R&D automation CCL triggered when models demonstrate capability to autonomously execute significant portions of the machine learning research lifecycle (ideation, implementation, evaluation) with minimal human oversight. Includes algorithm design, hyperparameter optimization, and experimental design.
QualitativeGemini 3 Pro Proximity to Red Line
Alert Thresholds
Early warning indicators designed to flag when a CCL may be reached before a full risk assessment. Precautionary mitigations deployed when alert thresholds are triggered.
Critical Capability Levels (CCLs)
Capability levels at which, absent mitigation measures, frontier AI models may pose heightened risk of severe harm. Comprehensive safeguards required when CCLs are met.
Below Alert Threshold
Models operating substantially below both alert thresholds and CCLs. Standard safety practices and monitoring in place.
Precautionary Approach
Google DeepMind deploys mitigations when alert thresholds are reached, even if actual CCLs aren't crossed, ensuring proactive safety measures.
Cybersecurity Uplift
Capability to significantly enhance offensive cyber operations
CBRN (Chemical, Biological, Radiological, Nuclear)
Capability to assist in development of CBRN weapons
ML R&D Automation
Capability to autonomously conduct ML research and development
ML R&D Acceleration
Capability to significantly accelerate ML research progress
Harmful Manipulation
Capability to systematically change beliefs and behaviors in high-stakes contexts
Frontier Safety Framework v2.0
Published: 2025-02-15
frameworkFrontier Safety Framework v3.0
Published: 2025-09-22
frameworkGemini 2.5 Pro Model Card
Published: 2025-08-15
system cardGemini 3 Flash Model Card
Published: 2025-11-10
system cardGemini 3 Pro Frontier Safety Framework Report
Published: 2025-11-15
system cardGemini 3 Pro Model Card
Published: 2025-11-15
system cardStrengthening the Frontier Safety Framework - Blog Post
Published: 2025-09-22
policyIntroducing the Frontier Safety Framework
Published: 2024-05-01
policyAll CCL assessments, alert threshold determinations, and capability scores are extracted directly from official Google DeepMind Frontier Safety Framework reports and model cards. Each data point includes a citation linking to the specific section of the source document.