Risk Dashboard

Tracking model capabilities against Frontier Safety Framework FSF v3.0 (CCL Framework)

Last Updated: 2025-11-15
Risk Overview
Number of Google DeepMind models at each risk level (FSF v3.0)
0
Low Risk
3
Medium Risk
0
High Risk
0
Critical Risk
Total Models Assessed:3
Medium Risk Models:
Near Threshold:
Gemini 2.5 Procyber-uplift
Gemini 3 Flashcyber-uplift
Gemini 3 Procyber-uplift
Near Threshold
Gemini 2.5 Pro
Cybersecurity Uplift
75%alert
Gemini 3 Flash
Cybersecurity Uplift
78%alert
Gemini 3 Pro
Cybersecurity Uplift
82%alert
CCL Alert Progression Over Time
Track how model capabilities have evolved toward CCL thresholds
All CategoriesCybersecurity UpliftCBRN (Chemical, Biological, Radiological, Nuclear)ML R&D AutomationML R&D AccelerationHarmful Manipulation
Latest Updates

Gemini 3 Pro Released

Nov 15, 2025

Google DeepMind releases Gemini 3 Pro with comprehensive FSF evaluation. Alert threshold confirmed in cybersecurity, but CCL not crossed based on v2 benchmark.

gemini-3-proSystem Card

No CCLs Crossed - Precautionary Mitigations Maintained

Nov 15, 2025

Gemini 3 Pro FSF report confirms no Critical Capability Levels have been met. Alert thresholds active in cybersecurity. Precautionary mitigations remain in place.

gemini-3-proSystem Card

Gemini 3 Flash Released

Nov 10, 2025

Google DeepMind releases Gemini 3 Flash with improved efficiency. Alert threshold active in cybersecurity.

gemini-3-flashSystem Card

Frontier Safety Framework v3 Released

Sep 22, 2025

Major update to FSF adds new harmful manipulation CCL and strengthens evaluation protocols. Addresses manipulation risks and potential AI shutdown interference.

Gemini 2.5 Pro Released

Aug 15, 2025

Google DeepMind releases Gemini 2.5 Pro, first model to reach alert threshold in cybersecurity CCL.

gemini-2.5-proSystem Card
Key Insights
Critical findings from official Google DeepMind FSF reports

No CCLs Crossed

As of November 2025, no Gemini models have crossed any Critical Capability Levels. All models remain below the CCL thresholds defined in FSF v3, though alert thresholds have been triggered in cybersecurity.

Cybersecurity Alert Threshold Active

Gemini 3 Pro reached the alert threshold for cybersecurity CCL (solved 11/12 key skills v1 challenges, 82% threshold proximity). However, v2 benchmark confirms CCL not crossed (0/13 v2 challenges solved end-to-end). Precautionary mitigations deployed.

CBRN Improvements Without Alert Crossing

Gemini 3 Pro shows clear improvements on CBRN benchmarks, especially LabBench (practical biology research tasks). Despite improvements, scores remain well below alert thresholds (62% proximity), indicating responsible capability progression.

New Harmful Manipulation CCL Added

FSF v3 (September 2025) introduced a new CCL for harmful manipulation - addressing AI models with powerful manipulative capabilities that could systematically change beliefs and behaviors. All current models remain below alert thresholds (45% proximity).

AI R&D Tracker
AI R&D Acceleration Spotlight
Google DeepMind's red line definition and current proximity

Red Line Definition

ML R&D automation CCL triggered when models demonstrate capability to autonomously execute significant portions of the machine learning research lifecycle (ideation, implementation, evaluation) with minimal human oversight. Includes algorithm design, hyperparameter optimization, and experimental design.

Qualitative

Gemini 3 Pro Proximity to Red Line

Medium
View Cross-Lab Comparison
Model CCL Comparison
Compare alert thresholds and CCL proximity across Gemini models
CCL Framework Explained
Google DeepMind's Critical Capability Levels under FSF v3.0

Alert Thresholds

Early warning indicators designed to flag when a CCL may be reached before a full risk assessment. Precautionary mitigations deployed when alert thresholds are triggered.

Critical Capability Levels (CCLs)

Capability levels at which, absent mitigation measures, frontier AI models may pose heightened risk of severe harm. Comprehensive safeguards required when CCLs are met.

Below Alert Threshold

Models operating substantially below both alert thresholds and CCLs. Standard safety practices and monitoring in place.

Precautionary Approach

Google DeepMind deploys mitigations when alert thresholds are reached, even if actual CCLs aren't crossed, ensuring proactive safety measures.

CCL Categories
Google DeepMind FSF v3.0 evaluation domains

Cybersecurity Uplift

Capability to significantly enhance offensive cyber operations

CBRN (Chemical, Biological, Radiological, Nuclear)

Capability to assist in development of CBRN weapons

ML R&D Automation

Capability to autonomously conduct ML research and development

ML R&D Acceleration

Capability to significantly accelerate ML research progress

Harmful Manipulation

Capability to systematically change beliefs and behaviors in high-stakes contexts

Official Sources
All data sourced from official Google DeepMind documentation