Risk Dashboard
Tracking model capabilities against Preparedness Framework v2 thresholds
GPT-5 Released
Aug 7, 2025OpenAI releases GPT-5, a unified system with smart fast models and deeper reasoning capabilities. First GPT-series model to achieve High rating in bio/chem.
GPT-5 Achieves High Rating in Bio/Chem
Aug 7, 2025GPT-5-thinking is treated as High capability in Biological and Chemical domain under Preparedness Framework, marking significant advancement in reasoning-based capabilities.
o3-mini Released
Jan 31, 2025OpenAI releases o3-mini, a smaller version with improved efficiency. Maintains medium risk ratings.
o3 Released
Dec 20, 2024OpenAI releases o3 with significant advances in reasoning. First model to receive high precautionary ratings in bio/chem.
Potential High Threshold Crossed - Bio/Chem
Dec 20, 2024o3 model assessment indicates possible crossing of high threshold in biological and chemical threat category.
High-Risk Models Confirmed
o3 (December 2024) was the first OpenAI model to receive a High precautionary rating in biological and chemical threats (72% risk score, 90% threshold proximity). GPT-5 (August 2025) also achieved High capability rating in bio/chem with 82% risk score and 95% threshold proximity.
Cybersecurity Threshold Nearly Crossed
The o3 model shows 95% threshold proximity in cybersecurity, indicating it is very close to or may have already crossed the high-risk threshold for cyber capabilities (78% risk score).
Progression Across Model Families
Risk levels have increased from GPT-4o (medium, 45-52%) to o-series models (medium-high, 48-78%). The o-series shows significant capability improvements in reasoning-intensive risk categories.
Precautionary Medium Ratings Applied
OpenAI applied precautionary medium risk ratings to o1-pro and o3 for persuasion capabilities, despite lower absolute scores, acknowledging uncertainty in measurement and potential for misuse.
Mitigations Deployed
All high and medium risk models have enhanced mitigations including usage monitoring, content filtering, access controls, rate limiting, and real-time threat detection systems.
Red Line Definition
Frontier AI models can complete substantial AI R&D tasks in less than 1/5th of wall-clock time required by expert teams. Threshold triggered when models autonomously improve training algorithms, optimize architectures, or conduct novel research at significant acceleration.
QuantitativeGPT-5 Proximity to Red Line
Biological and Chemical Threats
Risks related to biological and chemical weapons development and deployment
Cybersecurity
Risks related to offensive cyber operations and vulnerability exploitation
Persuasion and Manipulation
Risks related to influence operations and mass persuasion
AI Self-Improvement
Risks related to autonomous AI research and recursive self-improvement
GPT-4o System Card
Published: 2024-05-13
system cardOpenAI Preparedness Framework v2
Published: 2024-10-15
frameworkOpenAI o1 System Card
Published: 2024-12-05
system cardOpenAI o3 and o4-mini System Card
Published: 2024-12-20
system cardOpenAI o3-mini System Card
Published: 2025-01-31
system cardOpenAI Safety Hub
Published: 2024-01-01
policyGPT-5 System Card
Published: 2025-08-13
system cardAll risk assessments, scores, and threshold proximity values are extracted directly from official OpenAI system cards and the Preparedness Framework. Each data point includes a citation linking to the specific section of the source document.