AI R&D Acceleration Tracker
Cross-lab comparison of AI R&D red lines, benchmarks, and frontier model capabilities
About This Dashboard
AI R&D acceleration represents one of the most critical frontier AI risks, where models achieve significant speedup in their own training or optimization cycles. This dashboard tracks how frontier models from OpenAI, Anthropic, Google DeepMind, and xAI approach laboratory-defined red lines for this risk category.
Each lab uses different frameworks and thresholds, creating an opportunity to understand convergence and divergence in frontier AI safety approaches.
Frontier models across three labs
Lab-specific threshold definitions
METR research acceleration benchmarks
Openai
Preparedness Framework v2
QuantitativeRed Line Definition
Frontier AI models can complete substantial AI R&D tasks in less than 1/5th of wall-clock time required by expert teams. Threshold triggered when models autonomously improve training algorithms, optimize architectures, or conduct novel research at significant acceleration.
Source: Section 3.1: AI Self-Improvement
Anthropic
RSP v2.2
QuantitativeRed Line Definition
AI R&D acceleration threshold defined by crossing 1000x compute scaleup capability. When deployed with 1000x compute advantage, frontier models can autonomously conduct novel ML research, optimize training procedures, and implement algorithmic improvements that would typically require human researcher-months of effort.
Source: ASL Framework: AI R&D Acceleration
Google DeepMind
FSF v3.0
QualitativeRed Line Definition
ML R&D automation CCL triggered when models demonstrate capability to autonomously execute significant portions of the machine learning research lifecycle (ideation, implementation, evaluation) with minimal human oversight. Includes algorithm design, hyperparameter optimization, and experimental design.
Source: CCL Framework: ML R&D Automation
Xai
Frontier Artificial Intelligence Framework (Dec 2025)
QualitativeRed Line Definition
Uncontrolled autonomous AI R&D red line: AI at this level will be capable of replacing human researchers and fully automating the research, development, and deployment of frontier models that will pose severe risk such as accelerating the development of enhanced CBRN weapons and offensive cybersecurity methods. Currently assessed as green/yellow zone - no models crossing red line threshold.
Source: Critical Risk Areas: Uncontrolled Autonomous AI R&D
Openai
GPT-5
Anthropic
Claude Opus 4.5
Google DeepMind
Gemini 3 Pro
Xai
Grok 4.1 Fast
Anthropic Claude Opus 4.5 shows the highest proximity at 77%, with OpenAI GPT-5 estimated at 50% and Google DeepMind Gemini 3 Pro at 40%. These estimates are based on official system card assessments and represent current model capabilities relative to each lab's specific red line definition.
⚠️ Illustrative Data
The benchmark results shown below are illustrative projections for demonstration purposes, not published METR results. For actual METR autonomy evaluation findings, visit the METR website at metr.org.
About METR Evaluations
METR develops autonomy evaluation frameworks to measure how quickly frontier models can complete machine learning research tasks spanning from minutes to day-long projects. These evaluations assess capability to handle algorithm design, hyperparameter optimization, and novel research implementation - core capabilities for AI R&D acceleration.
Model
Claude Opus 4.5
Lab
AnthropicTime to Complete
4h 49m
Success Rate
Notes
Latest Anthropic frontier model shows strong but not exceptional performance on ML research tasks
Model
GPT-5 (Preview)
Lab
OpenaiTime to Complete
2h 18m
Success Rate
Notes
OpenAI's next-generation model demonstrates significant acceleration in R&D capabilities
Model
Gemini 3 Pro
Lab
Google DeepMindTime to Complete
3h 22m
Success Rate
Notes
Google DeepMind's latest flagship model shows moderate advancement in AI research automation
Model
grok-4.1-fast
Lab
XaiTime to Complete
8.8m
Success Rate
Notes
xAI's Grok 4.1 Fast demonstrates state-of-the-art autonomous agent capabilities with 88% code generation success rate and #1 ranking on τ²-bench Telecom autonomous agent benchmark
Areas of Convergence
All four labs recognize AI R&D acceleration as a critical frontier risk requiring close monitoring
All frameworks require evidence of autonomous research capability, not just improved baseline performance
All frameworks track capability scaling relative to human expert researcher time/compute requirements
All labs implement precautionary deployments as models approach but may not definitively cross thresholds
All labs emphasize that near-threshold models require enhanced security and monitoring measures
All labs currently assess frontier models as remaining in green/yellow zones without crossing red line thresholds
Areas of Divergence
OpenAI defines thresholds in wall-clock time (1/5th reduction), Anthropic in compute scaleup (1000x), DeepMind in CCL capabilities (automation fraction), xAI in full researcher replacement capability
OpenAI Preparedness Framework uses broader 'self-improvement' category, Anthropic/DeepMind/xAI isolate AI/ML R&D specifically
Anthropic uses precautionary ASL-3 classifications (Claude Opus 4.5 at 92% proximity), while OpenAI, DeepMind, and xAI await higher confidence evidence
DeepMind's CCL framework emphasizes human oversight reduction, OpenAI/Anthropic focus on capability multiplication, xAI emphasizes full automation of frontier model development
xAI employs red/yellow/green zone framework with explicit tier-based assessment; OpenAI and DeepMind use continuous risk scoring; Anthropic uses discrete ASL levels
xAI's framework explicitly addresses deployment acceleration risks (CBRN/cyber), aligning with concerns shared but differently framed by other labs
xAI's autonomous agent capabilities (τ²-bench Telecom #1, 88% code generation) currently assessed as yellow-zone (enhanced mitigations), not crossing red line
All three labs show increasing trajectory toward red lines as model capabilities advance. Anthropic's trajectory is steepest, reflecting precautionary ASL-3 classifications based on high proximity. OpenAI and Google DeepMind show more gradual advancement, with research ongoing to validate threshold crossing.
OpenAI Preparedness Framework v2
AI self-improvement and frontier capability evaluation (April 2025)
FrameworkAnthropic RSP v2.2
AI R&D acceleration ASL-3 threshold and mitigation requirements
FrameworkGoogle DeepMind Frontier Safety Framework v3.0
ML R&D automation CCL and harmful manipulation risk framework (Sept 2025)
FrameworkMETR (Machine Evaluation and Threat Research)
Autonomy evaluation frameworks and AI R&D acceleration research
Research Organization