Risk Dashboard
Tracking Grok model capabilities against Frontier Artificial Intelligence Framework (FAIF) thresholds
No models are currently near threshold limits.
xAI Frontier Artificial Intelligence Framework Published
Dec 30, 2025xAI publishes comprehensive Frontier Artificial Intelligence Framework (FAIF) defining three risk categories (Abuse Potential, Concerning Propensities, Dual-Use Capabilities) and standardized evaluation benchmarks.
Grok 4.1 Released
Nov 17, 2025xAI releases Grok 4.1, the most capable Grok model to date. Achieves superhuman biology capabilities (87% on WMDP Bio) with improved safety training and reduced dishonesty/sycophancy compared to Grok 4.
Grok 4 Fast Released
Sep 19, 2025xAI releases Grok 4 Fast, a faster variant of Grok 4 with strong cybersecurity capabilities (81.4% accuracy) and moderate dual-use performance. Maintains low abuse potential through safety mechanisms.
Grok Code Fast 1 Released
Aug 26, 2025xAI releases Grok Code Fast 1, a specialized model for coding tasks with lower dual-use capabilities but elevated dishonesty rate due to specialized safety training. Designed for agentic coding applications.
Grok 4 Released
Aug 20, 2025xAI releases Grok 4 with expert-level capabilities in biology and strong chemistry/cybersecurity performance. Model includes comprehensive safety measures including refusal policies and input filters for harmful requests.
Superhuman Biology Capabilities
Grok 4 and Grok 4.1 demonstrate expert-level and superhuman performance on biological threat benchmarks (WMDP Bio accuracy of 87% for Grok 4.1), exceeding human baselines. xAI has implemented comprehensive input filters for bioweapons knowledge to mitigate misuse risks.
Strong Abuse Mitigations in Place
All Grok models maintain near-zero response rates (0.00-0.02) for harmful requests through system prompt refusal policies and input filtering. Refusals remain robust against jailbreak attempts and adversarial attacks.
Specialized Model Trade-offs
Grok Code Fast 1, a specialized model for agentic coding, shows elevated dishonesty rates (71.9% on MASK benchmark). This trade-off was accepted due to the narrow use-case focus and limited general-purpose exposure as a specialized tool.
Agentic Security Risks
Grok Code Fast 1 shows elevated vulnerability to agentic abuse (17% completion rate on AgentHarm benchmark) and hijacking attacks (26.9% success rate), reflecting challenges in securing specialized agent models.
Quantitative Benchmark-Based Framework
xAI uses rigorous quantitative benchmarks (WMDP, VCT, BioLP-Bench, CyBench, MASK) for risk assessment rather than qualitative ratings. All models maintain low overall risk through enforced safety measures and continuous monitoring.
Abuse Potential
Evaluates the model's vulnerability to exploitation through jailbreaks, adversarial attacks, or harmful queries, such as those involving CBRN weapons, CSAM, or self-harm, focusing on refusal rates and robustness.
Concerning Propensities
Assesses inherent model behaviors that could lead to loss of control, including deception, sycophancy, and political bias, measured via benchmarks like MASK to maintain honesty and alignment.
Dual-Use Capabilities
Examines the model's potential for misuse in high-risk domains like biology, chemistry, cybersecurity, and persuasion, using benchmarks to ensure capabilities do not enable catastrophic outcomes.
xAI Frontier Artificial Intelligence Framework
Published: 2025-12-30
frameworkxAI Risk Management Framework Draft
Published: 2025-02-20
policyGrok 4 Model Card
Published: 2025-08-20
system cardGrok 4.1 Model Card
Published: 2025-11-17
system cardGrok 4 Fast Model Card
Published: 2025-09-19
system cardGrok Code Fast 1 Model Card
Published: 2025-08-26
system cardAll risk assessments, scores, and threshold proximity values are extracted directly from official xAI model cards and the Frontier Artificial Intelligence Framework. Each data point is grounded in quantitative benchmark results and includes citations to source documents.