RED30 Model Risk Analysis
30 Universal AI Red Line Indicators - Verified Comparison Across 16 Frontier Models
RED30: 30 Universal AI Red Line Indicators
These 30 Universal AI Red Line Indicators define the minimum, non-negotiable boundaries of acceptable AI behavior across jurisdictions, grounded in criminal law, human rights, data protection, and consumer protection frameworks. Together, the indicators comprehensively cover all major AI risk domains, forming an attempt to create a global baseline for AI governance and model capability assessment.
Immediate Severe Harm - 8 Indicators
Serious Systemic Discrimination - 8 Indicators
Significant Individual Rights - 8 Indicators
Transparency & Accountability - 6 Indicators
Anthropic (5)
- •Claude Opus 4.5
- •Claude Sonnet 4.5
- •Claude Opus 4
- •Claude Sonnet 4
- •Claude Haiku 4.5
OpenAI (4)
- •GPT-5
- •gpt-5-thinking
- •o3
- •o1-pro
Google DeepMind (3)
- •Gemini 3 Pro
- •Gemini 3 Flash
- •Gemini 2.5 Pro
xAI (4)
- •Grok 4.1
- •Grok 4
- •Grok 4 Fast
- •Grok Code Fast 1
[X] Strong Evidence
Tested & documented with quantitative evaluation
[~] Partial Evidence
Policy exists with testing, or qualitative evidence
[ ] No Evidence
Not tested/documented
[N/A] Not Applicable
Not applicable to lab's scope or design
| Lab | [X] Strong | [~] Partial | [ ] None | [N/A] |
|---|---|---|---|---|
| Anthropic | 2 | 8 | 10 | 10 |
| OpenAI | 2 | 7 | 11 | 10 |
| Google DeepMind | 2 | 7 | 11 | 10 |
| xAI | 3 | 7 | 10 | 10 |
What This Shows
- [X] Strong: Red line defined + quantitatively tested
- [~] Partial: Policy defined with some testing or qualitative evaluation
- [ ] None: Category not formally tested or documented
- [N/A]: Not applicable to lab's framework or design scope
Key Insight
All four labs have policies covering ~20 indicators (strong + partial evidence). Most gaps fall in testing-intensive categories like discrimination auditing and bias detection. xAI shows strongest quantitative evidence for harmful content prevention (CSAM, fraud).
Detailed Assessment by Category
Review each category below to see all indicators with lab-by-lab assessment
Direct severe harm, crimes, vulnerable populations
| # | Indicator | Anthropic | OpenAI | DeepMind | xAI |
|---|---|---|---|---|---|
| 1 | CSAM Prevention Generation prevention, detection, reporting mechanisms | Tested | Tested | Tested | Tested |
| 2 | Fraud/Identity Theft Deepfake fraud, impersonation, document forgery prevention | Partial | Partial | Partial | Partial |
| 3 | Violence/Terrorism Violent content generation, terrorist materials prevention | Partial | Partial | Partial | Partial |
| 4 | Harassment Doxxing, threatening content, stalking prevention | Partial | Partial | Partial | Partial |
| 5 | Children's Data Age verification, parental consent, child targeting prevention | Partial | Partial | Partial | Partial |
| 6 | Sensitive Data Health data, medical privacy, biometric data protection | Partial | Partial | Partial | Partial |
| 7 | Biometric Misuse Facial recognition, voice biometrics, biometric tracking prevention | No Evidence | No Evidence | No Evidence | No Evidence |
| 8 | Mass Surveillance Surveillance capabilities, pattern analysis, population tracking | Not Applicable | Not Applicable | Not Applicable | Not Applicable |
| Category 1: Score | 2/8 | 2/8 | 2/8 | 3/8 | |
Systemic discrimination, serious exploitation
| # | Indicator | Anthropic | OpenAI | DeepMind | xAI |
|---|---|---|---|---|---|
| 9 | Racial Discrim. Biased outputs in hiring, lending, services | Partial | Partial | Partial | Partial |
| 10 | Gender Discrim. Gender bias in recommendations, evaluations, hiring | No Evidence | No Evidence | No Evidence | No Evidence |
| 11 | Employment Discrim. Hiring algorithms, performance evaluation, termination decisions | Not Applicable | Not Applicable | Not Applicable | Not Applicable |
| 12 | Credit Discrim. Credit scoring, loan approval, pricing discrimination | Not Applicable | Not Applicable | Not Applicable | Not Applicable |
| 13 | Housing Discrim. Tenant screening, housing allocation discrimination | Not Applicable | Not Applicable | Not Applicable | Not Applicable |
| 14 | Defamation False statement generation, reputation harm prevention | Partial | Partial | Partial | Partial |
| 15 | Vulnerability Exploit Predatory targeting, manipulation, elder/child exploitation | Partial | Partial | Partial | Partial |
| 16 | Algorithmic Redlining Geographic exclusion, demographic service denial | Not Applicable | Not Applicable | Not Applicable | Not Applicable |
| Category 2: Score | 2/8 | 2/8 | 2/8 | 3/8 | |
Individual rights, fairness, due process
| # | Indicator | Anthropic | OpenAI | DeepMind | xAI |
|---|---|---|---|---|---|
| 17 | Unauthorized Data Training data collection, user data processing transparency | No Evidence | No Evidence | No Evidence | No Evidence |
| 18 | Disability Discrim. Interface barriers, accessibility compliance, service denial | No Evidence | No Evidence | No Evidence | No Evidence |
| 19 | Age Discrimination Age-based hiring bias, service recommendations | No Evidence | No Evidence | No Evidence | No Evidence |
| 20 | Lack of Explainability Black box decisions, opaque reasoning, lack of justification | Partial | Partial | Partial | Partial |
| 21 | No Human Review Fully automated high-impact decisions, no appeal mechanism | Not Applicable | Not Applicable | Not Applicable | Not Applicable |
| 22 | No Contestation No appeal process, no complaint mechanism for wrong decisions | Not Applicable | Not Applicable | Not Applicable | Not Applicable |
| 23 | Deceptive Marketing False advertising, misleading content generation | Partial | Partial | Partial | Partial |
| 24 | Dark Patterns Manipulative UI, deceptive design, unwanted decisions | Not Applicable | Not Applicable | Not Applicable | Not Applicable |
| Category 3: Score | 2/8 | 2/8 | 2/8 | 3/8 | |
Transparency, emerging requirements, accountability
| # | Indicator | Anthropic | OpenAI | DeepMind | xAI |
|---|---|---|---|---|---|
| 25 | Data Transparency Data collection disclosure, informed consent mechanisms | Partial | Partial | Partial | Partial |
| 26 | Data Deletion Right to be forgotten, data deletion mechanisms | No Evidence | No Evidence | No Evidence | No Evidence |
| 27 | Cross-Border Transfers International data flows, cloud storage safeguards | Partial | Partial | Partial | Partial |
| 28 | Synthetic Content Label Deepfake labels, AI disclosure, synthetic media watermarking | No Evidence | No Evidence | No Evidence | No Evidence |
| 29 | High-Risk Opacity Medical AI, legal AI, financial decision transparency | Not Applicable | Not Applicable | Not Applicable | Not Applicable |
| 30 | Audit Trail Decision logging, regulatory compliance, accountability | No Evidence | No Evidence | No Evidence | No Evidence |
| Category 4: Score | 2/6 | 2/6 | 2/6 | 3/6 | |
CATEGORY 1: CRITICAL HARM (Direct Victims, Universal Prohibition)
Crimes with severe harm to identifiable victims. Criminal prosecution in virtually all countries. No cultural variation in prohibition. Examples: CSAM, fraud, violence, mass surveillance.
CATEGORY 2: SYSTEMIC HARM (Discrimination, Protected Groups)
Systematic harm affecting entire demographic groups. Protected by international conventions (CERD, CEDAW, ILO). Strong civil liability. Examples: racial/gender/employment discrimination, defamation.
CATEGORY 3: INDIVIDUAL HARM (Fairness, Due Process)
Individual rights protections and procedural fairness. Growing legal recognition. Important but secondary harm. Examples: privacy violations, explainability, disability access, appeal mechanisms.
CATEGORY 4: EMERGING STANDARDS (Transparency & Accountability)
Newer requirements becoming best practices. Variable enforcement globally. Process-focused rather than outcome harm. Examples: data transparency, synthetic content labels, audit trails.
This framework tracks 30 universal AI red line indicators across 4 severity tiers. It represents the most important harms that AI systems could cause, from immediate severe harms (TIER 1) to emerging standards (TIER 4).
Best Documented:
CSAM prevention (especially xAI), violence/terrorism filtering, content policies
Needs Improvement:
Racial/gender discrimination testing, data deletion mechanisms, audit trails
Framework Based On: International conventions (UN, ILO, CEDAW, CERD), national laws, published system cards, and universal harm principles