2026年夏季全球_AI_安全指数报告_127页_2mb
报告摘要
AI Safety Index Summary - Summer 2026
Core Content
The AI Safety Index is an independent assessment by the Future of Life Institute (FLI) that evaluates the safety practices of nine leading AI companies across six critical domains: Risk Assessment, Current Harms, Safety Frameworks, Existential Safety, Governance & Accountability, and Information Sharing. The report uses a GPA-style grading system (A-F) and is based on 37 indicators, with evidence collected up to June 3, 2026.
Key Findings
- Anthropic, OpenAI, and Google DeepMind lead in safety practices, with Anthropic earning the highest overall grade (C+).
- Meta improves its ranking, while xAI and Mistral show significant decline.
- Existential Safety remains the weakest domain, with no company scoring above a C-.
- Z.ai, Alibaba Cloud, DeepSeek, and Mistral have no published safety frameworks, and their scores reflect regulatory compliance rather than independent safety leadership.
- Military AI use is a growing concern, with major companies like Anthropic, OpenAI, Google DeepMind, and Meta shifting from prior bans to active defense partnerships.
- Safety rhetoric often outpaces actual behavior, as some companies' public commitments do not align with their operational practices.
- Whistleblower protections and transparent safety policies are lacking across the industry, with Mistral and xAI scoring the lowest in this category.
Grading Overview
| Company | Overall Grade | Score | Grade Trend (Winter 25) | Risk Assessment | Current Harms | Safety Frameworks | Existential Safety | Governance & Accountability | Information Sharing |
|---|---|---|---|---|---|---|---|---|---|
| Anthropic | C+ | 2.66 | C+ | C+ | B- | B- | D+ | B | B+ |
| OpenAI | C | 2.28 | C+▼ | C+ | C | C+ | D+ | C | B- |
| Google DeepMind | C | 2.01 | C | C+ | C | C | D | C- | B- |
| Meta | D+ | 1.32 | D▲ | D+ | D- | C- | F | D+ | D+ |
| Z.ai | D- | 0.88 | - | F | C- | D- | F | D- | D |
| Alibaba Cloud | D- | 0.87 | - | F | C- | D- | F | D- | D |
| xAI | F | 0.65 | - | D- | F | D- | F | F | D |
| DeepSeek | F | 0.47 | - | F | D- | F | F | F | D- |
| Mistral | F | 0.33 | - | F | F | F | F | F | D- |
Main Domains and Findings
Risk Assessment
- Key indicators: Internal Testing, External Testing
- Findings:
- Anthropic and OpenAI show strong external testing and evaluation processes.
- Companies like Meta and Mistral have weak or undefined testing practices.
- The industry is moving away from prior commitments to pause development if redlines are approached.
Current Harms
- Key indicators: Safety Performance, Digital Responsibility, Major Safety Incidents & Response, Military Use of AI
- Findings:
- AI systems are increasingly causing real-world harm, including wrongful-death claims.
- Military use of AI is a growing risk, with several companies engaging in defense partnerships.
- Safety benchmarks like HELM and TrustLLM are used to measure model performance and robustness.
Safety Frameworks
- Key indicators: Risk Identification, Risk Analysis & Evaluation, Risk Treatment, Risk Governance
- Findings:
- Safety frameworks are being updated, but many lack quantitative thresholds and independent audits.
- Companies are not fully aligning their public policy positions with their safety commitments.
Existential Safety
- Key indicators: Existential Safety Strategy, Internal Monitoring and Control Interventions, Technical AI Safety Research, Supporting External Safety Research
- Findings:
- No company scores above a C- in this domain.
- Attempts at alignment and control, such as Anthropic's constitutional classifiers, are deemed inadequate by the panel.
- There is a lack of binding mitigations and preventive measures.
Governance & Accountability
- Key indicators: Company Structure & Mandate, Whistleblowing Protection, Policy Transparency, Policy Quality, Reporting Culture
- Findings:
- Companies are lacking in whistleblower protections and transparent governance structures.
- Mistral and xAI score the lowest in this domain due to poor policies and weak reporting cultures.
Information Sharing
- Key indicators: Technical Specifications, System Prompt Transparency, Behavior Specification Transparency, Voluntary Commitment, Public Policy Engagement
- Findings:
- Anthropic and OpenAI score well in information sharing.
- Companies like Mistral and xAI have poor transparency in their policies and practices.
Key Recommendations
-
Anthropic:
- Reverse the RSP 3.0 walk-back on pause commitments.
- Replace qualitative thresholds with quantitative and risk-tied ones.
- Treat prevention as seriously as interpretability and detection.
- Establish stronger safeguards for military use.
-
OpenAI:
- Remove leadership's ability to override the Safety Advisory Group.
- Make safety thresholds measurable, risk-tiered, and externally enforceable.
- Align public policy with safety commitments.
- Evaluate internal-deployment risks before broad use.
-
Google DeepMind & Meta:
- Establish clear decision-making authority and independent audit mechanisms.
- Reverse the backsliding on pause commitments.
- Strengthen whistleblower protections and align culture with policy.
-
Z.ai & Alibaba Cloud:
- Publish a full safety framework and governance structure.
- Move beyond passive deference to regulation and engage in proactive safety research.
- Establish and publicize a whistleblower policy.
-
xAI & DeepSeek:
- Publish a full safety framework and governance structure.
- Engage substantively with existential safety.
- Improve safety benchmark performance.
-
Mistral:
- Publish a full safety framework and governance structure.
- Engage substantively with existential safety.
- Improve weak safety benchmark performance.
Methodology Highlights
- The Index evaluates companies on 37 indicators across six domains.
- Companies are selected based on market influence and technological impact.
- Evidence is collected from publicly available materials and targeted surveys.
- Independent Review Panel of seven experts assigns domain-level grades (A-F) based on absolute performance standards.
Independent Review Panel
The panel includes notable figures such as:
- David Krueger: Expert in AI alignment and safety.
- Tegan Maharaj: Focuses on responsible AI development.
- Stuart Russell: Co-author of the standard AI textbook and advocate for ethical AI.
- Sharon Li: Researcher in safe and reliable AI.
- Sneha Revanur: Founder of Encode, an organization advocating for ethical AI regulation.
- Robert Trager: Expert in AI governance and international regulation.
- Yi Zeng: Chinese AI researcher and policy advisor.
Conclusion
The AI Safety Index underscores the urgent need for stronger safety practices in the AI industry, especially as capabilities continue to rise and real-world harms increase. While some companies like Anthropic and OpenAI have made progress, the overall trend is concerning, with many failing to meet basic safety standards. The Index emphasizes the importance of transparent policies, binding safety measures, and proactive research to ensure the responsible development and deployment of AI systems.
试读结束,高清完整版pdf/doc/ppt,请点下载