> **来源:[研报客](https://pc.yanbaoke.cn)** # AI Safety Index Summary - Summer 2026 ## Core Content The **AI Safety Index** is an independent assessment by the Future of Life Institute (FLI) that evaluates the safety practices of nine leading AI companies across six critical domains: **Risk Assessment**, **Current Harms**, **Safety Frameworks**, **Existential Safety**, **Governance & Accountability**, and **Information Sharing**. The report uses a GPA-style grading system (A-F) and is based on 37 indicators, with evidence collected up to June 3, 2026. ## Key Findings - **Anthropic, OpenAI, and Google DeepMind** lead in safety practices, with Anthropic earning the highest overall grade (C+). - **Meta** improves its ranking, while **xAI** and **Mistral** show significant decline. - **Existential Safety** remains the weakest domain, with no company scoring above a C-. - **Z.ai**, **Alibaba Cloud**, **DeepSeek**, and **Mistral** have **no published safety frameworks**, and their scores reflect **regulatory compliance rather than independent safety leadership**. - **Military AI use** is a growing concern, with major companies like Anthropic, OpenAI, Google DeepMind, and Meta shifting from prior bans to active defense partnerships. - **Safety rhetoric often outpaces actual behavior**, as some companies' public commitments do not align with their operational practices. - **Whistleblower protections** and **transparent safety policies** are lacking across the industry, with **Mistral** and **xAI** scoring the lowest in this category. ## Grading Overview | Company | Overall Grade | Score | Grade Trend (Winter 25) | Risk Assessment | Current Harms | Safety Frameworks | Existential Safety | Governance & Accountability | Information Sharing | |-----------------|---------------|--------|------------------------|------------------|--------------|-------------------|-------------------|-----------------------------|---------------------| | Anthropic | C+ | 2.66 | C+ | C+ | B- | B- | D+ | B | B+ | | OpenAI | C | 2.28 | C+▼ | C+ | C | C+ | D+ | C | B- | | Google DeepMind | C | 2.01 | C | C+ | C | C | D | C- | B- | | Meta | D+ | 1.32 | D▲ | D+ | D- | C- | F | D+ | D+ | | Z.ai | D- | 0.88 | - | F | C- | D- | F | D- | D | | Alibaba Cloud | D- | 0.87 | - | F | C- | D- | F | D- | D | | xAI | F | 0.65 | - | D- | F | D- | F | F | D | | DeepSeek | F | 0.47 | - | F | D- | F | F | F | D- | | Mistral | F | 0.33 | - | F | F | F | F | F | D- | ## Main Domains and Findings ### Risk Assessment - **Key indicators**: Internal Testing, External Testing - **Findings**: - Anthropic and OpenAI show strong external testing and evaluation processes. - Companies like Meta and Mistral have weak or undefined testing practices. - The industry is moving away from prior commitments to pause development if redlines are approached. ### Current Harms - **Key indicators**: Safety Performance, Digital Responsibility, Major Safety Incidents & Response, Military Use of AI - **Findings**: - AI systems are increasingly causing real-world harm, including wrongful-death claims. - Military use of AI is a growing risk, with several companies engaging in defense partnerships. - Safety benchmarks like HELM and TrustLLM are used to measure model performance and robustness. ### Safety Frameworks - **Key indicators**: Risk Identification, Risk Analysis & Evaluation, Risk Treatment, Risk Governance - **Findings**: - Safety frameworks are being updated, but many lack **quantitative thresholds** and **independent audits**. - Companies are not fully aligning their **public policy positions** with their **safety commitments**. ### Existential Safety - **Key indicators**: Existential Safety Strategy, Internal Monitoring and Control Interventions, Technical AI Safety Research, Supporting External Safety Research - **Findings**: - No company scores above a C- in this domain. - Attempts at alignment and control, such as Anthropic's constitutional classifiers, are deemed **inadequate** by the panel. - There is a lack of **binding mitigations** and **preventive measures**. ### Governance & Accountability - **Key indicators**: Company Structure & Mandate, Whistleblowing Protection, Policy Transparency, Policy Quality, Reporting Culture - **Findings**: - Companies are lacking in **whistleblower protections** and **transparent governance structures**. - **Mistral** and **xAI** score the lowest in this domain due to poor policies and weak reporting cultures. ### Information Sharing - **Key indicators**: Technical Specifications, System Prompt Transparency, Behavior Specification Transparency, Voluntary Commitment, Public Policy Engagement - **Findings**: - Anthropic and OpenAI score well in information sharing. - Companies like Mistral and xAI have **poor transparency** in their policies and practices. ## Key Recommendations - **Anthropic**: - Reverse the RSP 3.0 walk-back on pause commitments. - Replace qualitative thresholds with **quantitative and risk-tied** ones. - Treat **prevention** as seriously as **interpretability** and **detection**. - Establish stronger **safeguards for military use**. - **OpenAI**: - Remove leadership's ability to **override the Safety Advisory Group**. - Make safety thresholds **measurable, risk-tiered, and externally enforceable**. - Align public policy with safety commitments. - Evaluate internal-deployment risks **before** broad use. - **Google DeepMind & Meta**: - Establish **clear decision-making authority** and **independent audit mechanisms**. - Reverse the **backsliding on pause commitments**. - Strengthen **whistleblower protections** and **align culture with policy**. - **Z.ai & Alibaba Cloud**: - Publish a **full safety framework** and **governance structure**. - Move beyond **passive deference** to regulation and engage in **proactive safety research**. - Establish and publicize a **whistleblower policy**. - **xAI & DeepSeek**: - Publish a **full safety framework** and **governance structure**. - Engage substantively with **existential safety**. - Improve **safety benchmark performance**. - **Mistral**: - Publish a **full safety framework** and **governance structure**. - Engage substantively with **existential safety**. - Improve **weak safety benchmark performance**. ## Methodology Highlights - The Index evaluates companies on **37 indicators** across **six domains**. - Companies are selected based on **market influence and technological impact**. - Evidence is collected from **publicly available materials** and **targeted surveys**. - **Independent Review Panel** of seven experts assigns **domain-level grades** (A-F) based on **absolute performance standards**. ## Independent Review Panel The panel includes notable figures such as: - **David Krueger**: Expert in AI alignment and safety. - **Tegan Maharaj**: Focuses on responsible AI development. - **Stuart Russell**: Co-author of the standard AI textbook and advocate for ethical AI. - **Sharon Li**: Researcher in safe and reliable AI. - **Sneha Revanur**: Founder of Encode, an organization advocating for ethical AI regulation. - **Robert Trager**: Expert in AI governance and international regulation. - **Yi Zeng**: Chinese AI researcher and policy advisor. ## Conclusion The AI Safety Index underscores the **urgent need for stronger safety practices** in the AI industry, especially as capabilities continue to rise and real-world harms increase. While some companies like Anthropic and OpenAI have made progress, the **overall trend is concerning**, with many failing to meet basic safety standards. The Index emphasizes the importance of **transparent policies**, **binding safety measures**, and **proactive research** to ensure the responsible development and deployment of AI systems.