人工智能风险缓释的系统映射_证据扫描与初步AI风险缓释分类体系_21页_671kb
报告摘要
Summary of "Mapping AI Risk Mitigations: Evidence Scan and Preliminary AI Risk Mitigation Taxonomy"
Core Content
This paper presents a preliminary AI Risk Mitigation Taxonomy developed through a rapid evidence scan of 13 AI risk mitigation frameworks published between 2023 and 2025. These frameworks were analyzed to extract 831 distinct AI risk mitigations, which were then clustered and categorized into a structured taxonomy. The goal of the work is to provide a common frame of reference for AI risk mitigation, enabling better coordination and communication among stakeholders in the AI ecosystem.
The Taxonomy organizes mitigations into four main categories and 23 subcategories, and is publicly available on the AI Risk Initiative website (airisk.mit.edu) for further refinement and use.
Main Categories and Subcategories
1. Governance & Oversight
- Purpose: Establish human oversight mechanisms and decision protocols to ensure accountability, ethical conduct, and risk management.
- Subcategories:
- 1.1 Board Structure & Oversight
- 1.2 Risk Management
- 1.3 Conflict of Interest Protections
- 1.4 Whistleblower Reporting & Protection
- 1.5 Safety Decision Frameworks
- 1.6 Environmental Impact Management
- 1.7 Societal Impact Assessment
2. Technical & Security
- Purpose: Secure AI systems and constrain model behaviors to ensure safety, alignment with human values, and content integrity.
- Subcategories:
- 2.1 Model & Infrastructure Security
- 2.2 Model Alignment
- 2.3 Model Safety Engineering
- 2.4 Content Safety Controls
3. Operational Process
- Purpose: Govern AI system deployment, usage, monitoring, incident handling, and validation to promote safety and accountability.
- Subcategories:
- 3.1 Testing & Auditing
- 3.2 Data Governance
- 3.3 Access Management
- 3.4 Staged Deployment
- 3.5 Post-deployment Monitoring
- 3.6 Incident Response & Recovery
4. Transparency & Accountability
- Purpose: Enable external scrutiny, communication of AI system information, and ensure accountability to users, regulators, and the public.
- Subcategories:
- 4.1 System Documentation
- 4.2 Risk Disclosure
- 4.3 Incident Reporting
- 4.4 Governance Disclosure
- 4.5 Third-Party System Access
- 4.6 User Rights & Recourse
Key Findings
-
Mitigation Distribution:
- Operational Process Controls accounted for 36% of all mitigations (n=295).
- Governance & Oversight Controls accounted for 30% (n=248).
- Transparency & Accountability Controls accounted for 21% (n=171).
- Technical & Security Controls accounted for 12% (n=101).
- 16 mitigations (2%) could not be classified and are listed in Appendix B.
-
Inconsistencies in Terminology:
- The term "risk management" is widely used but inconsistently defined across frameworks.
- Some frameworks emphasize governance structures, while others focus on monitoring and evaluation after deployment.
-
Underrepresented Mitigations:
- Categories such as Conflict of Interest Protections, Whistleblower Reporting & Protection, and Staged Deployment were mentioned in less than 1% of the mitigations, indicating they may be neglected or underemphasized in current AI risk practices.
-
Methodology:
- The Taxonomy was developed using thematic synthesis and framework synthesis methods.
- Mitigations were manually extracted and iteratively clustered across three rounds of analysis.
- Large language models were used as assistive tools, but human coding was ultimately required for accurate classification.
Practical Implications
-
The preliminary AI Risk Mitigation Database and Taxonomy support risk management across various stakeholder groups, including:
- Technology Managers and AI Developers: To assess and implement mitigations effectively.
- Policymakers and Regulators: To translate general obligations into specific, measurable actions.
- Researchers and External Auditors: To engage in independent evaluation and validation of AI systems.
-
The work is part of a broader AI Risk Initiative (airisk.mit.edu) aimed at helping decision-makers identify, prioritize, and manage AI-related risks.
Conclusion
This paper introduces a structured and categorized framework for AI risk mitigations, addressing the fragmentation and inconsistency in current risk management approaches. The Taxonomy and Database offer a starting point for better coordination, synthesis, and implementation of AI risk mitigations. While the Taxonomy is preliminary, it represents a significant step toward creating a shared understanding of AI risk mitigation practices, which is essential for building trustworthy AI systems.
试读结束,高清完整版pdf/doc/ppt,请点下载