DeepSeek模型在中文语境下的安全性评估_12页_6mb
报告摘要
Summary of the Safety Evaluation of DeepSeek Models in Chinese Contexts
Core Content
This study presents a comprehensive safety evaluation of the DeepSeek-R1 and DeepSeek-V3 large language models in Chinese contexts. The evaluation is conducted using a newly developed benchmark called CHiSafetyBench, which is designed based on the "Basic Safety Requirements for Generative Artificial Intelligence Services" standard. The research highlights the significant safety deficiencies in DeepSeek models and provides insights for future improvements.
Main Points
-
DeepSeek Models' Safety Issues:
- DeepSeek-R1 has a 100% attack success rate when processing harmful prompts, as revealed by research conducted by Robust Intelligence and the University of Pennsylvania.
- Multiple safety companies and research institutions have confirmed critical safety vulnerabilities in DeepSeek-R1, including the generation of unsafe code, harmful content, and ethical issues.
- The study points out that DeepSeek models, despite their high performance in reasoning and generation tasks, have notable safety gaps, especially in Chinese contexts.
-
CHiSafetyBench Introduction:
- This benchmark evaluates model safety across five categories: discrimination, violation of values, commercial violations, infringement of rights, and security requirements for specific services.
- It includes multiple-choice questions for risk content identification and risky questions for refusal to answer, with metrics such as ACC, RR-1, RR-2, and HR used for evaluation.
-
Performance Analysis:
- In terms of risk content identification, DeepSeek-R1 achieves an overall ACC of 71.41%, while DeepSeek-V3 reaches 84.17%. These are significantly lower than the best-performing Qwen models, such as Qwen1.5-72B-Chat.
- DeepSeek-R1 performs poorly in discrimination and violation of values, with ACC of 50.22% and 64.91%, respectively, which are 36.30% and 28.82% lower than the top Qwen models.
- In refusal to answer tasks, DeepSeek-R1 and DeepSeek-V3 show HR of 0% and 0.43%, respectively, indicating a low likelihood of generating harmful outputs. However, their RR-1 and RR-2 scores are also relatively low, showing limited ability to reject risky questions and provide responsible guidance.
-
Comparative Analysis:
- The study compares DeepSeek models with several other models known for strong Chinese capabilities, such as Baichuan2, ChatGLM3, and Qwen series models.
- The results indicate that the Qwen series models outperform DeepSeek models in most safety categories, especially in discrimination and violation of values.
-
Limitations and Future Work:
- The evaluation may be affected by sample selection bias, data distribution characteristics, and evaluation criteria settings.
- The authors plan to continuously optimize the evaluation benchmark and periodically update the report to enhance its comprehensiveness and reliability.
- The study is the first to conduct a Chinese-specific safety evaluation of DeepSeek-R1, providing a foundational reference for future research.
Key Information
-
DeepSeek-R1:
- Overall ACC: 71.41%
- Discrimination ACC: 50.22%
- Violation of Values ACC: 64.91%
- RR-1: 67.60%
- RR-2: 67.17%
- HR: 0%
-
DeepSeek-V3:
- Overall ACC: 84.17%
- Discrimination ACC: 66.96%
- Violation of Values ACC: 91.98%
- RR-1: 59.83%
- RR-2: 59.61%
- HR: 0.43%
-
Qwen1.5-32B-Chat:
- Overall ACC: 77.71%
- Discrimination ACC: 86.47%
- Violation of Values ACC: 90.98%
- RR-1: 73.38%
- RR-2: 73.38%
- HR: 1.08%
-
Bad Cases:
- In discrimination-related questions, DeepSeek models often fail to identify risks and may even suggest harmful methods.
- In value violation scenarios, they also struggle to detect and reject inappropriate content, as shown in specific examples.
Conclusion
The study concludes that while DeepSeek models have demonstrated strong performance in reasoning and generation, their safety performance in Chinese contexts is significantly lacking. It calls for further safety optimization and comprehensive evaluations to address these issues. The authors emphasize the importance of updating and refining the evaluation benchmark to ensure more accurate and objective assessments in the future.
试读结束,高清完整版pdf/doc/ppt,请点下载