2025面向前沿大型语言模型生物学知识的综合基准测试研究报告_67页_4mb
报告摘要
Summary of Biological Knowledge Benchmarking Report
Issue
The report addresses the risk that advanced AI systems (frontier large language models-Large language models LLMs) could be misused for developing biological and chemical weapons. AI's demonstrated deep scientific knowledge may lower technical barriers for malicious actors, present et pass # High_risk threats.
Approach
Models & Benchmarks
- 39 LLMs were evaluated: 22 closed-weight and 17 open-weight/fine-tuned.
- Included unsafety-tuned models (custom Llama 3.1 405B) and biology-specific tuning to assess capability changes.
- 8 public benchmarks covering biology/chemistry knowledge and refusal rates for bioweapon-relevant prompts.
Key Findings
- Frontier LLM Performance: Latest reasoning models (e.g., OpenAI's
o3) exceed human expert performance on biology/chemistry benchmarks (e.g., BioLP-bench, GPQA). Models like Claude 3.7 approachPhD-levelexpertise in troubleshooting biological protocols. - Refusal Behavior: Reducing safety guardrails significantly lowers refusal rates (e.g., HarmBench from 100% to 2.4%), but also reduces knowledge performance.
- Benchmark Saturation: Benchmarks like
WMDP biologyare nearing performance ceilings (≈85% across models released after its March 2024 publication), limiting future evaluations. - Data Contamination Risk: Public benchmark data may pollute future model training, reducing benchmark utility.
Recommendations
- Enhance Human Baselines: Include expert baselines for contextualizing model performance, especially for subfields.
- Develop Challenging Benchmarks: Create
private/supersized datasetsto avoid contamination and assess tool-use capabilities (e.g., LAB-Bench with agent-task integration). - Improve Standardization: Report
full benchmark detailsto ensure reproducibility across implementations. - Hybrid Benchmarking Approach: Balance public accessibility with
private datasetsfor sensitive measurements relevant to biological risks.
Overall Contribution
The report highlights the growing capability of LLMs but emphasizes the need for robust benchmarking practices to detect misuse risks and guide policy.
展开完整摘要
试读结束,高清完整版pdf/doc/ppt,请点下载