2025-05-27-SuperCLUE-中文大模型基准测评2025年5月报告_39页_13mb
报告摘要
Summary of 2025 SuperCLUE Benchmark Report
This report evaluates progress toward artificial general intelligence (AGI) through Chinese large models, based on the 2025 May SuperCLUE benchmark. It highlights advancements in model capabilities, identifies gaps, and provides comprehensive analysis.
Key Findings
- Overall, Chinese models show significant improvement, with o4-mini(high) leading in total score (70.51) due to strong performance in reasoning and code generation.
- Domestic models like Doubao-1.5-thinking-pro-250415 excel in text-related tasks, achieving high scores, while performance in instruction following lags behind international models by a large margin.
- Small-parameter models, such as Qwen3 series, demonstrate competitive performance, often surpassing larger closed-source models in inference tasks.
- A notable gap remains in AGI potential, with international models leading in comprehensive abilities, but domestic models narrowing the divide.
Trends in 2025 Model Development
- 2025 saw rapid innovation, with models evolving through phases from preparation to fusion, characterized by continuous capability breakthroughs.
- China's models achieved catch-up in many areas, including multi-modal applications and industry-specific uses, narrowing the gap with international counterparts over the past 25 months.
- Open-source and collaborative ecosystems drove progress, with models like DeepSeek-R1 and Qwen series gaining prominence in affordability and accessibility.
SuperCLUE Benchmark Overview
- The benchmark is an independent, dynamic evaluation system updated every two months, covering six primary dimensions: mathematics, science, coding, agent capabilities, instruction following, and text processing.
- Each iteration uses original questions to prevent data contamination, aligning closely with real-world user interactions.
- Evaluation methods include automated assessments based on correctness, execution tests, and human-annotated scoring, ensuring reliable comparisons across models.
Comprehensive Evaluation Results
- Top Performers: o4-mini(high) dominated with high scores, while Chinese models like SenseNova V6 Reasoner and Doubao showed strength in specialized tasks.
- Agent and Capability Analysis: International models outperformed domestically developed agents; however, domestic models showed potential in emerging areas like edge computing.
- Efficiency and Cost-Effectiveness: Chinese models offered better value, with high scores at lower costs, especially in smaller parameter sizes.
- Human-AI Consistency: Results from Chatbot Arena show high correlation (Pearson 0.86, Spearman 0.89) with human evaluations in English benchmarks, validating the benchmark's reliability.
- Maturity Index: Text-related abilities are most mature, while instruction following, mathematical, and scientific reasoning remain critical areas for advancement.
- 10B and Edge Models: Small models achieved competitive results, particularly in text tasks, indicating their suitability for resource-constrained applications.
The report underscores the global acceleration of AGI development, with China making substantial strides but highlighting the need for further enhancements in foundational skills to compete fully.
展开完整摘要
试读结束,高清完整版pdf/doc/ppt,请点下载