2025大语言模型推理能力榜_中文语境下_最强大脑_测评揭晓_15页_2mb
报告摘要
Summary
The report evaluates the reasoning capabilities of 36 large language models (LLMs) in Chinese-language contexts, focusing on basic logical reasoning and contextual reasoning tasks. The evaluation framework uses criteria like accuracy, logical coherence, and conciseness, with questions drawn from new creation and benchmark datasets. Key findings reveal that GPT-o3 leads in basic logical reasoning with a score of 97, while Gemini 2.5 Flash excels in contextual reasoning. Doubao 1.5 Pro (Thinking) secures the top composite score of 93, followed closely by GPT-5 (Auto). Chinese-developed models perform strongly overall, demonstrating competitive reasoning abilities.
Analysis of efficiency shows that models with superior reasoning often have higher costs in token efficiency, response time, and API usage. Chinese models like Yi-Lightning offer cost advantages, and Doubao 1.5 Pro balances high reasoning performance with efficiency in token use, response time, and API costs. The study concludes that China's AI landscape in reasoning tasks is advancing rapidly, with potential for future enhancements in speed and cost-effectiveness to improve real-world applications.
试读结束,高清完整版pdf/doc/ppt,请点下载