剑桥大学+GPT-4+技术报告-英-100页_4mb
报告摘要
-
Model Details: GPT-4 is a large-scale, multimodal Transformer-based model that accepts text and image inputs to produce text outputs. It was developed using reinforcement learning from human feedback (RLHF) and scaled alignment processes.
-
Capabilities:
- GPT-4 exhibits human-level performance on professional benchmarks (e.g., passing a simulated bar exam with top 10% accuracy) and professional exams like LSAT, GRE, and AP tests.
- On NLP benchmarks, it outperforms previous models and state-of-the-art systems by considerable margins, including multilingual MMLU tasks.
- It demonstrates strong visual input processing and multimodal problem-solving capabilities.
-
Limitations:
- GPT-4 still suffers from hallucinations, lacks full reliability, and knowledge beyond mid-2021 data.
- Performance does not consistently improve with scale (e.g., Inverse Scaling Prize tasks).
- Adversarial testing revealed potential harms, though mitigations were applied.
-
Safety & Mitigations:
- GPT-4 employs adversarial testing by domain experts and model-assisted safety pipelines (including RLHF and rule-based reward models).
- Toxicity and safety metrics improved, with GPT-4 generating toxic responses 0.73% of the time compared to GPT-3.5’s 6.48%.
- Deployment risks remain due to potential jailbreaks, requiring continuous safety monitoring.
-
Predictable Scaling:
- The project developed infrastructure for reliable performance predictions from smaller models with reduced compute, enabling efficient training alignment.
- Calibration decreases post-RestoringHF, though RLHF improves certain behaviors.
展开完整摘要
试读结束,高清完整版pdf/doc/ppt,请点下载