2024-01-21-OPPO-2023多模态预训练模型在OPPO端云场景的落地实践报告_46页_10mb
报告摘要
Analysis Summary
The provided content describes OPPO's research and practice on multi-modal pretraining models, focusing on end-side (or "edge-side") application of advanced AI technologies. Key points include:
-
Edge-Side Image-Text Retrieval Technology:
- Purpose: Enhance intelligent search for user queries (e.g., "和女朋3友去迪士尼").
- Key Features:
- Natural language query processing for image search.
- Efficient search speed by deploying large models on devices.
- Privacy protection by processing search locally.
- Challenges:
- Lightweight deployment after model compression.
- Implementation of an efficient vector retrieval engine.
- Benefits: Improved efficiency in information search, cost savings, and enhanced user privacy.
-
Implementation and Framework:
- Use of multi-modal alignment technologies like CLIP and ALBEF.
- Model optimization techniques such as orthogonal finetuning and multi-teacher distillation.
- Progressive distillation for faster sampling of diffusion models.
-
Domain-Specific Applications:
- Person Beauty Vertical: Focus on high-quality fine-tuning and domain-specific adjustments.
- Text Rendering: Specialized handling for Chinese text rendering in text-to-image models.
- Personalized Generation: Use of SDD dataset and Subject-Diffusion for personalized image generation.
-
Experimental Results:
- Performance comparisons of models (e.g., ALBEF vs. CLI
compact specific text-to-image models). - Quantitative data such as precision@k metrics for image-text retrieval.
- Benchmark results on edge hardware like high-end smartphones (e.g., Snapdragon 8 Gen2).
- Performance comparisons of models (e.g., ALBEF vs. CLI
-
Technical Challenges:
- Balancing model size reduction and performance.
- Ensuring fast inference on edge devices by optimizing sampling techniques.
Summary Conclusion
The report explores the application of large language models and multi-modal techniques in on-device scenarios, emphasizing lightweight model deployment and domain-specific optimizations for various applications like image-text retrieval, text-image generation, and personalized content creation. It highlights the trade-offs between model fidelity and deployment efficiency on hardware like Snapdragon-based devices, along with empirical results discussing achievable performance and user benefits.
试读结束,高清完整版pdf/doc/ppt,请点下载