常扬-文档解析技术加速大模型训练与应用_42页_11mb
报告摘要
Digital Summit Speech on Development: Document Parsing Technology Overview
Part 1: Large Model Training and Application Challenges
The speaker discusses key hurdles in large model (LLM) training and application, particularly document parsing issues. Current challenges include insufficient quality of training data, high token costs, and inaccurate parsing of documents in varied formats like PDFs and scanned pages. This impacts applications such as Retrieval-Augmented Generation (RAG), where precise extraction from sources is vital. Demand exists for higher precision, efficiency, and handling of complex layouts (e.g., multi-column, merged cells) to support industries like finance and manufacturing.
Part 2: Document Parsing Research and Development Progress
Document parsing technology has evolved through multiple stages, from optical methods in the 1920s to deep learning dominance starting in 2014. Research focuses on parsing unstructured documents, including handling challenges like element overlap, diversity in document types, and advanced OCR techniques. Key areas include version analysis, text recognition, and applications in industries that rely on digitized content, with ongoing efforts to improve accuracy and adapt to real-world layouts.
Part 3: TextIn Document Parsing Algorithm Framework
TextIn employs a comprehensive algorithm pipeline with physical and logical page layout analysis. Physical analysis uses detection models to group elements visually, while logical analysis leverages transformers to understand semantic relationships, forming a document tree structure. The system processes multi-format documents using AI for text, table, and formula recognition, with optimizations for speed and accuracy, outperforming alternatives in benchmarks for efficiency and coverage.
Part 4: Large Model Applications Using Document Parsing
TextIn enables applications such as open-domain information extraction from business documents, improving efficiency in extracting key facts from varied sources. It also supports analyst-focused AI systems for financial reports, providing reliable answers through Natural Language Query (NLQ) and RAG integration, reducing manual effort and enhancing decision-making by leveraging accurate document parsing.
Part 5: Summary and Future Outlook
The presentation summarizes TextIn's role in advancing LLM applications, highlighting its features like high precision handling of diverse document types (e.g., tables, formulas), faster processing times, and competitive performance. The TextIn platform matrix showcases integrated applications for industries, emphasizing its contribution to automated document handling and AI-driven insights. Future directions include expanding software capabilities and integrating deep learning advancements for broader AI application contexts.
试读结束,高清完整版pdf/doc/ppt,请点下载