Everything+you+need+to+know+about+Multilingual+LLMs-+Towards+fair,+performant+and+reliable+models+for+languages+of+the-144页_9mb
报告摘要
ACL 2023 Tutorial Summary: Multilingual LLMs
This summary covers the key aspects of the ACL 2023 tutorial on Multilingual LLMs, aimed at creating fair, performant, and reliable models for diverse languages.
Introduction
- Highlights the need for linguistic inclusivity in NLP, noting that 88% of the world's languages lack significant language technology benefits.
- Emphasizes the tutorial's goal of advancing multilingual AI by building upon English-centric models and addressing global language disparities.
Data Collection and Training
- Data Challenges: Significant gaps in data quantity and quality across languages (e.g., CommonCrawl shows 57 languages account for <0.001% of web data). Quality issues include incorrect language identification and machine-generated data.
- Model Architectures: Covers encoder-only (e.g., mBERT, XLM-R), decoder-only (e.g., mT5, BART), and encoder-decoder models (e.g., T5). Discusses training objectives like masked language modeling (MLM) and electra-style paradigms for better performance.
- Tokenization and Sampling: Importance of tokenization strategies with Unicode support to improve multilingual coverage. Sampling techniques like temperature sampling aim for fair language coverage, though unimax sampling showed better results.
Prompting Strategies
- Explores various prompting methods for multilingual tasks, including monolingual, translate-test, cross-lingual, and chain-of-thought prompting.
- Results show that cross-lingual prompting and chain-of-thought techniques often outperform direct prompting, with English-centric examples providing advantages even in low-resource languages.
Evaluation and Interpretability
- Benchmarking Issues: Most benchmarks favor high-resource Indo-European languages, leading to poor performance on low-resource languages. Models like fine-tuned mBERT outperform GPT-3.5 on some tasks but struggle elsewhere.
- Calibration and Probing: Multilingual models are often mis-calibrated, overestimating confidence. Structural and intrinsic probing techniques reveal shared neural patterns for linguistic properties across languages, aiding interpretability. Causal probing helps understand cross-lingual transfer mechanisms.
Responsible AI
- Ethical Concerns: Addresses harms like biased outputs, offensive language, hallucinations, and performance disparities. Gender bias is prevalent across languages, with typological differences affecting mitigation strategies.
- Cultural and Distributive Justice: Highlights the need for culturally sensitive AI, decolonizing RAI discourse, and ensuring fair model selection based on principles like Rawlsian justice (maximizing minimum performance across languages).
Community Engagement
- Stressing collaboration with language communities to address their needs, drawing from case studies like AI4Bharat and community-led initiatives.
- Emphasizes building tools in partnership with stakeholders, enabling income opportunities via platforms like Project Karya, while acknowledging challenges in representation and infrastructure.
Conclusion and Open Questions
- Summarizes key open research areas, such as optimizing data mixtures, improving sample efficiency, and enhancing post-training methods for multilingual models.
- Notes unaddressed topics, including transliteration, efficient fine-tuning, and multimodal multilingual AI, providing tutorial resources for further reading.
展开完整摘要
试读结束,高清完整版pdf/doc/ppt,请点下载