中文
相关论文

相关论文: Krutrim LLM: Multilingual Foundational Model for o…

200 篇论文

Indic languages like Hindi and Tamil are underrepresented in the natural language processing (NLP) field compared to languages like English. Due to this underrepresentation, performance on NLP tasks (such as search algorithms) in Indic…

计算与语言 · 计算机科学 2022-10-13 Adhitya Thirumala , Elisa Ferracane

In this paper, we conduct one of the very first studies for cross-corpora performance evaluation in the spoken language identification (LID) problem. Cross-corpora evaluation was not explored much in LID research, especially for the Indian…

音频与语音处理 · 电气工程与系统科学 2021-05-13 Spandan Dey , Goutam Saha , Md Sahidullah

The development of Large Language Models (LLMs) remains heavily skewed towards English and a few other high-resource languages. This linguistic disparity is particularly evident for Bangla - the 5th most spoken language. A few initiatives…

计算与语言 · 计算机科学 2025-07-01 Nishat Raihan , Marcos Zampieri

Several recent papers have published good solutions for language identification (LID) for about 300 high-resource and medium-resource languages. However, there is no LID available that (i) covers a wide range of low-resource languages, (ii)…

计算与语言 · 计算机科学 2024-07-04 Amir Hossein Kargaran , Ayyoob Imani , François Yvon , Hinrich Schütze

Large Language Models (LLMs) have shown strong generalization across tasks in high-resource languages; however, their linguistic competence in low-resource and morphologically rich languages such as Tamil remains largely unexplored.…

AI-driven education, particularly Large Language Models (LLMs), has the potential to address learning disparities in rural K-12 schools. However, research on AI adoption in rural India remains limited, with existing studies focusing…

计算机与社会 · 计算机科学 2025-05-07 Harshita Goyal , Garima Garg , Prisha Mordia , Veena Ramachandran , Dhruv Kumar , Jagat Sesh Challa

This paper introduces PMIndiaSum, a multilingual and massively parallel summarization corpus focused on languages in India. Our corpus provides a training and testing ground for four language families, 14 languages, and the largest to date…

计算与语言 · 计算机科学 2023-10-23 Ashok Urlana , Pinzhen Chen , Zheng Zhao , Shay B. Cohen , Manish Shrivastava , Barry Haddow

Recent research in multilingual language models (LM) has demonstrated their ability to effectively handle multiple languages in a single model. This holds promise for low web-resource languages (LRL) as multilingual models can enable…

计算与语言 · 计算机科学 2021-06-10 Yash Khemchandani , Sarvesh Mehtani , Vaidehi Patil , Abhijeet Awasthi , Partha Talukdar , Sunita Sarawagi

Large Language Models (LLMs) exhibit significant disparities in performance across languages, primarily benefiting high-resource languages while marginalizing underrepresented ones. Continual Pretraining (CPT) has emerged as a promising…

计算与语言 · 计算机科学 2025-10-09 Zihao Li , Shaoxiong Ji , Hengyu Luo , Jörg Tiedemann

Large language models (LLMs) are reported to be partial to certain cultures owing to the training data dominance from the English corpora. Since multilingual cultural data are often expensive to collect, existing efforts handle this by…

计算与语言 · 计算机科学 2024-12-04 Cheng Li , Mengzhou Chen , Jindong Wang , Sunayana Sitaram , Xing Xie

We investigate a surprising limitation of LLMs: their inability to consistently generate text in a user's desired language. We create the Language Confusion Benchmark (LCB) to evaluate such failures, covering 15 typologically diverse…

计算与语言 · 计算机科学 2025-04-07 Kelly Marchisio , Wei-Yin Ko , Alexandre Bérard , Théo Dehaze , Sebastian Ruder

Vision-language models score well on mathematical, scientific, and spatial reasoning benchmarks, yet these evaluations are overwhelmingly English. I present the first cross-lingual visual reasoning audit for Indian languages. 980 questions…

计算与语言 · 计算机科学 2026-03-31 Swastik R

In this paper, we introduce SUTRA, multilingual Large Language Model architecture capable of understanding, reasoning, and generating text in over 50 languages. SUTRA's design uniquely decouples core conceptual understanding from…

计算与语言 · 计算机科学 2024-05-14 Abhijit Bendale , Michael Sapienza , Steven Ripplinger , Simon Gibbs , Jaewon Lee , Pranav Mistry

Large language models (LLMs) increasingly mediate human communication, decision support, content creation, and information retrieval. Despite impressive fluency, these systems frequently produce biased or stereotypical content, especially…

计算与语言 · 计算机科学 2025-12-11 Muneeb Ur Raheem Khan

Large language models (LLMs) have demonstrated remarkable capabilities across a range of natural language processing (NLP) tasks, capturing the attention of both practitioners and the broader public. A key question that now preoccupies the…

计算与语言 · 计算机科学 2025-06-04 Yahan Li , Yi Wang , Yi Chang , Yuan Wu

Large Language Models (LLMs) exhibit emerging in-context learning abilities through prompt engineering. The recent progress in large-scale generative models has further expanded their use in real-world language applications. However, the…

计算与语言 · 计算机科学 2024-04-12 Linyi Yang , Shuibai Zhang , Zhuohao Yu , Guangsheng Bao , Yidong Wang , Jindong Wang , Ruochen Xu , Wei Ye , Xing Xie , Weizhu Chen , Yue Zhang

India has a rich linguistic landscape with languages from 4 major language families spoken by over a billion people. 22 of these languages are listed in the Constitution of India (referred to as scheduled languages) are the focus of this…

Existing cultural commonsense benchmarks treat nations as monolithic, assuming uniform practices within national boundaries. But does cultural commonsense hold uniformly within a nation, or does it vary at the sub-national level? We…

计算与语言 · 计算机科学 2026-04-16 Sangmitra Madhusudan , Trush Shashank More , Steph Buongiorno , Renata Dividino , Jad Kabbara , Ali Emami

This research investigates biases in text-to-image (TTI) models for the Indic languages widely spoken across India. It evaluates and compares the generative performance and cultural relevance of leading TTI models in these languages against…

计算与语言 · 计算机科学 2024-08-02 Surbhi Mittal , Arnav Sudan , Mayank Vatsa , Richa Singh , Tamar Glaser , Tal Hassner

In this paper, we explore the utility of translationese as synthetic data created using machine translation for pre-training language models (LMs) for low-resource languages (LRLs). Our simple methodology consists of translating large…

计算与语言 · 计算机科学 2025-07-08 Meet Doshi , Raj Dabre , Pushpak Bhattacharyya