中文
相关论文

相关论文: No Language Data Left Behind: A Comparative Study …

200 篇论文

Solving complicated AI tasks with different domains and modalities is a key step toward artificial general intelligence. While there are numerous AI models available for various domains and modalities, they cannot handle complicated AI…

计算与语言 · 计算机科学 2023-12-05 Yongliang Shen , Kaitao Song , Xu Tan , Dongsheng Li , Weiming Lu , Yueting Zhuang

Multilingual Large Language Models (LLMs) can process many languages, yet how they internally represent this diversity remains unclear. Do they form shared multilingual representations with language-specific decoding, and if so, why does…

计算与语言 · 计算机科学 2026-02-10 Abir Harrasse , Florent Draye , Punya Syon Pandey , Zhijing Jin , Bernhard Schölkopf

The global deployment of large language models (LLMs) has raised concerns about cultural misalignment, yet the linguistic properties of fine-tuning datasets used for cultural adaptation remain poorly understood. We adopt a dataset-centric…

We present CCI3.0-HQ (https://huggingface.co/datasets/BAAI/CCI3-HQ), a high-quality 500GB subset of the Chinese Corpora Internet 3.0 (CCI3.0)(https://huggingface.co/datasets/BAAI/CCI3-Data), developed using a novel two-stage hybrid…

计算与语言 · 计算机科学 2024-10-28 Liangdong Wang , Bo-Wen Zhang , Chengwei Wu , Hanyu Zhao , Xiaofeng Shi , Shuhao Gu , Jijie Li , Quanyue Ma , TengFei Pan , Guang Liu

The HuggingFace Datasets Hub hosts thousands of datasets, offering exciting opportunities for language model training and evaluation. However, datasets for a specific task type often have different schemas, making harmonization challenging.…

计算与语言 · 计算机科学 2023-05-17 Damien Sileo

To ensure equitable access to the benefits of large language models (LLMs), it is essential to evaluate their capabilities across the world's languages. We introduce the AI Language Proficiency Monitor, a comprehensive multilingual…

计算与语言 · 计算机科学 2025-07-14 David Pomerenke , Jonas Nothnagel , Simon Ostermann

We introduce Multi-SimLex, a large-scale lexical resource and evaluation benchmark covering datasets for 12 typologically diverse languages, including major languages (e.g., Mandarin Chinese, Spanish, Russian) as well as less-resourced ones…

Spoken language datasets are vital for advancing linguistic research, Natural Language Processing, and speech technology. However, resources dedicated to Italian, a linguistically rich and diverse Romance language, remain underexplored…

计算与语言 · 计算机科学 2025-03-13 Marco Giordano , Claudia Rinaldi

We study to what extend Chinese, Japanese and Korean faces can be classified and which facial attributes offer the most important cues. First, we propose a novel way of obtaining large numbers of facial images with nationality labels. Then…

计算机视觉与模式识别 · 计算机科学 2016-10-25 Yu Wang , Haofu Liao , Yang Feng , Xiangyang Xu , Jiebo Luo

This study introduces KPoEM (Korean Poetry Emotion Mapping), a novel dataset that serves as a foundation for both emotion-centered analysis and generative applications in modern Korean poetry. Despite advancements in NLP, poetry remains…

计算与语言 · 计算机科学 2026-01-15 Iro Lim , Haein Ji , Byungjun Kim

Large Language Models (LLMs) have shown remarkable abilities across various tasks, yet their development has predominantly centered on high-resource languages like English and Chinese, leaving low-resource languages underserved. To address…

As ChatGPT and GPT-4 spearhead the development of Large Language Models (LLMs), more researchers are investigating their performance across various tasks. But more research needs to be done on the interpretability capabilities of LLMs, that…

计算与语言 · 计算机科学 2023-10-27 Dongfang Li , Jindi Yu , Baotian Hu , Zhenran Xu , Min Zhang

We introduce GECKO, a bilingual large language model (LLM) optimized for Korean and English, along with programming languages. GECKO is pretrained on the balanced, high-quality corpus of Korean and English employing LLaMA architecture. In…

计算与语言 · 计算机科学 2024-05-27 Sungwoo Oh , Donggyu Kim

With the development of large language models (LLMs), social biases in these LLMs have become a pressing issue. Although there are various benchmarks for social biases across languages, the extent to which Japanese LLMs exhibit social…

计算与语言 · 计算机科学 2025-06-16 Hitomi Yanaka , Namgi Han , Ryoma Kumon , Jie Lu , Masashi Takeshita , Ryo Sekizawa , Taisei Kato , Hiromi Arai

Multilingual pretrained language models (mPLMs) acquire valuable, generalizable linguistic information during pretraining and have advanced the state of the art on task-specific finetuning. To date, only ~31 out of ~2,000 African languages…

计算与语言 · 计算机科学 2023-05-30 Ife Adebara , AbdelRahim Elmadany , Muhammad Abdul-Mageed , Alcides Alcoba Inciarte

The advancements of neural dialogue generation models show promising results on modeling short-text conversations. However, training such models usually needs a large-scale high-quality dialogue corpus, which is hard to access. In this…

计算与语言 · 计算机科学 2022-04-27 Yida Wang , Pei Ke , Yinhe Zheng , Kaili Huang , Yong Jiang , Xiaoyan Zhu , Minlie Huang

Understanding how developers combine programming languages in practice reveals the hidden structure of the software ecosystem: which languages are used as complements, which define coherent technology stacks, and which bridge disparate…

软件工程 · 计算机科学 2026-04-16 Bachan Ghimire , Nitin Gupta

Language models (LM or LLM) are increasingly deployed in the field of artificial intelligence (AI) and its applications, but the question arises as to whether they can be a common resource managed and maintained by a community of users.…

计算机与社会 · 计算机科学 2024-03-20 Robin Quillivic , Salma Mesmoudi

Multilingual NLP is often treated as a route to global inclusion, but linguistic coverage and cultural competence frequently diverge. This paper synthesizes over 50 papers spanning multilingual performance inequality, cross-lingual…

计算与语言 · 计算机科学 2026-05-05 Sina Bagheri Nezhad

Historical records in Korea before the 20th century were primarily written in Hanja, an extinct language based on Chinese characters and not understood by modern Korean or Chinese speakers. Historians with expertise in this time period have…

计算与语言 · 计算机科学 2022-10-12 Haneul Yoo , Jiho Jin , Juhee Son , JinYeong Bak , Kyunghyun Cho , Alice Oh