中文
相关论文

相关论文: Massively Multi-Cultural Knowledge Acquisition & L…

200 篇论文

Numerous recent studies have shown that Large Language Models (LLMs) are biased towards a Western and Anglo-centric worldview, which compromises their usefulness in non-Western cultural settings. However, "culture" is a complex,…

计算机与社会 · 计算机科学 2025-02-17 Sougata Saha , Saurabh Kumar Pandey , Monojit Choudhury

Metaphors are pervasive in communication, making them crucial for natural language processing (NLP). Previous research on automatic metaphor processing predominantly relies on training data consisting of English samples, which often reflect…

计算与语言 · 计算机科学 2025-06-10 Senqi Yang , Dongyu Zhang , Jing Ren , Ziqi Xu , Xiuzhen Zhang , Yiliao Song , Hongfei Lin , Feng Xia

Large language models (LLMs) are now deployed worldwide, inspiring a surge of benchmarks that measure their multilingual and multicultural abilities. However, these benchmarks prioritize generic language understanding or superficial…

Online platforms, particularly Wikipedia, have become critical infrastructures for providing diverse linguistic and cultural contexts. This human-curated knowledge now forms the foundation for modern AI. However, we have not yet fully…

计算机与社会 · 计算机科学 2025-07-31 Akira Matsui , Fujio Toriumi , Mitsuo Yoshida , Taichi Murayama , Shiori Hironaka

Large language models (LLMs) are used globally across many languages, but their English-centric pretraining raises concerns about cross-lingual disparities for cultural awareness, often resulting in biased outputs. However, comprehensive…

计算与语言 · 计算机科学 2025-09-23 Raoyuan Zhao , Beiduo Chen , Barbara Plank , Michael A. Hedderich

Text-to-image generation models have achieved strong performance in culturally homogeneous settings, yet their ability to generate multicultural scenes, where people and landmarks originate from different cultures, remains largely…

计算机视觉与模式识别 · 计算机科学 2026-04-20 Parth Bhalerao , Mounika Yalamarty , Brian Trinh , Oana Ignat

Existing commonsense reasoning datasets for AI and NLP tasks fail to address an important aspect of human life: cultural differences. We introduce an approach that extends prior work on crowdsourcing commonsense knowledge by incorporating…

人工智能 · 计算机科学 2020-12-22 Anurag Acharya , Kartik Talamadupula , Mark A Finlayson

Large Language Models (LLMs) have demonstrated remarkable capabilities across various tasks, yet they often exhibit a specific cultural biases, neglecting the values and linguistic diversity of low-resource regions. This cultural bias not…

计算与语言 · 计算机科学 2025-05-28 Ruixiang Feng , Shen Gao , Xiuying Chen , Lisi Chen , Shuo Shang

Cultural competence, defined as the ability to understand and adapt to multicultural contexts, is increasingly vital for large language models (LLMs) in global environments. While several cultural benchmarks exist to assess LLMs' cultural…

计算与语言 · 计算机科学 2025-09-16 Xinyu Zhang , Pei Zhang , Shuang Luo , Jialong Tang , Yu Wan , Baosong Yang , Fei Huang

Recent research has taken advantage of Wikipedia's multilingualism as a resource for cross-language information retrieval and machine translation, as well as proposed techniques for enriching its cross-language structure. The availability…

数据库 · 计算机科学 2011-11-01 Thanh Nguyen , Viviane Moreira , Huong Nguyen , Hoa Nguyen , Juliana Freire

To enhance language models' cultural awareness, we design a generalizable pipeline to construct cultural knowledge bases from different online communities on a massive scale. With the pipeline, we construct CultureBank, a knowledge base…

计算与语言 · 计算机科学 2024-04-24 Weiyan Shi , Ryan Li , Yutong Zhang , Caleb Ziems , Chunhua yu , Raya Horesh , Rogério Abreu de Paula , Diyi Yang

Language is a cornerstone of cultural identity, yet globalization and the dominance of major languages have placed nearly 3,000 languages at risk of extinction. Existing AI-driven translation models prioritize efficiency but often fail to…

计算与语言 · 计算机科学 2025-06-10 Mahfuz Ahmed Anik , Abdur Rahman , Azmine Toushik Wasi , Md Manjurul Ahsan

As large language models (LLMs) become increasingly accessible in many countries, it is essential to align them to serve pluralistic human values across cultures. However, pluralistic culture alignment in LLMs remain an open problem. In…

计算与语言 · 计算机科学 2024-10-18 Shaoyang Xu , Yongqi Leng , Linhao Yu , Deyi Xiong

To serve global users safely and productively, LLMs need culture-specific knowledge that might not be learned during pre-training. How do we find such knowledge that is (1) salient to in-group users, but (2) unknown to LLMs? The most common…

计算与语言 · 计算机科学 2025-11-03 Caleb Ziems , William Held , Jane Yu , Amir Goldberg , David Grusky , Diyi Yang

We present a comprehensive three-phase study to examine (1) the cultural understanding of Large Multimodal Models (LMMs) by introducing DalleStreet, a large-scale dataset generated by DALL-E 3 and validated by humans, containing 9,935…

计算与语言 · 计算机科学 2024-10-21 Anjishnu Mukherjee , Ziwei Zhu , Antonios Anastasopoulos

Research has shown that while large language models (LLMs) can generate their responses based on cultural context, they are not perfect and tend to generalize across cultures. However, when evaluating the cultural bias of a language…

计算与语言 · 计算机科学 2025-12-29 Vitthal Bhandari

Knowledge built culturally across generations allows humans to learn far more than an individual could glean from their own experience in a lifetime. Cultural knowledge in turn rests on language: language is the richest record of what…

Existing benchmarks that measure cultural adaptation in LLMs are misaligned with the actual challenges these models face when interacting with users from diverse cultural backgrounds. In this work, we introduce the first framework and…

计算与语言 · 计算机科学 2025-10-14 Shreya Havaldar , Sunny Rai , Young-Min Cho , Lyle Ungar

Large Language Models (LLMs) are predominantly trained and aligned in ways that reinforce Western-centric epistemologies and socio-cultural norms, leading to cultural homogenization and limiting their ability to reflect global…

计算与语言 · 计算机科学 2025-05-15 Abdullah Mushtaq , Imran Taj , Rafay Naeem , Ibrahim Ghaznavi , Junaid Qadir

Generative models are known to have reduced performance in different global cultural contexts and languages. While continual data updates have been commonly conducted to improve overall model performance, bolstering and evaluating this…