English
Related papers

Related papers: SDUs DAISY: A Benchmark for Danish Culture

200 papers

AI is flattening culture. Evaluations of "culture" are showing the myriad ways in which large AI models are homogenizing language and culture, averaging out rich linguistic differences into generic expressions. I call this phenomenon…

Human-Computer Interaction · Computer Science 2025-07-02 Daniel Mwesigwa

This work presents a maturity model for assessing catalogues of semantic artefacts, one of the keystones that permit semantic interoperability of systems. We defined the dimensions and related features to include in the maturity model by…

Digital Libraries · Computer Science 2024-06-06 Oscar Corcho , Fajar J. Ekaputra , Ivan Heibi , Clement Jonquet , Andras Micsik , Silvio Peroni , Emanuele Storti

The amount of sequence data obtained from ancient samples has dramatically expanded in the last decade, and so have the types of questions that can now be addressed using ancient DNA. In the field of human history, while ancient DNA has…

Populations and Evolution · Quantitative Biology 2020-01-08 Fernando Racimo , Martin Sikora , Hannes Schroeder , Carles Lalueza-Fox

There is a lack of empirical evidence about global attitudes around whether and how GenAI should represent cultures. This paper assesses understandings and beliefs about culture as it relates to GenAI from a large-scale global survey. We…

Computation and Language · Computer Science 2026-03-09 Erin van Liemt , Renee Shelby , Andrew Smart , Sinchana Kumbale , Richard Zhang , Neha Dixit , Qazi Mamunur Rashid , Jamila Smith-Loud

The language technology moonshot moment of Generative Large Language Models (GLLMs) was not limited to English: These models brought a surge of technological applications, investments, and hype to low-resource languages as well. However,…

Computation and Language · Computer Science 2025-03-05 Søren Vejlgaard Holm , Lars Kai Hansen , Martin Carsten Nielsen

The context of this paper is the creation of large uniform archaeological datasets from heterogeneous published resources, such as find catalogues - with the help of AI and Big Data. The paper is concerned with the challenge of consistent…

Computer Vision and Pattern Recognition · Computer Science 2025-03-24 Kevin Klein , Antoine Muller , Alyssa Wohde , Alexander V. Gorelik , Volker Heyd , Ralf Lämmel , Yoan Diekmann , Maxime Brami

Although the cultural (mis)alignment of Large Language Models (LLMs) has attracted increasing attention -- often framed in terms of cultural bias -- until recently there has been limited work on the design and development of datasets for…

Human culture research has witnessed an opportunity of revolution thanks to the big data and social network revolution. Websites such as Douban.com, Goodreads.com, Pandora and IMDB become the new gold mine for cultural researchers. In 2021…

Computers and Society · Computer Science 2023-07-27 Hao Wang

Communicating about some vital topics -- such as sexuality and health -- is treated as taboo and subjected to censorship. How can we construct knowledge about these topics? Wikipedia is home to numerous high-quality knowledge artifacts…

Computers and Society · Computer Science 2026-03-10 Kaylea Champion , Benjamin Mako Hill

Recent breakthroughs in generative AI have opened the door to new research perspectives in the domain of art and cultural heritage, where a large number of artifacts have been digitized. There is a need for innovation to ease the access and…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Valentine Bernasconi , Gustavo Marfia

A `state of the art' model A surpasses humans in a benchmark B, but fails on similar benchmarks C, D, and E. What does B have that the other benchmarks do not? Recent research provides the answer: spurious bias. However, developing A to…

Computation and Language · Computer Science 2020-08-11 Swaroop Mishra , Anjana Arunkumar , Bhavdeep Sachdeva , Chris Bryan , Chitta Baral

We introduce DS-1000, a code generation benchmark with a thousand data science problems spanning seven Python libraries, such as NumPy and Pandas. Compared to prior works, DS-1000 incorporates three core features. First, our problems…

Software Engineering · Computer Science 2022-11-22 Yuhang Lai , Chengxi Li , Yiming Wang , Tianyi Zhang , Ruiqi Zhong , Luke Zettlemoyer , Scott Wen-tau Yih , Daniel Fried , Sida Wang , Tao Yu

Despite the rapid development of large language models (LLMs) for the Korean language, there remains an obvious lack of benchmark datasets that test the requisite Korean cultural and linguistic knowledge. Because many existing Korean…

Computation and Language · Computer Science 2024-07-08 Eunsu Kim , Juyoung Suk , Philhoon Oh , Haneul Yoo , James Thorne , Alice Oh

Novelty modeling and detection is a core topic in Natural Language Processing (NLP), central to numerous tasks such as recommender systems and automatic summarization. It involves identifying pieces of text that deviate in some way from…

Computation and Language · Computer Science 2025-05-14 Florian Carichon , Romain Rampa , Golnoosh Farnadi

The Audio Mostly (AM) conference has long been a platform for exploring the intersection of sound, technology, and culture. Despite growing interest in sonic cultures, discussions on the role of cultural diversity in sound design and…

Physics and Society · Physics 2026-01-16 Rubén García-Benito

Existing commonsense reasoning datasets for AI and NLP tasks fail to address an important aspect of human life: cultural differences. We introduce an approach that extends prior work on crowdsourcing commonsense knowledge by incorporating…

Artificial Intelligence · Computer Science 2020-12-22 Anurag Acharya , Kartik Talamadupula , Mark A Finlayson

While recent years have seen remarkable progress in music generation models, research on their biases across countries, languages, cultures, and musical genres remains underexplored. This gap is compounded by the lack of datasets and…

Sound · Computer Science 2025-10-03 Ahmet Solak , Florian Grötschla , Luca A. Lanzendörfer , Roger Wattenhofer

We present a fully reproducible demonstration of an AI-assisted scientific workflow designed for a broad physics, mathematics, and computer-science readership. The initial project artifact stack was generated from one single user prompt and…

Other Condensed Matter · Physics 2026-03-17 Kin Hung Fung

Large Language Models (LLMs) often exhibit cultural biases due to training data dominated by high-resource languages like English and Chinese. This poses challenges for accurately representing and evaluating diverse cultural contexts,…

Computation and Language · Computer Science 2025-08-11 Zhong Ken Hew , Jia Xin Low , Sze Jue Yang , Chee Seng Chan

We propose a new class of "grand challenge" AI problems that we call creative captioning---generating clever, interesting, or abstract captions for images, as well as understanding such captions. Creative captioning draws on core AI…

Artificial Intelligence · Computer Science 2020-10-02 Maithilee Kunda , Irina Rabkina