English
Related papers

Related papers: AI4D -- African Language Dataset Challenge

200 papers

New deep-learning architectures are created every year, achieving state-of-the-art results in image recognition and leading to the belief that, in a few years, complex tasks such as sign language translation will be considerably easier,…

Computer Vision and Pattern Recognition · Computer Science 2021-04-05 Alvaro Leandro Cavalcante Carneiro , Lucas de Brito Silva , Denis Henrique Pinheiro Salvadeo

Despite comprising one-third of global languages, African languages are critically underrepresented in Artificial Intelligence (AI), threatening linguistic diversity and cultural heritage. Ghanaian languages, in particular, face an alarming…

Computation and Language · Computer Science 2024-05-14 Sheriff Issaka , Zhaoyi Zhang , Mihir Heda , Keyi Wang , Yinka Ajibola , Ryan DeMar , Xuefeng Du

We present an analysis of the performance of machine learning classifiers on discriminating between similar languages and language varieties. We carried out a number of experiments using the results of the two editions of the Discriminating…

Computation and Language · Computer Science 2016-10-04 Cyril Goutte , Serge Léger , Shervin Malmasi , Marcos Zampieri

Automatic Speech Recognition (ASR) technology has witnessed significant advancements in recent years, revolutionizing human-computer interactions. While major languages have benefited from these developments, lesser-resourced languages like…

Computation and Language · Computer Science 2024-11-25 Muhammad Sharif , Zeeshan Abbas , Jiangyan Yi , Chenglin Liu

With language models becoming increasingly ubiquitous, it has become essential to address their inequitable treatment of diverse demographic groups and factors. Most research on evaluating and mitigating fairness harms has been concentrated…

Computation and Language · Computer Science 2023-03-01 Krithika Ramesh , Sunayana Sitaram , Monojit Choudhury

Biased associations have been a challenge in the development of classifiers for detecting toxic language, hindering both fairness and accuracy. As potential solutions, we investigate recently introduced debiasing methods for text…

Computation and Language · Computer Science 2021-02-02 Xuhui Zhou , Maarten Sap , Swabha Swayamdipta , Noah A. Smith , Yejin Choi

Social media datasets are essential for research on disinformation, influence operations, social sensing, hate speech detection, cyberbullying, and other significant topics. However, access to these datasets is often restricted due to costs…

Computers and Society · Computer Science 2024-07-12 Henry Tari , Danial Khan , Justus Rutten , Darian Othman , Rishabh Kaushal , Thales Bertaglia , Adriana Iamnitchi

Recent advances in word embeddings and language models use large-scale, unlabelled data and self-supervised learning to boost NLP performance. Multilingual models, often trained on web-sourced data like Wikipedia, face challenges: few…

Computation and Language · Computer Science 2025-07-02 David Ifeoluwa Adelani

In recent years, large language models (e.g., Open AI's GPT-4, Meta's LLaMa, Google's PaLM) have become the dominant approach for building AI systems to analyze and generate language online. However, the automated systems that increasingly…

Computation and Language · Computer Science 2023-06-14 Gabriel Nicholas , Aliya Bhatia

In the field of speaker diarization, the development of technology is constrained by two problems: insufficient data resources and poor generalization ability of deep learning models. To address these two problems, firstly, we propose an…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-01 Shilong Wu

According to interviews with people who work with speech impaired persons, speech impaired people have difficulties in communicating with other people around them who do not know the sign language, and this situation may cause them to…

Computer Vision and Pattern Recognition · Computer Science 2021-02-24 Arda Mavi

Generative large-scale language models create the fifth paradigm of scientific research, organically combine data science and computational intelligence, transform the research paradigm of natural language processing and multimodal…

Digital Libraries · Computer Science 2024-04-30 Jiangfeng Liu , Ziyi Wang , Jing Xie , Lei Pei

In the evolving landscape of deep learning, there is a pressing need for more comprehensive datasets capable of training models across multiple modalities. Concurrently, in digital humanities, there is a growing demand to leverage…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Peter Grönquist , Deblina Bhattacharjee , Bahar Aydemir , Baran Ozaydin , Tong Zhang , Mathieu Salzmann , Sabine Süsstrunk

Researchers often rely on humans to code (label, annotate, etc.) large sets of texts. This kind of human coding forms an important part of social science research, yet the coding process is both resource intensive and highly variable from…

Artificial Intelligence · Computer Science 2023-06-06 Christopher Michael Rytting , Taylor Sorensen , Lisa Argyle , Ethan Busby , Nancy Fulda , Joshua Gubler , David Wingate

Building a dialogue system that can communicate naturally with humans is a challenging yet interesting problem of agent-based computing. The rapid growth in this area is usually hindered by the long-standing problem of data scarcity as…

Computation and Language · Computer Science 2021-04-23 Munazza Zaib , Quan Z. Sheng , Wei Emma Zhang

Despite advancements in speech recognition, accented speech remains challenging. While previous approaches have focused on modeling techniques or creating accented speech datasets, gathering sufficient data for the multitude of accents,…

Computation and Language · Computer Science 2024-02-06 Abraham Toluwase Owodunni , Aditya Yadavalli , Chris Chinenye Emezue , Tobi Olatunji , Clinton C Mbataku

Progress in AI is driven largely by the scale and quality of training data. Despite this, there is a deficit of empirical analysis examining the attributes of well-established datasets beyond text. In this work we conduct the largest and…

Recent improvements in multilingual ASR have not been equally distributed across languages and language varieties. To advance state-of-the-art (SOTA) ASR models, we present the Interspeech 2025 ML-SUPERB 2.0 Challenge. We construct a new…