English
Related papers

Related papers: NusaBERT: Teaching IndoBERT to be Multilingual and…

200 papers

Being less resource languages, Indian-Indian and English-Indian language MT system developments faces the difficulty to translate various lexical phenomena. In this paper, we present our work on a comparative study of 440 phrase-based…

Computation and Language · Computer Science 2017-10-09 Sreelekha S , Pushpak Bhattacharyya

In this paper we focus on constructing useful embeddings of textual information in vacancies and resumes, which we aim to incorporate as features into job to job seeker matching models alongside other features. We explain our task where…

Computation and Language · Computer Science 2021-09-15 Dor Lavi , Volodymyr Medentsiy , David Graus

India's linguistic landscape, spanning 22 scheduled languages and hundreds of marginalized dialects, has driven rapid growth in NLP datasets, benchmarks, and pretrained models. However, no dedicated survey consolidates resources developed…

Computation and Language · Computer Science 2026-04-21 Raghvendra Kumar , Devankar Raj , Sriparna Saha

As language-specific training data tends to be sparsely available compared to English, document retrieval in many languages has been largely relying on multilingual models. In Japanese, the best performing deep-learning based retrieval…

Computation and Language · Computer Science 2024-09-24 Benjamin Clavié

In this paper, we introduce Neural Information Retrieval resources for 11 widely spoken Indian Languages (Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Oriya, Punjabi, Tamil, and Telugu) from two major Indian language…

Information Retrieval · Computer Science 2023-12-18 Saiful Haq , Ashutosh Sharma , Pushpak Bhattacharyya

India is a multilingual society with 1369 rationalized languages and dialects being spoken across the country (INDIA, 2011). Of these, the 22 scheduled languages have a staggering total of 1.17 billion speakers and 121 languages have more…

The multilingual BERT model is trained on 104 languages and meant to serve as a universal language model and tool for encoding sentences. We explore how well the model performs on several languages across several tasks: a diagnostic…

Computation and Language · Computer Science 2019-10-10 Samuel Rönnqvist , Jenna Kanerva , Tapio Salakoski , Filip Ginter

The development of high-performing, robust, and reliable speech technologies depends on large, high-quality datasets. However, African languages -- including our focus, Igbo, Hausa, and Yoruba -- remain under-represented due to insufficient…

We present SwissBERT, a masked language model created specifically for processing Switzerland-related text. SwissBERT is a pre-trained model that we adapted to news articles written in the national languages of Switzerland -- German,…

Computation and Language · Computer Science 2024-01-17 Jannis Vamvas , Johannes Graën , Rico Sennrich

Currently, text-to-image synthesis uses text encoder and image generator architecture. Research on this topic is challenging. This is because of the domain gap between natural language and vision. Nowadays, most research on this topic only…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Made Raharja Surya Mahadi , Nugraha Priya Utama

We present IndoNLI, the first human-elicited NLI dataset for Indonesian. We adapt the data collection protocol for MNLI and collect nearly 18K sentence pairs annotated by crowd workers and experts. The expert-annotated data is used…

Computation and Language · Computer Science 2022-03-30 Rahmad Mahendra , Alham Fikri Aji , Samuel Louvan , Fahrurrozi Rahman , Clara Vania

Multi-lingual contextualized embeddings, such as multilingual-BERT (mBERT), have shown success in a variety of zero-shot cross-lingual tasks. However, these models are limited by having inconsistent contextualized representations of…

Computation and Language · Computer Science 2020-07-14 Libo Qin , Minheng Ni , Yue Zhang , Wanxiang Che

Machine translation in low-resource language pairs faces significant challenges due to the scarcity of parallel corpora and linguistic resources. This study focuses on the case of English-Marathi language pairs, where existing datasets are…

Computation and Language · Computer Science 2024-09-05 Nidhi Kowtal , Tejas Deshpande , Raviraj Joshi

Speech translation for Indian languages remains a challenging task due to the scarcity of large-scale, publicly available datasets that capture the linguistic diversity and domain coverage essential for real-world applications. Existing…

We introduce a multilabel probing task to assess the morphosyntactic representations of word embeddings from multilingual language models. We demonstrate this task with multilingual BERT (Devlin et al., 2018), training probes for seven…

Computation and Language · Computer Science 2021-04-20 Naomi Tachikawa Shapiro , Amandalynne Paullada , Shane Steinert-Threlkeld

We probe the layers in multilingual BERT (mBERT) for phylogenetic and geographic language signals across 100 languages and compute language distances based on the mBERT representations. We 1) employ the language distances to infer and…

Computation and Language · Computer Science 2020-11-05 Taraka Rama , Lisa Beinborn , Steffen Eger

Yor\`ub\'a an African language with roughly 47 million speakers encompasses a continuum with several dialects. Recent efforts to develop NLP technologies for African languages have focused on their standard dialects, resulting in…

Computation and Language · Computer Science 2024-07-01 Orevaoghene Ahia , Anuoluwapo Aremu , Diana Abagyan , Hila Gonen , David Ifeoluwa Adelani , Daud Abolade , Noah A. Smith , Yulia Tsvetkov

Multilingual BERT (M-BERT) has been a huge success in both supervised and zero-shot cross-lingual transfer learning. However, this success has focused only on the top 104 languages in Wikipedia that it was trained on. In this paper, we…

Computation and Language · Computer Science 2020-04-29 Zihan Wang , Karthikeyan K , Stephen Mayhew , Dan Roth

Although the prediction of dialects is an important language processing task, with a wide range of applications, existing work is largely limited to coarse-grained varieties. Inspired by geolocation research, we propose the novel task of…

Computation and Language · Computer Science 2020-12-08 Muhammad Abdul-Mageed , Chiyu Zhang , AbdelRahim Elmadany , Lyle Ungar

Lakota, a critically endangered language of the Sioux people in North America, faces significant challenges due to declining fluency among younger generations. This paper introduces LakotaBERT, the first large language model (LLM) tailored…

Computation and Language · Computer Science 2025-03-25 Kanishka Parankusham , Rodrigue Rizk , KC Santosh
‹ Prev 1 3 4 5 6 7 10 Next ›