English
Related papers

Related papers: ViSoBERT: A Pre-Trained Language Model for Vietnam…

200 papers

Transformers are the most eminent architectures used for a vast range of Natural Language Processing tasks. These models are pre-trained over a large text corpus and are meant to serve state-of-the-art results over tasks like text…

Computation and Language · Computer Science 2022-11-15 Abhishek Velankar , Hrushikesh Patil , Raviraj Joshi

Natural Language Understanding (NLU) for low-resource languages remains a major challenge in NLP due to the scarcity of high-quality data and language-specific models. Maithili, despite being spoken by millions, lacks adequate computational…

Computation and Language · Computer Science 2026-02-03 Sumit Yadav , Raju Kumar Yadav , Utsav Maskey , Gautam Siddharth Kashyap , Ganesh Gautam , Usman Naseem

With the burgeoning amount of data of image-text pairs and diversity of Vision-and-Language (V\&L) tasks, scholars have introduced an abundance of deep learning models in this research domain. Furthermore, in recent years, transfer learning…

Computation and Language · Computer Science 2024-12-12 Thong Nguyen , Cong-Duy Nguyen , Xiaobao Wu , See-Kiong Ng , Anh Tuan Luu

We present VietNormalizer1, an open-source, zero-dependency Python library for Vietnamese text normalization targeting Text-to-Speech (TTS) and Natural Language Processing (NLP) applications. Vietnamese text normalization is a critical yet…

Computation and Language · Computer Science 2026-03-05 Hung Vu Nguyen , Loan Do , Thanh Ngoc Nguyen , Ushik Shrestha Khwakhali , Thanh Pham , Vinh Do , Charlotte Nguyen , Hien Nguyen

Spoken Named Entity Recognition (NER) aims to extract named entities from speech and categorise them into types like person, location, organization, etc. In this work, we present VietMed-NER - the first spoken NER dataset in the medical…

Audio and Speech Processing · Electrical Eng. & Systems 2025-04-03 Khai Le-Duc , David Thulke , Hung-Phong Tran , Long Vo-Dang , Khai-Nguyen Nguyen , Truong-Son Hy , Ralf Schlüter

Large-scale auto-regressive language models pretrained on massive text have demonstrated their impressive ability to perform new natural language tasks with only a few text examples, without the need for fine-tuning. Recent studies further…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-15 Heting Gao , Junrui Ni , Kaizhi Qian , Yang Zhang , Shiyu Chang , Mark Hasegawa-Johnson

Mental health is a critical issue in modern society, and mental disorders could sometimes turn to suicidal ideation without adequate treatment. Early detection of mental disorders and suicidal ideation from social content provides a…

Computation and Language · Computer Science 2022-07-19 Shaoxiong Ji , Tianlin Zhang , Luna Ansari , Jie Fu , Prayag Tiwari , Erik Cambria

This paper presents XLS-R, a large-scale model for cross-lingual speech representation learning based on wav2vec 2.0. We train models with up to 2B parameters on nearly half a million hours of publicly available speech audio in 128…

We introduce FaBERT, a Persian BERT-base model pre-trained on the HmBlogs corpus, encompassing both informal and formal Persian texts. FaBERT is designed to excel in traditional Natural Language Understanding (NLU) tasks, addressing the…

Computation and Language · Computer Science 2024-02-12 Mostafa Masumi , Seyed Soroush Majd , Mehrnoush Shamsfard , Hamid Beigy

Recent advances in self-supervised speech models have shown significant improvement in many downstream tasks. However, these models predominantly centered on frame-level training objectives, which can fall short in spoken language…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-09 Hung-Chieh Fang , Nai-Xuan Ye , Yi-Jen Shih , Puyuan Peng , Hsuan-Fu Wang , Layne Berry , Hung-yi Lee , David Harwath

Large Language Models pre-trained with self-supervised learning have demonstrated impressive zero-shot generalization capabilities on a wide spectrum of tasks. In this work, we present WeLM: a well-read pre-trained language model for…

Computation and Language · Computer Science 2023-05-17 Hui Su , Xiao Zhou , Houjin Yu , Xiaoyu Shen , Yuwen Chen , Zilin Zhu , Yang Yu , Jie Zhou

The Arabic language is a morphologically rich language with relatively few resources and a less explored syntax compared to English. Given these limitations, Arabic Natural Language Processing (NLP) tasks like Sentiment Analysis (SA), Named…

Computation and Language · Computer Science 2021-03-09 Wissam Antoun , Fady Baly , Hazem Hajj

French language models, such as CamemBERT, have been widely adopted across industries for natural language processing (NLP) tasks, with models like CamemBERT seeing over 4 million downloads per month. However, these models face challenges…

Computation and Language · Computer Science 2024-11-14 Wissam Antoun , Francis Kulumba , Rian Touchent , Éric de la Clergerie , Benoît Sagot , Djamé Seddah

This paper presents UniBERT, a compact multilingual language model that uses an innovative training framework that integrates three components: masked language modeling, adversarial training, and knowledge distillation. Pre-trained on a…

Computation and Language · Computer Science 2025-09-03 Andrei-Marius Avram , Marian Lupaşcu , Dumitru-Clementin Cercel , Ionuţ Mironică , Ştefan Trăuşan-Matu

Recently, fine-tuning large pre-trained Transformer models using downstream datasets has received a rising interest. Despite their success, it is still challenging to disentangle the benefits of large-scale datasets and Transformer…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-19 Junyi Peng , Oldřich Plchot , Themos Stafylakis , Ladislav Mošner , Lukáš Burget , Jan Černocký

Data is a cornerstone for fine-tuning large language models, yet acquiring suitable data remains challenging. Challenges encompassed data scarcity, linguistic diversity, and domain-specific content. This paper presents lessons learned while…

Computation and Language · Computer Science 2023-11-03 Thanh Nguyen Ngoc , Quang Nhat Tran , Arthur Tang , Bao Nguyen , Thuy Nguyen , Thanh Pham

To the best of our knowledge, this paper made the first attempt to answer whether word segmentation is necessary for Vietnamese sentiment classification. To do this, we presented five pre-trained monolingual S4- based language models for…

Computation and Language · Computer Science 2023-01-03 Duc-Vu Nguyen , Ngan Luu-Thuy Nguyen

In recent years, pre-trained language models (PLMs) achieve the best performance on a wide range of natural language processing (NLP) tasks. While the first models were trained on general domain data, specialized ones have emerged to more…

Computation and Language · Computer Science 2023-05-08 Yanis Labrak , Adrien Bazoge , Richard Dufour , Mickael Rouvier , Emmanuel Morin , Béatrice Daille , Pierre-Antoine Gourraud

Biomedical text mining is becoming increasingly important as the number of biomedical documents rapidly grows. With the progress in natural language processing (NLP), extracting valuable information from biomedical literature has gained…

Computation and Language · Computer Science 2019-10-21 Jinhyuk Lee , Wonjin Yoon , Sungdong Kim , Donghyeon Kim , Sunkyu Kim , Chan Ho So , Jaewoo Kang

We investigate the reasoning ability of pretrained vision and language (V&L) models in two tasks that require multimodal integration: (1) discriminating a correct image-sentence pair from an incorrect one, and (2) counting entities in an…

Computer Vision and Pattern Recognition · Computer Science 2021-06-18 Letitia Parcalabescu , Albert Gatt , Anette Frank , Iacer Calixto