English
Related papers

Related papers: A Baseline Readability Model for Cebuano

200 papers

Language modelling and machine translation tasks mostly use subword or character inputs, but syllables are seldom used. Syllables provide shorter sequences than characters, require less-specialised extracting rules than morphemes, and their…

Computation and Language · Computer Science 2022-10-07 Arturo Oncevay , Kervy Dante Rivas Rojas , Liz Karen Chavez Sanchez , Roberto Zariquiey

At present Automatic Speaker Recognition system is a very important issue due to its diverse applications. Hence, it becomes absolutely necessary to obtain models that take into consideration the speaking style of a person, vocal tract…

Vision-language models have progressed rapidly, but Tibetan remains a severely underserved low-resource language due to the lack of reproducible training and evaluation infrastructure. To fill this gap, we introduce FTibSuite, a…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Guixian Xu , Yide Liang , Zeli Su , Xuexian Song , Ziyin Zhang , Yushuang Dong , Ting Zhang , Xu Han

The latest work on language representations carefully integrates contextualized features into language model training, which enables a series of success especially in various machine reading comprehension and natural language inference…

Computation and Language · Computer Science 2020-02-05 Zhuosheng Zhang , Yuwei Wu , Hai Zhao , Zuchao Li , Shuailiang Zhang , Xi Zhou , Xiang Zhou

Evaluating text comprehension in educational settings is critical for understanding student performance and improving curricular effectiveness. This study investigates the capability of state-of-the-art language models-RoBERTa Base,…

Computation and Language · Computer Science 2024-12-25 Abdullah Khondoker , Enam Ahmed Taufik , Md Iftekhar Islam Tashik , S M Ishtiak mahmud , Antara Firoz Parsa

We present foundation language models developed to power Apple Intelligence features, including a ~3 billion parameter model designed to run efficiently on devices and a large server-based language model designed for Private Cloud Compute.…

Artificial Intelligence · Computer Science 2026-05-28 Tom Gunter , Zirui Wang , Chong Wang , Ruoming Pang , Andy Narayanan , Aonan Zhang , Bowen Zhang , Chen Chen , Chung-Cheng Chiu , David Qiu , Deepak Gopinath , Dian Ang Yap , Dong Yin , Feng Nan , Floris Weers , Guoli Yin , Haoshuo Huang , Jianyu Wang , Jiarui Lu , John Peebles , Ke Ye , Mark Lee , Nan Du , Qibin Chen , Quentin Keunebroek , Sam Wiseman , Syd Evans , Tao Lei , Vivek Rathod , Xiang Kong , Xianzhi Du , Yanghao Li , Yongqiang Wang , Yuan Gao , Zaid Ahmed , Zhaoyang Xu , Zhiyun Lu , Al Rashid , Albin Madappally Jose , Alec Doane , Alfredo Bencomo , Allison Vanderby , Andrew Hansen , Ankur Jain , Anupama Mann Anupama , Areeba Kamal , Bugu Wu , Carolina Brum , Charlie Maalouf , Chinguun Erdenebileg , Chris Dulhanty , Daniel Parilla , Dominik Moritz , Doug Kang , Eduardo Jimenez , Evan Ladd , Fangping Shi , Felix Bai , Frank Chu , Fred Hohman , Hadas Kotek , Hannah Gillis Coleman , Jane Li , Jeffrey Bigham , Jeffery Cao , Jeff Lai , Jessica Cheung , Jiulong Shan , Joe Zhou , John Li , Jun Qin , Karanjeet Singh , Karla Vega , Kelvin Zou , Laura Heckman , Lauren Gardiner , Margit Bowler , Maria Cordell , Meng Cao , Nicole Hay , Nilesh Shahdadpuri , Otto Godwin , Pranay Dighe , Pushyami Rachapudi , Ramsey Tantawi , Roman Frigg , Sam Davarnia , Sanskruti Shah , Saptarshi Guha , Sasha Sirovica , Shen Ma , Shuang Ma , Simon Wang , Sulgi Kim , Suma Jayaram , Vaishaal Shankar , Varsha Paidi , Vivek Kumar , Xin Wang , Xin Zheng , Walker Cheng , Yael Shrager , Yang Ye , Yasu Tanaka , Yihao Guo , Yunsong Meng , Zhao Tang Luo , Zhi Ouyang , Alp Aygar , Alvin Wan , Andrew Walkingshaw , Andy Narayanan , Antonie Lin , Arsalan Farooq , Brent Ramerth , Colorado Reed , Chris Bartels , Chris Chaney , David Riazati , Eric Liang Yang , Erin Feldman , Gabriel Hochstrasser , Guillaume Seguin , Irina Belousova , Joris Pelemans , Karen Yang , Keivan Alizadeh Vahid , Liangliang Cao , Mahyar Najibi , Marco Zuliani , Max Horton , Minsik Cho , Nikhil Bhendawade , Patrick Dong , Piotr Maj , Pulkit Agrawal , Qi Shan , Qichen Fu , Regan Poston , Sam Xu , Shuangning Liu , Sushma Rao , Tashweena Heeramun , Thomas Merth , Uday Rayala , Victor Cui , Vivek Rangarajan Sridhar , Wencong Zhang , Wenqi Zhang , Wentao Wu , Xingyu Zhou , Xinwen Liu , Yang Zhao , Yin Xia , Zhile Ren , Zhongzheng Ren

Large Language Models (LLMs) achieve remarkable performance across various tasks, but their tendency to produce hallucinations limits reliable adoption. Benchmarks such as TruthfulQA have been developed to measure truthfulness, yet they are…

Computation and Language · Computer Science 2025-09-09 Lorenzo Alfred Nery , Ronald Dawson Catignas , Thomas James Tiam-Lee

While cross-lingual word embeddings have been studied extensively in recent years, the qualitative differences between the different algorithms remain vague. We observe that whether or not an algorithm uses a particular feature set…

Computation and Language · Computer Science 2017-01-11 Omer Levy , Anders Søgaard , Yoav Goldberg

Increasing model size when pretraining natural language representations often results in improved performance on downstream tasks. However, at some point further model increases become harder due to GPU/TPU memory limitations and longer…

Computation and Language · Computer Science 2020-02-11 Zhenzhong Lan , Mingda Chen , Sebastian Goodman , Kevin Gimpel , Piyush Sharma , Radu Soricut

We address a notable gap in Natural Language Processing (NLP) by introducing a collection of resources designed to improve Machine Translation (MT) for low-resource languages, with a specific focus on African languages. First, we introduce…

Computation and Language · Computer Science 2024-07-15 AbdelRahim Elmadany , Ife Adebara , Muhammad Abdul-Mageed

We propose that small pretrained foundational generative language models with millions of parameters can be utilized as a general learning framework for sequence-based tasks. Our proposal overcomes the computational resource, skill set, and…

Computation and Language · Computer Science 2024-02-09 Ben Fauber

Language Models (LMs) have been ubiquitously leveraged in various tasks including spoken language understanding (SLU). Spoken language requires careful understanding of speaker interactions, dialog states and speech induced multimodal…

Computation and Language · Computer Science 2021-09-22 Ayush Kumar , Mukuntha Narayanan Sundararaman , Jithendra Vepa

Several recent studies have shown that strong natural language understanding (NLU) models are prone to relying on unwanted dataset biases without learning the underlying task, resulting in models that fail to generalize to out-of-domain…

Computation and Language · Computer Science 2020-04-27 Rabeeh Karimi Mahabadi , Yonatan Belinkov , James Henderson

In this paper, we study pre-trained sequence-to-sequence models for a group of related languages, with a focus on Indic languages. We present IndicBART, a multilingual, sequence-to-sequence pre-trained model focusing on 11 Indic languages…

Computation and Language · Computer Science 2022-10-28 Raj Dabre , Himani Shrotriya , Anoop Kunchukuttan , Ratish Puduppully , Mitesh M. Khapra , Pratyush Kumar

BERT-based models are currently used for solving nearly all Natural Language Processing (NLP) tasks and most often achieve state-of-the-art results. Therefore, the NLP community conducts extensive research on understanding these models, but…

Computation and Language · Computer Science 2021-05-06 Robert Mroczkowski , Piotr Rybak , Alina Wróblewska , Ireneusz Gawlik

Large language models have made tremendous progress in recent years, but low-resource languages, like Tibetan, remain significantly underrepresented in their evaluation. Despite Tibetan being spoken by over seven million people, it has…

Transformer-based language models are now widely used in Natural Language Processing (NLP). This statement is especially true for English language, in which many pre-trained models utilizing transformer-based architecture have been…

Computation and Language · Computer Science 2020-06-11 Sławomir Dadas , Michał Perełkiewicz , Rafał Poświata

Knowledge base construction entails acquiring structured information to create a knowledge base of factual and relational data, facilitating question answering, information retrieval, and semantic understanding. The challenge called…

Computation and Language · Computer Science 2023-10-13 Dong Yang , Xu Wang , Remzi Celebi

Structured language models for speech recognition have been shown to remedy the weaknesses of n-gram models. All current structured language models are, however, limited in that they do not take into account dependencies between…

Computation and Language · Computer Science 2007-05-23 Rens Bod

Machine translation systems achieve near human-level performance on some languages, yet their effectiveness strongly relies on the availability of large amounts of parallel sentences, which hinders their applicability to the majority of…

Computation and Language · Computer Science 2018-08-15 Guillaume Lample , Myle Ott , Alexis Conneau , Ludovic Denoyer , Marc'Aurelio Ranzato