English
Related papers

Related papers: LANS: Large-scale Arabic News Summarization Corpus

200 papers

Millions of people around the world can not access content on the Web because most of the content is not readily available in their language. Machine translation (MT) systems have the potential to change this for many languages. Current MT…

Computation and Language · Computer Science 2021-12-16 Asmelash Teka Hadgu , Abel Aregawi , Adam Beaudoin

Automatic Text Summarization (ATS) is becoming relevant with the growth of textual data; however, with the popularization of public large-scale datasets, some recent machine learning approaches have focused on dense models and architectures…

Computation and Language · Computer Science 2023-03-07 Vinícius Camargo da Silva , João Paulo Papa , Kelton Augusto Pontara da Costa

In recent years, Large Language Models (LLMs) have become widely used in medical applications, such as clinical decision support, medical education, and medical question answering. Yet, these models are often English-centric, limiting their…

Computation and Language · Computer Science 2026-02-06 Chaimae Abouzahir , Congbo Ma , Nizar Habash , Farah E. Shamout

Dialectal Arabic (DA) speech data vary widely in domain coverage, dialect labeling practices, and recording conditions, complicating cross-dataset comparison and model evaluation. To characterize this landscape, we conduct a computational…

Computation and Language · Computer Science 2026-01-30 Peter Sullivan , AbdelRahim Elmadany , Alcides Alcoba Inciarte , Muhammad Abdul-Mageed

The role of predicting sarcasm in the text is known as automatic sarcasm detection. Given the prevalence and challenges of sarcasm in sentiment-bearing text, this is a critical phase in most sentiment analysis tasks. With the increasing…

Computation and Language · Computer Science 2021-08-04 Bashar Talafha , Muhy Eddin Za'ter , Samer Suleiman , Mahmoud Al-Ayyoub , Mohammed N. Al-Kabi

While natural language processing tools have been developed extensively for some of the world's languages, a significant portion of the world's over 7000 languages are still neglected. One reason for this is that evaluation datasets do not…

Computation and Language · Computer Science 2024-06-05 Chunlan Ma , Ayyoob ImaniGooghari , Haotian Ye , Renhao Pei , Ehsaneddin Asgari , Hinrich Schütze

Computer-assisted reading and analysis of text has various applications in the humanities and social sciences. The increasing size of many electronic text archives has the advantage of a more complete analysis but the disadvantage of taking…

Databases · Computer Science 2007-05-23 Steven Keith , Owen Kaser , Daniel Lemire

Automatic text summarization (ATS) is an emerging technology to assist clinicians in providing continuous and coordinated care. This study presents an approach to summarize doctor-patient dialogues using generative large language models…

Computation and Language · Computer Science 2024-03-21 Mengxian Lyu , Cheng Peng , Xiaohan Li , Patrick Balian , Jiang Bian , Yonghui Wu

Arabic Optical Character Recognition (OCR) is essential for converting vast amounts of Arabic print media into digital formats. However, training modern OCR models, especially powerful vision-language models, is hampered by the lack of…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Omer Nacar , Yasser Al-Habashi , Serry Sibaee , Adel Ammar , Wadii Boulila

We present AraLingBench: a fully human annotated benchmark for evaluating the Arabic linguistic competence of large language models (LLMs). The benchmark spans five core categories: grammar, morphology, spelling, reading comprehension, and…

Computation and Language · Computer Science 2025-12-10 Mohammad Zbeeb , Hasan Abed Al Kader Hammoud , Sina Mukalled , Nadine Rizk , Fatima Karnib , Issam Lakkis , Ammar Mohanna , Bernard Ghanem

Previous research in multi-document news summarization has typically concentrated on collating information that all sources agree upon. However, the summarization of diverse information dispersed across multiple articles about an event…

Computation and Language · Computer Science 2024-03-26 Kung-Hsiang Huang , Philippe Laban , Alexander R. Fabbri , Prafulla Kumar Choubey , Shafiq Joty , Caiming Xiong , Chien-Sheng Wu

In the past few years, neural abstractive text summarization with sequence-to-sequence (seq2seq) models have gained a lot of popularity. Many interesting techniques have been proposed to improve seq2seq models, making them capable of…

Computation and Language · Computer Science 2020-09-22 Tian Shi , Yaser Keneshloo , Naren Ramakrishnan , Chandan K. Reddy

The application of handwritten text recognition to historical works is highly dependant on accurate text line retrieval. A number of systems utilizing a robust baseline detection paradigm have emerged recently but the advancement of layout…

Computer Vision and Pattern Recognition · Computer Science 2019-07-10 Benjamin Kiessling , Daniel Stökl Ben Ezra , Matthew Thomas Miller

With the emergence of Web 2.0 technology and the expansion of on-line social networks, current Internet users have the ability to add their reviews, ratings and opinions on social media and on commercial and news web sites. Sentiment…

Computation and Language · Computer Science 2018-09-18 Mo'ath Alrefai , Hossam Faris , Ibrahim Aljarah

We present an ongoing initiative to provide open, very large, high-quality, and richly annotated textual datasets for almost 200 languages. At 30 trillion tokens, this is likely the largest generally available multilingual collection of LLM…

Analyzing vast textual data and summarizing key information from electronic health records imposes a substantial burden on how clinicians allocate their time. Although large language models (LLMs) have shown promise in natural language…

Large language models (LLMs) have emerged as a widely-used tool for information seeking, but their generated outputs are prone to hallucination. In this work, our aim is to allow LLMs to generate text with citations, improving their factual…

Computation and Language · Computer Science 2023-11-01 Tianyu Gao , Howard Yen , Jiatong Yu , Danqi Chen

This survey offers a comprehensive overview of Large Language Models (LLMs) designed for Arabic language and its dialects. It covers key architectures, including encoder-only, decoder-only, and encoder-decoder models, along with the…

Computation and Language · Computer Science 2026-05-20 Malak Mashaabi , Shahad Al-Khalifa , Hend Al-Khalifa

We present the results and findings of the First Nuanced Arabic Dialect Identification Shared Task (NADI). This Shared Task includes two subtasks: country-level dialect identification (Subtask 1) and province-level sub-dialect…

Computation and Language · Computer Science 2020-11-11 Muhammad Abdul-Mageed , Chiyu Zhang , Houda Bouamor , Nizar Habash

The rapid spread of fake news presents a significant global challenge, particularly in low-resource languages like Bangla, which lack adequate datasets and detection tools. Although manual fact-checking is accurate, it is expensive and slow…

Computation and Language · Computer Science 2025-02-18 Hrithik Majumdar Shibu , Shrestha Datta , Md. Sumon Miah , Nasrullah Sami , Mahruba Sharmin Chowdhury , Md. Saiful Islam