中文
相关论文

相关论文: SiDiaC-v.2.0: Sinhala Diachronic Corpus Version 2.…

200 篇论文

In the recent years, transformer-based models have lead to significant advances in language modelling for natural language processing. However, they require a vast amount of data to be (pre-)trained and there is a lack of corpora in…

This draft is a working document, having a summary of nighty-four (94) papers with additional sections on Traceability of Software Requirements (Section 4), Formal Methods and Its Tools (Section 5), Unifying Theories of Programming (UTP)…

软件工程 · 计算机科学 2025-06-24 Arshad Beg , Diarmuid O'Donoghue , Rosemary Monahan

This paper introduces a new corpus of Mandarin-English code-switching speech recognition--TALCS corpus, suitable for training and evaluating code-switching speech recognition systems. TALCS corpus is derived from real online one-to-one…

计算与语言 · 计算机科学 2022-06-28 Chengfei Li , Shuhao Deng , Yaoping Wang , Guangjing Wang , Yaguang Gong , Changbin Chen , Jinfeng Bai

This paper presents the "Speak & Improve Challenge 2025: Spoken Language Assessment and Feedback" -- a challenge associated with the ISCA SLaTE 2025 Workshop. The goal of the challenge is to advance research on spoken language assessment…

计算与语言 · 计算机科学 2024-12-18 Mengjie Qian , Kate Knill , Stefano Banno , Siyuan Tang , Penny Karanasou , Mark J. F. Gales , Diane Nicholls

Recent years showed a strong increase in biomedical sciences and an inherent increase in publication volume. Extraction of specific information from these sources requires highly sophisticated text mining and information extraction tools.…

计算与语言 · 计算机科学 2020-04-09 Johannes Kirschnick , Philippe Thomas , Roland Roller , Leonhard Hennig

We present PashtoCorp, a 1.25-billion-word corpus for Pashto, a language spoken by 60 million people that remains severely underrepresented in NLP. The corpus is assembled from 39 sources spanning seven HuggingFace datasets and 32…

计算与语言 · 计算机科学 2026-03-18 Hanif Rahman

Scientific articles published prior to the "age of digitization" (~1997) require Optical Character Recognition (OCR) to transform scanned documents into machine-readable text, a process that often produces errors. We develop a pipeline for…

数字图书馆 · 计算机科学 2023-09-22 Jill P. Naiman , Morgan G. Cosillo , Peter K. G. Williams , Alyssa Goodman

Language modeling has witnessed remarkable advancements in recent years, with Large Language Models (LLMs) like ChatGPT setting unparalleled benchmarks in human-like text generation. However, a prevailing limitation is the…

计算与语言 · 计算机科学 2023-11-13 Abhinand Balachandran

Large Language Models (LLMs) demonstrate remarkable fluency across high-resource languages yet consistently fail to generate coherent text in Kashmiri, a language spoken by approximately seven million people. This performance disparity…

计算与语言 · 计算机科学 2026-01-06 Haq Nawaz Malik

This paper presents the Norwegian Review Corpus (NoReC), created for training and evaluating models for document-level sentiment analysis. The full-text reviews have been collected from major Norwegian news sources and cover a range of…

This paper introduces AFRIDOC-MT, a document-level multi-parallel translation dataset covering English and five African languages: Amharic, Hausa, Swahili, Yor\`ub\'a, and Zulu. The dataset comprises 334 health and 271 information…

We present an expanded version of our previously released Kazakh text-to-speech (KazakhTTS) synthesis corpus. In the new KazakhTTS2 corpus, the overall size has increased from 93 hours to 271 hours, the number of speakers has risen from two…

音频与语音处理 · 电气工程与系统科学 2022-04-21 Saida Mussakhojayeva , Yerbolat Khassanov , Huseyin Atakan Varol

This research provides the first comprehensive analysis of the performance of pre-trained language models for Sinhala text classification. We test on a set of different Sinhala text classification tasks and our analysis shows that out of…

计算与语言 · 计算机科学 2022-08-18 Vinura Dhananjaya , Piyumal Demotte , Surangika Ranathunga , Sanath Jayasena

This paper presents two significant contributions: First, it introduces a novel dataset of 19th-century Latin American newspaper texts, addressing a critical gap in specialized corpora for historical and linguistic analysis in this region.…

计算与语言 · 计算机科学 2025-03-31 Laura Manrique-Gómez , Tony Montes , Arturo Rodríguez-Herrera , Rubén Manrique

Speech is a hierarchical collection of text, prosody, emotions, dysfluencies, etc. Automatic transcription of speech that goes beyond text (words) is an underexplored problem. We focus on transcribing speech along with non-fluencies…

音频与语音处理 · 电气工程与系统科学 2024-12-03 Jiachen Lian , Xuanru Zhou , Zoe Ezzes , Jet Vonk , Brittany Morin , David Baquirin , Zachary Mille , Maria Luisa Gorno Tempini , Gopala Krishna Anumanchipalli

The need for large text corpora has increased with the advent of pretrained language models and, in particular, the discovery of scaling laws for these models. Most available corpora have sufficient data only for languages with large…

计算与语言 · 计算机科学 2025-03-05 Amir Hossein Kargaran , François Yvon , Hinrich Schütze

The visual quality of an image is confounded by a number of intertwined factors including its semantic content, distortion characteristics and appearance properties such as brightness, contrast, sharpness, and colourfulness. Distilling high…

计算机视觉与模式识别 · 计算机科学 2024-12-11 Fei Zhou , Tianhao Gu , Zhicong Huang , Guoping Qiu

We introduce the Self-Annotated Reddit Corpus (SARC), a large corpus for sarcasm research and for training and evaluating systems for sarcasm detection. The corpus has 1.3 million sarcastic statements -- 10 times more than any previous…

计算与语言 · 计算机科学 2018-03-26 Mikhail Khodak , Nikunj Saunshi , Kiran Vodrahalli

In this work, we present to the NLP community, and to the wider research community as a whole, an application for the diachronic analysis of research corpora. We open source an easy-to-use tool coined: DRIFT, which allows researchers to…

计算与语言 · 计算机科学 2021-09-13 Abheesht Sharma , Gunjan Chhablani , Harshit Pandey , Rajaswa Patil

The use of Project Gutenberg (PG) as a text corpus has been extremely popular in statistical analysis of language for more than 25 years. However, in contrast to other major linguistic datasets of similar importance, no consensual full…

计算与语言 · 计算机科学 2018-12-20 Martin Gerlach , Francesc Font-Clos