中文
相关论文

相关论文: AGRR-2019: A Corpus for Gapping Resolution in Russ…

200 篇论文

Analyzing how humans revise their writings is an interesting research question, not only from an educational perspective but also in terms of artificial intelligence. Better understanding of this process could facilitate many NLP…

计算与语言 · 计算机科学 2022-06-06 Omid Kashefi , Tazin Afrin , Meghan Dale , Christopher Olshefski , Amanda Godley , Diane Litman , Rebecca Hwa

We present a new corpus with coreference annotation, Russian Coreference Corpus (RuCoCo). The goal of RuCoCo is to obtain a large number of annotated texts while maintaining high inter-annotator agreement. RuCoCo contains news texts in…

计算与语言 · 计算机科学 2022-06-13 Vladimir Dobrovolskii , Mariia Michurina , Alexandra Ivoylova

AI generated content (AIGC) presents considerable challenge to educators around the world. Instructors need to be able to detect such text generated by large language models, either with the naked eye or with the help of some tools. There…

计算与语言 · 计算机科学 2023-09-26 Yikang Liu , Ziyin Zhang , Wanyang Zhang , Shisen Yue , Xiaojing Zhao , Xinyuan Cheng , Yiwen Zhang , Hai Hu

In this paper, we present a system for information extraction from scientific texts in the Russian language. The system performs several tasks in an end-to-end manner: term recognition, extraction of relations between terms, and term…

计算与语言 · 计算机科学 2021-09-15 Elena Bruches , Anastasia Mezentseva , Tatiana Batura

Automatic speech recognition (ASR) systems are designed to transcribe spoken language into written text and find utility in a variety of applications including voice assistants and transcription services. However, it has been observed that…

计算与语言 · 计算机科学 2023-07-21 Anand Kumar Rai , Siddharth D Jaiswal , Animesh Mukherjee

Warning: this work contains upsetting or disturbing content. Large language models (LLMs) tend to learn the social and cultural biases present in the raw pre-training data. To test if an LLM's behavior is fair, functional datasets are…

计算与语言 · 计算机科学 2024-03-27 Veronika Grigoreva , Anastasiia Ivanova , Ilseyar Alimova , Ekaterina Artemova

Retrieval-Augmented Generation (RAG) has become the standard approach for grounding large language models in information that was not available during training. While existing datasets and benchmarks focus on web or other public sources,…

信息检索 · 计算机科学 2026-05-21 Yuhong Sun , Joachim Rahmfeld , Chris Weaver , Weijia Chen , Roshan Desai , Wenxi Huang , Mark H. Butler

Large language models (LLMs) have been widely applied in question answering over scientific research papers. To enhance the professionalism and accuracy of responses, many studies employ external knowledge augmentation. However, existing…

计算与语言 · 计算机科学 2025-02-21 Jiayin Lan , Jiaqi Li , Baoxin Wang , Ming Liu , Dayong Wu , Shijin Wang , Bing Qin

Large pre-trained language models are capable of generating varied and fluent texts. Starting from the prompt, these models generate a narrative that can develop unpredictably. The existing methods of controllable text generation, which…

计算与语言 · 计算机科学 2022-06-22 Sergey Vychegzhanin , Evgeny Kotelnikov

We introduce Balalaika, an open-source, data-centric pipeline for processing audio and producing prosody-aware annotations. It combines semantic VAD for context-preserving segmentation, multi-ASR ensembling with ROVER consensus decoding,…

计算与语言 · 计算机科学 2026-05-05 Kirill Borodin , Nikita Vasiliev , Vasiliy Kudryavtsev , Maxim Maslov , Mikhail Gorodnichev , Grach Mkrtchian

Neural machine translation is the current state-of-the-art in machine translation. Although it is successful in a resource-rich setting, its applicability for low-resource language pairs is still debatable. In this paper, we explore the…

计算与语言 · 计算机科学 2019-10-02 Aidar Valeev , Ilshat Gibadullin , Albina Khusainova , Adil Khan

Retrieval-Augmented Generation (RAG) has emerged as a standard framework for knowledge-intensive NLP tasks, combining large language models (LLMs) with document retrieval from external corpora. Despite its widespread use, most RAG pipelines…

信息检索 · 计算机科学 2025-08-26 Mandeep Rathee , V Venktesh , Sean MacAvaney , Avishek Anand

We describe the AGReE system, which takes user-submitted passages as input and automatically generates grammar practice exercises that can be completed while reading. Multiple-choice practice items are generated for a variety of different…

计算与语言 · 计算机科学 2022-11-07 Sophia Chan , Swapna Somasundaran , Debanjan Ghosh , Mengxuan Zhao

Many errors in student essays can be explained by influence from the native language (L1). L1 interference refers to errors influenced by a speaker's first language, such as using stadion instead of stadium, reflecting lexical…

计算与语言 · 计算机科学 2026-03-10 Darya Kharlamova , Irina Proskurina

The review process is essential to ensure the quality of publications. Recently, the increase of submissions for top venues in machine learning and NLP has caused a problem of excessive burden on reviewers and has often caused concerns…

计算与语言 · 计算机科学 2021-09-30 Ana Sabina Uban , Cornelia Caragea

For a researcher, writing a good research statement is crucial but costs a lot of time and effort. To help researchers, in this paper, we propose the research statement generation (RSG) task which aims to summarize one's research…

信息检索 · 计算机科学 2021-04-30 Wenhao Wu , Sujian Li

This paper describes an English audio and textual dataset of debating speeches, a unique resource for the growing research field of computational argumentation and debating technologies. We detail the process of speech recording by…

Natural language processing (NLP) utilizes text data augmentation to overcome sample size constraints. Increasing the sample size is a natural and widely used strategy for alleviating these challenges. In this study, we chose Arabic to…

计算与语言 · 计算机科学 2024-11-08 Ahlam Alrehili , Areej Alhothali

Wiktionary is a unique, peculiar, valuable and original resource for natural language processing (NLP). The paper describes an open-source Wiktionary parser: its architecture and requirements followed by a description of Wiktionary features…

信息检索 · 计算机科学 2010-06-28 A. A. Krizhanovsky

Recent progress in speech processing has highlighted that high-quality performance across languages requires substantial training data for each individual language. While existing multilingual datasets cover many languages, they often…

计算与语言 · 计算机科学 2025-10-28 Samuel Pfisterer , Florian Grötschla , Luca A. Lanzendörfer , Florian Yan , Roger Wattenhofer