中文
相关论文

相关论文: RuCoCo: a new Russian corpus with coreference anno…

200 篇论文

This paper presents NorNE, a manually annotated corpus of named entities which extends the annotation of the existing Norwegian Dependency Treebank. Comprising both of the official standards of written Norwegian (Bokm{\aa}l and Nynorsk),…

计算与语言 · 计算机科学 2020-03-09 Fredrik Jørgensen , Tobias Aasmoe , Anne-Stine Ruud Husevåg , Lilja Øvrelid , Erik Velldal

We compiled a new sentence splitting corpus that is composed of 203K pairs of aligned complex source and simplified target sentences. Contrary to previously proposed text simplification corpora, which contain only a small number of split…

计算与语言 · 计算机科学 2019-09-27 Christina Niklaus , Andre Freitas , Siegfried Handschuh

Recognizing non-standard entity types and relations, such as B2B products, product classes and their producers, in news and forum texts is important in application areas such as supply chain monitoring and market research. However, there is…

计算与语言 · 计算机科学 2020-04-08 Saskia Schön , Veselina Mironova , Aleksandra Gabryszak , Leonhard Hennig

Argumentation mining is a field of computational linguistics that is devoted to extracting from texts and classifying arguments and relations between them, as well as constructing an argumentative structure. A significant obstacle to…

计算与语言 · 计算机科学 2021-06-29 Irina Fishcheva , Valeriya Goloviznina , Evgeny Kotelnikov

TACO is an open image dataset for litter detection and segmentation, which is growing through crowdsourcing. Firstly, this paper describes this dataset and the tools developed to support it. Secondly, we report instance segmentation…

计算机视觉与模式识别 · 计算机科学 2020-03-18 Pedro F Proença , Pedro Simões

We introduce the Self-Annotated Reddit Corpus (SARC), a large corpus for sarcasm research and for training and evaluating systems for sarcasm detection. The corpus has 1.3 million sarcastic statements -- 10 times more than any previous…

计算与语言 · 计算机科学 2018-03-26 Mikhail Khodak , Nikunj Saunshi , Kiran Vodrahalli

In this study we address the problem of automated word stress detection in Russian using character level models and no part-speech-taggers. We use a simple bidirectional RNN with LSTM nodes and achieve the accuracy of 90% or higher. We…

计算与语言 · 计算机科学 2019-07-15 Maria Ponomareva , Kirill Milintsevich , Ekaterina Chernyak , Anatoly Starostin

Automatic text summarization aims to produce a brief but crucial summary for the input documents. Both extractive and abstractive methods have witnessed great success in English datasets in recent years. However, there has been a minimal…

计算与语言 · 计算机科学 2021-10-22 Danqing Wang , Jiaze Chen , Xianze Wu , Hao Zhou , Lei Li

Automatic text summarization is widely regarded as the highly difficult problem, partially because of the lack of large text summarization data set. Due to the great challenge of constructing the large scale summaries for full text, in this…

计算与语言 · 计算机科学 2016-02-22 Baotian Hu , Qingcai Chen , Fangze Zhu

SberQuAD -- a large scale analog of Stanford SQuAD in the Russian language - is a valuable resource that has not been properly presented to the scientific community. We fill this gap by providing a description, a thorough analysis, and…

计算与语言 · 计算机科学 2020-11-02 Pavel Efimov , Andrey Chertok , Leonid Boytsov , Pavel Braslavski

Information Extraction is a well-researched area of Natural Language Processing with applications in web search and question answering concerned with identifying entities and relationships between them as expressed in a given context,…

信息检索 · 计算机科学 2020-11-17 Erin Macdonald , Denilson Barbosa

We present a new corpus comprising annotations of medical entities in case reports, originating from PubMed Central's open access library. In the case reports, we annotate cases, conditions, findings, factors and negation modifiers.…

计算与语言 · 计算机科学 2020-03-31 Sarah Schulz , Jurica Ševa , Samuel Rodriguez , Malte Ostendorff , Georg Rehm

This paper introduces the Balanced Arabic Readability Evaluation Corpus (BAREC), a large-scale, fine-grained dataset for Arabic readability assessment. BAREC consists of 69,441 sentences spanning 1+ million words, carefully curated to cover…

计算与语言 · 计算机科学 2025-06-17 Khalid N. Elmadani , Nizar Habash , Hanada Taha-Thomure

We report on a language resource consisting of 2000 annotated bibliography entries, which is being analyzed as part of our research on indicative document summarization. We show how annotated bibliographies cover certain aspects of…

计算与语言 · 计算机科学 2007-05-23 Min-Yen Kan , Judith L. Klavans , Kathleen R. McKeown

This paper presents an extension to a very low-resource parallel corpus collected in an endangered language, Griko, making it useful for computational research. The corpus consists of 330 utterances (about 20 minutes of speech) which have…

We recorded and preprocessed ZuCo 2.0, a new dataset of simultaneous eye-tracking and electroencephalography during natural reading and during annotation. This corpus contains gaze and brain activity data of 739 sentences, 349 in a normal…

计算与语言 · 计算机科学 2020-03-10 Nora Hollenstein , Marius Troendle , Ce Zhang , Nicolas Langer

We present SciDMT, an enhanced and expanded corpus for scientific mention detection, offering a significant advancement over existing related resources. SciDMT contains annotated scientific documents for datasets (D), methods (M), and tasks…

人工智能 · 计算机科学 2024-06-24 Huitong Pan , Qi Zhang , Cornelia Caragea , Eduard Dragut , Longin Jan Latecki

This paper presents a large-scale corpus for non-task-oriented dialogue response selection, which contains over 27K distinct prompts more than 82K responses collected from social media. To annotate this corpus, we define a 5-grade rating…

计算与语言 · 计算机科学 2018-05-16 Jing Li , Yan Song , Haisong Zhang , Shuming Shi

Coreference resolution aims to identify words and phrases which refer to same entity in a text, a core task in natural language processing. In this paper, we extend this task to resolving coreferences in long-form narrations of visual…

计算机视觉与模式识别 · 计算机科学 2023-03-20 Arushi Goel , Basura Fernando , Frank Keller , Hakan Bilen

Abstract. When writing an academic paper, researchers often spend considerable time reviewing and summarizing papers to extract relevant citations and data to compose the Introduction and Related Work sections. To address this problem, we…

信息检索 · 计算机科学 2023-06-22 Juan Ramirez-Orta , Eduardo Xamena , Ana Maguitman , Axel J. Soto , Flavia P. Zanoto , Evangelos Milios