中文
相关论文

相关论文: The EcoLexicon English Corpus as an open corpus in…

200 篇论文

Many projects have applied knowledge patterns (KPs) to the retrieval of specialized information. Yet terminologists still rely on manual analysis of concordance lines to extract semantic information, since there are no user-friendly…

计算与语言 · 计算机科学 2018-04-17 P. León-Araúz , A. San Martín

We introduce the Emergent Language Corpus Collection (ELCC): a collection of corpora generated from open source implementations of emergent communication systems across the literature. These systems include a variety of signalling game…

计算与语言 · 计算机科学 2024-12-05 Brendon Boldt , David Mortensen

We present OpenGloss, a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic definitions, encyclopedic context, etymological histories, and semantic relationships in a unified resource.…

计算与语言 · 计算机科学 2025-11-25 Michael J. Bommarito

In this paper we describe the Japanese-English Subtitle Corpus (JESC). JESC is a large Japanese-English parallel corpus covering the underrepresented domain of conversational dialogue. It consists of more than 3.2 million examples, making…

计算与语言 · 计算机科学 2018-02-22 Reid Pryzant , Yongjoo Chung , Dan Jurafsky , Denny Britz

The Architecture, Engineering, and Construction (AEC) industry is undergoing rapid digital transformation, producing diverse digital assets such as datasets, computational models, use cases, and educational materials across the built…

计算机与社会 · 计算机科学 2026-03-23 Ruoxin Xiong , Yanyu Wang , Jiannan Cai , Kaijian Liu , Yuansheng Zhu , Pingbo Tang , Nora El-Gohary , George Edward Gibson

We present DepCC, the largest-to-date linguistically analyzed corpus in English including 365 million documents, composed of 252 billion tokens and 7.5 billion of named entity occurrences in 14.3 billion sentences from a web-scale crawl of…

计算与语言 · 计算机科学 2018-03-01 Alexander Panchenko , Eugen Ruppert , Stefano Faralli , Simone Paolo Ponzetto , Chris Biemann

Semantic knowledge can be a great asset to natural language processing systems, but it is usually hand-coded for each application. Although some semantic information is available in general-purpose knowledge bases such as WordNet and Cyc,…

cmp-lg · 计算机科学 2008-02-03 Ellen Riloff , Jessica Shepherd

Automatic machine learning systems can inadvertently accentuate and perpetuate inappropriate human biases. Past work on examining inappropriate biases has largely focused on just individual systems. Further, there is no benchmark dataset…

计算与语言 · 计算机科学 2018-05-14 Svetlana Kiritchenko , Saif M. Mohammad

Open information extraction (OIE) systems extract relations and their arguments from natural language text in an unsupervised manner. The resulting extractions are a valuable resource for downstream tasks such as knowledge base…

计算与语言 · 计算机科学 2019-04-30 Kiril Gashteovski , Sebastian Wanner , Sven Hertling , Samuel Broscheit , Rainer Gemulla

The Eye Movements on Machine-Generated Texts Corpus (EMTeC) is a naturalistic eye-movements-while-reading corpus of 107 native English speakers reading machine-generated texts. The texts are generated by three large language models using…

This paper presents a syntactic lexicon for English that was originally derived from the Oxford Advanced Learner's Dictionary and the Oxford Dictionary of Current Idiomatic English, and then modified and augmented by hand. There are more…

cmp-lg · 计算机科学 2008-02-03 Dania Egedi , Patrick Martin

This paper introduces a new open-source speech corpus named "speechocean762" designed for pronunciation assessment use, consisting of 5000 English utterances from 250 non-native speakers, where half of the speakers are children. Five…

计算与语言 · 计算机科学 2021-06-03 Junbo Zhang , Zhiwen Zhang , Yongqing Wang , Zhiyong Yan , Qiong Song , Yukai Huang , Ke Li , Daniel Povey , Yujun Wang

CODEC is a document and entity ranking benchmark that focuses on complex research topics. We target essay-style information needs of social science researchers, i.e. "How has the UK's Open Banking Regulation benefited Challenger Banks?".…

信息检索 · 计算机科学 2022-05-18 Iain Mackie , Paul Owoicho , Carlos Gemmell , Sophie Fischer , Sean MacAvaney , Jeffrey Dalton

Service-based IT infrastructures are today's trend and the future for every enterprise willing to support dynamic and agile business to contend with the ever changing e-demands and requirements. A digital ecosystem is an emerging business…

软件工程 · 计算机科学 2012-04-03 Youssef Bassil

The need for large text corpora has increased with the advent of pretrained language models and, in particular, the discovery of scaling laws for these models. Most available corpora have sufficient data only for languages with large…

计算与语言 · 计算机科学 2025-03-05 Amir Hossein Kargaran , François Yvon , Hinrich Schütze

We release a corpus of 43 million atomic edits across 8 languages. These edits are mined from Wikipedia edit history and consist of instances in which a human editor has inserted a single contiguous phrase into, or deleted a single…

计算与语言 · 计算机科学 2018-08-29 Manaal Faruqui , Ellie Pavlick , Ian Tenney , Dipanjan Das

The Universal Knowledge Core (UKC) is a large multilingual lexical database with a focus on language diversity and covering over a thousand languages. The aim of the database, as well as its tools and data catalogue, is to make the somewhat…

We present a freely available, genre-balanced English web corpus totaling 4M tokens and featuring a large number of high-quality automatic annotation layers, including dependency trees, non-named entity annotations, coreference resolution,…

计算与语言 · 计算机科学 2020-06-19 Luke Gessler , Siyao Peng , Yang Liu , Yilun Zhu , Shabnam Behzad , Amir Zeldes

We describe the design of Comlex Syntax, a computational lexicon providing detailed syntactic information for approximately 38,000 English headwords. We consider the types of errors which arise in creating such a lexicon, and how such…

cmp-lg · 计算机科学 2009-09-25 Ralph Grishman , Catherine Macleod , Adam Meyers

We present Lit2Vec, a reproducible workflow for constructing and validating a chemistry corpus from the Semantic Scholar Open Research Corpus using conservative, metadata-based license screening. Using this workflow, we assembled an…

数据库 · 计算机科学 2026-04-15 Mahmoud Amiri , Jamile Mohammad Jafari , Sara Mostafapour , Thomas Bocklitz
‹ 上一页 1 2 3 10 下一页 ›