中文
相关论文

相关论文: NYTWIT: A Dataset of Novel Words in the New York T…

200 篇论文

There are a few prominent practices for conducting reviews of academic literature, including searching for specific keywords on Google Scholar or checking citations from some initial seed paper(s). These approaches serve a critical purpose…

人机交互 · 计算机科学 2021-10-01 Arpit Narechania , Alireza Karduni , Ryan Wesslen , Emily Wall

Millions of news articles from hundreds of thousands of sources around the globe appear in news aggregators every day. Consuming such a volume of news presents an almost insurmountable challenge. For example, a reader searching on…

计算与语言 · 计算机科学 2020-06-02 Joshua Bambrick , Minjie Xu , Andy Almonte , Igor Malioutov , Guim Perarnau , Vittorio Selo , Iat Chong Chan

We present NewsQA, a challenging machine comprehension dataset of over 100,000 human-generated question-answer pairs. Crowdworkers supply questions and answers based on a set of over 10,000 news articles from CNN, with answers consisting of…

计算与语言 · 计算机科学 2017-02-08 Adam Trischler , Tong Wang , Xingdi Yuan , Justin Harris , Alessandro Sordoni , Philip Bachman , Kaheer Suleman

Novelpy (v1.2) is an open-source Python package designed to compute bibliometrics indicators. The package aims to provide a tool to the scientometrics community that centralizes different measures of novelty and disruptiveness, enables…

数字图书馆 · 计算机科学 2022-11-21 Pierre Pelletier , Kevin Wirtz

In order to simplify a sentence, human editors perform multiple rewriting transformations: they split it into several shorter sentences, paraphrase words (i.e. replacing complex words or phrases by simpler synonyms), reorder components,…

计算与语言 · 计算机科学 2020-05-04 Fernando Alva-Manchego , Louis Martin , Antoine Bordes , Carolina Scarton , Benoît Sagot , Lucia Specia

Existing paraphrase identification datasets lack sentence pairs that have high lexical overlap without being paraphrases. Models trained on such data fail to distinguish pairs like flights from New York to Florida and flights from Florida…

计算与语言 · 计算机科学 2019-04-03 Yuan Zhang , Jason Baldridge , Luheng He

Synthetic data sets are used across linguistic domains and NLP tasks, particularly in scenarios where authentic data is limited (or even non-existent). One such domain is that of clinical (healthcare) contexts, where there exist significant…

计算与语言 · 计算机科学 2026-03-17 Steven Bedrick , A. Seza Doğruöz , Sergiu Nisioi

Advancements in Large Language Models (LLMs) have significantly enhanced instruction-following capabilities. However, most Instruction Fine-Tuning (IFT) datasets are predominantly in English, limiting model performance in other languages.…

NLP models often degrade in performance when real world data distributions differ markedly from training data. However, existing dataset drift metrics in NLP have generally not considered specific dimensions of linguistic drift that affect…

计算与语言 · 计算机科学 2023-05-29 Tyler A. Chang , Kishaloy Halder , Neha Anna John , Yogarshi Vyas , Yassine Benajiba , Miguel Ballesteros , Dan Roth

A word embedding is a low-dimensional, dense and real- valued vector representation of a word. Word embeddings have been used in many NLP tasks. They are usually gener- ated from a large text corpus. The embedding of a word cap- tures both…

计算与语言 · 计算机科学 2017-08-15 Quanzhi Li , Sameena Shah , Xiaomo Liu , Armineh Nourbakhsh

Recent studies have evaluated the creativity/novelty of large language models (LLMs) primarily from a semantic perspective, using benchmarks from cognitive science. However, accessing the novelty in scholarly publications is a largely…

计算与语言 · 计算机科学 2024-09-26 Ethan Lin , Zhiyuan Peng , Yi Fang

Most research on emotion analysis from text focuses on the task of emotion classification or emotion intensity regression. Fewer works address emotions as a phenomenon to be tackled with structured learning, which can be explained by the…

计算与语言 · 计算机科学 2020-03-04 Laura Bostan , Evgeny Kim , Roman Klinger

This paper describes a test suite submission providing detailed statistics of linguistic performance for the state-of-the-art German-English systems of the Fifth Conference of Machine Translation (WMT20). The analysis covers 107 phenomena…

计算与语言 · 计算机科学 2020-10-16 Eleftherios Avramidis , Vivien Macketanz , Ursula Strohriegel , Aljoscha Burchardt , Sebastian Möller

Neologism-aware machine translation aims to translate source sentences containing neologisms into target languages. This field remains underexplored compared with general machine translation (MT). In this paper, we propose an agentic…

计算与语言 · 计算机科学 2026-05-26 Zhongtao Miao , Kaiyan Zhao , Masaaki Nagata , Yoshimasa Tsuruoka

Over the past years, deep learning methods allowed for new state-of-the-art results in ad-hoc information retrieval. However such methods usually require large amounts of annotated data to be effective. Since most standard ad-hoc…

信息检索 · 计算机科学 2020-03-18 Jibril Frej , Didier Schwab , Jean-Pierre Chevallet

We present a new machine learning and text information extraction approach to detection of cyber threat events in Twitter that are novel (previously non-extant) and developing (marked by significance with respect to similarity with a…

信息检索 · 计算机科学 2019-07-19 Avishek Bose , Vahid Behzadan , Carlos Aguirre , William H. Hsu

Novelty is a core requirement in academic publishing and a central focus of peer review, yet the growing volume of submissions has placed increasing pressure on human reviewers. While large language models (LLMs), including those fine-tuned…

计算与语言 · 计算机科学 2026-04-14 Wenqing Wu , Yi Zhao , Yuzhuo Wang , Siyou Li , Juexi Shao , Yunfei Long , Chengzhi Zhang

The prevalence of social media presents a growing opportunity to collect and analyse examples of English varieties. Whilst usage of these varieties was - and, in many cases, still is - used only in spoken contexts or hard-to-access private…

计算与语言 · 计算机科学 2024-01-23 Nhi Pham , Lachlan Pham , Adam L. Meyers

As events progress, news articles often update with new information: if we are not cautious, we risk propagating outdated facts. In this work, we hypothesize that linguistic features indicate factual fluidity, and that we can predict which…

计算与语言 · 计算机科学 2024-12-02 Alexander Spangher , Kung-Hsiang Huang , Hyundong Cho , Jonathan May

Accurate lexical entailment (LE) and natural language inference (NLI) often require large quantities of costly annotations. To alleviate the need for labeled data, we introduce WikiNLI: a resource for improving model performance on NLI and…

计算与语言 · 计算机科学 2020-10-06 Mingda Chen , Zewei Chu , Karl Stratos , Kevin Gimpel