中文
相关论文

相关论文: LScDC-new large scientific dictionary

200 篇论文

Citation information in scholarly data is an important source of insight into the reception of publications and the scholarly discourse. Outcomes of citation analyses and the applicability of citation based machine learning approaches…

数字图书馆 · 计算机科学 2022-01-12 Tarek Saier , Michael Färber , Tornike Tsereteli

Word embeddings are an essential instrument in many NLP tasks. Most available resources are trained on general language from Web corpora or Wikipedia dumps. However, word embeddings for domain-specific language are rare, in particular for…

计算与语言 · 计算机科学 2023-02-14 Ricardo Schiffers , Dagmar Kern , Daniel Hienert

Latent Semantic Analysis (LSA) and Word2vec are some of the most widely used word embeddings. Despite the popularity of these techniques, the precise mechanisms by which they acquire new semantic relations between words remain unclear. In…

计算与语言 · 计算机科学 2019-01-23 Edgar Altszyler , Mariano Sigman , Diego Fernandez Slezak

Dense word embeddings, which encode semantic meanings of words to low dimensional vector spaces have become very popular in natural language processing (NLP) research due to their state-of-the-art performances in many NLP tasks. Word…

计算与语言 · 计算机科学 2018-07-20 Lutfi Kerem Senel , Ihsan Utlu , Veysel Yucesoy , Aykut Koc , Tolga Cukur

Scientific literature is typically dense, requiring significant background knowledge and deep comprehension for effective engagement. We introduce SciDQA, a new dataset for reading comprehension that challenges LLMs for a deep understanding…

计算与语言 · 计算机科学 2024-11-11 Shruti Singh , Nandan Sarkar , Arman Cohan

The scientific literature's exponential growth makes it increasingly challenging to navigate and synthesize knowledge across disciplines. Large language models (LLMs) are powerful tools for understanding scientific text, but they fail to…

计算与语言 · 计算机科学 2025-05-30 Abhipsha Das , Nicholas Lourie , Siavash Golkar , Mariel Pettee

The CL-SciSumm Shared Task is the first medium-scale shared task on scientific document summarization in the computational linguistics~(CL) domain. In 2019, it comprised three tasks: (1A) identifying relationships between citing documents…

计算与语言 · 计算机科学 2019-07-24 Muthu Kumar Chandrasekaran , Michihiro Yasunaga , Dragomir Radev , Dayne Freitag , Min-Yen Kan

The EcoLexicon English Corpus (EEC) is a 23.1-million-word corpus of contemporary environmental texts. It was compiled by the LexiCon research group for the development of EcoLexicon (Faber, Leon-Arauz & Reimerink 2016; San Martin et al.…

计算与语言 · 计算机科学 2018-07-17 Pilar Leon-Arauz , Antonio San Martin , Arianne Reimerink

This paper analysed author-assigned and title keywords into constituent words collected from 769 articles published in the journal Low Temperature Physics since the year 2006 to 2010. The total number of distinct keywords over the said time…

信息检索 · 计算机科学 2021-08-05 Bidyarthi Dutta

Scientific multi-label text classification suffers from extreme class imbalance, where specialized terminology exhibits severe power-law distributions that challenge standard classification approaches. Existing scientific corpora lack…

Large language models (LLMs) have grown in their usage to provide support for question answering across numerous disciplines. The models on their own have already shown promise for answering basic questions, however fail quickly where…

信息检索 · 计算机科学 2025-04-15 David Brett , Anniek Myatt

Scientific articles are long text documents organized into sections, each describing aspects of the research. Analyzing scientific production has become progressively challenging due to the increase in the number of available articles.…

计算与语言 · 计算机科学 2024-04-02 Gustavo Bartz Guedes , Ana Estela Antunes da Silva

Standing at the forefront of knowledge dissemination, digital libraries curate vast collections of scientific literature. However, these scholarly writings are often laden with jargon and tailored for domain experts rather than the general…

计算与语言 · 计算机科学 2024-08-08 Haining Wang , Jason Clark

Large Language Models (LLMs) show great promise as a powerful tool for scientific literature exploration. However, their effectiveness in providing scientifically accurate and comprehensive answers to complex questions within specialized…

Automatic abstractive text summarization is an important and challenging research topic of natural language processing. Among many widely used languages, the Chinese language has a special property that a Chinese character contains rich…

计算与语言 · 计算机科学 2018-09-11 Chieh-Teng Chang , Chi-Chia Huang , Chih-Yuan Yang , Jane Yung-Jen Hsu

The computational study of lexical semantic change (LSC) has taken off in the past few years and we are seeing increasing interest in the field, from both computational sciences and linguistics. Most of the research so far has focused on…

计算与语言 · 计算机科学 2021-01-20 Simon Hengchen , Nina Tahmasebi , Dominik Schlechtweg , Haim Dubossarsky

This paper describes a corpus of about 3000 English literary texts with about 250 million words extracted from the Gutenberg project that span a range of genres from both fiction and non-fiction written by more than 130 authors (e.g.,…

计算与语言 · 计算机科学 2018-01-09 Arthur M. Jacobs

Scholarly data are largely fragmented across siloed databases with divergent metadata and missing linkages among them. We present the Science Data Lake, a locally-deployable infrastructure built on DuckDB and simple Parquet files that…

数字图书馆 · 计算机科学 2026-03-04 Jonas Wilinski

In some areas of computing, natural language processing and information science, progress is made by sharing datasets and challenging the community to design the best algorithm for an associated task. This article introduces a shared…

数字图书馆 · 计算机科学 2026-01-27 Mike Thelwall

This paper presents an approach to classify documents in any language into an English topical label space, without any text categorization training data. The approach, Cross-Lingual Dataless Document Classification (CLDDC) relies on mapping…

计算与语言 · 计算机科学 2016-11-15 Yangqiu Song , Stephen Mayhew , Dan Roth