中文
相关论文

相关论文: A framework for constructing a huge name disambigu…

200 篇论文

Human annotated data plays a crucial role in machine learning (ML) research and development. However, the ethical considerations around the processes and decisions that go into dataset annotation have not received nearly enough attention.…

Scholarly data is growing continuously containing information about the articles from a plethora of venues including conferences, journals, etc. Many initiatives have been taken to make scholarly data available as Knowledge Graphs (KGs).…

人工智能 · 计算机科学 2022-06-02 Cristian Santini , Genet Asefa Gesese , Silvio Peroni , Aldo Gangemi , Harald Sack , Mehwish Alam

Name disambiguation aims to identify unique authors with the same name. Existing name disambiguation methods always exploit author attributes to enhance disambiguation results. However, some discriminative author attributes (e.g., email and…

数字图书馆 · 计算机科学 2021-01-21 Qingyun Sun , Hao Peng , Jianxin Li , Senzhang Wang , Xiangyu Dong , Liangxuan Zhao , Philip S. Yu , Lifang He

As digital collections of scientific literature are widespread and used frequently in knowledge-intense working environments, it has become a challenge to identify author names correctly. The treatment of homonyms is crucial for the…

数字图书馆 · 计算机科学 2017-03-07 Thomas Krämer , Fakhri Momeni , Philipp Mayr

Massive-scale historical document collections are crucial for social science research. Despite increasing digitization, these documents typically lack unique cross-document identifiers for individuals mentioned within the texts, as well as…

计算与语言 · 计算机科学 2024-06-25 Abhishek Arora , Emily Silcock , Leander Heldring , Melissa Dell

Acronyms are the short forms of phrases that facilitate conveying lengthy sentences in documents and serve as one of the mainstays of writing. Due to their importance, identifying acronyms and corresponding phrases (i.e., acronym…

计算与语言 · 计算机科学 2020-10-29 Amir Pouran Ben Veyseh , Franck Dernoncourt , Quan Hung Tran , Thien Huu Nguyen

Data is the engine of modern computer vision, which necessitates collecting large-scale datasets. This is expensive, and guaranteeing the quality of the labels is a major challenge. In this paper, we investigate efficient annotation…

计算机视觉与模式识别 · 计算机科学 2021-04-27 Yuan-Hong Liao , Amlan Kar , Sanja Fidler

The HuggingFace Datasets Hub hosts thousands of datasets, offering exciting opportunities for language model training and evaluation. However, datasets for a specific task type often have different schemas, making harmonization challenging.…

计算与语言 · 计算机科学 2023-05-17 Damien Sileo

Patents and scientific papers provide an essential source for measuring science and technology output, to be used as a basis for the most varied scientometric analyzes. Authors' and inventors' names are the key identifiers to carry out…

信息检索 · 计算机科学 2024-02-27 David Reymond , Heman Khouilla , Sandrine Wolff , Manuel Durand-Barthez

Human annotators typically provide annotated data for training machine learning models, such as neural networks. Yet, human annotations are subject to noise, impairing generalization performances. Methodological research on approaches…

计算机视觉与模式识别 · 计算机科学 2024-07-31 Marek Herde , Denis Huseljic , Lukas Rauch , Bernhard Sick

Patent data represent a significant source of information on innovation and the evolution of technology through networks of citations, co-invention and co-assignment of new patents. A major obstacle to extracting useful information from…

数字图书馆 · 计算机科学 2016-01-11 Greg Morrison , Massimo Riccaboni , Fabio Pammolli

Semantic annotation of long texts, such as novels, remains an open challenge in Natural Language Processing (NLP). This research investigates the problem of detecting person entities and assigning them unique identities, i.e., recognizing…

计算与语言 · 计算机科学 2021-10-05 Weronika Łajewska , Anna Wróblewska

With the growing prevalence of large language models, it is increasingly common to annotate datasets for machine learning using pools of crowd raters. However, these raters often work in isolation as individual crowdworkers. In this work,…

计算机与社会 · 计算机科学 2024-08-05 Sonja Schmer-Galunder , Ruta Wheelock , Scott Friedman , Alyssa Chvasta , Zaria Jalan , Emily Saltz

As large language models (LLMs) rapidly advance and integrate into daily life, the privacy risks they pose are attracting increasing attention. We focus on a specific privacy risk where LLMs may help identify the authorship of anonymous…

计算与语言 · 计算机科学 2024-11-21 Zichen Wen , Dadi Guo , Huishuai Zhang

Reference texts such as encyclopedias and news articles can manifest biased language when objective reporting is substituted by subjective writing. Existing methods to detect bias mostly rely on annotated data to train machine learning…

计算与语言 · 计算机科学 2021-12-20 Timo Spinde , David Krieger , Manuel Plank , Bela Gipp

The increasing diversity of languages used on the web introduces a new level of complexity to Information Retrieval (IR) systems. We can no longer assume that textual content is written in one language or even the same language family. In…

计算与语言 · 计算机科学 2014-10-15 Rami Al-Rfou , Vivek Kulkarni , Bryan Perozzi , Steven Skiena

Large-scale datasets are essential to modern day deep learning. Advocates argue that understanding these methods requires dataset transparency (e.g. "dataset curation, motivation, composition, collection process, etc..."). However, almost…

计算机视觉与模式识别 · 计算机科学 2023-06-27 Nadine Chang , Francesco Ferroni , Michael J. Tarr , Martial Hebert , Deva Ramanan

We present a resource for the task of FrameNet semantic frame disambiguation of over 5,000 word-sentence pairs from the Wikipedia corpus. The annotations were collected using a novel crowdsourcing approach with multiple workers per sentence…

计算与语言 · 计算机科学 2020-06-15 Anca Dumitrache , Lora Aroyo , Chris Welty

In multi-label classification, each example in a dataset may be annotated as belonging to one or more classes (or none of the classes). Example applications include image (or document) tagging where each possible tag either applies to a…

机器学习 · 计算机科学 2022-11-28 Aditya Thyagarajan , Elías Snorrason , Curtis Northcutt , Jonas Mueller

Sequence-to-sequence models have recently gained the state of the art performance in summarization. However, not too many large-scale high-quality datasets are available and almost all the available ones are mainly news articles with…

计算与语言 · 计算机科学 2018-10-23 Mahnaz Koupaee , William Yang Wang