中文
相关论文

相关论文: Can inferred provenance and its visualisation be u…

200 篇论文

Motivation: Annotations are a key feature of many biological databases, used to convey our knowledge of a sequence to the reader. Ideally, annotations are curated manually, however manual curation is costly, time consuming and requires…

计算工程、金融与科学 · 计算机科学 2013-08-22 Michael J. Bell , Colin S. Gillespie , Daniel Swan , Phillip Lord

The primary mission of UniProt is to support biological research by maintaining a stable, comprehensive, fully classified, richly and accurately annotated protein sequence knowledgebase, with extensive cross-references to external…

数据库 · 计算机科学 2010-12-09 Jerven Bolleman , Alain Gateau , Sebastien Gehant , Nicole Redaschi

As the quantity of data being depositing into biological databases continues to increase, it becomes ever more vital to develop methods that enable us to understand this data and ensure that the knowledge is correct. It is widely-held that…

数字图书馆 · 计算机科学 2017-05-25 Michael J Bell , Phillip Lord

Research in Computational Linguistics is dependent on text corpora for training and testing new tools and methodologies. While there exists a plethora of annotated linguistic information, these corpora are often not interoperable without…

计算与语言 · 计算机科学 2020-11-03 Timo Lek , Anna de Groot , Tobias Kuhn , Roser Morante

Understanding protein function is one of the keys to understanding life at the molecular level. It is also important in several scenarios including human disease and drug discovery. In this age of rapid and affordable biological sequencing,…

分布式、并行与集群计算 · 计算机科学 2017-08-24 Sabeur Aridhi , Seyed Ziaeddin Alborzi , Malika Smaïl-Tabbone , Marie-Dominique Devignes , David Ritchie

Provenance is information about the origin, derivation, ownership, or history of an object. It has recently been studied extensively in scientific databases and other settings due to its importance in helping scientists judge data validity,…

编程语言 · 计算机科学 2008-12-03 James Cheney , Umut Acar , Amal Ahmed

Many real-world applications require automated data annotation, such as identifying tissue origins based on gene expressions and classifying images into semantic categories. Annotation classes are often numerous and subject to changes over…

计算与语言 · 计算机科学 2018-07-03 Maxim Grechkin , Hoifung Poon , Bill Howe

Long-term data-driven studies have become indispensable in many areas of science. Often, the data formats, structures and semantics of data change over time, the data sets evolve. Therefore, studies over several decades in particular have…

数据库 · 计算机科学 2022-11-28 Tanja Auge , Andreas Heuer

Knowledge Graphs are repositories of information that gather data from a multitude of domains and sources in the form of semantic triples, serving as a source of structured data for various crucial applications in the modern web landscape,…

计算与语言 · 计算机科学 2022-10-27 Gabriel Amaral , Odinaldo Rodrigues , Elena Simperl

We study in this paper provenance information for queries with aggregation. Provenance information was studied in the context of various query languages that do not allow for aggregation, and recent work has suggested to capture provenance…

数据库 · 计算机科学 2015-03-17 Yael Amsterdamer , Daniel Deutch , Val Tannen

Science is conducted collaboratively, often requiring the sharing of knowledge about computational experiments. When experiments include only datasets, they can be shared using Uniform Resource Identifiers (URIs) or Digital Object…

数据库 · 计算机科学 2018-06-19 Zhihao Yuan , Dai Hai Ton That , Siddhant Kothari , Gabriel Fils , Tanu Malik

Even though computational reproducibility is widely accepted as necessary for research validation and reuse, it is often not considered during the research process. This is because reproducibility tools are typically stand-alone and require…

分布式、并行与集群计算 · 计算机科学 2020-03-04 Ana Trisovic , Chris R. Jones , Ben Couturier , Marco Clemencic

Provenance is an increasing concern due to the ongoing revolution in sharing and processing scientific data on the Web and in other computer systems. It is proposed that many computer systems will need to become provenance-aware in order to…

编程语言 · 计算机科学 2014-01-06 Umut A. Acar , Amal Ahmed , James Cheney , Roly Perera

Factual correctness is often the limiting factor in practical applications of natural language generation in high-stakes domains such as healthcare. An essential requirement for maintaining factuality is the ability to deal with rare…

计算与语言 · 计算机科学 2023-06-07 Maksim Eremeev , Ilya Valmianski , Xavier Amatriain , Anitha Kannan

Provenance, or information about the origin or derivation of data, is important for assessing the trustworthiness of data and identifying and correcting mistakes. Most prior implementations of data provenance have involved heavyweight…

编程语言 · 计算机科学 2017-08-23 Stefan Fehrenbach , James Cheney

Scientific progress increasingly depends on data management, particularly to clean and curate data so that it can be systematically analyzed and reused. A wealth of techniques for managing and curating data (and its provenance) have been…

编程语言 · 计算机科学 2018-03-28 Jan Stolarek , James Cheney

LLMs have achieved remarkable fluency and coherence in text generation, yet their widespread adoption has raised concerns about content reliability and accountability. In high-stakes domains, it is crucial to understand where and how the…

计算与语言 · 计算机科学 2025-06-02 Junnan Zhu , Min Xiao , Yining Wang , Feifei Zhai , Yu Zhou , Chengqing Zong

The power of natural language generation models has provoked a flurry of interest in automatic methods to detect if a piece of text is human or machine-authored. The problem so far has been framed in a standard supervised way and consists…

计算与语言 · 计算机科学 2021-11-05 Matthias Gallé , Jos Rozen , Germán Kruszewski , Hady Elsahar

We present a dataset that contains every instance of all tokens (~ words) ever written in undeleted, non-redirect English Wikipedia articles until October 2016, in total 13,545,349,787 instances. Each token is annotated with (i) the article…

计算与语言 · 计算机科学 2017-03-27 Fabian Flöck , Kenan Erdogan , Maribel Acosta

Nowadays, with the booming development of the Internet, people benefit from its convenience due to its open and sharing nature. A large volume of natural language texts is being generated by users in various forms, such as search queries,…

计算与语言 · 计算机科学 2019-08-07 Chenwei Zhang
‹ 上一页 1 2 3 10 下一页 ›