中文
相关论文

相关论文: Croissant Baker: Metadata Generation for Discovera…

200 篇论文

Data prefetching--loading data into the cache before it is requested--is essential for reducing I/O overhead and improving database performance. While traditional prefetchers focus on sequential patterns, recent learning-based approaches,…

数据库 · 计算机科学 2025-10-14 Farzaneh Zirak , Farhana Choudhury , Renata Borovica-Gajic

Machine Learning (ML) has been widely adopted in design exploration using high level synthesis (HLS) to give a better and faster performance, and resource and power estimation at very early stages for FPGA-based design. To perform…

硬件体系结构 · 计算机科学 2023-08-22 Zhigang Wei , Aman Arora , Ruihao Li , Lizy K. John

Generating datasets that "look like" given real ones is an interesting tasks for healthcare applications of ML and many other fields of science and engineering. In this paper we propose a new method of general application to binary datasets…

机器学习 · 统计学 2018-07-05 Laura Aviñó , Matteo Ruffini , Ricard Gavaldà

Commonsense generation aims to generate a realistic sentence describing a daily scene under the given concepts, which is very challenging, since it requires models to have relational reasoning and compositional generalization capabilities.…

计算与语言 · 计算机科学 2022-10-24 Xingwei He , Yeyun Gong , A-Long Jin , Weizhen Qi , Hang Zhang , Jian Jiao , Bartuer Zhou , Biao Cheng , SM Yiu , Nan Duan

Synthetic data has gained significant momentum thanks to sophisticated machine learning tools that enable the synthesis of high-dimensional datasets. However, many generation techniques do not give the data controller control over what…

Literature analysis facilitates researchers to acquire a good understanding of the development of science and technology. The traditional literature analysis focuses largely on the literature metadata such as topics, authors, abstracts,…

人工智能 · 计算机科学 2021-01-29 Linlin Hou , Ji Zhang , Ou Wu , Ting Yu , Zhen Wang , Zhao Li , Jianliang Gao , Yingchun Ye , Rujing Yao

Misinformation is becoming increasingly prevalent on social media and in news articles. It has become so widespread that we require algorithmic assistance utilising machine learning to detect such content. Training these machine learning…

机器学习 · 计算机科学 2022-03-09 Dan Saattrup Nielsen , Ryan McConville

The selection, development, or comparison of machine learning methods in data mining can be a difficult task based on the target problem and goals of a particular study. Numerous publicly available real-world and simulated benchmark…

机器学习 · 计算机科学 2017-03-03 Randal S. Olson , William La Cava , Patryk Orzechowski , Ryan J. Urbanowicz , Jason H. Moore

Search engines these days can serve datasets as search results. Datasets get picked up by search technologies based on structured descriptions on their official web pages, informed by metadata ontologies such as the Dataset content type of…

The Model Context Protocol (MCP) has recently emerged as a standardized interface for connecting language models with external tools and data. As the ecosystem rapidly expands, the lack of a structured, comprehensive view of existing MCP…

密码学与安全 · 计算机科学 2025-07-01 Zhiwei Lin , Bonan Ruan , Jiahao Liu , Weibo Zhao

The first and maybe the most important step in designing a model-based predictive controller is to develop a model that is as accurate as possible and that is valid under a wide range of operating conditions. The sugar boiling process is a…

化学物理 · 物理学 2012-12-24 Alfred Jean Philippe Lauret , Harry Boyer , Jean Claude Gatina

Traditional keyphrase prediction methods predict a single set of keyphrases per document, failing to cater to the diverse needs of users and downstream applications. To bridge the gap, we introduce on-demand keyphrase generation, a novel…

计算与语言 · 计算机科学 2024-10-07 Di Wu , Xiaoxian Shen , Kai-Wei Chang

Causal discovery is fundamental to scientific understanding and reliable decision-making. Existing approaches face critical limitations: purely data-driven methods suffer from statistical indistinguishability and modeling assumptions, while…

计算与语言 · 计算机科学 2026-01-21 Bo Peng , Sirui Chen , Lei Xu , Chaochao Lu

Synthetic datasets are important for evaluating and testing machine learning models. When evaluating real-life recommender systems, high-dimensional categorical (and sparse) datasets are often considered. Unfortunately, there are not many…

信息检索 · 计算机科学 2024-12-11 Miha Malenšek , Blaž Škrlj , Blaž Mramor , Jure Demšar

A lack of accessible data has historically restricted malware analysis research, and practitioners have relied heavily on datasets provided by industry sources to advance. Existing public datasets are limited by narrow scope - most include…

密码学与安全 · 计算机科学 2025-06-06 Robert J. Joyce , Gideon Miller , Phil Roth , Richard Zak , Elliott Zaresky-Williams , Hyrum Anderson , Edward Raff , James Holt

Programming language-design and run-time-implementation require detailed knowledge about the programs that users want to implement. Acquiring this knowledge is hard, and there is little tool support to effectively estimate whether a…

编程语言 · 计算机科学 2017-04-03 Stephan Brandauer , Tobias Wrigstad

We focus on Multimodal Machine Reading Comprehension (M3C) where a model is expected to answer questions based on given passage (or context), and the context and the questions can be in different modalities. Previous works such as RecipeQA…

计算机视觉与模式识别 · 计算机科学 2021-10-25 Pritish Sahu , Karan Sikka , Ajay Divakaran

This paper describes EMBER: a labeled benchmark dataset for training machine learning models to statically detect malicious Windows portable executable files. The dataset includes features extracted from 1.1M binary files: 900K training…

密码学与安全 · 计算机科学 2018-04-18 Hyrum S. Anderson , Phil Roth

We introduce a new dataset, MELINDA, for Multimodal biomEdicaL experImeNt methoD clAssification. The dataset is collected in a fully automated distant supervision manner, where the labels are obtained from an existing curated database, and…

计算与语言 · 计算机科学 2020-12-18 Te-Lin Wu , Shikhar Singh , Sayan Paul , Gully Burns , Nanyun Peng

In the current environment of data generation and publication, there is an ever-growing number of datasets available for download. This growth precipitates an existing challenge: sourcing and integrating relevant datasets for analysis is…

数字图书馆 · 计算机科学 2024-10-15 Mark S. Fox , Bart Gajderowicz , Dishu Lyu