中文
相关论文

相关论文: Data Motif-based Proxy Benchmarks for Big Data and…

200 篇论文

ProMoAI is a novel tool that leverages Large Language Models (LLMs) to automatically generate process models from textual descriptions, incorporating advanced prompt engineering, error handling, and code generation techniques. Beyond…

数据库 · 计算机科学 2024-08-09 Humam Kourani , Alessandro Berti , Daniel Schuster , Wil M. P. van der Aalst

This work proposes a time series prediction method based on the kernel view of linear reservoirs. In particular, the time series motifs of the reservoir kernel are used as representational basis on which general readouts are constructed. We…

机器学习 · 计算机科学 2024-12-05 Peter Tino , Robert Simon Fong , Roberto Fabio Leonarduzzi

Robust causal discovery in time series datasets depends on reliable benchmark datasets with known ground-truth causal relationships. However, such datasets remain scarce, and existing synthetic alternatives often overlook critical temporal…

机器学习 · 计算机科学 2025-06-03 Muhammad Hasan Ferdous , Emam Hossain , Md Osman Gani

Large language models are increasingly used as proxies for human subjects in social science research, yet external validity requires that synthetic agents faithfully reflect the preferences of target human populations. We introduce…

人工智能 · 计算机科学 2026-01-30 Bingchen Wang , Zi-Yu Khoo , Jingtan Wang

Predictive modelling is vital to guide preventive efforts. Whilst large-scale prospective cohort studies and a diverse toolkit of available machine learning (ML) algorithms have facilitated such survival task efforts, choosing the…

Power is the primary design objective of large-scale integrated circuits (ICs), especially for complex modern processors (i.e., CPUs). Accurate CPU power evaluation requires designers to go through the whole time-consuming IC implementation…

硬件体系结构 · 计算机科学 2025-12-09 Qijun Zhang , Yao Lu , Mengming Li , Shang Liu , Zhiyao Xie

Existing high-dimensional statistical methods are largely established for analyzing individual-level data. In this work, we study estimation and inference for high-dimensional linear models where we only observe "proxy data", which include…

统计方法学 · 统计学 2022-01-12 Sai Li , T. Tony Cai , Hongzhe Li

We introduce a fairness-aware dataset for job recommendations in advertising, designed to foster research in algorithmic fairness within real-world scenarios. It was collected and prepared to comply with privacy standards and business…

机器学习 · 计算机科学 2024-11-05 Mariia Vladimirova , Federico Pavone , Eustache Diemert

Data series motif discovery represents one of the most useful primitives for data series mining, with applications to many domains, such as robotics, entomology, seismology, medicine, and climatology, and others. The state-of-the-art motif…

数据库 · 计算机科学 2020-09-01 Michele Linardi , Yan Zhu , Themis Palpanas , Eamonn Keogh

Many practical applications, ranging from paper-reviewer assignment in peer review to job-applicant matching for hiring, require human decision makers to identify relevant matches by combining their expertise with predictions from machine…

机器学习 · 计算机科学 2023-02-17 Joon Sik Kim , Valerie Chen , Danish Pruthi , Nihar B. Shah , Ameet Talwalkar

Big data analytics applications play a significant role in data centers, and hence it has become increasingly important to understand their behaviors in order to further improve the performance of data center computer systems, in which…

分布式、并行与集群计算 · 计算机科学 2015-04-21 Zhen Jia , Lei Wang , Jianfeng Zhan , Lixin Zhang , Chunjie Luo , Ninghui Sun

Learning causal relationships from time series data is an important but challenging problem. Existing synthetic datasets often contain hidden artifacts that can be exploited by causal discovery methods, reducing their usefulness for…

机器学习 · 计算机科学 2026-03-23 Xiaoyu He , Petr Ryšavý , Jakub Mareček

Motivated by applications in social network community analysis, we introduce a new clustering paradigm termed motif clustering. Unlike classical clustering, motif clustering aims to minimize the number of clustering errors associated with…

社会与信息网络 · 计算机科学 2017-01-31 Pan Li , Hoang Dau , Gregory Puleo , Olgica Milenkovic

Compound AI applications, composed from interactions between Large Language Models (LLMs), Machine Learning (ML) models, external tools and data sources are quickly becoming an integral workload in datacenters. Their diverse sub-components…

分布式、并行与集群计算 · 计算机科学 2026-04-14 Paramuth Samuthrsindh , Angel Cervantes , Varun Gohil , Gohar Irfan Chaudhry , Christina Delimitrou , Adam Belay

Recent progress in deep learning has been driven by increasingly larger models. However, their computational and energy demands have grown proportionally, creating significant barriers to their deployment and to a wider adoption of deep…

机器学习 · 计算机科学 2025-09-16 Pedro Savarese

Time Series Motif Discovery (TSMD), which aims at finding recurring patterns in time series, is an important task in numerous application domains, and many methods for this task exist. These methods are usually evaluated qualitatively. A…

机器学习 · 计算机科学 2024-12-13 Daan Van Wesenbeeck , Aras Yurtman , Wannes Meert , Hendrik Blockeel

Fraud detection and anti-money-laundering (AML) compliance are high-value domains for large language models (LLMs), but their serving requirements differ sharply from generic chat workloads. Compliance prompts are often prefix-heavy,…

人工智能 · 计算机科学 2026-05-13 Prathamesh Vasudeo Naik , Naresh Dintakurthi , Yue Wang

AI-generated content has evolved from monolithic models to modular workflows, particularly on platforms like ComfyUI, enabling customization in creative pipelines. However, crafting effective workflows requires great expertise to…

计算与语言 · 计算机科学 2025-06-12 Zhenran Xu , Yiyu Wang , Xue Yang , Longyue Wang , Weihua Luo , Kaifu Zhang , Baotian Hu , Min Zhang

The rapidly growing demand for high-quality data in Large Language Models (LLMs) has intensified the need for scalable, reliable, and semantically rich data preparation pipelines. However, current practices remain dominated by ad-hoc…

Estimating a causal query from observational data is an essential task in the analysis of biomolecular networks. Estimation takes as input a network topology, a query estimation method, and observational measurements on the network…