中文
相关论文

相关论文: Introducing Data Primitives: Data Formats for the …

200 篇论文

Big data programming frameworks have become increasingly important for the development of applications for which performance and scalability are critical. In those complex frameworks, optimizing code by hand is hard and time-consuming,…

计算机科学中的逻辑 · 计算机科学 2023-06-14 Sarah Chlyah , Nils Gesbert , Pierre Geneves , Nabil Layaida

Data analysis based on information from several sources is common in economic and biomedical studies. This setting is often referred to as the data fusion problem, which differs from traditional missing data problems since no complete data…

统计方法学 · 统计学 2022-04-07 Wei Li , Shanshan Luo , Wangli Xu

We propose new ensemble models for multivariate functional data classification as combinations of semi-metric-based weak learners. Our models extend current semi-metric-type methods from the univariate to the multivariate case, propose new…

In this manuscript a unified framework for conducting inference on complex aggregated data in high dimensional settings is proposed. The data are assumed to be a collection of multiple non-Gaussian realizations with underlying undirected…

应用统计 · 统计学 2013-10-14 Fang Han , Han Liu , Brian Caffo

How do latent and inference time computations enable large language models (LLMs) to solve multi-step reasoning? We introduce a framework for tracing and steering algorithmic primitives that underlie model reasoning. Our approach links…

机器学习 · 计算机科学 2026-02-17 Samuel Lippl , Thomas McGee , Kimberly Lopez , Ziwen Pan , Pierce Zhang , Salma Ziadi , Oliver Eberle , Ida Momennejad

Empirical claims often rely on one population, design, and analysis. Many-analysts, multiverse, and robustness studies expose how results can vary across plausible analytic choices. Synthesizing these results, however, is nontrivial as all…

统计方法学 · 统计学 2025-11-24 František Bartoš , Suzanne Hoogeveen , Alexandra Sarafoglou , Samuel Pawel

Partial Least Squares (PLS) methods have been heavily exploited to analyse the association between two blocs of data. These powerful approaches can be applied to data sets where the number of variables is greater than the number of…

机器学习 · 统计学 2017-02-24 Pierre Lafaye de Micheaux , Benoit Liquet , Matthew Sutton

Models of the spatial distribution of animals provide useful tools to help ecologists quantify species-environment relationships, and they are increasingly being used to help determine the impacts of climate and habitat changes on species.…

统计方法学 · 统计学 2020-09-04 Joe Watson , Ruth Joy , Dominic Tollit , Sheila J Thornton , Marie Auger-Méthé

Multivariate spatially-oriented data sets are prevalent in the environmental and physical sciences. Scientists seek to jointly model multiple variables, each indexed by a spatial location, to capture any underlying spatial association for…

统计方法学 · 统计学 2021-08-19 Lu Zhang , Sudipto Banerjee

Multiple sets of measurements on the same objects obtained from different platforms may reflect partially complementary information of the studied system. The integrative analysis of such data sets not only provides us with the opportunity…

统计方法学 · 统计学 2020-10-15 Yipeng Song , Johan A. Westerhuis , Age K. Smilde

Individuals and organizations cope with an always-growing amount of data, which is heterogeneous in its contents and formats. An adequate data management process yielding data quality and control over its lifecycle is a prerequisite to…

The organization and mining of malaria genomic and post-genomic data is highly motivated by the necessity to predict and characterize new biological targets and new drugs. Biological targets are sought in a biological space designed from…

Compositional data, also referred to as simplicial data, naturally arise in many scientific domains such as geochemistry, microbiology, and economics. In such domains, obtaining sensible lower-dimensional representations and modes of…

统计方法学 · 统计学 2025-04-15 Hyeon Lee , Kassel Liam Hingee , Janice L. Scealy , Andrew T. A. Wood , Eric Grunsky , J. S. Marron

Recent technological advancements have led to the rapid generation of high-throughput biological data, which can be used to address novel scientific questions in broad areas of research. These data can be thought of as a large matrix with…

统计计算 · 统计学 2021-03-01 Jane W. Liang , Saunak Sen

Many research questions can be answered quickly and efficiently using data already collected for previous research. This practice is called secondary data analysis (SDA), and has gained popularity due to lower costs and improved research…

数字图书馆 · 计算机科学 2020-04-07 Yasith Jayawardana , Sampath Jayarathna

Today's data-heavy research environment requires the integration of different sources of information into structured data sets that can not be analyzed as simple matrices. We introduce an old technique, known in the European data analyses…

应用统计 · 统计学 2012-02-27 Omar De la Cruz , Susan Holmes

A large class of traditional graph and data mining algorithms can be concisely expressed in Datalog, and other Logic-based languages, once aggregates are allowed in recursion. In fact, for most BigData algorithms, the difficult semantic…

编程语言 · 计算机科学 2019-07-25 Ariyam Das , Carlo Zaniolo

Data segmentation a.k.a. multiple change point analysis has received considerable attention due to its importance in time series analysis and signal processing, with applications in a variety of fields including natural and social sciences,…

统计方法学 · 统计学 2021-07-09 Haeran Cho , Claudia Kirch

Understanding geometric properties of natural language processing models' latent spaces allows the manipulation of these properties for improved performance on downstream tasks. One such property is the amount of data spread in a model's…

机器学习 · 计算机科学 2023-08-02 Anna C. Marbut , Katy McKinney-Bock , Travis J. Wheeler

Curating, processing, and combining large-scale medical imaging datasets from national studies is a non-trivial task due to the intense computation and data throughput required, variability of acquired data, and associated financial…