中文
相关论文

相关论文: Finding Groups of Cross-Correlated Features in Bi-…

200 篇论文

We study the problem of linear feature selection when features are highly correlated. Such settings pose two fundamental challenges. First, how should model similarity be defined? Simply counting features in common can be misleading: two…

统计方法学 · 统计学 2026-03-24 Xiaozhu Zhang , Jacob Bien , Armeen Taeb

The notion of a bimodule herd is introduced and studied. A bimodule herd consists of a $B$-$A$ bimodule, its formal dual, called a pen, and a map, called a shepherd, which satisfies untiality and coassociativity conditions. It is shown that…

环与代数 · 数学 2008-06-10 Tomasz Brzezinski , Joost Vercruysse

Biclustering is a method for detecting homogeneous submatrices in a given observed matrix, and it is an effective tool for relational data analysis. Although there are many studies that estimate the underlying bicluster structure of a…

统计方法学 · 统计学 2021-07-16 Chihiro Watanabe , Taiji Suzuki

Genome Wide Association Studies (GWAS) and eQTL analyses have produced a large and growing number of genetic associations linked to a wide range of human phenotypes. As of 2013, there were more than 11,000 SNPs associated with a trait as…

基因组学 · 定量生物学 2016-09-28 John Platig , Peter Castaldi , Dawn DeMeo , John Quackenbush

The goal of coreset selection in supervised learning is to produce a weighted subset of data, so that training only on the subset achieves similar performance as training on the entire dataset. Existing methods achieved promising results in…

机器学习 · 计算机科学 2023-01-25 Xiao Zhou , Renjie Pi , Weizhong Zhang , Yong Lin , Tong Zhang

In genome wide association studies (GWAS), researchers are often dealing with non-normally distributed traits or a mixture of discrete-continuous traits. However, most of the current region-based methods rely on multivariate linear mixed…

统计方法学 · 统计学 2021-09-30 Julien St-Pierre , Karim Oualkacha

Few-shot learning for fine-grained image classification has gained recent attention in computer vision. Among the approaches for few-shot learning, due to the simplicity and effectiveness, metric-based methods are favorably state-of-the-art…

计算机视觉与模式识别 · 计算机科学 2021-02-03 Xiaoxu Li , Jijie Wu , Zhuo Sun , Zhanyu Ma , Jie Cao , Jing-Hao Xue

Similarity is a fundamental measure in network analyses and machine learning algorithms, with wide applications ranging from personalized recommendation to socio-economic dynamics. We argue that an effective similarity measurement should…

物理与社会 · 物理学 2015-12-07 Jian-Guo Liu , Lei Hou , Xue Pan , Qiang Guo , Tao Zhou

Community detection in complex networks is a topic of high interest in many fields. Bipartite networks are a special type of complex networks in which nodes are decomposed into two disjoint sets, and only nodes between the two sets can be…

社会与信息网络 · 计算机科学 2015-04-01 Zhenping Li , Rui-Sheng Wang , Shihua Zhang , Xiang-Sun Zhang

Graph clustering is a challenging pattern recognition problem whose goal is to identify vertex partitions with high intra-group connectivity. This paper investigates a bi-objective problem that maximizes the number of intra-cluster edges of…

社会与信息网络 · 计算机科学 2019-09-10 Camila P. S. Tautenhain , Mariá C. V. Nascimento

Genome-wide eQTL mapping explores the relationship between gene expression values and DNA variants to understand genetic causes of human disease. Due to the large number of genes and DNA variants that need to be assessed simultaneously,…

应用统计 · 统计学 2018-04-10 Jacob Rhyne , Jung-Ying Tzeng , Teng Zhang , X. Jessie Jeng

Feature selection techniques have been used as the workhorse in biomarker discovery applications for a long time. Surprisingly, the stability of feature selection with respect to sampling variations has long been under-considered. It is…

计算工程、金融与科学 · 计算机科学 2010-01-07 Zengyou He , Weichuan Yu

We consider popular matching problems in both bipartite and non-bipartite graphs with strict preference lists. It is known that every stable matching is a min-size popular matching. A subclass of max-size popular matchings called dominant…

离散数学 · 计算机科学 2018-06-13 Yuri Faenza , Telikepalli Kavitha , Vladlena Powers , Xingyu Zhang

In recent years, with the development of microarray technique, discovery of useful knowledge from microarray data has become very important. Biclustering is a very useful data mining technique for discovering genes which have similar…

计算工程、金融与科学 · 计算机科学 2009-09-09 Mohsen lashkargir , S. Amirhassan Monadjemi , Ahmad Baraani Dastjerdi

Interactions among multiple genes across the genome may contribute to the risks of many complex human diseases. Whole-genome single nucleotide polymorphisms (SNPs) data collected for many thousands of SNP markers from thousands of…

应用统计 · 统计学 2011-11-28 Yu Zhang , Jing Zhang , Jun S. Liu

Multimodal manifold modeling methods extend the spectral geometry-aware data analysis to learning from several related and complementary modalities. Most of these methods work based on two major assumptions: 1) there are the same number of…

机器学习 · 计算机科学 2021-05-13 Maysam Behmanesh , Peyman Adibi , Jocelyn Chanussot , Sayyed Mohammad Saeed Ehsani

Biclustering is a problem in machine learning and data mining that seeks to group together rows and columns of a dataset according to certain criteria. In this work, we highlight the natural relation that quantum computing models like boson…

量子物理 · 物理学 2024-05-30 Ajinkya Borle , Ameya Bhave

A wide range of data that appear in scientific experiments and simulations are multivariate or multifield in nature, consisting of multiple scalar fields. Topological feature search of such data aims to reveal important properties useful to…

计算几何 · 计算机科学 2024-06-06 Tripti Agarwal , Amit Chattopadhyay , Vijay Natarajan

Causal modeling has long been an attractive topic for many researchers and in recent decades there has seen a surge in theoretical development and discovery algorithms. Generally discovery algorithms can be divided into two approaches:…

机器学习 · 统计学 2017-02-06 Ridho Rahmadi , Perry Groot , Marianne Heins , Hans Knoop , Tom Heskes

Traditional clustering methods are limited when dealing with huge and heterogeneous groups of gene expression data, which motivates the development of bi-clustering methods. Bi-clustering methods are used to mine bi-clusters whose subsets…

计算机视觉与模式识别 · 计算机科学 2020-05-13 Kaijie Xu , Witold Pedrycz , Zhiwu Li , Yinghui Quan , Weike Nie