English
Related papers

Related papers: Data reduction strategy in the PandaX-4T experimen…

200 papers

Dataset pruning has been widely studied for 2D images to remove redundancy and accelerate training, while particular pruning methods for 3D data remain largely unexplored. In this work, we study dataset pruning for 3D data, where its…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Xiaohan Zhao , Xinyi Shang , Jiacheng Liu , Zhiqiang Shen

Recently, tremendous interest has been devoted to develop data fusion strategies for energy efficiency in buildings, where various kinds of information can be processed. However, applying the appropriate data fusion strategy to design an…

Computers and Society · Computer Science 2020-09-15 Yassine Himeur , Abdullah Alsalemi , Ayman Al-Kababji , Faycal Bensaali , Abbes Amira

Dataset distillation (DD) aims to compress large-scale datasets into compact synthetic counterparts for efficient model training. However, existing DD methods exhibit substantial performance degradation on long-tailed datasets. We identify…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Ruixi Wu , Shaobo Wang , Jiahuan Chen , Zhiyuan Liu , Yicun Yang , Zhaorun Chen , Zekai Li , Kaixin Li , Xinming Wang , Hongzhu Yi , Kai Wang , Linfeng Zhang

Positive-Unlabeled (PU) learning aims to learn a model with rare positive samples and abundant unlabeled samples. Compared with classical binary classification, the task of PU learning is much more challenging due to the existence of many…

Computer Vision and Pattern Recognition · Computer Science 2022-12-01 Chengming Xu , Chen Liu , Siqian Yang , Yabiao Wang , Shijie Zhang , Lijie Jia , Yanwei Fu

Motivation: Although principal component analysis is frequently applied to reduce the dimensionality of matrix data, the method is sensitive to noise and bias and has difficulty with comparability and interpretation. These issues are…

Methodology · Statistics 2012-12-27 Tomokazu Konishi

The high-dimensional data setting, in which p >> n, is a challenging statistical paradigm that appears in many real-world problems. In this setting, learning a compact, low-dimensional representation of the data can substantially help…

Machine Learning · Computer Science 2018-08-07 Micol Marchetti-Bowick , Benjamin J. Lengerich , Ankur P. Parikh , Eric P. Xing

Dynamic data selection aims to accelerate training with lossless performance. However, reducing training data inherently limits data diversity, potentially hindering generalization. While data augmentation is widely used to enhance…

Machine Learning · Computer Science 2025-05-13 Suorong Yang , Peng Ye , Furao Shen , Dongzhan Zhou

The randomized subspace Newton convex methods for the sensor selection problem are proposed. The randomized subspace Newton algorithm is straightforwardly applied to the convex formulation, and the customized method in which the part of the…

Systems and Control · Electrical Eng. & Systems 2021-05-03 Taku Nonomura , Shunsuke Ono , Kumi Nakai , Yuji Saito

We report on a blinded analysis of low-energy electronic-recoil data from the first science run of the XENONnT dark matter experiment. Novel subsystems and the increased 5.9 tonne liquid xenon target reduced the background in the (1, 30)…

High Energy Physics - Experiment · Physics 2022-11-16 E. Aprile , K. Abe , F. Agostini , S. Ahmed Maouloud , L. Althueser , B. Andrieu , E. Angelino , J. R. Angevaare , V. C. Antochi , D. Antón Martin , F. Arneodo , L. Baudis , A. L. Baxter , L. Bellagamba , R. Biondi , A. Bismark , A. Brown , S. Bruenner , G. Bruno , R. Budnik , T. K. Bui , C. Cai , C. Capelli , J. M. R. Cardoso , D. Cichon , M. Clark , A. P. Colijn , J. Conrad , J. J. Cuenca-García , J. P. Cussonneau , V. D'Andrea , M. P. Decowski , P. Di Gangi , S. Di Pede , A. Di Giovanni , R. Di Stefano , S. Diglio , K. Eitel , A. Elykov , S. Farrell , A. D. Ferella , C. Ferrari , H. Fischer , W. Fulgione , P. Gaemers , R. Gaior , A. Gallo Rosso , M. Galloway , F. Gao , R. Gardner , R. Glade-Beucke , L. Grandi , J. Grigat , M. Guida , R. Hammann , A. Higuera , C. Hils , L. Hoetzsch , J. Howlett , M. Iacovacci , Y. Itow , J. Jakob , F. Joerg , A. Joy , N. Kato , M. Kara , P. Kavrigin , S. Kazama , M. Kobayashi , G. Koltman , A. Kopec , F. Kuger , H. Landsman , R. F. Lang , L. Levinson , I. Li , S. Li , S. Liang , S. Lindemann , M. Lindner , K. Liu , J. Loizeau , F. Lombardi , J. Long , J. A. M. Lopes , Y. Ma , C. Macolino , J. Mahlstedt , A. Mancuso , L. Manenti , F. Marignetti , T. Marrodán Undagoitia , K. Martens , J. Masbou , D. Masson , E. Masson , S. Mastroianni , M. Messina , K. Miuchi , K. Mizukoshi , A. Molinario , S. Moriyama , K. Morå , Y. Mosbacher , M. Murra , J. Müller , K. Ni , U. Oberlack , B. Paetsch , J. Palacio , P. Paschos , R. Peres , C. Peters , J. Pienaar , M. Pierre , V. Pizzella , G. Plante , J. Qi , J. Qin , D. Ramírez García , S. Reichard , A. Rocchetti , N. Rupp , L. Sanchez , J. M. F. dos Santos , I. Sarnoff , G. Sartorelli , J. Schreiner , D. Schulte , P. Schulte , H. Schulze Eißing , M. Schumann , L. Scotto Lavina , M. Selvi , F. Semeria , P. Shagin , S. Shi , E. Shockley , M. Silva , H. Simgen , J. Stephen , A. Takeda , P. -L. Tan , A. Terliuk , D. Thers , F. Toschi , G. Trinchero , C. Tunnell , F. Tönnies , K. Valerius , G. Volta , Y. Wei , C. Weinheimer , M. Weiss , D. Wenz , C. Wittweg , T. Wolf , D. Xu , Z. Xu , M. Yamashita , L. Yang , J. Ye , L. Yuan , G. Zavattini , M. Zhong , T. Zhu

Dataset distillation creates a small distilled set that enables efficient training by capturing key information from the full dataset. While existing dataset distillation methods perform well on balanced datasets, they struggle under…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Xiao Cui , Yulei Qin , Xinyue Li , Wengang Zhou , Hongsheng Li , Houqiang Li

To solve key biomedical problems, experimentalists now routinely measure millions or billions of features (dimensions) per sample, with the hope that data science techniques will be able to build accurate data-driven inferences. Because…

The great success of deep learning heavily relies on increasingly larger training data, which comes at a price of huge computational and infrastructural costs. This poses crucial questions that, do all training data contribute to model's…

Machine Learning · Computer Science 2023-02-28 Shuo Yang , Zeke Xie , Hanyu Peng , Min Xu , Mingming Sun , Ping Li

In semantic segmentation, training data down-sampling is commonly performed due to limited resources, the need to adapt image size to the model input, or improve data augmentation. This down-sampling typically employs different strategies…

Computer Vision and Pattern Recognition · Computer Science 2024-02-20 Roberto Alcover-Couso , Marcos Escudero-Vinolo , Juan C. SanMiguel , Jose M. Martinez

Machine learning techniques are increasingly being applied in high-energy nuclear physics data analysis thanks to their outstanding performance. One key challenge in such applications is the construction of training samples that can…

Nuclear Experiment · Physics 2025-11-14 Yan Wang , Rangrong Ma , Kaifeng Shen , Zebo Tang , Wangmei Zha

Stochastic resetting, a method for accelerating target search in random processes, often incurs temporal and energetic costs. For a diffusing particle, a lower bound exists for the energetic cost of reaching the target, which is attained at…

Statistical Mechanics · Physics 2024-09-17 Ofir Tal-Friedman , Tommer D. Keidar , Shlomi Reuveni , Yael Roichman

Storage-efficient privacy-preserving learning is crucial due to increasing amounts of sensitive user data required for modern learning tasks. We propose a framework for reducing the storage cost of user data while at the same time providing…

Information Theory · Computer Science 2023-03-23 Berivan Isik , Tsachy Weissman

Spectro-microscopy is an experimental technique which can be used to observe spatial variations in chemical state and changes in chemical state over time or under experimental conditions. As a result it has broad applications across areas…

Medical Physics · Physics 2026-02-05 Maike Meier , Lorenzo Lazzarino , Boris Shustin , Hussam Al Daas , Paul Quinn

Faced with the scarcity of clean label data in real scenarios, seismic denoising methods based on supervised learning (SL) often encounter performance limitations. Specifically, when a model trained on synthetic data is directly applied to…

Geophysics · Physics 2023-11-07 Shijun Cheng , Zhiyao Cheng , Chao Jiang , Weijian Mao , Qingchen Zhang

Noise is one of the primary sources of interference in seismic exploration. Many authors have proposed various methods to remove noise from seismic data; however, in the face of strong noise conditions, satisfactory results are often not…

Geophysics · Physics 2024-04-04 Junheng Peng , Yong Li , Yingtian Liu , Zhangquan Liao