English
Related papers

Related papers: Enhancing Protein Predictive Models via Proteins D…

200 papers

Data augmentation in time series forecasting plays a crucial role in enhancing model performance by introducing variability while maintaining the underlying temporal patterns. However, time series data offers fewer augmentation strategies…

Machine Learning · Computer Science 2025-11-12 Dang Nha Nguyen , Hai Dang Nguyen , Khoa Tho Anh Nguyen

Data augmentations are effective in improving the invariance of learning machines. We argue that the core challenge of data augmentations lies in designing data transformations that preserve labels. This is relatively straightforward for…

Machine Learning · Computer Science 2023-03-01 Youzhi Luo , Michael McThrow , Wing Yee Au , Tao Komikado , Kanji Uchino , Koji Maruhashi , Shuiwang Ji

Quantitative measurements produced by mass spectrometry proteomics experiments offer a direct way to explore the role of proteins in molecular mechanisms. However, analysis of such data is challenging due to the large proportion of missing…

Methodology · Statistics 2025-01-22 Haeun Moon , Jin-Hong Du , Jing Lei , Kathryn Roeder

Objective: The use of deep learning for electroencephalography (EEG) classification tasks has been rapidly growing in the last years, yet its application has been limited by the relatively small size of EEG datasets. Data augmentation,…

Machine Learning · Computer Science 2022-11-16 Cédric Rommel , Joseph Paillard , Thomas Moreau , Alexandre Gramfort

Text classification is a representative downstream task of natural language processing, and has exhibited excellent performance since the advent of pre-trained language models based on Transformer architecture. However, in pre-trained…

Computation and Language · Computer Science 2022-04-07 Byeong-Cheol Jo , Tak-Sung Heo , Yeongjoon Park , Yongmin Yoo , Won Ik Cho , Kyungsun Kim

Data augmentation is an essential technique in natural language processing (NLP) for enriching training datasets by generating diverse samples. This process is crucial for improving the robustness and generalization capabilities of NLP…

Computation and Language · Computer Science 2025-10-16 Zaitian Wang , Jinghan Zhang , Xinhao Zhang , Kunpeng Liu , Pengfei Wang , Yuanchun Zhou

Computational protein design is experiencing a transformation driven by AI/ML. However, the range of potential protein sequences and structures is astronomically vast, even for moderately sized proteins. Hence, achieving convergence between…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-10-09 Aymen Alsaadi , Jonathan Ash , Mikhail Titov , Matteo Turilli , Andre Merzky , Shantenu Jha , Sagar Khare

Learning recipe and food image representation in common embedding space is non-trivial but crucial for cross-modal recipe retrieval. In this paper, we propose a new perspective for this problem by utilizing foundation models for data…

Information Retrieval · Computer Science 2024-07-18 Fangzhou Song , Bin Zhu , Yanbin Hao , Shuo Wang

Manipulating data, such as weighting data examples or augmenting with new instances, has been increasingly used to improve model training. Previous work has studied various rule- or learning-based approaches designed for specific types of…

Machine Learning · Computer Science 2019-10-29 Zhiting Hu , Bowen Tan , Ruslan Salakhutdinov , Tom Mitchell , Eric P. Xing

Recent studies on semi-supervised semantic segmentation (SSS) have seen fast progress. Despite their promising performance, current state-of-the-art methods tend to increasingly complex designs at the cost of introducing more network…

Computer Vision and Pattern Recognition · Computer Science 2022-12-12 Zhen Zhao , Lihe Yang , Sifan Long , Jimin Pi , Luping Zhou , Jingdong Wang

Data Augmentation (DA) -- generating extra training samples beyond original training set -- has been widely-used in today's unbiased VQA models to mitigate the language biases. Current mainstream DA strategies are synthetic-based methods,…

Computer Vision and Pattern Recognition · Computer Science 2022-09-16 Long Chen , Yuhang Zheng , Jun Xiao

In semi-supervised semantic segmentation (SSSS), data augmentation plays a crucial role in the weak-to-strong consistency regularization framework, as it enhances diversity and improves model generalization. Recent strong augmentation…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Lingyan Ran , Yali Li , Tao Zhuo , Shizhou Zhang , Yanning Zhang

Augmenting training datasets has been shown to improve the learning effectiveness for several computer vision tasks. A good augmentation produces an augmented dataset that adds variability while retaining the statistical properties of the…

Computer Vision and Pattern Recognition · Computer Science 2020-04-01 Tom Ching LingChen , Ava Khonsari , Amirreza Lashkari , Mina Rafi Nazari , Jaspreet Singh Sambee , Mario A. Nascimento

Background: The increasing volume and variety of genotypic and phenotypic data is a major defining characteristic of modern biomedical sciences. At the same time, the limitations in technology for generating data and the inherently…

Quantitative Methods · Quantitative Biology 2016-12-07 Yuxiang Jiang , Tal Ronnen Oron , Wyatt T Clark , Asma R Bankapur , Daniel D'Andrea , Rosalba Lepore , Christopher S Funk , Indika Kahanda , Karin M Verspoor , Asa Ben-Hur , Emily Koo , Duncan Penfold-Brown , Dennis Shasha , Noah Youngs , Richard Bonneau , Alexandra Lin , Sayed ME Sahraeian , Pier Luigi Martelli , Giuseppe Profiti , Rita Casadio , Renzhi Cao , Zhaolong Zhong , Jianlin Cheng , Adrian Altenhoff , Nives Skunca , Christophe Dessimoz , Tunca Dogan , Kai Hakala , Suwisa Kaewphan , Farrokh Mehryary , Tapio Salakoski , Filip Ginter , Hai Fang , Ben Smithers , Matt Oates , Julian Gough , Petri Törönen , Patrik Koskinen , Liisa Holm , Ching-Tai Chen , Wen-Lian Hsu , Kevin Bryson , Domenico Cozzetto , Federico Minneci , David T Jones , Samuel Chapman , Dukka B K. C. , Ishita K Khan , Daisuke Kihara , Dan Ofer , Nadav Rappoport , Amos Stern , Elena Cibrian-Uhalte , Paul Denny , Rebecca E Foulger , Reija Hieta , Duncan Legge , Ruth C Lovering , Michele Magrane , Anna N Melidoni , Prudence Mutowo-Meullenet , Klemens Pichler , Aleksandra Shypitsyna , Biao Li , Pooya Zakeri , Sarah ElShal , Léon-Charles Tranchevent , Sayoni Das , Natalie L Dawson , David Lee , Jonathan G Lees , Ian Sillitoe , Prajwal Bhat , Tamás Nepusz , Alfonso E Romero , Rajkumar Sasidharan , Haixuan Yang , Alberto Paccanaro , Jesse Gillis , Adriana E Sedeño-Cortés , Paul Pavlidis , Shou Feng , Juan M Cejuela , Tatyana Goldberg , Tobias Hamp , Lothar Richter , Asaf Salamov , Toni Gabaldon , Marina Marcet-Houben , Fran Supek , Qingtian Gong , Wei Ning , Yuanpeng Zhou , Weidong Tian , Marco Falda , Paolo Fontana , Enrico Lavezzo , Stefano Toppo , Carlo Ferrari , Manuel Giollo , Damiano Piovesan , Silvio Tosatto , Angela del Pozo , José M Fernández , Paolo Maietta , Alfonso Valencia , Michael L Tress , Alfredo Benso , Stefano Di Carlo , Gianfranco Politano , Alessandro Savino , Hafeez Ur Rehman , Matteo Re , Marco Mesiti , Giorgio Valentini , Joachim W Bargsten , Aalt DJ van Dijk , Branislava Gemovic , Sanja Glisic , Vladmir Perovic , Veljko Veljkovic , Nevena Veljkovic , Danillo C Almeida-e-Silva , Ricardo ZN Vencio , Malvika Sharan , Jörg Vogel , Lakesh Kansakar , Shanshan Zhang , Slobodan Vucetic , Zheng Wang , Michael JE Sternberg , Mark N Wass , Rachael P Huntley , Maria J Martin , Claire O'Donovan , Peter N Robinson , Yves Moreau , Anna Tramontano , Patricia C Babbitt , Steven E Brenner , Michal Linial , Christine A Orengo , Burkhard Rost , Casey S Greene , Sean D Mooney , Iddo Friedberg , Predrag Radivojac

The excellent performance of deep neural networks is usually accompanied by a large number of parameters and computations, which have limited their usage on the resource-limited edge devices. To address this issue, abundant methods such as…

Computer Vision and Pattern Recognition · Computer Science 2023-05-23 Muzhou Yu , Linfeng Zhang , Kaisheng Ma

In the realm of medical imaging, the training of machine learning models necessitates a large and varied training dataset to ensure robustness and interoperability. However, acquiring such diverse and heterogeneous data can be difficult due…

Image and Video Processing · Electrical Eng. & Systems 2023-03-03 Manuel Cossio

Deep learning has revolutionized the performance of classification, but meanwhile demands sufficient labeled data for training. Given insufficient data, while many techniques have been developed to help combat overfitting, the challenge…

Computer Vision and Pattern Recognition · Computer Science 2018-09-05 Xiaofeng Zhang , Zhangyang Wang , Dong Liu , Qing Ling

Data augmentation, a technique in which a training set is expanded with class-preserving transformations, is ubiquitous in modern machine learning pipelines. In this paper, we seek to establish a theoretical framework for understanding data…

Machine Learning · Computer Science 2019-03-21 Tri Dao , Albert Gu , Alexander J. Ratner , Virginia Smith , Christopher De Sa , Christopher Ré

Data augmentation is used extensively to improve model generalisation. However, reliance on external libraries to implement augmentation methods introduces a vulnerability into the machine learning pipeline. It is well known that backdoors…

Machine Learning · Computer Science 2022-10-03 Joseph Rance , Yiren Zhao , Ilia Shumailov , Robert Mullins

Medical image analysis suffers from a lack of labeled data due to several challenges including patient privacy and lack of experts. Although some AI models only perform well with large amounts of data, we will move to data augmentation…

Image and Video Processing · Electrical Eng. & Systems 2025-11-26 Khadija Rais , Mohamed Amroune , Mohamed Yassine Haouam , Abdelmadjid Benmachiche