English
Related papers

Related papers: A Spark ML driven preprocessing approach for deep …

200 papers

In this paper, we primarily focus on understanding the data preprocessing pipeline for DNN Training in the public cloud. First, we run experiments to test the performance implications of the two major data preprocessing methods using either…

Machine Learning · Computer Science 2023-04-19 Ping Gong , Yuxin Ma , Cheng Li , Xiaosong Ma , Sam H. Noh

The use of large pretrained neural networks to create contextualized word embeddings has drastically improved performance on several natural language processing (NLP) tasks. These computationally expensive models have begun to be applied to…

Computers and Society · Computer Science 2019-12-03 Benjamin Clavié , Kobi Gal

We introduce Microsoft Machine Learning for Apache Spark (MMLSpark), an ecosystem of enhancements that expand the Apache Spark distributed computing library to tackle problems in Deep Learning, Micro-Service Orchestration, Gradient…

This study investigates the automation of meta-analysis in scientific documents using large language models (LLMs). Meta-analysis is a robust statistical method that synthesizes the findings of multiple studies support articles to provide a…

Computation and Language · Computer Science 2024-11-19 Jawad Ibn Ahad , Rafeed Mohammad Sultan , Abraham Kaikobad , Fuad Rahman , Mohammad Ruhul Amin , Nabeel Mohammed , Shafin Rahman

Extracting information from big data sets, both real and simulated, is a modern hallmark of the physical sciences. In practice, students face barriers to learning ``Big Data'' methods in undergraduate physics and astronomy curricula. As an…

Physics Education · Physics 2025-09-12 Stéphane Delorme , Leon Mach , Hubert Paszkiewicz , Richard Ruiz

Spark is an in-memory analytics platform that targets commodity server environments today. It relies on the Hadoop Distributed File System (HDFS) to persist intermediate checkpoint states and final processing results. In Spark, immutable…

Distributed, Parallel, and Cluster Computing · Computer Science 2017-08-22 Mijung Kim , Jun Li , Haris Volos , Manish Marwah , Alexander Ulanov , Kimberly Keeton , Joseph Tucek , Lucy Cherkasova , Le Xu , Pradeep Fernando

Analyzing and evaluating students' progress in any learning environment is stressful and time consuming if done using traditional analysis methods. This is further exasperated by the increasing number of students due to the shift of focus…

Computers and Society · Computer Science 2024-02-06 Abdallah Moubayed , MohammadNoor Injadat , Nouh Alhindawi , Ghassan Samara , Sara Abuasal , Raed Alazaidah

Deploying Machine Learning (ML) algorithms within databases is a challenge due to the varied computational footprints of modern ML algorithms and the myriad of database technologies each with its own restrictive syntax. We introduce an…

In this community review report, we discuss applications and techniques for fast machine learning (ML) in science -- the concept of integrating power ML methods into the real-time experimental data processing loop to accelerate scientific…

Machine Learning · Computer Science 2023-02-07 Allison McCarn Deiana , Nhan Tran , Joshua Agar , Michaela Blott , Giuseppe Di Guglielmo , Javier Duarte , Philip Harris , Scott Hauck , Mia Liu , Mark S. Neubauer , Jennifer Ngadiuba , Seda Ogrenci-Memik , Maurizio Pierini , Thea Aarrestad , Steffen Bahr , Jurgen Becker , Anne-Sophie Berthold , Richard J. Bonventre , Tomas E. Muller Bravo , Markus Diefenthaler , Zhen Dong , Nick Fritzsche , Amir Gholami , Ekaterina Govorkova , Kyle J Hazelwood , Christian Herwig , Babar Khan , Sehoon Kim , Thomas Klijnsma , Yaling Liu , Kin Ho Lo , Tri Nguyen , Gianantonio Pezzullo , Seyedramin Rasoulinezhad , Ryan A. Rivera , Kate Scholberg , Justin Selig , Sougata Sen , Dmitri Strukov , William Tang , Savannah Thais , Kai Lukas Unger , Ricardo Vilalta , Belinavon Krosigk , Thomas K. Warburton , Maria Acosta Flechas , Anthony Aportela , Thomas Calvet , Leonardo Cristella , Daniel Diaz , Caterina Doglioni , Maria Domenica Galati , Elham E Khoda , Farah Fahim , Davide Giri , Benjamin Hawks , Duc Hoang , Burt Holzman , Shih-Chieh Hsu , Sergo Jindariani , Iris Johnson , Raghav Kansal , Ryan Kastner , Erik Katsavounidis , Jeffrey Krupa , Pan Li , Sandeep Madireddy , Ethan Marx , Patrick McCormack , Andres Meza , Jovan Mitrevski , Mohammed Attia Mohammed , Farouk Mokhtar , Eric Moreno , Srishti Nagu , Rohin Narayan , Noah Palladino , Zhiqiang Que , Sang Eon Park , Subramanian Ramamoorthy , Dylan Rankin , Simon Rothman , Ashish Sharma , Sioni Summers , Pietro Vischia , Jean-Roch Vlimant , Olivia Weng

Automatic machine learning, or AutoML, holds the promise of truly democratizing the use of machine learning (ML), by substantially automating the work of data scientists. However, the huge combinatorial search space of candidate pipelines…

Machine Learning · Computer Science 2022-04-21 Ripon K. Saha , Akira Ura , Sonal Mahajan , Chenguang Zhu , Linyi Li , Yang Hu , Hiroaki Yoshida , Sarfraz Khurshid , Mukul R. Prasad

Data is crucial for machine learning (ML) applications, yet acquiring large datasets can be costly and time-consuming, especially in complex, resource-intensive fields like biopharmaceuticals. A key process in this industry is upstream…

Machine Learning · Computer Science 2025-06-23 Johnny Peng , Thanh Tung Khuat , Katarzyna Musial , Bogdan Gabrys

In the contemporary information era, significantly accelerated by the advent of Large-scale Language Models, the proliferation of scientific literature is reaching unprecedented levels. Researchers urgently require efficient tools for…

Computation and Language · Computer Science 2024-01-18 Feng Jiang , Kuang Wang , Haizhou Li

The reasoning capabilities of Large Language Models (LLMs) play a critical role in many downstream tasks, yet depend strongly on the quality of training data. Despite various proposed data construction methods, their practical utility in…

Computation and Language · Computer Science 2025-10-09 Yike Zhao , Simin Guo , Ziqing Yang , Shifan Han , Dahua Lin , Fei Tan

Recently, large language models (LLMs) have shown promising abilities to generate novel research ideas in science, a direction which coincides with many foundational principles in computational creativity (CC). In light of these…

Artificial Intelligence · Computer Science 2025-05-23 Aishik Sanyal , Samuel Schapiro , Sumuk Shashidhar , Royce Moon , Lav R. Varshney , Dilek Hakkani-Tur

Large language models (LLMs) have demonstrated significant potential in code generation tasks. However, there remains a performance gap between open-source and closed-source models. To address this gap, existing approaches typically…

Computation and Language · Computer Science 2025-04-18 Weijie Lv , Xuan Xia , Sheng-Jun Huang

Data science and machine learning algorithms running on big data infrastructure are increasingly important in activities ranging from business intelligence and analytics to cybersecurity, smart city management, and many fields of science…

Distributed, Parallel, and Cluster Computing · Computer Science 2017-10-10 Eduardo Rodrigues , Ricardo Morla

This paper presents novel prompting techniques to improve the performance of automatic summarization systems for scientific articles. Scientific article summarization is highly challenging due to the length and complexity of these…

Computation and Language · Computer Science 2023-12-18 Aldan Creo , Manuel Lama , Juan C. Vidal

In the Big Data era, the community of PAM faces strong challenges, including the need for more standardized processing tools accross its different applications in oceanography, and for more scalable and high-performance computing systems to…

Audio and Speech Processing · Electrical Eng. & Systems 2019-03-17 D. Cazau

High-quality textual training data is essential for the success of multimodal data processing tasks, yet outputs from image captioning models like BLIP and GIT often contain errors and anomalies that are difficult to rectify using…

Computation and Language · Computer Science 2025-02-25 Elyas Meguellati , Nardiena Pratama , Shazia Sadiq , Gianluca Demartini

Recent work shows that post-training datasets for LLMs can be substantially downsampled without noticeably deteriorating performance. However, data selection often incurs high computational costs or is limited to narrow domains. In this…

Computation and Language · Computer Science 2025-09-25 Paramita Mirza , Lucas Weber , Fabian Küch