English
Related papers

Related papers: Oasis: Data Curation and Assessment System for Pre…

200 papers

Data curation is a field with origins in librarianship and archives, whose scholarship and thinking on data issues go back centuries, if not millennia. The field of machine learning is increasingly observing the importance of data curation…

Computers and Society · Computer Science 2025-01-06 Eshta Bhardwaj , Harshit Gujral , Siyi Wu , Ciara Zogheib , Tegan Maharaj , Christoph Becker

The quality of training data has a huge impact on the efficiency, accuracy and complexity of machine learning tasks. Various tools and techniques are available that assess data quality with respect to general cleaning and profiling checks.…

Vision-language models (VLMs) are trained for thousands of GPU hours on carefully curated web datasets. In recent times, data curation has gained prominence with several works developing strategies to retain 'high-quality' subsets of 'raw'…

Machine Learning · Computer Science 2024-04-11 Sachin Goyal , Pratyush Maini , Zachary C. Lipton , Aditi Raghunathan , J. Zico Kolter

Though preceding work in computational argument quality (AQ) mostly focuses on assessing overall AQ, researchers agree that writers would benefit from feedback targeting individual dimensions of argumentation theory. However, a large-scale…

Computation and Language · Computer Science 2020-11-04 Anne Lauscher , Lily Ng , Courtney Napoles , Joel Tetreault

Large language models are commonly trained on a mixture of filtered web data and curated high-quality corpora, such as social media conversations, books, or technical papers. This curation process is believed to be necessary to produce…

Entity resolution (ER) presents unique challenges for evaluation methodology. While crowdsourcing platforms acquire ground truth, sound approaches to sampling must drive labelling efforts. In ER, extreme class imbalance between matching and…

Machine Learning · Computer Science 2017-06-27 Neil G. Marchant , Benjamin I. P. Rubinstein

The performance of deep neural networks is strongly influenced by the quality of their training data. However, mitigating dataset bias by manually curating challenging edge cases remains a major bottleneck. To address this, we propose an…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Kyeongryeol Go

Synthetic data is widely used in healthcare to create datasets that are similar to original data but without the privacy concerns. Generating and evaluating synthetic data across privacy, utility and fairness is crucial for facilitating…

The advancement of artificial intelligence (AI) for organ segmentation and tumor detection is propelled by the growing availability of computed tomography (CT) datasets with detailed, per-voxel annotations. However, these AI models often…

Image and Video Processing · Electrical Eng. & Systems 2024-05-29 Jie Liu , Yixiao Zhang , Kang Wang , Mehmet Can Yavuz , Xiaoxi Chen , Yixuan Yuan , Haoliang Li , Yang Yang , Alan Yuille , Yucheng Tang , Zongwei Zhou

Recent breakthroughs in natural language processing and computer vision, driven by efficient pre-training on large datasets, have enabled foundation models to excel on a wide range of tasks. However, this potential has not yet been fully…

Machine Learning · Computer Science 2025-02-03 Özgün Turgut , Philip Müller , Martin J. Menten , Daniel Rueckert

Recent work has demonstrated that increased training dataset diversity improves general cross-domain knowledge and downstream generalization capability for large-scale language models. With this in mind, we present \textit{the Pile}: an 825…

This paper presents TTSOps, a fully automated closed-loop framework for constructing multi-speaker text-to-speech (TTS) systems from noisy, uncurated web-scale speech data, often referred to as ``dark data,'' such as online videos.…

Sound · Computer Science 2025-11-12 Kentaro Seki , Shinnosuke Takamichi , Takaaki Saeki , Hiroshi Saruwatari

The availability of large data sets is providing an impetus for driving current artificial intelligent developments. There are, however, challenges for developing solutions with small data sets due to practical and cost-effective deployment…

Machine Learning · Computer Science 2024-05-17 Luca Gherardini , Varun Ravi Varma , Karol Capala , Roger Woods , Jose Sousa

Recent advances in conversational AI have been substantial, but developing real-time systems for perceptual task guidance remains challenging. These systems must provide interactive, proactive assistance based on streaming visual inputs,…

Artificial Intelligence · Computer Science 2025-06-09 Yichi Zhang , Xin Luna Dong , Zhaojiang Lin , Andrea Madotto , Anuj Kumar , Babak Damavandi , Joyce Chai , Seungwhan Moon

Reliable automatic evaluation of summarization systems is challenging due to the multifaceted and subjective nature of the task. This is especially the case for languages other than English, where human evaluations are scarce. In this work,…

Multilingual document understanding remains limited for low-resource languages due to scarce training data and model-based annotation pipelines that perpetuate existing biases. We introduce DocAtlas, a framework that constructs…

We present Autonomous Data Selection (AutoDS), a method that leverages base language models themselves as zero-shot "generative classifiers" to automatically curate high-quality mathematical texts. Unlike prior approaches that require human…

Computation and Language · Computer Science 2025-07-23 Yifan Zhang , Yifan Luo , Yang Yuan , Andrew C Yao

The introduction of large language models and other influential developments in AI-based language processing have led to an evolution in the methods available to quantitatively analyse language data. With the resultant growth of attention…

Computation-Enabled Object Storage (COS) systems, such as MinIO and Ceph, have recently emerged as promising storage solutions for post hoc, SQL-based analysis on large-scale datasets in High-Performance Computing (HPC) environments. By…

This work investigates the selection of high-quality pre-training data from massive corpora to enhance LMs' capabilities for downstream usage. We formulate data selection as a generalized Optimal Control problem, which can be solved…

Computation and Language · Computer Science 2025-03-20 Yuxian Gu , Li Dong , Hongning Wang , Yaru Hao , Qingxiu Dong , Furu Wei , Minlie Huang
‹ Prev 1 4 5 6 7 8 10 Next ›