English
Related papers

Related papers: Expectation-Maximization as the Engine of Scalable…

200 papers

Electronic health records contain inconsistently structured or free-text data, requiring efficient preprocessing to enable predictive health care models. Although artificial intelligence-driven natural language processing tools show promise…

AI systems in high-consequence domains such as defense, intelligence, and disaster response must detect rare, high-impact events while operating under tight resource constraints. Traditional annotation strategies that prioritize label…

Machine Learning · Computer Science 2025-05-22 Dave Cook , Tim Klawa

The lack of sufficient annotated image data is a common issue in medical image segmentation. For some organs and densities, the annotation may be scarce, leading to poor model training convergence, while other organs have plenty of…

Image and Video Processing · Electrical Eng. & Systems 2021-09-22 Anastasia Makarevich , Azade Farshad , Vasileios Belagiannis , Nassir Navab

Traditional image annotation tasks rely heavily on human effort for object selection and label assignment, making the process time-consuming and prone to decreased efficiency as annotators experience fatigue after extensive work. This paper…

Computer Vision and Pattern Recognition · Computer Science 2025-03-17 He Zhang , Xinyi Fu , John M. Carroll

The advancement of artificial intelligence (AI) for organ segmentation and tumor detection is propelled by the growing availability of computed tomography (CT) datasets with detailed, per-voxel annotations. However, these AI models often…

Image and Video Processing · Electrical Eng. & Systems 2024-05-29 Jie Liu , Yixiao Zhang , Kang Wang , Mehmet Can Yavuz , Xiaoxi Chen , Yixuan Yuan , Haoliang Li , Yang Yang , Alan Yuille , Yucheng Tang , Zongwei Zhou

There is a dire need for medical imaging datasets with accompanying annotations to perform downstream patient analysis. However, it is difficult to manually generate these annotations, due to the time-consuming nature, and the variability…

Image and Video Processing · Electrical Eng. & Systems 2024-06-21 Deepa Krishnaswamy , Vamsi Krishna Thiriveedhi , Cosmin Ciausu , David Clunie , Steve Pieper , Ron Kikinis , Andrey Fedorov

Patient recruitment remains a major bottleneck in clinical trials, calling for scalable and automated solutions. We present TrialMatchAI, an AI-powered recommendation system that automates patient-to-trial matching by processing…

Automated text annotation is a compelling use case for generative large language models (LLMs) in social media research. Recent work suggests that LLMs can achieve strong performance on annotation tasks; however, these studies evaluate LLMs…

Computation and Language · Computer Science 2024-09-24 Nicholas Pangakis , Samuel Wolken

Supervised deep learning methods for segmentation require large amounts of labelled training data, without which they are prone to overfitting, not generalizing well to unseen images. In practice, obtaining a large number of annotations…

Computer Vision and Pattern Recognition · Computer Science 2019-03-01 Krishna Chaitanya , Neerav Karani , Christian Baumgartner , Olivio Donati , Anton Becker , Ender Konukoglu

While large language models (LLMs) demonstrate reasonable zero-shot capability across many downstream tasks, fine-tuning is a common practice to improve their performance. However, a task's data efficiency--i.e., the number of fine-tuning…

Machine Learning · Computer Science 2026-01-01 Gyung Hyun Je , Colin Raffel

The translation of AI-generated brain metastases (BM) segmentation into clinical practice relies heavily on diverse, high-quality annotated medical imaging datasets. The BraTS-METS 2023 challenge has gained momentum for testing and…

Other Quantitative Biology · Quantitative Biology 2024-12-10 Ahmed W. Moawad , Anastasia Janas , Ujjwal Baid , Divya Ramakrishnan , Rachit Saluja , Nader Ashraf , Nazanin Maleki , Leon Jekel , Nikolay Yordanov , Pascal Fehringer , Athanasios Gkampenis , Raisa Amiruddin , Amirreza Manteghinejad , Maruf Adewole , Jake Albrecht , Udunna Anazodo , Sanjay Aneja , Syed Muhammad Anwar , Timothy Bergquist , Veronica Chiang , Verena Chung , Gian Marco Conte , Farouk Dako , James Eddy , Ivan Ezhov , Nastaran Khalili , Keyvan Farahani , Juan Eugenio Iglesias , Zhifan Jiang , Elaine Johanson , Anahita Fathi Kazerooni , Florian Kofler , Kiril Krantchev , Dominic LaBella , Koen Van Leemput , Hongwei Bran Li , Marius George Linguraru , Xinyang Liu , Zeke Meier , Bjoern H Menze , Harrison Moy , Klara Osenberg , Marie Piraud , Zachary Reitman , Russell Takeshi Shinohara , Chunhao Wang , Benedikt Wiestler , Walter Wiggins , Umber Shafique , Klara Willms , Arman Avesta , Khaled Bousabarah , Satrajit Chakrabarty , Nicolo Gennaro , Wolfgang Holler , Manpreet Kaur , Pamela LaMontagne , MingDe Lin , Jan Lost , Daniel S. Marcus , Ryan Maresca , Sarah Merkaj , Gabriel Cassinelli Pedersen , Marc von Reppert , Aristeidis Sotiras , Oleg Teytelboym , Niklas Tillmans , Malte Westerhoff , Ayda Youssef , Devon Godfrey , Scott Floyd , Andreas Rauschecker , Javier Villanueva-Meyer , Irada Pfluger , Jaeyoung Cho , Martin Bendszus , Gianluca Brugnara , Justin Cramer , Gloria J. Guzman Perez-Carillo , Derek R. Johnson , Anthony Kam , Benjamin Yin Ming Kwan , Lillian Lai , Neil U. Lall , Fatima Memon , Mark Krycia , Satya Narayana Patro , Bojan Petrovic , Tiffany Y. So , Gerard Thompson , Lei Wu , E. Brooke Schrickel , Anu Bansal , Frederik Barkhof , Cristina Besada , Sammy Chu , Jason Druzgal , Alexandru Dusoi , Luciano Farage , Fabricio Feltrin , Amy Fong , Steve H. Fung , R. Ian Gray , Ichiro Ikuta , Michael Iv , Alida A. Postma , Amit Mahajan , David Joyner , Chase Krumpelman , Laurent Letourneau-Guillon , Christie M. Lincoln , Mate E. Maros , Elka Miller , Fanny Moron , Esther A. Nimchinsky , Ozkan Ozsarlak , Uresh Patel , Saurabh Rohatgi , Atin Saha , Anousheh Sayah , Eric D. Schwartz , Robert Shih , Mark S. Shiroishi , Juan E. Small , Manoj Tanwar , Jewels Valerie , Brent D. Weinberg , Matthew L. White , Robert Young , Vahe M. Zohrabian , Aynur Azizova , Melanie Maria Theresa Bruseler , Mohanad Ghonim , Mohamed Ghonim , Abdullah Okar , Luca Pasquini , Yasaman Sharifi , Gagandeep Singh , Nico Sollmann , Theodora Soumala , Mahsa Taherzadeh , Philipp Vollmuth , Martha Foltyn-Dumitru , Ajay Malhotra , Aly H. Abayazeed , Francesco Dellepiane , Philipp Lohmann , Victor M. Perez-Garcia , Hesham Elhalawani , Maria Correia de Verdier , Sanaria Al-Rubaiey , Rui Duarte Armindo , Kholod Ashraf , Moamen M. Asla , Mohamed Badawy , Jeroen Bisschop , Nima Broomand Lomer , Jan Bukatz , Jim Chen , Petra Cimflova , Felix Corr , Alexis Crawley , Lisa Deptula , Tasneem Elakhdar , Islam H. Shawali , Shahriar Faghani , Alexandra Frick , Vaibhav Gulati , Muhammad Ammar Haider , Fatima Hierro , Rasmus Holmboe Dahl , Sarah Maria Jacobs , Kuang-chun Jim Hsieh , Sedat G. Kandemirli , Katharina Kersting , Laura Kida , Sofia Kollia , Ioannis Koukoulithras , Xiao Li , Ahmed Abouelatta , Aya Mansour , Ruxandra-Catrinel Maria-Zamfirescu , Marcela Marsiglia , Yohana Sarahi Mateo-Camacho , Mark McArthur , Olivia McDonnell , Maire McHugh , Mana Moassefi , Samah Mostafa Morsi , Alexander Munteanu , Khanak K. Nandolia , Syed Raza Naqvi , Yalda Nikanpour , Mostafa Alnoury , Abdullah Mohamed Aly Nouh , Francesca Pappafava , Markand D. Patel , Samantha Petrucci , Eric Rawie , Scott Raymond , Borna Roohani , Sadeq Sabouhi , Laura M. Sanchez-Garcia , Zoe Shaked , Pokhraj P. Suthar , Talissa Altes , Edvin Isufi , Yaseen Dhemesh , Jaime Gass , Jonathan Thacker , Abdul Rahman Tarabishy , Benjamin Turner , Sebastiano Vacca , George K. Vilanilam , Daniel Warren , David Weiss , Fikadu Worede , Sara Yousry , Wondwossen Lerebo , Alejandro Aristizabal , Alexandros Karargyris , Hasan Kassem , Sarthak Pati , Micah Sheller , Katherine E. Link , Evan Calabrese , Nourel hoda Tahon , Ayman Nada , Yuri S. Velichko , Spyridon Bakas , Jeffrey D. Rudie , Mariam Aboian

Understanding how segmentation performance scales with training data is fundamental for developing data-efficient medical AI systems. In this study, we systematically revisit data scaling behavior across 15 anatomical segmentation tasks…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Yuetan Chu , Zhongyi Han , Gongning Luo , Xin Gao

In safety-critical applications like medical diagnosis, certainty associated with a model's prediction is just as important as its accuracy. Consequently, uncertainty estimation and reduction play a crucial role. Uncertainty in predictions…

Image and Video Processing · Electrical Eng. & Systems 2023-09-12 Abhishek Singh Sambyal , Narayanan C. Krishnan , Deepti R. Bathula

The rise of Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) has rapidly increased the need for high-quality, curated information retrieval datasets. These datasets, however, are currently created with off-the-shelf…

Information Retrieval · Computer Science 2026-02-05 Sameh Khattab , Marie Bauer , Lukas Heine , Till Rostalski , Jens Kleesiek , Julian Friedrich

Recent Artificial Intelligence (AI) models have matched or exceeded human experts in several benchmarks of biomedical task performance, but surgical benchmarks in particular are often missing from prominent medical benchmark suites. Since…

Public imaging datasets are critical for the development and evaluation of automated tools in cancer imaging. Unfortunately, many do not include annotations or image-derived features, complicating their downstream analysis. Artificial…

Computer Vision and Pattern Recognition · Computer Science 2023-06-02 Deepa Krishnaswamy , Dennis Bontempi , Vamsi Thiriveedhi , Davide Punzo , David Clunie , Christopher P Bridge , Hugo JWL Aerts , Ron Kikinis , Andrey Fedorov

The quality of the dataset is crucial for ensuring optimal performance and reliability of downstream task models. However, datasets often contain noisy data inadvertently included during the construction process. Numerous attempts have been…

Computation and Language · Computer Science 2024-09-25 Juhwan Choi , Jungmin Yun , Kyohoon Jin , YoungBin Kim

Radiotherapy (RT) planning is complex, subjective, and time-intensive. Advances with artificial intelligence (AI) promise to improve its precision and efficiency, but progress is often limited by the scarcity of large, standardized…

Learning from human feedback has become a pivot technique in aligning large language models (LLMs) with human preferences. However, acquiring vast and premium human feedback is bottlenecked by time, labor, and human capability, resulting in…

Computation and Language · Computer Science 2024-07-17 Ganqu Cui , Lifan Yuan , Ning Ding , Guanming Yao , Bingxiang He , Wei Zhu , Yuan Ni , Guotong Xie , Ruobing Xie , Yankai Lin , Zhiyuan Liu , Maosong Sun

Multimodal (MM) learning is emerging as a promising paradigm in biomedical artificial intelligence (AI) applications, integrating complementary modality, which highlight different aspects of patient health. The scarcity of large…

Artificial Intelligence · Computer Science 2025-12-01 Niccolo Marini , Zhaohui Liang , Sivaramakrishnan Rajaraman , Zhiyun Xue , Sameer Antani