English
Related papers

Related papers: Whombat: An open-source annotation tool for machin…

200 papers

Speaker anonymization is the task of modifying a speech recording such that the original speaker cannot be identified anymore. Since the first Voice Privacy Challenge in 2020, along with the release of a framework, the popularity of this…

Sound · Computer Science 2023-12-25 Sarina Meyer , Xiaoxiao Miao , Ngoc Thang Vu

We present a lightweight annotation tool, the Data AnnotatoR Tool (DART), for the general task of labeling structured data with textual descriptions. The tool is implemented as an interactive application that reduces human efforts in…

Computation and Language · Computer Science 2020-12-02 Ernie Chang , Jeriah Caplinger , Alex Marin , Xiaoyu Shen , Vera Demberg

The marine ecosystem is changing at an alarming rate, exhibiting biodiversity loss and the migration of tropical species to temperate basins. Monitoring the underwater environments and their inhabitants is of fundamental importance to…

Sound · Computer Science 2022-01-17 Michele Mancusi , Nicola Zonca , Emanuele Rodolà , Silvia Zuffi

Detecting the presence of animal vocalisations in nature is essential to study animal populations and their behaviors. A recent development in the field is the introduction of the task known as few-shot bioacoustic sound event detection,…

Audio and Speech Processing · Electrical Eng. & Systems 2024-03-28 Jinhua Liang , Ines Nolasco , Burooj Ghani , Huy Phan , Emmanouil Benetos , Dan Stowell

Recent advances in artificial intelligence, including the development of highly sophisticated large language models (LLM), have proven beneficial in many real-world applications. However, evidence of inherent bias encoded in these LLMs has…

Computation and Language · Computer Science 2023-09-19 Vithya Yogarajan , Gillian Dobbie , Timothy Pistotti , Joshua Bensemann , Kobe Knowles

Seamless interaction between AI agents and humans using natural language remains a key goal in AI research. This paper addresses the challenges of developing interactive agents capable of understanding and executing grounded natural…

Artificial Intelligence · Computer Science 2024-07-15 Shrestha Mohanty , Negar Arabzadeh , Andrea Tupini , Yuxuan Sun , Alexey Skrynnik , Artem Zholus , Marc-Alexandre Côté , Julia Kiseleva

Bioacoustic sound event detection (BioSED) is crucial for biodiversity conservation but faces practical challenges during model development and training: limited amounts of annotated data, sparse events, species diversity, and class…

Sound · Computer Science 2025-05-30 Shiqi Zhang , Tuomas Virtanen

Data annotation, the practice of assigning descriptive labels to raw data, is pivotal in optimizing the performance of machine learning models. However, it is a resource-intensive process susceptible to biases introduced by annotators. The…

The value of quality treebanks is steadily increasing due to the crucial role they play in the development of natural language processing tools. The creation of such treebanks is enormously labor-intensive and time-consuming. Especially…

Computation and Language · Computer Science 2024-02-06 Salih Furkan Akkurt , Büşra Marşan , Susan Uskudarli

Passive acoustic monitoring can be an effective way of monitoring wildlife populations that are acoustically active but difficult to survey visually. Digital recorders allow surveyors to gather large volumes of data at low cost, but…

Sound · Computer Science 2023-08-25 Yuheng Wang , Juan Ye , David L. Borchers

The traditional data annotation process is often labor-intensive, time-consuming, and susceptible to human bias, which complicates the management of increasingly complex datasets. This study explores the potential of large language models…

Computation and Language · Computer Science 2024-09-17 Jianfei Wu , Xubin Wang , Weijia Jia

The advent of Large Language Models (LLMs) has advanced the benchmark in various Natural Language Processing (NLP) tasks. However, large amounts of labelled training data are required to train LLMs. Furthermore, data annotation and training…

Computation and Language · Computer Science 2024-03-05 Sargam Yadav , Abhishek Kaushik , Kevin McDaid

Linguistic annotation of transcribed speech is essential for research in language acquisition, language disorders, and sociolinguistics, yet remains labor-intensive and time-consuming. While Large Language Models (LLMs) have shown promise…

Computation and Language · Computer Science 2026-05-19 Qingwen Zhao , Hongao Zhu , Yunqi He , Rui Wang , Aijun Huang , Hai Hu

Ecological and conservation studies monitoring bird communities typically rely on species classification based on bird vocalizations. Historically, this has been based on expert volunteers going into the field and making lists of the bird…

Methodology · Statistics 2026-05-29 Haoxuan Wang , Patrik Lauha , David B. Dunson

Evaluating production-level retrieval systems at scale is a crucial yet challenging task due to the limited availability of a large pool of well-trained human annotators. Large Language Models (LLMs) have the potential to address this…

Information Retrieval · Computer Science 2024-09-19 Kasra Hosseini , Thomas Kober , Josip Krapac , Roland Vollgraf , Weiwei Cheng , Ana Peleteiro Ramallo

Despite a growing ecosystem of tools supporting Systematic Literature Reviews (SLRs), integrating them into user-friendly workflows remains challenging. The Streamlined Workflow for Automating Machine-Actionable Systematic Literature…

Digital Libraries · Computer Science 2026-03-06 Tim Wittenborg , Allard Oelen , Manuel Prinz

Neural approaches have become very popular in Question Answering (QA), however, they require a large amount of annotated data. In this work, we propose a novel approach that combines data augmentation via question-answer generation with…

Computation and Language · Computer Science 2024-09-16 Maximilian Kimmich , Andrea Bartezzaghi , Jasmina Bogojeska , Cristiano Malossi , Ngoc Thang Vu

Advanced omics technologies and facilities generate a wealth of valuable data daily; however, the data often lacks the essential metadata required for researchers to find and search them effectively. The lack of metadata poses a significant…

We present HowSumm, a novel large-scale dataset for the task of query-focused multi-document summarization (qMDS), which targets the use-case of generating actionable instructions from a set of sources. This use-case is different from the…

Computation and Language · Computer Science 2021-10-12 Odellia Boni , Guy Feigenblat , Guy Lev , Michal Shmueli-Scheuer , Benjamin Sznajder , David Konopnicki

Surgical AI often involves multiple tasks within a single procedure, like phase recognition or assessing the Critical View of Safety in laparoscopic cholecystectomy. Traditional models, built for one task at a time, lack flexibility,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-11 Soham Walimbe , Britty Baby , Vinkle Srivastav , Nicolas Padoy
‹ Prev 1 8 9 10 Next ›