English
Related papers

Related papers: Sound Search by Text Description or Vocal Imitatio…

200 papers

Achieving diverse and high-quality audio transformations from text prompts remains challenging, as existing methods are fundamentally constrained by their reliance on a limited set of differentiable audio effects. This paper proposes…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-21 Hojoon Ki , Jongsuk Kim , Minchan Kwon , Junmo Kim

Person search by natural language aims at retrieving a specific person in a large-scale image pool that matches the given textual descriptions. While most of the current methods treat the task as a holistic visual and textual feature…

Computer Vision and Pattern Recognition · Computer Science 2020-07-31 Zhe Wang , Zhiyuan Fang , Jun Wang , Yezhou Yang

Due to the excellent capacities of large language models (LLMs), it becomes feasible to develop LLM-based agents for reliable user simulation. Considering the scarcity and limit (e.g., privacy issues) of real user data, in this paper, we…

Information Retrieval · Computer Science 2024-02-28 Ruiyang Ren , Peng Qiu , Yingqi Qu , Jing Liu , Wayne Xin Zhao , Hua Wu , Ji-Rong Wen , Haifeng Wang

Video Multimethod Assessment Fusion (VMAF) [1], [2], [3] is a popular tool in the industry for measuring coded video quality. In this study, we propose an auditory-inspired frontend in existing VMAF for creating videos of reference and…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-08 Arijit Biswas , Harald Mundt

Audio-based multimedia retrieval tasks may identify semantic information in audio streams, i.e., audio concepts (such as music, laughter, or a revving engine). Conventional Gaussian-Mixture-Models have had some success in classifying a…

Audio and Speech Processing · Electrical Eng. & Systems 2017-10-13 Mirco Ravanelli , Benjamin Elizalde , Karl Ni , Gerald Friedland

Text-based person search is the task of finding person images that are the most relevant to the natural language text description given as query. The main challenge of this task is a large gap between the target images and text queries,…

Computer Vision and Pattern Recognition · Computer Science 2025-01-03 Jicheol Park , Boseung Jeong , Dongwon Kim , Suha Kwak

We introduce an approach to identifying speaker names in dialogue transcripts, a crucial task for enhancing content accessibility and searchability in digital media archives. Despite the advancements in speech recognition, the task of…

Computation and Language · Computer Science 2024-07-18 Minh Nguyen , Franck Dernoncourt , Seunghyun Yoon , Hanieh Deilamsalehy , Hao Tan , Ryan Rossi , Quan Hung Tran , Trung Bui , Thien Huu Nguyen

Speech is a natural interface for humans to interact with robots. Yet, aligning a robot's voice to its appearance is challenging due to the rich vocabulary of both modalities. Previous research has explored a few labels to describe robots…

Human-Computer Interaction · Computer Science 2024-02-09 Pol van Rijn , Silvan Mertes , Kathrin Janowski , Katharina Weitz , Nori Jacoby , Elisabeth André

Sample selection approaches are popular in robust learning from noisy labels. However, how to properly control the selection process so that deep networks can benefit from the memorization effect is a hard problem. In this paper, motivated…

Machine Learning · Computer Science 2020-09-21 Quanming Yao , Hansi Yang , Bo Han , Gang Niu , James Kwok

The proliferation of speech technologies and rising privacy legislation calls for the development of privacy preservation solutions for speech applications. These are essential since speech signals convey a wealth of rich, personal and…

Audio and Speech Processing · Electrical Eng. & Systems 2020-09-01 Paul-Gauthier Noé , Jean-François Bonastre , Driss Matrouf , Natalia Tomashenko , Andreas Nautsch , Nicholas Evans

Though recent technological advances have enabled note-taking through different modalities (e.g., keyboard, digital ink, voice), there is still a lack of understanding of the effect of the modality choice on learning. In this paper, we…

Human-Computer Interaction · Computer Science 2020-12-08 Anam Ahmad Khan , Sadia Nawaz , Joshua Newn , Jason M. Lodge , James Bailey , Eduardo Velloso

Large language models (LLMs) have shown great promise for capturing contextual information in natural language processing tasks. We propose a novel approach to speaker diarization that incorporates the prowess of LLMs to exploit contextual…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-15 Tae Jin Park , Kunal Dhawan , Nithin Koluguri , Jagadeesh Balam

Accurately describing images with text is a foundation of explainable AI. Vision-Language Models (VLMs) like CLIP have recently addressed this by aligning images and texts in a shared embedding space, expressing semantic similarities…

Computer Vision and Pattern Recognition · Computer Science 2024-12-20 Pingchuan Ma , Lennart Rietdorf , Dmytro Kotovenko , Vincent Tao Hu , Björn Ommer

Recognizing human non-speech vocalizations is an important task and has broad applications such as automatic sound transcription and health condition monitoring. However, existing datasets have a relatively small number of vocal sound…

Sound · Computer Science 2022-06-22 Yuan Gong , Jin Yu , James Glass

In-home IoT devices play a major role in healthcare systems as smart personal assistants. They usually come with a voice-enabled feature to add an extra level of usability and convenience to elderly, disabled people, and patients. In this…

Cryptography and Security · Computer Science 2018-09-13 Mohammad Hadian , Thamer Altuwaiyan , Xiaohui Liang , Wei Li

The Text Extraction of the Audio from the Video plays an important role in multimedia editing and processing. As a popular open source toolkit, Whisper performs fast in human voice recognition. However, the recognition performance is…

Sound · Computer Science 2024-07-16 Jinwei Lin

How does textual representation of audio relate to the Large Language Model's (LLMs) learning about the audio world? This research investigates the extent to which LLMs can be prompted to generate audio, despite their primary training in…

Environmental sounds like footsteps, keyboard typing, or dog barking carry rich information and emotional context, making them valuable for designing haptics in user applications. Existing audio-to-vibration methods, however, rely on…

Human-Computer Interaction · Computer Science 2026-01-27 Yinan Li , Hasti Seifi

Audio captioning aims at describing the content of audio clips with human language. Due to the ambiguity of audio, different people may perceive the same audio differently, resulting in caption disparities (i.e., one audio may correlate to…

Sound · Computer Science 2022-04-19 Yiming Zhang , Hong Yu , Ruoyi Du , Zhanyu Ma , Yuan Dong

To automatically test web applications, crawling-based techniques are usually adopted to mine the behavior models, explore the state spaces or detect the violated invariants of the applications. However, in existing crawlers, rules for…

Software Engineering · Computer Science 2016-08-24 Jun-Wei Lin , Farn Wang