English
Related papers

Related papers: Rescribe: Authoring and Automatically Editing Audi…

200 papers

Our brains combine vision and hearing to create a more elaborate interpretation of the world. When the visual input is insufficient, a rich panoply of sounds can be used to describe our surroundings. Since more than 1,000 hours of videos…

Computer Vision and Pattern Recognition · Computer Science 2019-12-30 Rohan Mahadev , Hongyu Lu

In recent years, automatic speech recognition (ASR) systems have significantly improved, especially in languages with a vast amount of transcribed speech data. However, ASR systems tend to perform poorly for low-resource languages with…

Computation and Language · Computer Science 2024-06-04 Ara Yeroyan , Nikolay Karpov

Podcast episodes often contain material extraneous to the main content, such as advertisements, interleaved within the audio and the written descriptions. We present classifiers that leverage both textual and listening patterns in order to…

Computation and Language · Computer Science 2021-06-15 Sravana Reddy , Yongze Yu , Aasish Pappu , Aswin Sivaraman , Rezvaneh Rezapour , Rosie Jones

Video storytelling is engaging multimedia content that utilizes video and its accompanying narration to attract the audience, where a key challenge is creating narrations for recorded visual scenes. Previous studies on dense video…

Multimedia · Computer Science 2024-12-31 Dingyi Yang , Chunru Zhan , Ziheng Wang , Biao Wang , Tiezheng Ge , Bo Zheng , Qin Jin

Screen recordings of mobile applications are easy to capture and include a wealth of information, making them a popular mechanism for users to inform developers of the problems encountered in the bug reports. However, watching the bug…

Software Engineering · Computer Science 2023-02-03 Sidong Feng , Mulong Xie , Yinxing Xue , Chunyang Chen

In this paper, the problem of describing visual contents of a video sequence with natural language is addressed. Unlike previous video captioning work mainly exploiting the cues of video contents to make a language description, we propose a…

Computer Vision and Pattern Recognition · Computer Science 2019-06-05 Wei Zhang , Bairui Wang , Lin Ma , Wei Liu

Voice assistants provide users a new way of interacting with digital products, allowing them to retrieve information and complete tasks with an increased sense of control and flexibility. Such products are comprised of several machine…

Audio and Speech Processing · Electrical Eng. & Systems 2021-08-27 Shachaf Poran , Gil Amsalem , Amit Beka , Dmitri Goldenberg

Despite the impressive progress of multimodal generative models, video-to-audio generation still suffers from limited performance and limits the flexibility to prioritize sound synthesis for specific objects within the scene. Conversely,…

Computer Vision and Pattern Recognition · Computer Science 2024-12-30 Yujin Jeong , Yunji Kim , Sanghyuk Chun , Jiyoung Lee

Repository summarization is a crucial research question in development and maintenance for software engineering. Existing repository summarization techniques primarily focus on summarizing code according to the directory tree, which is…

Software Engineering · Computer Science 2025-10-14 Yifeng Zhu , Xianlin Zhao , Xutian Li , Yanzhen Zou , Haizhuo Yuan , Yue Wang , Bing Xie

Despite recent breakthroughs, audio foundation models struggle in processing complex multi-source acoustic scenes. We refer to this challenging domain as audio stories, which can have multiple speakers and background/foreground sound…

Many speech segments in movies are re-recorded in a studio during postproduction, to compensate for poor sound quality as recorded on location. Manual alignment of the newly-recorded speech with the original lip movements is a tedious task.…

Computer Vision and Pattern Recognition · Computer Science 2018-08-21 Tavi Halperin , Ariel Ephrat , Shmuel Peleg

Video summarization is a crucial research area that aims to efficiently browse and retrieve relevant information from the vast amount of video content available today. With the exponential growth of multimedia data, the ability to extract…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Hai-Dang Huynh-Lam , Ngoc-Phuong Ho-Thi , Minh-Triet Tran , Trung-Nghia Le

In video production, inserting B-roll is a widely used technique to enrich the story and make a video more engaging. However, determining the right content and positions of B-roll and actually inserting it within the main footage can be…

Human-Computer Interaction · Computer Science 2019-03-01 Bernd Huber , Hijung Valentina Shin , Bryan Russell , Oliver Wang , Gautham J. Mysore

We describe a system for large-scale audiovisual translation and dubbing, which translates videos from one language to another. The source language's speech content is transcribed to text, translated, and automatically synthesized into…

Computer Vision and Pattern Recognition · Computer Science 2020-11-09 Yi Yang , Brendan Shillingford , Yannis Assael , Miaosen Wang , Wendi Liu , Yutian Chen , Yu Zhang , Eren Sezener , Luis C. Cobo , Misha Denil , Yusuf Aytar , Nando de Freitas

Visual dubbing, the synchronization of facial movements with new speech, is crucial for making content accessible across different languages, enabling broader global reach. However, current methods face significant limitations. Existing…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Binyamin Manela , Sharon Gannot , Ethan Fetyaya

aTrain is an open-source and offline tool for transcribing audio data in multiple languages with CPU and NVIDIA GPU support. It is specifically designed for researchers using qualitative data generated from various forms of speech…

Sound · Computer Science 2023-10-19 Armin Haberl , Jürgen Fleiß , Dominik Kowald , Stefan Thalmann

We introduce an approach to identifying speaker names in dialogue transcripts, a crucial task for enhancing content accessibility and searchability in digital media archives. Despite the advancements in speech recognition, the task of…

Computation and Language · Computer Science 2024-07-18 Minh Nguyen , Franck Dernoncourt , Seunghyun Yoon , Hanieh Deilamsalehy , Hao Tan , Ryan Rossi , Quan Hung Tran , Trung Bui , Thien Huu Nguyen

Podcast summary, an important factor affecting end-users' listening decisions, has often been considered a critical feature in podcast recommendation systems, as well as many downstream applications. Existing abstractive summarization…

Computation and Language · Computer Science 2020-08-27 Chujie Zheng , Harry Jiannan Wang , Kunpeng Zhang , Ling Fan

With the exponential growth of video content, the need for automated video highlight detection to extract key moments or highlights from lengthy videos has become increasingly pressing. This technology has the potential to enhance user…

Computer Vision and Pattern Recognition · Computer Science 2025-05-16 Zahidul Islam , Sujoy Paul , Mrigank Rochan

Creating an animated data video enriched with audio narration takes a significant amount of time and effort and requires expertise. Users not only need to design complex animations, but also turn written text scripts into audio narrations…

Human-Computer Interaction · Computer Science 2024-06-10 Yun Wang , Leixian Shen , Zhengxin You , Xinhuan Shu , Bongshin Lee , John Thompson , Haidong Zhang , Dongmei Zhang