English
Related papers

Related papers: A Synchronized Audio-Visual Multi-View Capture Sys…

200 papers

This paper addresses the gap in predicting turn-taking and backchannel actions in human-machine conversations using multi-modal signals (linguistic, acoustic, and visual). To overcome the limitation of existing datasets, we propose an…

Computation and Language · Computer Science 2025-05-21 Yuxin Lin , Yinglin Zheng , Ming Zeng , Wangzheng Shi

With advances in information acquisition technologies, multi-view data become ubiquitous. Multi-view learning has thus become more and more popular in machine learning and data mining fields. Multi-view unsupervised or semi-supervised…

Machine Learning · Computer Science 2018-04-04 Guoqing Chao , Shiliang Sun , Jinbo Bi

The dominant paradigm for video chat employs a single camera at each end of the conversation, but some conversations can be greatly enhanced by using multiple cameras at one or both ends. This paper provides the first rigorous investigation…

Multimedia · Computer Science 2012-09-10 John MacCormick

To date, little attention has been given to multi-view 3D human mesh estimation, despite real-life applicability (e.g., motion capture, sport analysis) and robustness to single-view ambiguities. Existing solutions typically suffer from poor…

Computer Vision and Pattern Recognition · Computer Science 2022-12-13 Xuan Gong , Liangchen Song , Meng Zheng , Benjamin Planche , Terrence Chen , Junsong Yuan , David Doermann , Ziyan Wu

We propose to synthesize high-quality and synchronized audio, given video and optional text conditions, using a novel multimodal joint training framework MMAudio. In contrast to single-modality training conditioned on (limited) video data…

Computer Vision and Pattern Recognition · Computer Science 2025-04-09 Ho Kei Cheng , Masato Ishii , Akio Hayakawa , Takashi Shibuya , Alexander Schwing , Yuki Mitsufuji

A major challenge in text-video and text-audio retrieval is the lack of large-scale training data. This is unlike image-captioning, where datasets are in the order of millions of samples. To close this gap we propose a new video mining…

Computer Vision and Pattern Recognition · Computer Science 2022-04-05 Arsha Nagrani , Paul Hongsuck Seo , Bryan Seybold , Anja Hauth , Santiago Manen , Chen Sun , Cordelia Schmid

While diffusion model for audio-driven avatar video generation have achieved notable process in synthesizing long sequences with natural audio-visual synchronization and identity consistency, the generation of music-performance videos with…

Computer Vision and Pattern Recognition · Computer Science 2025-12-03 Jiahui Chen , Weida Wang , Runhua Shi , Huan Yang , Chaofan Ding , Zihao Chen

Humans have the ability to utilize visual cues, such as lip movements and visual scenes, to enhance auditory perception, particularly in noisy environments. However, current Automatic Speech Recognition (ASR) or Audio-Visual Speech…

Computation and Language · Computer Science 2025-04-11 Lakshmipathi Balaji , Karan Singla

Fluorescence microscopy is widely employed for the analysis of living biological samples; however, the utility of the resulting recordings is frequently constrained by noise, temporal variability, and inconsistent visualisation of signals…

Computer Vision and Pattern Recognition · Computer Science 2026-01-16 Hassan Eshkiki , Sarah Costa , Mostafa Mohammadpour , Farinaz Tanhaei , Christopher H. George , Fabio Caraffini

Video dubbing aims to generate high-fidelity speech that is precisely temporally aligned with the visual content. Existing methods still suffer from limitations in speech naturalness and audio-visual synchronization, and are limited to…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-08 Kaidi Wang , Yi He , Wenhao Guan , Weijie Wu , Hongwu Ding , Xiong Zhang , Di Wu , Meng Meng , Jian Luan , Lin Li , Qingyang Hong

Automatic speech recognition can potentially benefit from the lip motion patterns, complementing acoustic speech to improve the overall recognition performance, particularly in noise. In this paper we propose an audio-visual fusion strategy…

Audio and Speech Processing · Electrical Eng. & Systems 2019-05-02 George Sterpu , Christian Saam , Naomi Harte

Most lip-to-speech (LTS) synthesis models are trained and evaluated under the assumption that the audio-video pairs in the dataset are perfectly synchronized. In this work, we show that the commonly used audio-visual datasets, such as GRID,…

Sound · Computer Science 2023-03-02 Zhe Niu , Brian Mak

Speech sounds convey a great deal of information about the scenes, resulting in a variety of effects ranging from reverberation to additional ambient sounds. In this paper, we manipulate input speech to sound as though it was recorded…

Computer Vision and Pattern Recognition · Computer Science 2024-09-24 Tingle Li , Renhao Wang , Po-Yao Huang , Andrew Owens , Gopala Anumanchipalli

Multi-view image acquisition systems with two or more cameras can be rather costly due to the number of high resolution image sensors that are required. Recently, it has been shown that by covering a low resolution sensor with a non-regular…

Image and Video Processing · Electrical Eng. & Systems 2022-04-11 Markus Jonscher , Jürgen Seiler , Thomas Richter , Michel Bätz , André Kaup

General audio understanding is a fundamental goal for large audio-language models, with audio captioning serving as a cornerstone task for their development. However, progress in this domain is hindered by existing datasets, which lack the…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-26 Yadong Niu , Tianzi Wang , Heinrich Dinkel , Xingwei Sun , Jiahao Zhou , Gang Li , Jizhong Liu , Junbo Zhang , Jian Luan

Multi-modal based speech separation has exhibited a specific advantage on isolating the target character in multi-talker noisy environments. Unfortunately, most of current separation strategies prefer a straightforward fusion based on…

Sound · Computer Science 2022-03-08 Junwen Xiong , Peng Zhang , Lei Xie , Wei Huang , Yufei Zha , Yanning Zhang

We propose a novel method for spatiotemporal multi-camera calibration using freely moving people in multiview videos. Since calibrating multiple cameras and finding matches across their views are inherently interdependent, performing both…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Sang-Eun Lee , Ko Nishino , Shohei Nobuhara

Humans use multiple senses to comprehend the environment. Vision and language are two of the most vital senses since they allow us to easily communicate our thoughts and perceive the world around us. There has been a lot of interest in…

Computation and Language · Computer Science 2026-05-13 Thong Nguyen , Yi Bin , Junbin Xiao , Leigang Qu , Yicong Li , Jay Zhangjie Wu , Cong-Duy Nguyen , See-Kiong Ng , Luu Anh Tuan

In this paper, we propose a quality-aware end-to-end audio-visual neural speaker diarization framework, which comprises three key techniques. First, our audio-visual model takes both audio and visual features as inputs, utilizing a series…

Multimedia · Computer Science 2024-10-31 Mao-Kui He , Jun Du , Shu-Tong Niu , Qing-Feng Liu , Chin-Hui Lee

This paper proposes a practical multimodal video summarization task setting and a dataset to train and evaluate the task. The target task involves summarizing a given video into a predefined number of keyframe-caption pairs and displaying…

Computation and Language · Computer Science 2023-12-05 Keito Kudo , Haruki Nagasawa , Jun Suzuki , Nobuyuki Shimizu
‹ Prev 1 8 9 10 Next ›