English
Related papers

Related papers: MAVFlow: Preserving Paralinguistic Elements with C…

200 papers

Facial expression recognition is an essential task for various applications, including emotion detection, mental health analysis, and human-machine interactions. In this paper, we propose a multi-modal facial expression recognition method…

Computer Vision and Pattern Recognition · Computer Science 2023-03-21 Jun-Hwa Kim , Namho Kim , Chee Sun Won

Vision-guided speech generation aims to produce authentic speech from facial appearance or lip motions without relying on auditory signals, offering significant potential for applications such as dubbing in filmmaking and assisting…

Computer Vision and Pattern Recognition · Computer Science 2025-03-20 Jiaxin Ye , Hongming Shan

Natural language understanding inherently depends on integrating multiple complementary perspectives spanning from surface syntax to deep semantics and world knowledge. However, current Aspect-Based Sentiment Analysis (ABSA) systems…

Computation and Language · Computer Science 2026-03-20 Smitha Muthya Sudheendra , Mani Deep Cherukuri , Jaideep Srivastava

Audio-visual speech recognition (AVSR) combines audio-visual modalities to improve speech recognition, especially in noisy environments. However, most existing methods deploy the unidirectional enhancement or symmetric fusion manner, which…

Multimedia · Computer Science 2025-08-12 Junxiao Xue , Xiaozhen Liu , Xuecheng Wu , Xinyi Yin , Danlei Huang , Fei Yu

Existing works have made strides in video generation, but the lack of sound effects (SFX) and background music (BGM) hinders a complete and immersive viewer experience. We introduce a novel semantically consistent v ideo-to-audio generation…

Multimedia · Computer Science 2024-04-29 Gehui Chen , Guan'an Wang , Xiaowen Huang , Jitao Sang

Flow matching (FM) trains a time-dependent vector field that transports samples from a simple prior to a complex data distribution. However, for high-dimensional images, each training sample supervises only a single trajectory and…

Computer Vision and Pattern Recognition · Computer Science 2026-05-04 George Stoica , Sayak Paul , Matthew Wallingford , Vivek Ramanujan , Abhay Nori , Winson Han , Ali Farhadi , Ranjay Krishna , Judy Hoffman

Image-to-image translation and voice conversion enable the generation of a new facial image and voice while maintaining some of the semantics such as a pose in an image and linguistic content in audio, respectively. They can aid in the…

Computer Vision and Pattern Recognition · Computer Science 2023-03-02 Naoya Takahashi , Mayank K. Singh , Yuki Mitsufuji

Direct speech-to-speech translation (S2ST) aims to convert speech from one language into another, and has demonstrated significant progress to date. Despite the recent success, current S2ST models still suffer from distinct degradation in…

Computation and Language · Computer Science 2023-05-25 Rongjie Huang , Huadai Liu , Xize Cheng , Yi Ren , Linjun Li , Zhenhui Ye , Jinzheng He , Lichao Zhang , Jinglin Liu , Xiang Yin , Zhou Zhao

Generating high-quality time-series data is challenging because real-world signals often exhibit multimodal patterns and multiscale dynamics, including oscillations and high-frequency variations. Flow Matching (FM) offers an efficient…

Machine Learning · Computer Science 2026-05-29 Junru Zhang , Lang Feng , Jinbo Wang , Xu Guo , Yucheng Wang , Han Yu , Min Wu , Yabo Dong , Duanqing Xu

Training supervised video captioning model requires coupled video-caption pairs. However, for many targeted languages, sufficient paired data are not available. To this end, we introduce the unpaired video captioning task aiming to train…

Computer Vision and Pattern Recognition · Computer Science 2022-11-23 Fenglin Liu , Xian Wu , Chenyu You , Shen Ge , Yuexian Zou , Xu Sun

Pre-trained Vision Foundation Models (VFMs) provide strong visual representations for a wide range of applications. In this paper, we continually pre-train prevailing VFMs in a multimodal manner such that they can effortlessly process…

Computer Vision and Pattern Recognition · Computer Science 2025-05-19 Yitong Chen , Lingchen Meng , Wujian Peng , Zuxuan Wu , Yu-Gang Jiang

Vision-language models offer strong few-shot capability through prompt tuning but remain vulnerable to noisy labels, which can corrupt prompts and degrade cross-modal alignment. Existing approaches struggle because they often lack the…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Lu Niu , Cheng Xue

Self-supervised feed-forward methods for scene flow estimation offer real-time efficiency, but their supervision from two-frame point correspondences is unreliable and often breaks down under occlusions. Multi-frame supervision has the…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Qingwen Zhang , Chenhan Jiang , Xiaomeng Zhu , Yunqi Miao , Yushan Zhang , Olov Andersson , Patric Jensfelt

Modern text-to-image diffusion models encode rich visual priors, but expose them only through one-way text-conditioned generation. Existing unified vision--language models derived from them recover bidirectional capability through…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Eric Tillmann Bill , Enis Simsar , Alessio Tonioni , Thomas Hofmann

Cross-lingual cross-modal retrieval has garnered increasing attention recently, which aims to achieve the alignment between vision and target language (V-T) without using any annotated V-T data pairs. Current methods employ machine…

Computer Vision and Pattern Recognition · Computer Science 2024-02-02 Yabing Wang , Fan Wang , Jianfeng Dong , Hao Luo

Zero-shot scene understanding in real-world settings presents major challenges due to the complexity and variability of natural scenes, where models must recognize new objects, actions, and contexts without prior labeled examples. This work…

Computer Vision and Pattern Recognition · Computer Science 2025-10-30 Manjunath Prasad Holenarasipura Rajiv , B. M. Vidyavathi

Humans have the ability to utilize visual cues, such as lip movements and visual scenes, to enhance auditory perception, particularly in noisy environments. However, current Automatic Speech Recognition (ASR) or Audio-Visual Speech…

Computation and Language · Computer Science 2025-04-11 Lakshmipathi Balaji , Karan Singla

Language-aligned vision foundation models (VFMs) enable versatile visual understanding for always-on contextual AI, but their deployment on edge devices is hindered by strict latency and power constraints. We present AdaVFM, an adaptive…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Yiwei Zhao , Yi Zheng , Huapeng Su , Jieyu Lin , Stefano Ambrogio , Cijo Jose , Michael Ramamonjisoa , Patrick Labatut , Barbara De Salvo , Chiao Liu , Phillip B. Gibbons , Ziyun Li

Emotion Recognition in Conversation (ERC) is essential for effective human-machine interaction, aiming to identify speakers' emotional states in multi-turn dialogues. Early text-based methods struggle with complex scenarios like sarcasm…

Artificial Intelligence · Computer Science 2026-05-19 Linan ZHU , Zihao Zhai , Xiao Han , Yuqian Fu , Xiangfan Chen , Xiangjie Kong , Guojiang Shen

As Large Language Models (LLMs) expand beyond text, integrating speech as a native modality has given rise to SpeechLLMs, which directly process spoken language and enable speech-to-text translation (ST) and other downstream tasks,…