English
Related papers

Related papers: Dual Audio-Centric Modality Coupling for Talking H…

200 papers

Dynamic NeRFs have recently garnered growing attention for 3D talking portrait synthesis. Despite advances in rendering speed and visual quality, challenges persist in enhancing efficiency and effectiveness. We present R2-Talker, an…

Computer Vision and Pattern Recognition · Computer Science 2023-12-12 Zhiling Ye , LiangGuo Zhang , Dingheng Zeng , Quan Lu , Ning Jiang

In this paper, we propose a novel audio-driven talking head method capable of simultaneously generating highly expressive facial expressions and hand gestures. Unlike existing methods that focus on generating full-body or half-body poses,…

Computer Vision and Pattern Recognition · Computer Science 2025-01-22 Linrui Tian , Siqi Hu , Qi Wang , Bang Zhang , Liefeng Bo

Recently, more and more personalized speech enhancement systems (PSE) with excellent performance have been proposed. However, two critical issues still limit the performance and generalization ability of the model: 1) Acoustic environment…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-23 Xiaofeng Ge , Jiangyu Han , Haixin Guan , Yanhua Long

Automatic speaker naming is the problem of localizing as well as identifying each speaking character in a TV/movie/live show video. This is a challenging problem mainly attributes to its multimodal nature, namely face cue alone is…

Computer Vision and Pattern Recognition · Computer Science 2015-07-20 Yongtao Hu , Jimmy Ren , Jingwen Dai , Chang Yuan , Li Xu , Wenping Wang

Audio-visual information fusion enables a performance improvement in speech recognition performed in complex acoustic scenarios, e.g., noisy environments. It is required to explore an effective audio-visual fusion strategy for audiovisual…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-07 Liangfa Wei , Jie Zhang , Junfeng Hou , Lirong Dai

We propose EnCLAP, a novel framework for automated audio captioning. EnCLAP employs two acoustic representation models, EnCodec and CLAP, along with a pretrained language model, BART. We also introduce a new training objective called masked…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-01 Jaeyeon Kim , Jaeyoon Jung , Jinjoo Lee , Sang Hoon Woo

Animating high-fidelity video portrait with speech audio is crucial for virtual reality and digital entertainment. While most previous studies rely on accurate explicit structural information, recent works explore the implicit scene…

Computer Vision and Pattern Recognition · Computer Science 2022-02-11 Xian Liu , Yinghao Xu , Qianyi Wu , Hang Zhou , Wayne Wu , Bolei Zhou

Speech-driven facial animation is the process which uses speech signals to automatically synthesize a talking character. The majority of work in this domain creates a mapping from audio features to visual features. This often requires…

Audio and Speech Processing · Electrical Eng. & Systems 2018-07-20 Konstantinos Vougioukas , Stavros Petridis , Maja Pantic

Audio-driven 3D facial animation aims to generate synchronized lip movements and vivid facial expressions from arbitrary audio clips. While existing methods can produce synchronized lip motions, they often rely on predefined identity or…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Xuangeng Chu , Yuan Gan , Ziteng Cui , Shuhong Liu , Jian Wang , Bing Zhou , Tatsuya Harada

Large Language Models (LLMs) have recently been extended to the video domain, enabling sophisticated video-language understanding. However, existing Video LLMs often exhibit limitations in fine-grained temporal reasoning, restricting their…

Computer Vision and Pattern Recognition · Computer Science 2026-01-29 Bo-Cheng Chiu , Jen-Jee Chen , Yu-Chee Tseng , Feng-Chi Chen , An-Zi Yen

In recent years, the talking head generation has become a focal point for researchers. Considerable effort is being made to refine lip-sync motion, capture expressive facial expressions, generate natural head poses, and achieve high-quality…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Farzaneh Jafari , Stefano Berretti , Anup Basu

Automated audio captioning (AAC) aims to generate informative descriptions for various sounds from nature and/or human activities. In recent years, AAC has quickly attracted research interest, with state-of-the-art systems now relying on a…

Recent advances in video diffusion models have unlocked new potential for realistic audio-driven talking video generation. However, achieving seamless audio-lip synchronization, maintaining long-term identity consistency, and producing…

Computer Vision and Pattern Recognition · Computer Science 2024-12-06 Longtao Zheng , Yifan Zhang , Hanzhong Guo , Jiachun Pan , Zhenxiong Tan , Jiahao Lu , Chuanxin Tang , Bo An , Shuicheng Yan

We propose the Multi-Head Density Adaptive Attention Mechanism (DAAM), a novel probabilistic attention framework that can be used for Parameter-Efficient Fine-tuning (PEFT), and the Density Adaptive Transformer (DAT), designed to enhance…

Machine Learning · Computer Science 2024-10-01 Georgios Ioannides , Aman Chadha , Aaron Elkins

Current region feature-based image captioning methods have progressed rapidly and achieved remarkable performance. However, they are still prone to generating irrelevant descriptions due to the lack of contextual information and the…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Jun Wan , Jun Liu , Zhihui lai , Jie Zhou

Significant progress has been made in talking-face video generation research; however, precise lip-audio synchronization and high visual quality remain challenging in editing lip shapes based on input audio. This paper introduces JoyGen, a…

Computer Vision and Pattern Recognition · Computer Science 2025-01-06 Qili Wang , Dajiang Wu , Zihang Xu , Junshi Huang , Jun Lv

By integrating recent advances in large language models (LLMs) and generative models into the emerging semantic communication (SC) paradigm, in this article we put forward to a novel framework of language-oriented semantic communication…

Signal Processing · Electrical Eng. & Systems 2023-09-21 Hyelin Nam , Jihong Park , Jinho Choi , Mehdi Bennis , Seong-Lyun Kim

Referring video segmentation aims to segment the corresponding video object described by the language expression. To address this task, we first design a two-stream encoder to extract CNN-based visual features and transformer-based…

Computer Vision and Pattern Recognition · Computer Science 2022-03-31 Guang Feng , Lihe Zhang , Zhiwei Hu , Huchuan Lu

Large Audio Language Models struggle to disentangle overlapping events in complex acoustic scenes, yielding temporally inconsistent captions and frequent hallucinations. We introduce Timestamped Audio Captioner (TAC), a model that produces…

Audio-driven portrait animation aims to synthesize portrait videos that are conditioned by given audio. Animating high-fidelity and multimodal video portraits has a variety of applications. Previous methods have attempted to capture…

Computer Vision and Pattern Recognition · Computer Science 2023-07-20 Yunfei Liu , Lijian Lin , Fei Yu , Changyin Zhou , Yu Li