中文
相关论文

相关论文: PAEFF: Precise Alignment and Enhanced Gated Featur…

200 篇论文

Text-to-video retrieval requires precise alignment between language and temporally rich audio-video signals. However, existing methods often emphasize visual cues while underutilizing audio semantics or relying on coarse fusion strategies,…

计算机视觉与模式识别 · 计算机科学 2025-11-19 Bowen Yang , Yun Cao , Chen He , Xiaosu Su

The advancements of technology have led to the use of multimodal systems in various real-world applications. Among them, the audio-visual systems are one of the widely used multimodal systems. In the recent years, associating face and voice…

Recent advances in deep learning and automatic speech recognition (ASR) have enabled the end-to-end (E2E) ASR system and boosted the accuracy to a new level. The E2E systems implicitly model all conventional ASR components, such as the…

Existing facial editing methods have achieved remarkable results, yet they often fall short in supporting multimodal conditional local facial editing. One of the significant evidences is that their output image quality degrades dramatically…

计算机视觉与模式识别 · 计算机科学 2024-12-30 Wanglong Lu , Jikai Wang , Xiaogang Jin , Xianta Jiang , Hanli Zhao

Forced alignment refers to a technology that time-aligns a given transcription with a corresponding speech. However, as the forced alignment technologies have developed using speech audio, they might fail in alignment when the input speech…

计算机视觉与模式识别 · 计算机科学 2023-03-16 Minsu Kim , Chae Won Kim , Yong Man Ro

Recent advancements in multimodal fusion have witnessed the remarkable success of vision-language (VL) models, which excel in various multimodal applications such as image captioning and visual question answering. However, building VL…

计算机视觉与模式识别 · 计算机科学 2024-10-24 Zhiwei Hao , Jianyuan Guo , Li Shen , Yong Luo , Han Hu , Yonggang Wen

In recent years, with the advent of deep-learning, face recognition has achieved exceptional success. However, many of these deep face recognition models perform much better in handling frontal faces compared to profile faces. The major…

计算机视觉与模式识别 · 计算机科学 2021-07-30 Fariborz Taherkhani , Veeru Talreja , Jeremy Dawson , Matthew C. Valenti , Nasser M. Nasrabadi

Face recognition technology has been deployed in various real-life applications. The most sophisticated deep learning-based face recognition systems rely on training millions of face images through complex deep neural networks to achieve…

计算机视觉与模式识别 · 计算机科学 2024-01-25 Dong Han , Yong Li , Joachim Denzler

DeepFake Audio, unlike DeepFake images and videos, has been relatively less explored from detection perspective, and the solutions which exist for the synthetic speech classification either use complex networks or dont generalize to…

声音 · 计算机科学 2022-10-24 Vardhan Dongre , Abhinav Thimma Reddy , Nikhitha Reddeddy

When we hear the word "house", we don't just process sound, we imagine walls, doors, memories. The brain builds meaning through layers, moving from raw acoustics to rich, multimodal associations. Inspired by this, we build on recent work…

机器学习 · 计算机科学 2025-11-11 Kateryna Shapovalenko , Quentin Auster

Combining face swapping with lip synchronization technology offers a cost-effective solution for customized talking face generation. However, directly cascading existing models together tends to introduce significant interference between…

计算机视觉与模式识别 · 计算机科学 2024-05-10 Zeren Zhang , Haibo Qin , Jiayu Huang , Yixin Li , Hui Lin , Yitao Duan , Jinwen Ma

Deep metric learning aims to learn an embedding function, modeled as deep neural network. This embedding function usually puts semantically similar images close while dissimilar images far from each other in the learned embedding space.…

计算机视觉与模式识别 · 计算机科学 2018-09-03 Wonsik Kim , Bhavya Goyal , Kunal Chawla , Jungmin Lee , Keunjoo Kwon

Brain computer interface (BCI) is the only way for some special patients to communicate with the outside world and provide a direct control channel between brain and the external devices. As a non-invasive interface, the scalp…

定量方法 · 定量生物学 2018-08-15 Chuanqi Tan , Fuchun Sun , Wenchang Zhang , Shaobo Liu , Chunfang Liu

Face image synthesis is gaining more attention in computer security due to concerns about its potential negative impacts, including those related to fake biometrics. Hence, building models that can detect the synthesized face images is an…

计算机视觉与模式识别 · 计算机科学 2024-01-10 Roberto Leyva , Victor Sanchez , Gregory Epiphaniou , Carsten Maple

Speech-driven 3D facial animation has garnered lots of attention thanks to its broad range of applications. Despite recent advancements in achieving realistic lip motion, current methods fail to capture the nuanced emotional undertones…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Jisoo Kim , Jungbin Cho , Joonho Park , Soonmin Hwang , Da Eun Kim , Geon Kim , Youngjae Yu

In this paper, we present a deep learning based image feature extraction method designed specifically for face images. To train the feature extraction model, we construct a large scale photo-realistic face image dataset with ground-truth…

计算机视觉与模式识别 · 计算机科学 2018-03-13 Boyi Jiang , Juyong Zhang , Bailin Deng , Yudong Guo , Ligang Liu

Face recognition has already been well studied under the visible light and the infrared,in both intra-spectral and cross-spectral cases. However, how to fuse different light bands, i.e., hyperspectral face recognition, is still an open…

计算机视觉与模式识别 · 计算机科学 2020-09-15 Zhicheng Cao , Xi Cen , Liaojun Pang

The rapid evolution of generative AI has increased the threat of realistic audio-visual deepfakes, demanding robust detection methods. Existing solutions primarily address unimodal (audio or visual) forgeries but struggle with multimodal…

计算机视觉与模式识别 · 计算机科学 2025-06-17 Jian Wang , Baoyuan Wu , Li Liu , Qingshan Liu

Acoustic word embeddings (AWEs) are discriminative representations of speech segments, and learned embedding space reflects the phonetic similarity between words. With multi-view learning, where text labels are considered as supplementary…

音频与语音处理 · 电气工程与系统科学 2022-06-28 Myunghun Jung , Hoirin Kim

Recent advances in diffusion-based lip-syncing generative models have demonstrated their ability to produce highly synchronized talking face videos for visual dubbing. Although these models excel at lip synchronization, they often struggle…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Yanyu Zhu , Lichen Bai , Jintao Xu , Hai-tao Zheng