中文
相关论文

相关论文: Universal Activation Verbalizer: A Unified Framewo…

200 篇论文

Speech-driven 3D facial animation is important for many multimedia applications. Recent work has shown promise in using either Diffusion models or Transformer architectures for this task. However, their mere aggregation does not lead to…

计算机视觉与模式识别 · 计算机科学 2024-02-09 Zhiyuan Ma , Xiangyu Zhu , Guojun Qi , Chen Qian , Zhaoxiang Zhang , Zhen Lei

Audio-visual active speaker detection (AV-ASD) aims to identify which visible face is speaking in a scene with one or more persons. Most existing AV-ASD methods prioritize capturing speech-lip correspondence. However, there is a noticeable…

音频与语音处理 · 电气工程与系统科学 2024-04-02 Ruijie Tao , Xinyuan Qian , Rohan Kumar Das , Xiaoxue Gao , Jiadong Wang , Haizhou Li

Recent advances in image generation have made diffusion models powerful tools for creating high-quality images. However, their iterative denoising process makes understanding and interpreting their semantic latent spaces more challenging…

计算与语言 · 计算机科学 2024-11-06 E. Zhixuan Zeng , Yuhao Chen , Alexander Wong

Recently, deep learning becomes the main focus of machine learning research and has greatly impacted many important fields. However, deep learning is criticized for lack of interpretability. As a successful unsupervised model in deep…

机器学习 · 计算机科学 2021-01-06 Fenglei Fan , Mengzhou Li , Yueyang Teng , Ge Wang

We propose UniT, a Unified Transformer model to simultaneously learn the most prominent tasks across different domains, ranging from object detection to natural language understanding and multimodal reasoning. Based on the transformer…

计算机视觉与模式识别 · 计算机科学 2021-08-19 Ronghang Hu , Amanpreet Singh

When humans perceive the world, they naturally integrate multiple audio-visual tasks within dynamic, real-world scenes. However, current works such as event localization, parsing, segmentation and question answering are mostly explored…

计算机视觉与模式识别 · 计算机科学 2026-03-09 Guangyao Li , Xin Wang , Wenwu Zhu

Explanations have gained an increasing level of interest in the AI and Machine Learning (ML) communities in order to improve model transparency and allow users to form a mental model of a trained ML model. However, explanations can go…

机器学习 · 计算机科学 2022-10-11 Stefano Teso , Öznur Alkan , Wolfang Stammer , Elizabeth Daly

Self-supervised learning methods such as wav2vec 2.0 have shown promising results in learning speech representations from unlabelled and untranscribed speech data that are useful for speech recognition. Since these representations are…

音频与语音处理 · 电气工程与系统科学 2022-03-22 Shehzeen Hussain , Van Nguyen , Shuhua Zhang , Erik Visser

With recent advances in multi-modal foundation models, the previously text-only large language models (LLM) have evolved to incorporate visual input, opening up unprecedented opportunities for various applications in visualization. Our work…

人机交互 · 计算机科学 2023-12-08 Shusen Liu , Haichao Miao , Zhimin Li , Matthew Olson , Valerio Pascucci , Peer-Timo Bremer

Due to the lack of effective cross-modal modeling, existing open-source audio-video generation methods often exhibit compromised lip synchronization and insufficient semantic consistency. To mitigate these drawbacks, we propose UniAVGen, a…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Guozhen Zhang , Zixiang Zhou , Teng Hu , Ziqiao Peng , Youliang Zhang , Yi Chen , Yuan Zhou , Qinglin Lu , Limin Wang

Explaining the decisions made by audio spoofing detection models is crucial for fostering trust in detection outcomes. However, current research on the interpretability of detection models is limited to applying XAI tools to post-trained…

声音 · 计算机科学 2025-07-28 Menglu Li , Xiao-Ping Zhang

The development of many vision models mainly focuses on improving their performance using metrics such as accuracy, IoU, and mAP, with less attention to explainability due to the complexity of applying xAI methods to provide a meaningful…

计算机视觉与模式识别 · 计算机科学 2025-08-29 Phu-Vinh Nguyen , Tan-Hanh Pham , Chris Ngo , Truong Son Hy

Scaling humanoid foundation models is bottlenecked by the scarcity of robotic data. While massive egocentric human data offers a scalable alternative, bridging the cross-embodiment chasm remains a fundamental challenge due to kinematic…

机器人学 · 计算机科学 2026-04-22 Boyu Chen , Yi Chen , Lu Qiu , Jerry Bai , Yuying Ge , Yixiao Ge

We present a universal framework to model contextualized sentence representations with visual awareness that is motivated to overcome the shortcomings of the multimodal parallel data with manual annotations. For each sentence, we first…

计算与语言 · 计算机科学 2019-11-12 Zhuosheng Zhang , Rui Wang , Kehai Chen , Masao Utiyama , Eiichiro Sumita , Hai Zhao

In this study, our goal is to create interactive avatar agents that can autonomously plan and animate nuanced facial movements realistically, from both visual and behavioral perspectives. Given high-level inputs about the environment and…

计算机视觉与模式识别 · 计算机科学 2023-12-05 Duomin Wang , Bin Dai , Yu Deng , Baoyuan Wang

Canonical emotions, such as happy, sad, and fearful, are easy to understand and annotate. However, emotions are often compound, e.g. happily surprised, and can be mapped to the action units (AUs) used for expressing emotions, and trivially…

计算机视觉与模式识别 · 计算机科学 2024-04-02 Reni Paskaleva , Mykyta Holubakha , Andela Ilic , Saman Motamed , Luc Van Gool , Danda Paudel

Active speaker detection and speech enhancement have become two increasingly attractive topics in audio-visual scenario understanding. According to their respective characteristics, the scheme of independently designed architecture has been…

声音 · 计算机科学 2022-07-08 Junwen Xiong , Yu Zhou , Peng Zhang , Lei Xie , Wei Huang , Yufei Zha

Unified multimodal models can encode visual understanding and image generation within a shared backbone, yet understanding does not automatically translate into control: models may infer objects, relations, or knowledge cues but fail to…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Fuxiang Zhai , Sixiang Chen , Yingjin Li , Shuaibo Li , Jianyu Lai , Tengjun Huang , Lei Zhu

Video-driven 3D facial animation transfer aims to drive avatars to reproduce the expressions of actors. Existing methods have achieved remarkable results by constraining both geometric and perceptual consistency. However, geometric…

图形学 · 计算机科学 2024-10-10 Feng Qiu , Wei Zhang , Chen Liu , Rudong An , Lincheng Li , Yu Ding , Changjie Fan , Zhipeng Hu , Xin Yu

Attributing the behavior of Transformer models to internal computations is a central challenge in mechanistic interpretability. We introduce DePass, a unified framework for feature attribution based on a single decomposed forward pass.…

计算与语言 · 计算机科学 2025-10-27 Xiangyu Hong , Che Jiang , Kai Tian , Biqing Qi , Youbang Sun , Ning Ding , Bowen Zhou