中文
相关论文

相关论文: Kefa: A Knowledge Enhanced and Fine-grained Aligne…

200 篇论文

Emotionally talking head video generation aims to generate expressive portrait videos with accurate lip synchronization and emotional facial expressions. Current methods rely on simple emotional labels, leading to insufficient semantic…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Yahui Li , Yinfeng Yu , Liejun Wang , Shengjie Shen

Knowledge editing aims to modify outdated knowledge in language models efficiently while retaining their original capabilities. Mainstream datasets for knowledge editing are predominantly static and fail to keep in pace with the evolving…

计算与语言 · 计算机科学 2026-04-24 Chenming Tang , Yutong Yang , Kexue Wang , Yunfang Wu

Existing deep learning based speech enhancement mainly employ a data-driven approach, which leverage large amounts of data with a variety of noise types to achieve noise removal from noisy signal. However, the high dependence on the data…

声音 · 计算机科学 2024-01-24 Huaying Xue , Xiulian Peng , Yan Lu

Modern machine learning methods are critical to the development of large-scale personalized learning systems that cater directly to the needs of individual learners. The recently developed SPARse Factor Analysis (SPARFA) framework provides…

机器学习 · 统计学 2013-05-13 Andrew S. Lan , Christoph Studer , Andrew E. Waters , Richard G. Baraniuk

Speech enhancement models should meet very low latency requirements typically smaller than 5 ms for hearing assistive devices. While various low-latency techniques have been proposed, comparing these methods in a controlled setup using DNNs…

音频与语音处理 · 电气工程与系统科学 2024-09-17 Haibin Wu , Sebastian Braun

Audio-visual Navigation refers to an agent utilizing visual and auditory information in complex 3D environments to accomplish target localization and path planning, thereby achieving autonomous navigation. The core challenge of this task…

声音 · 计算机科学 2026-04-06 Xinyu Zhou , Yinfeng Yu

Referring expressions are natural language constructions used to identify particular objects within a scene. In this paper, we propose a unified framework for the tasks of referring expression comprehension and generation. Our model is…

计算机视觉与模式识别 · 计算机科学 2017-04-19 Licheng Yu , Hao Tan , Mohit Bansal , Tamara L. Berg

Traditional autonomous driving pipelines decouple camera design from downstream perception, relying on fixed optics and handcrafted ISPs that prioritize human viewable imagery rather than machine semantics. This separation discards…

计算机视觉与模式识别 · 计算机科学 2025-12-29 Reeshad Khan , John Gauch

A primary challenge when deploying speaker recognition systems in real-world applications is performance degradation caused by environmental mismatch. We propose a diffusion-based method that takes speaker embeddings extracted from a…

音频与语音处理 · 电气工程与系统科学 2025-05-23 KiHyun Nam , Jungwoo Heo , Jee-weon Jung , Gangin Park , Chaeyoung Jung , Ha-Jin Yu , Joon Son Chung

Existing methods on audio-visual deepfake detection mainly focus on high-level features for modeling inconsistencies between audio and visual data. As a result, these approaches usually overlook finer audio-visual artifacts, which are…

计算机视觉与模式识别 · 计算机科学 2024-10-15 Marcella Astrid , Enjie Ghorbel , Djamila Aouada

Audio-driven human animation technology is widely used in human-computer interaction, and the emergence of diffusion models has further advanced its development. Currently, most methods rely on multi-stage generation and intermediate…

计算机视觉与模式识别 · 计算机科学 2025-06-13 S. Z. Zhou , Y. B. Wang , J. F. Wu , T. Hu , J. N. Zhang

Existing 3D visual grounding methods rely on precise text prompts to locate objects within 3D scenes. Speech, as a natural and intuitive modality, offers a promising alternative. Real-world speech inputs, however, often suffer from…

计算机视觉与模式识别 · 计算机科学 2025-06-18 Yu Qi , Lipeng Gu , Honghua Chen , Liangliang Nan , Mingqiang Wei

With the proliferation of edge computing, efficient AI inference on edge devices has become essential for intelligent applications such as autonomous vehicles and VR/AR. In this context, we address the problem of efficient remote object…

信息论 · 计算机科学 2023-12-01 Xiangyu Gao , Yaping Sun , Dongyu Wei , Xiaodong Xu , Hao Chen , Hao Yin , Shuguang Cui

This paper introduces the retrieval-augmented large language model with Definite Finite Automaton (DFA-RAG), a novel framework designed to enhance the capabilities of conversational agents using large language models (LLMs). Traditional…

计算与语言 · 计算机科学 2024-06-04 Yiyou Sun , Junjie Hu , Wei Cheng , Haifeng Chen

In realistic speech enhancement settings for end-user devices, we often encounter only a few speakers and noise types that tend to reoccur in the specific acoustic environment. We propose a novel personalized speech enhancement method to…

音频与语音处理 · 电气工程与系统科学 2021-05-11 Sunwoo Kim , Minje Kim

Learning deep neural networks that are generalizable across different domains remains a challenge due to the problem of domain shift. Unsupervised domain adaptation is a promising avenue which transfers knowledge from a source domain to a…

机器学习 · 计算机科学 2020-08-20 Qingjie Meng , Daniel Rueckert , Bernhard Kainz

In this paper, we present a novel framework that jointly performs three tasks: speaker diarization, speech separation, and speaker counting. Our proposed framework integrates speaker diarization based on end-to-end neural diarization (EEND)…

音频与语音处理 · 电气工程与系统科学 2022-12-19 Soumi Maiti , Yushi Ueda , Shinji Watanabe , Chunlei Zhang , Meng Yu , Shi-Xiong Zhang , Yong Xu

This paper introduces a novel Self-supervised Fine-grained Dialogue Evaluation framework (SelF-Eval). The core idea is to model the correlation between turn quality and the entire dialogue quality. We first propose a novel automatic data…

计算与语言 · 计算机科学 2022-09-19 Longxuan Ma , Ziyu Zhuang , Weinan Zhang , Mingda Li , Ting Liu

While recent advances in machine learning have equipped Weather Foundation Models (WFMs) with substantial generalization capabilities across diverse downstream tasks, the escalating computational requirements associated with their expanding…

State-of-the-art speaker diarization systems utilize knowledge from external data, in the form of a pre-trained distance metric, to effectively determine relative speaker identities to unseen data. However, much of recent focus has been on…