中文
相关论文

相关论文: Weakly-supervised Audio Temporal Forgery Localizat…

200 篇论文

Recent multi-modal audio-language models (ALMs) excel at text-audio retrieval but struggle with frame-wise audio understanding. Prior works use temporal-aware labels or unsupervised training to improve frame-wise capabilities, but they…

Voice spoofing attacks pose a significant threat to automated speaker verification systems. Existing anti-spoofing methods often simulate specific attack types, such as synthetic or replay attacks. However, in real-world scenarios, the…

声音 · 计算机科学 2023-09-19 Awais Khan , Khalid Mahmood Malik , Shah Nawaz

Artefacts that serve to distinguish bona fide speech from spoofed or deepfake speech are known to reside in specific subbands and temporal segments. Various approaches can be used to capture and model such artefacts, however, none works…

音频与语音处理 · 电气工程与系统科学 2021-08-24 Hemlata Tak , Jee-weon Jung , Jose Patino , Madhu Kamble , Massimiliano Todisco , Nicholas Evans

Conventional federated learning (FL) heavily depends on high-quality labels, which are often impractical in the real world, leading to the federated label-noise (F-LN) problem. Worse still, the F-LN problem is exacerbated by the…

机器学习 · 计算机科学 2026-05-29 Yuxin Tian , Mouxing Yang , Yuhao Zhou , Jian Wang , Qing Ye , Tongliang Liu , Gang Niu , Jiancheng Lv

Given a text description, Temporal Language Grounding (TLG) aims to localize temporal boundaries of the segments that contain the specified semantics in an untrimmed video. TLG is inherently a challenging task, as it requires comprehensive…

计算机视觉与模式识别 · 计算机科学 2021-12-07 Fan Luo , Shaoxiang Chen , Jingjing Chen , Zuxuan Wu , Yu-Gang Jiang

Detecting maliciously falsified facial images and videos has attracted extensive attention from digital-forensics and computer-vision communities. An important topic in manipulation detection is the localization of the fake regions.…

计算机视觉与模式识别 · 计算机科学 2023-04-18 Weinan Guan , Wei Wang , Jing Dong , Bo Peng , Tieniu Tan

Recent anti-spoofing systems focus on spoofing detection, where the task is only to determine whether the test audio is fake. However, there are few studies putting attention to identifying the methods of generating fake speech. Common…

声音 · 计算机科学 2022-12-19 Tinglong Zhu , Xingming Wang , Xiaoyi Qin , Ming Li

Face forgery detection is raising ever-increasing interest in computer vision since facial manipulation technologies cause serious worries. Though recent works have reached sound achievements, there are still unignorable problems: a)…

计算机视觉与模式识别 · 计算机科学 2021-03-17 Jiaming Li , Hongtao Xie , Jiahong Li , Zhongyuan Wang , Yongdong Zhang

The rapid advancement of generative technologies has made synthetic images nearly indistinguishable from real ones, thereby creating an urgent need for robust detectors to counter misinformation. However, existing methods mainly rely on…

计算机视觉与模式识别 · 计算机科学 2026-05-07 Yutong Xiao , Ran Ran , Jiwei Wei , Shuchang Zhou , Ke Liu , Zheng Ziqiang , Caiyan Qin

When labeled data is insufficient, semi-supervised learning with the pseudo-labeling technique can significantly improve the performance of automatic speech recognition. However, pseudo-labels are often noisy, containing numerous incorrect…

音频与语音处理 · 电气工程与系统科学 2023-08-15 Han Zhu , Dongji Gao , Gaofeng Cheng , Daniel Povey , Pengyuan Zhang , Yonghong Yan

Quantum Federated Learning (QFL) is an emerging paradigm that combines quantum computing and federated learning (FL) to enable decentralized model training while maintaining data privacy over quantum networks. However, quantum noise remains…

量子物理 · 物理学 2025-07-18 Ratun Rahman , Atit Pokharel , Dinh C. Nguyen

Weakly supervised temporal action localization is a challenging task as only the video-level annotation is available during the training process. To address this problem, we propose a two-stage approach to fully exploit multi-resolution…

计算机视觉与模式识别 · 计算机科学 2025-06-24 Rui Su , Dong Xu , Luping Zhou , Wanli Ouyang

With the continuous improvement in the computational capabilities of edge devices such as intelligent sensors in the Industrial Internet of Things, these sensors are no longer limited to mere data collection but are increasingly capable of…

机器学习 · 计算机科学 2025-01-06 Heqiang Wang , Xiaoxiong Zhong , Kang Liu , Fangming Liu , Weizhe Zhang

Recently, Visual Foundation Models (VFMs) have shown a remarkable generalization performance in 3D perception tasks. However, their effectiveness in large-scale outdoor datasets remains constrained by the scarcity of accurate supervision…

计算机视觉与模式识别 · 计算机科学 2024-12-25 Pufan Zou , Shijia Zhao , Weijie Huang , Qiming Xia , Chenglu Wen , Wei Li , Cheng Wang

Recent advances in face forgery techniques produce nearly visually untraceable deepfake videos, which could be leveraged with malicious intentions. As a result, researchers have been devoted to deepfake detection. Previous studies have…

计算机视觉与模式识别 · 计算机科学 2022-10-13 Jiazhi Guan , Hang Zhou , Zhibin Hong , Errui Ding , Jingdong Wang , Chengbin Quan , Youjian Zhao

Large Audio-Language Models (LALMs) enable general audio understanding and demonstrate remarkable performance across various audio tasks. However, these models still face challenges in temporal perception (e.g., inferring event onset and…

声音 · 计算机科学 2026-04-16 Yanfeng Shi , Pengfei Cai , Jun Liu , Qing Gu , Nan Jiang , Lirong Dai , Ian McLoughlin , Yan Song

Voice anti-spoofing systems are crucial auxiliaries for automatic speaker verification (ASV) systems. A major challenge is caused by unseen attacks empowered by advanced speech synthesis technologies. Our previous research on one-class…

音频与语音处理 · 电气工程与系统科学 2022-11-08 Siwen Ding , You Zhang , Zhiyao Duan

The goal of weakly-supervised video moment retrieval is to localize the video segment most relevant to the given natural language query without access to temporal annotations during training. Prior strongly- and weakly-supervised approaches…

计算机视觉与模式识别 · 计算机科学 2020-03-31 Reuben Tan , Huijuan Xu , Kate Saenko , Bryan A. Plummer

Online federated learning (FL) enables geographically distributed devices to learn a global shared model from locally available streaming data. Most online FL literature considers a best-case scenario regarding the participating clients and…

机器学习 · 计算机科学 2023-10-31 Francois Gauthier , Vinay Chakravarthi Gogineni , Stefan Werner , Yih-Fang Huang , Anthony Kuh

Sound source localization aims to localize objects emitting the sound in visual scenes. Recent works obtaining impressive results typically rely on contrastive learning. However, the common practice of randomly sampling negatives in prior…

计算机视觉与模式识别 · 计算机科学 2024-08-30 Zengjie Song , Jiangshe Zhang , Yuxi Wang , Junsong Fan , Zhaoxiang Zhang