中文
相关论文

相关论文: MVAD: A Benchmark Dataset for Multimodal AI-Genera…

200 篇论文

The rapid advancement of generative AI has raised concerns about the authenticity of digital images, as highly realistic fake images can now be generated at low cost, potentially increasing societal risks. In response, several datasets have…

计算机视觉与模式识别 · 计算机科学 2026-02-12 Hanzhe Yu , Yun Ye , Jintao Rong , Qi Xuan , Chen Ma

In the field of speaker diarization, the development of technology is constrained by two problems: insufficient data resources and poor generalization ability of deep learning models. To address these two problems, firstly, we propose an…

音频与语音处理 · 电气工程与系统科学 2025-07-01 Shilong Wu

Most existing multimodal machine translation (MMT) datasets are predominantly composed of static images or short video clips, lacking extensive video data across diverse domains and topics. As a result, they fail to meet the demands of…

计算与语言 · 计算机科学 2025-05-12 Jinze Lv , Jian Chen , Zi Long , Xianghua Fu , Yin Chen

The rapid rise of video content on platforms such as TikTok and YouTube has transformed information dissemination, but it has also facilitated the spread of harmful content, particularly hate videos. Despite significant efforts to combat…

多媒体 · 计算机科学 2025-05-20 Yinghui Zhang , Tailin Chen , Yuchen Zhang , Zeyu Fu

Deepfakes represent a growing concern across domains such as disinformation, fraud, and non-consensual media. In particular, the rise of video conference and identity-driven attacks in high-stakes scenarios--such as impostor hiring--demands…

计算机视觉与模式识别 · 计算机科学 2026-04-08 Sarah Barrington , Maty Bohacek , Hany Farid

The surge of highly realistic synthetic videos produced by contemporary generative systems has significantly increased the risk of malicious use, challenging both humans and existing detectors. Against this backdrop, we take a…

计算机视觉与模式识别 · 计算机科学 2026-03-11 Youngseo Kim , Kwan Yun , Seokhyeon Hong , Sihun Cha , Colette Suhjung Koo , Junyong Noh

While hate speech detection (HSD) has been extensively studied in text, existing multi-modal approaches remain limited, particularly in videos. As modalities are not always individually informative, simple fusion methods fail to fully…

Currently, dialogue systems have achieved high performance in processing text-based communication. However, they have not yet effectively incorporated visual information, which poses a significant challenge. Furthermore, existing models…

计算与语言 · 计算机科学 2023-12-19 Viktor Moskvoretskii , Anton Frolov , Denis Kuznetsov

Recent advances in audio generation led to an increasing number of deepfakes, making the general public more vulnerable to financial scams, identity theft, and misinformation. Audio deepfake detectors promise to alleviate this issue, with…

Existing deepfake detectors face several challenges in achieving robustness and generalization. One of the primary reasons is their limited ability to extract relevant information from forgery videos, especially in the presence of various…

计算机视觉与模式识别 · 计算机科学 2023-05-01 Zhiyuan Yan , Peng Sun , Yubo Lang , Shuo Du , Shanzhuo Zhang , Wei Wang , Lei Liu

As tools for content editing mature, and artificial intelligence (AI) based algorithms for synthesizing media grow, the presence of manipulated content across online media is increasing. This phenomenon causes the spread of misinformation,…

计算机视觉与模式识别 · 计算机科学 2022-12-09 Trisha Mittal , Ritwik Sinha , Viswanathan Swaminathan , John Collomosse , Dinesh Manocha

To open up new possibilities to assess the multimodal perceptual quality of omnidirectional media formats, we proposed a novel open source 360 audiovisual (AV) quality dataset. The dataset consists of high-quality 360 video clips in…

The rapid advancement of generative models in creating highly realistic images poses substantial risks for misinformation dissemination. For instance, a synthetic image, when shared on social media, can mislead extensive audiences and erode…

计算机视觉与模式识别 · 计算机科学 2025-06-27 Zhenglin Huang , Jinwei Hu , Xiangtai Li , Yiwei He , Xingyu Zhao , Bei Peng , Baoyuan Wu , Xiaowei Huang , Guangliang Cheng

The tremendous recent advances in generative artificial intelligence techniques have led to significant successes and promise in a wide range of different applications ranging from conversational agents and textual content generation to…

计算机视觉与模式识别 · 计算机科学 2024-05-24 Hossein Aboutalebi , Dayou Mao , Rongqi Fan , Carol Xu , Chris He , Alexander Wong

The growing sophistication of speech generated by Artificial Intelligence (AI) has introduced new challenges in audio deepfake detection. Text-to-speech (TTS) and voice conversion (VC) technologies can create highly convincing synthetic…

声音 · 计算机科学 2026-03-17 Vamshi Nallaguntla , Aishwarya Fursule , Shruti Kshirsagar , Anderson R. Avila

Towards open-ended Video Anomaly Detection (VAD), existing methods often exhibit biased detection when faced with challenging or unseen events and lack interpretability. To address these drawbacks, we propose Holmes-VAD, a novel framework…

计算机视觉与模式识别 · 计算机科学 2024-07-02 Huaxin Zhang , Xiaohao Xu , Xiang Wang , Jialong Zuo , Chuchu Han , Xiaonan Huang , Changxin Gao , Yuehuan Wang , Nong Sang

The gaming and entertainment industry is rapidly evolving, driven by immersive experiences and the integration of generative AI (GAI) technologies. Training such models effectively requires large-scale datasets that capture the diversity…

计算机视觉与模式识别 · 计算机科学 2025-11-03 Yuanzhi Li , Lebin Zhou , Nam Ling , Zhenghao Chen , Wei Wang , Wei Jiang

The emergence of contemporary deepfakes has attracted significant attention in machine learning research, as artificial intelligence (AI) generated synthetic media increases the incidence of misinterpretation and is difficult to distinguish…

计算机视觉与模式识别 · 计算机科学 2024-11-12 Ammarah Hashmi , Sahibzada Adil Shahzad , Chia-Wen Lin , Yu Tsao , Hsin-Min Wang

Artificial intelligence (AI) in media has advanced rapidly over the last decade. The introduction of Generative Adversarial Networks (GANs) improved the quality of photorealistic image generation. Diffusion models later brought a new era of…

计算机视觉与模式识别 · 计算机科学 2025-11-17 Redwan Hussain , Mizanur Rahman , Prithwiraj Bhattacharjee

Video anomaly detection (VAD) with weak supervision has achieved remarkable performance in utilizing video-level labels to discriminate whether a video frame is normal or abnormal. However, current approaches are inherently limited to a…

计算机视觉与模式识别 · 计算机科学 2024-03-14 Peng Wu , Xuerong Zhou , Guansong Pang , Yujia Sun , Jing Liu , Peng Wang , Yanning Zhang