English
Related papers

Related papers: Toward Real-world Text Image Forgery Localization:…

200 papers

In this paper, we present an empirical study introducing a nuanced evaluation framework for text-to-image (T2I) generative models, applied to human image synthesis. Our framework categorizes evaluations into two distinct groups: first,…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Muxi Chen , Yi Liu , Jian Yi , Changran Xu , Qiuxia Lai , Hongliang Wang , Tsung-Yi Ho , Qiang Xu

The internet is filled with fake face images and videos synthesized by deep generative models. These realistic DeepFakes pose a challenge to determine the authenticity of multimedia content. As countermeasures, artifact-based detection…

Computer Vision and Pattern Recognition · Computer Science 2022-03-31 Gaojian Wang , Qian Jiang , Xin Jin , Xiaohui Cui

Numerous studies have shown that existing Face Recognition Systems (FRS), including commercial ones, often exhibit biases toward certain ethnicities due to under-represented data. In this work, we explore ethnicity alteration and skin tone…

Computer Vision and Pattern Recognition · Computer Science 2024-05-08 Praveen Kumar Chandaliya , Kiran Raja , Raghavendra Ramachandra , Zahid Akhtar , Christoph Busch

The rapid evolution of AIGC technology enables misleading viewers by tampering mere small segments within a video, rendering video-level detection inaccurate and unpersuasive. Consequently, temporal forgery localization (TFL), which aims to…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Boyang Zhao , Xin Liao , Jiaxin Chen , Xiaoshuai Wu , Yufeng Wu

Stuttering is a speech disorder where the natural flow of speech is interrupted by blocks, repetitions or prolongations of syllables, words and phrases. The majority of existing automatic speech recognition (ASR) interfaces perform poorly…

Score distillation sampling (SDS) demonstrates a powerful capability for text-conditioned 2D image and 3D object generation by distilling the knowledge from learned score functions. However, SDS often suffers from blurriness caused by noisy…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 SeonHwa Kim , Jiwon Kim , Soobin Park , Donghoon Ahn , Jiwon Kang , Seungryong Kim , Kyong Hwan Jin , Eunju Cha

This paper introduces F5-TTS, a fully non-autoregressive text-to-speech system based on flow matching with Diffusion Transformer (DiT). Without requiring complex designs such as duration model, text encoder, and phoneme alignment, the text…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-21 Yushen Chen , Zhikang Niu , Ziyang Ma , Keqi Deng , Chunhui Wang , Jian Zhao , Kai Yu , Xie Chen

Scene Text Editing (STE) aims to naturally modify text in images while preserving visual consistency, the decisive factors of which can be divided into three parts, i.e., text style, text content, and background. Previous methods have…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Yuchen Bao , Yiting Wang , Wenjian Huang , Haowei Wang , Shen Chen , Taiping Yao , Shouhong Ding , Jianguo Zhang

Collecting high-quality studio recordings of audio is challenging, which limits the language coverage of text-to-speech (TTS) systems. This paper proposes a framework for scaling a multilingual TTS model to 100+ languages using found data…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-17 Takaaki Saeki , Gary Wang , Nobuyuki Morioka , Isaac Elias , Kyle Kastner , Fadi Biadsy , Andrew Rosenberg , Bhuvana Ramabhadran , Heiga Zen , Françoise Beaufays , Hadar Shemtov

With diverse presentation forgery methods emerging continually, detecting the authenticity of images has drawn growing attention. Although existing methods have achieved impressive accuracy in training dataset detection, they still perform…

Computer Vision and Pattern Recognition · Computer Science 2024-03-20 Yingxin Lai , Guoqing Yang Yifan He , Zhiming Luo , Shaozi Li

The emergence of deepfake technologies has become a matter of social concern as they pose threats to individual privacy and public security. It is now of great significance to develop reliable deepfake detectors. However, with numerous face…

Computer Vision and Pattern Recognition · Computer Science 2023-03-16 Liang Shi , Jie Zhang , Shiguang Shan

Detecting AI-generated images, particularly deepfakes, has become increasingly crucial, with the primary challenge being the generalization to previously unseen manipulation methods. This paper tackles this issue by leveraging the forgery…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Wentang Song , Zhiyuan Yan , Yuzhen Lin , Taiping Yao , Changsheng Chen , Shen Chen , Yandan Zhao , Shouhong Ding , Bin Li

Most image super-resolution (SR) methods are developed on synthetic low-resolution (LR) and high-resolution (HR) image pairs that are constructed by a predetermined operation, e.g., bicubic downsampling. As existing methods typically learn…

Image and Video Processing · Electrical Eng. & Systems 2021-09-09 Sanghyun Son , Jaeha Kim , Wei-Sheng Lai , Ming-Husan Yang , Kyoung Mu Lee

The lack of large-scale datasets has been impeding the advance of deep learning approaches to the problem of F-formation detection. Moreover, most research works on this problem rely on input sensor signals of object location and…

Computer Vision and Pattern Recognition · Computer Science 2022-11-22 Giang Hoang , Tuan Nguyen Dinh , Tung Cao Hoang , Son Le Duy , Keisuke Hihara , Yumeka Utada , Akihiko Torii , Naoki Izumi , Long Tran Quoc

Controlling the style and characteristics of speech synthesis is crucial for adapting the output to specific contexts and user requirements. Previous Text-to-speech (TTS) works have focused primarily on the technical aspects of producing…

Sound · Computer Science 2025-09-04 Jiawei Zhang , Tian-Hao Zhang , Jun Wang , Jiaran Gao , Xinyuan Qian , Xu-Cheng Yin

Real life signals are in general non--stationary and non--linear. The development of methods able to extract their hidden features in a fast and reliable way is of high importance in many research fields. In this work we tackle the problem…

Numerical Analysis · Mathematics 2018-10-26 Antonio Cicone , Haomin Zhou

Diffusion models have achieved remarkable success in image synthesis, but the generated high-quality images raise concerns about potential malicious use. Existing detectors often struggle to capture discriminative clues across different…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Daichi Zhang , Tong Zhang , Shiming Ge , Sabine Süsstrunk

Despite the fact that DeepFake forgery detection algorithms have achieved impressive performance on known manipulations, they often face disastrous performance degradation when generalized to an unseen manipulation. Some recent works show…

Computer Vision and Pattern Recognition · Computer Science 2023-04-03 Chuer Yu , Xuhong Zhang , Yuxuan Duan , Senbo Yan , Zonghui Wang , Yang Xiang , Shouling Ji , Wenzhi Chen

Pre-trained text-to-image (T2I) diffusion models have shown strong potential for real-world image super-resolution (Real-ISR), owing to their noise-started generation process that enables realistic texture synthesis and captures the…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Wei Zhu , Kai Zhang , Yu Zheng , Lei Luo , Yong Guo , Jian Yang

Generalized few-shot semantic segmentation (GFSS) is fundamentally limited by the coverage of novel-class appearances under scarce annotations. While diffusion models can synthesize novel-class images at scale, practical gains are often…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Guohuan Xie , Xin He , Dingying Fan , Le Zhang , Ming-Ming Cheng , Yun Liu
‹ Prev 1 8 9 10 Next ›