中文
相关论文

相关论文: Toward Real-world Text Image Forgery Localization:…

200 篇论文

Edit-based approaches have recently shown promising results on multiple monolingual sequence transduction tasks. In contrast to conventional sequence-to-sequence (Seq2Seq) models, which learn to generate text from scratch as they are…

计算与语言 · 计算机科学 2022-05-11 Kostiantyn Omelianchuk , Vipul Raheja , Oleksandr Skurzhanskyi

Current temporal forgery localization (TFL) approaches typically rely on temporal boundary regression or continuous frame-level anomaly detection paradigms to derive candidate forgery proposals. However, they suffer not only from feature…

计算机视觉与模式识别 · 计算机科学 2026-03-09 Tianyi Wang , Xi Shao , Harry Cheng , Yinglong Wang , Mohan Kankanhalli

This paper explores multi-modal controllable Text-to-Speech Synthesis (TTS) where the voice can be generated from face image, and the characteristics of output speech (e.g., pace, noise level, distance, tone, place) can be controllable with…

音频与语音处理 · 电气工程与系统科学 2025-05-27 Minsu Kim , Pingchuan Ma , Honglie Chen , Stavros Petridis , Maja Pantic

Most recent style-transfer techniques based on generative architectures are able to obtain synthetic multimedia contents, or commonly called deepfakes, with almost no artifacts. Researchers already demonstrated that synthetic images contain…

计算机视觉与模式识别 · 计算机科学 2022-06-08 Luca Guarnera , Oliver Giudice , Sebastiano Battiato

Packing for Supervised Fine-Tuning (SFT) in autoregressive models involves concatenating data points of varying lengths until reaching the designed maximum length to facilitate GPU processing. However, randomly concatenating data points can…

机器学习 · 计算机科学 2025-02-27 Jiancheng Dong , Lei Jiang , Wei Jin , Lu Cheng

Although significant advances have been made in face recognition (FR), FR in unconstrained environments remains challenging due to the domain gap between the semi-constrained training datasets and unconstrained testing scenarios. To address…

计算机视觉与模式识别 · 计算机科学 2022-07-22 Feng Liu , Minchul Kim , Anil Jain , Xiaoming Liu

As a special case of common object removal, image person removal is playing an increasingly important role in social media and criminal investigation domains. Due to the integrity of person area and the complexity of human posture, person…

计算机视觉与模式识别 · 计算机科学 2022-09-30 Yunliang Jiang , Chenyang Gu , Zhenfeng Xue , Xiongtao Zhang , Yong Liu

Most of researches on image forensics have been mainly focused on detection of artifacts introduced by a single processing tool. They lead in the development of many specialized algorithms looking for one or more particular footprints under…

计算机视觉与模式识别 · 计算机科学 2017-01-31 Habib Ghaffari Hadigheh , Ghazali bin sulong

Facial images disclose many hidden personal traits such as age, gender, race, health, emotion, and psychology. Understanding these traits will help to classify the people in different attributes. In this paper, we have presented a novel…

计算机视觉与模式识别 · 计算机科学 2022-01-26 Rahul Goel , Modar Sulaiman , Kimia Noorbakhsh , Mahdi Sharifi , Rajesh Sharma , Pooyan Jamshidi , Kallol Roy

Neural Text-to-Speech (TTS) systems find broad applications in voice assistants, e-learning, and audiobook creation. The pursuit of modern models, like Diffusion Models (DMs), holds promise for achieving high-fidelity, real-time speech…

声音 · 计算机科学 2024-04-02 Xiang Li , Fan Bu , Ambuj Mehrish , Yingting Li , Jiale Han , Bo Cheng , Soujanya Poria

Well-designed prompts have demonstrated the potential to guide text-to-image models in generating amazing images. Although existing prompt engineering methods can provide high-level guidance, it is challenging for novice users to achieve…

多媒体 · 计算机科学 2026-03-27 Nailei Hei , Qianyu Guo , Zihao Wang , Yan Wang , Haofen Wang , Wenqiang Zhang

Modern deepfakes have evolved into localized and intermittent manipulations that require fine-grained temporal localization to mitigate severe digital security risks. The prohibitive cost of frame-level annotation makes weakly supervised…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Midou Guo , Qilin Yin , Wei Lu , Rui Yang

Typographic attacks exploit the interplay between text and visual content in multimodal foundation models, causing misclassifications when misleading text is embedded within images. Existing datasets are limited in size and diversity,…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Justus Westerhoff , Erblina Purelku , Jakob Hackstein , Jonas Loos , Leo Pinetzki , Erik Rodner , Lorenz Hufe

Text-to-image diffusion models have achieved unprecedented success but still struggle to produce high-quality results under limited sampling budgets. Existing training-free sampling acceleration methods are typically developed…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Zhenyu Zhou , Defang Chen , Siwei Lyu , Chun Chen , Can Wang

Collecting amounts of distorted/clean image pairs in the real world is non-trivial, which seriously limits the practical applications of these supervised learning-based methods on real-world image super-resolution (RealSR). Previous works…

计算机视觉与模式识别 · 计算机科学 2023-04-19 Xin Li , Xin Jin , Jun Fu , Xiaoyuan Yu , Bei Tong , Zhibo Chen

Existing deepfake detection methods have reported promising in-distribution results, by accessing published large-scale dataset. However, due to the non-smooth synthesis method, the fake samples in this dataset may expose obvious artifacts…

计算机视觉与模式识别 · 计算机科学 2020-12-22 Xinwei Sun , Botong Wu , Wei Chen

An unsupervised text-to-speech synthesis (TTS) system learns to generate speech waveforms corresponding to any written sentence in a language by observing: 1) a collection of untranscribed speech waveforms in that language; 2) a collection…

音频与语音处理 · 电气工程与系统科学 2022-08-17 Junrui Ni , Liming Wang , Heting Gao , Kaizhi Qian , Yang Zhang , Shiyu Chang , Mark Hasegawa-Johnson

The rapid advancement of generative models in creating highly realistic images poses substantial risks for misinformation dissemination. For instance, a synthetic image, when shared on social media, can mislead extensive audiences and erode…

计算机视觉与模式识别 · 计算机科学 2025-06-27 Zhenglin Huang , Jinwei Hu , Xiangtai Li , Yiwei He , Xingyu Zhao , Bei Peng , Baoyuan Wu , Xiaowei Huang , Guangliang Cheng

A commonly used evaluation metric for text-to-image synthesis is the Inception score (IS) \cite{inceptionscore}, which has been shown to be a quality metric that correlates well with human judgment. However, IS does not reveal properties of…

机器学习 · 计算机科学 2019-11-04 William Lund Sommer , Alexandros Iosifidis

Accurately estimating the point spread function (PSF) of an optical system requires solving free-space wave propagation, which entails evaluating a diffraction integral. This integral is traditionally computed numerically using Fast Fourier…

图像与视频处理 · 电气工程与系统科学 2026-05-22 Nicholas Ganino , Qi Guo