English
Related papers

Related papers: CelebV-Text: A Large-Scale Facial Text-Video Datas…

200 papers

The recent advances in deep learning have made it possible to generate photo-realistic images by using neural networks and even to extrapolate video frames from an input video clip. In this paper, for the sake of both furthering this…

Computer Vision and Pattern Recognition · Computer Science 2018-08-10 Lijie Fan , Wenbing Huang , Chuang Gan , Junzhou Huang , Boqing Gong

In this paper, we consider a novel and practical case for talking face video generation. Specifically, we focus on the scenarios involving multi-people interactions, where the talking context, such as audience or surroundings, is present.…

Computer Vision and Pattern Recognition · Computer Science 2024-02-29 Meidai Xuanyuan , Yuwang Wang , Honglei Guo , Qionghai Dai

Humans often specify and create through visual artifacts: typography sheets, sketches, reference images, and annotated scenes. Yet modern visual generators still ask users to serialize this intent into text, a bottleneck that compresses…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Yaofang Liu , Kangning Cui , Meng Chu , Zhaoqing Li , Suiyun Zhang , Jean-Michel Morel , Xiaodong Cun , Haoxuan Che , Rui Liu , Raymond H. Chan

A large number of annotated training images is crucial for training successful scene text recognition models. However, collecting sufficient datasets can be a labor-intensive and costly process, particularly for low-resource languages. To…

Computer Vision and Pattern Recognition · Computer Science 2023-06-28 Yangchen Xie , Xinyuan Chen , Hongjian Zhan , Palaiahankote Shivakum , Bing Yin , Cong Liu , Yue Lu

Diffusion models have demonstrated great success in text-to-video (T2V) generation. However, existing methods may face challenges when handling complex (long) video generation scenarios that involve multiple objects or dynamic changes in…

Computer Vision and Pattern Recognition · Computer Science 2024-10-15 Ye Tian , Ling Yang , Haotian Yang , Yuan Gao , Yufan Deng , Jingmin Chen , Xintao Wang , Zhaochen Yu , Xin Tao , Pengfei Wan , Di Zhang , Bin Cui

Video generation has increasingly gained interest in both academia and industry. Although commercial tools can generate plausible videos, there is a limited number of open-source models available for researchers and engineers. In this work,…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Haoxin Chen , Menghan Xia , Yingqing He , Yong Zhang , Xiaodong Cun , Shaoshu Yang , Jinbo Xing , Yaofang Liu , Qifeng Chen , Xintao Wang , Chao Weng , Ying Shan

Automated data visualization plays a crucial role in simplifying data interpretation, enhancing decision-making, and improving efficiency. While large language models (LLMs) have shown promise in generating visualizations from natural…

Computation and Language · Computer Science 2025-07-29 Mizanur Rahman , Md Tahmid Rahman Laskar , Shafiq Joty , Enamul Hoque

Efficient video-language modeling should consider the computational cost because of a large, sometimes intractable, number of video frames. Parametric approaches such as the attention mechanism may not be ideal since its computational cost…

Computer Vision and Pattern Recognition · Computer Science 2023-01-30 Sungdong Kim , Jin-Hwa Kim , Jiyoung Lee , Minjoon Seo

Text-to-video diffusion models enable the generation of high-quality videos that follow text instructions, making it easy to create diverse and individual content. However, existing approaches mostly focus on high-quality short video…

Computer Vision and Pattern Recognition · Computer Science 2025-04-17 Roberto Henschel , Levon Khachatryan , Hayk Poghosyan , Daniil Hayrapetyan , Vahram Tadevosyan , Zhangyang Wang , Shant Navasardyan , Humphrey Shi

Generating face image with specific gaze information has attracted considerable attention. Existing approaches typically input gaze values directly for face generation, which is unnatural and requires annotated gaze datasets for training,…

Computer Vision and Pattern Recognition · Computer Science 2024-10-01 Hengfei Wang , Zhongqun Zhang , Yihua Cheng , Hyung Jin Chang

Vision-Language Models pre-trained on large-scale image-text datasets have shown superior performance in downstream tasks such as image retrieval. Most of the images for pre-training are presented in the form of open domain common-sense…

Computer Vision and Pattern Recognition · Computer Science 2024-01-26 Xiangshuo Qiao , Xianxin Li , Xiaozhe Qu , Jie Zhang , Yang Liu , Yu Luo , Cihang Jin , Jin Ma

The Text to Audible-Video Generation (TAVG) task involves generating videos with accompanying audio based on text descriptions. Achieving this requires skillful alignment of both audio and video elements. To support research in this field,…

Computer Vision and Pattern Recognition · Computer Science 2024-04-23 Yuxin Mao , Xuyang Shen , Jing Zhang , Zhen Qin , Jinxing Zhou , Mochu Xiang , Yiran Zhong , Yuchao Dai

The objective of this paper is to jointly synthesize interactive videos and conversational speech from text and reference images. With the ultimate goal of building human-like conversational systems, recent studies have explored talking or…

Computer Vision and Pattern Recognition · Computer Science 2025-12-24 Ji-Hoon Kim , Junseok Ahn , Doyeop Kwak , Joon Son Chung , Shinji Watanabe

Text-guided image generation aimed to generate desired images conditioned on given texts, while text-guided image manipulation refers to semantically edit parts of a given image based on specified texts. For these two similar tasks, the key…

Computer Vision and Pattern Recognition · Computer Science 2023-09-22 Xiaozhou You , Jian Zhang

Text-conditioned image generation has gained significant attention in recent years and are processing increasingly longer and comprehensive text prompt. In everyday life, dense and intricate text appears in contexts like advertisements,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Alex Jinpeng Wang , Dongxing Mao , Jiawei Zhang , Weiming Han , Zhuobai Dong , Linjie Li , Yiqi Lin , Zhengyuan Yang , Libo Qin , Fuwei Zhang , Lijuan Wang , Min Li

With the advance of deep learning technology, automatic video generation from audio or text has become an emerging and promising research topic. In this paper, we present a novel approach to synthesize video from the text. The method builds…

Computer Vision and Pattern Recognition · Computer Science 2022-01-25 Sibo Zhang , Jiahong Yuan , Miao Liao , Liangjun Zhang

The growth of Social Networks has fueled the habit of people logging their day-to-day activities, and long First-Person Videos (FPVs) are one of the main tools in this new habit. Semantic-aware fast-forward methods are able to decrease the…

Computer Vision and Pattern Recognition · Computer Science 2020-01-01 Washington L. S. Ramos , Michel M. Silva , Edson R. Araujo , Alan C. Neves , Erickson R. Nascimento

Unlike existing methods that rely on source images as appearance references and use source speech to generate motion, this work proposes a novel approach that directly extracts information from the speech, addressing key challenges in…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-03 Jinting Wang , Jun Wang , Hei Victor Cheng , Li Liu

The increasing use of machine learning models has amplified the demand for high-quality, large-scale multimodal datasets. However, the availability of such datasets, especially those combining acoustic, visual and textual data, remains…

Multimedia · Computer Science 2025-09-09 Jorge E. León , Miguel Carrasco

We review research on generating visual data from text from the angle of "cross-modal generation." This point of view allows us to draw parallels between various methods geared towards working on input text and producing visual output,…

Computer Vision and Pattern Recognition · Computer Science 2024-01-23 Maciej Żelaszczyk , Jacek Mańdziuk