English
Related papers

Related papers: PolyGen: Fully Synthetic Vision-Language Training …

200 papers

Shortage of labeled seismic field data poses a significant challenge for deep-learning related applications in seismology. One approach to mitigate this issue is to use synthetic waveforms as a complement to field data. However, traditional…

Geophysics · Physics 2023-10-03 Guoyi Chen , Junlun Li , Hao Guo

We propose a framework for the automatic one-shot segmentation of synthetic images generated by a StyleGAN. Our framework is based on the observation that the multi-scale hidden features in the GAN generator hold useful semantic information…

Computer Vision and Pattern Recognition · Computer Science 2023-10-24 Ankit Manerikar , Avinash C. Kak

Current volumetric biomedical foundation models struggle to generalize as public 3D datasets are small and do not cover the broad diversity of medical procedures, conditions, anatomical regions, and imaging protocols. We address this by…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Neel Dey , Benjamin Billot , Hallee E. Wong , Clinton J. Wang , Mengwei Ren , P. Ellen Grant , Adrian V. Dalca , Polina Golland

In real-world images, slanted or curved texts, especially those on cans, banners, or badges, appear as frequently, if not more so, than flat texts due to artistic design or layout constraints. While high-quality visual text generation has…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Minxing Luo , Zixun Xia , Liaojun Chen , Zhenhang Li , Weichao Zeng , Jianye Wang , Wentao Cheng , Yaxing Wang , Yu Zhou , Jian Yang

While today's large language models exhibit impressive abilities in generating human-like text, they require massive amounts of data during training. We here take inspiration from human cognitive development to train models in limited data…

Computer Vision and Pattern Recognition · Computer Science 2024-11-05 Badr AlKhamissi , Yingtian Tang , Abdülkadir Gökce , Johannes Mehrer , Martin Schrimpf

Large-scale and categorical-balanced text data is essential for training effective Scene Text Recognition (STR) models, which is hard to achieve when collecting real data. Synthetic data offers a cost-effective and perfectly labeled…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Xingsong Ye , Yongkun Du , JiaXin Zhang , Chen Li , Jing Lyu , Zhineng Chen

Modern computer vision systems increasingly encounter performance limitations in data-scarce domains, where collecting large-scale, high-quality labeled data is costly or impractical. While controllable diffusion models enable scalable…

Image and Video Processing · Electrical Eng. & Systems 2026-05-12 Yukang Shen

Typical methods for text-to-image synthesis seek to design effective generative architecture to model the text-to-image mapping directly. It is fairly arduous due to the cross-modality translation. In this paper we circumvent this problem…

Computer Vision and Pattern Recognition · Computer Science 2020-07-14 Jiadong Liang , Wenjie Pei , Feng Lu

Sparse-view 3D modeling represents a fundamental tension between reconstruction fidelity and generative plausibility. While feed-forward reconstruction excels in efficiency and input alignment, it often lacks the global priors needed for…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Zhisheng Huang , Jiahao Chen , Cheng Lin , Chenyu Hu , Hanzhuo Huang , Zhengming Yu , Mengfei Li , Yuheng Liu , Zekai Gu , Zibo Zhao , Yuan Liu , Xin Li , Wenping Wang

While the accuracy of face recognition systems has improved significantly in recent years, the datasets used to train these models are often collected through web crawling without the explicit consent of users, raising ethical and privacy…

Computer Vision and Pattern Recognition · Computer Science 2025-07-23 Anjith George , Sebastien Marcel

A recent strand of work in view synthesis uses deep learning to generate multiplane images (a camera-centric, layered 3D representation) given two or more input images at known viewpoints. We apply this representation to single-view view…

Computer Vision and Pattern Recognition · Computer Science 2020-04-24 Richard Tucker , Noah Snavely

Synthetic data generation overcomes limitations of real-world machine learning. Traditional methods are valuable for augmenting costly datasets but only optimize one criterion: realism. In this paper, we tackle the problem of generating…

Machine Learning · Computer Science 2021-11-16 Chance N DeSmet , Diane J Cook

An effective perception system is a fundamental component for farming robots, as it enables them to properly perceive the surrounding environment and to carry out targeted operations. The most recent methods make use of state-of-the-art…

Computer Vision and Pattern Recognition · Computer Science 2021-09-07 Mulham Fawakherji , Ciro Potena , Alberto Pretto , Domenico D. Bloisi , Daniele Nardi

Medical Vision-Language Pre-training (VLP) learns representations jointly from medical images and paired radiology reports. It typically requires large-scale paired image-text datasets to achieve effective pre-training for both the image…

Computer Vision and Pattern Recognition · Computer Science 2024-05-01 Che Liu , Anand Shah , Wenjia Bai , Rossella Arcucci

Synthetic image datasets offer unmatched advantages for designing and evaluating deep neural networks: they make it possible to (i) render as many data samples as needed, (ii) precisely control each scene and yield granular ground truth…

Computer Vision and Pattern Recognition · Computer Science 2023-12-14 Florian Bordes , Shashank Shekhar , Mark Ibrahim , Diane Bouchacourt , Pascal Vincent , Ari S. Morcos

Synthetic polyp generation is a good alternative to overcome the privacy problem of medical data and the lack of various polyp samples. In this study, we propose a deep learning-based polyp image generation framework that generates…

Image and Video Processing · Electrical Eng. & Systems 2023-02-21 Hemin Ali Qadir , Ilangko Balasingham , Younghak Shin

Conversational agents are required to respond to their users not only with high quality (i.e. commonsense bearing) responses, but also considering multiple plausible alternative scenarios, reflecting the diversity in their responses.…

Computation and Language · Computer Science 2026-04-21 Tianhui Zhang , Bei Peng , Danushka Bollegala

Training data is an essential resource for creating capable and robust vision systems which are integral to the proper function of many robotic systems. Synthesized training data has been shown in recent years to be a viable alternative to…

Robotics · Computer Science 2024-11-12 Peter Gavriel , Adam Norton , Kenneth Kimble , Megan Zimmerman

We present LingGen, a controlled text generation model that allows fine-grained control over a large number of real-valued linguistic attributes. It encodes target attribute values with a dedicated linguistic attribute encoder and…

Computation and Language · Computer Science 2026-01-27 Mohamed Elgaar , Hadi Amiri

Large-scale synthetic datasets are beneficial to stereo matching but usually introduce known domain bias. Although unsupervised image-to-image translation networks represented by CycleGAN show great potential in dealing with domain gap, it…

Computer Vision and Pattern Recognition · Computer Science 2020-05-06 Rui Liu , Chengxi Yang , Wenxiu Sun , Xiaogang Wang , Hongsheng Li