中文
相关论文

相关论文: MMO-IG: Multi-Class and Multi-Scale Object Image G…

200 篇论文

The objective optimization of medical imaging systems requires full characterization of all sources of randomness in the measured data, which includes the variability within the ensemble of objects to-be-imaged. This can be accomplished by…

图像与视频处理 · 电气工程与系统科学 2020-01-28 Weimin Zhou , Sayantan Bhadra , Frank J. Brooks , Hua Li , Mark A. Anastasio

We propose a novel AutoRegressive Generation-based paradigm for image Segmentation (ARGenSeg), achieving multimodal understanding and pixel-level perception within a unified framework. Prior works integrating image segmentation into…

计算机视觉与模式识别 · 计算机科学 2025-10-24 Xiaolong Wang , Lixiang Ru , Ziyuan Huang , Kaixiang Ji , Dandan Zheng , Jingdong Chen , Jun Zhou

In autonomous driving scenarios, accurate perception is becoming an even more critical task for safe navigation. While LiDAR provides precise spatial data, its inherent sparsity makes it difficult to detect small or distant objects.…

计算机视觉与模式识别 · 计算机科学 2025-12-09 Minseung Lee , Seokha Moon , Seung Joon Lee , Reza Mahjourian , Jinkyu Kim

Recent advancements in generative models have revolutionized the field of artificial intelligence, enabling the creation of highly-realistic and detailed images. In this study, we propose a novel Mask Conditional Text-to-Image Generative…

计算机视觉与模式识别 · 计算机科学 2024-10-02 Rami Skaik , Leonardo Rossi , Tomaso Fontanini , Andrea Prati

We present MVMO (Multi-View, Multi-Object dataset): a synthetic dataset of 116,000 scenes containing randomly placed objects of 10 distinct classes and captured from 25 camera locations in the upper hemisphere. MVMO comprises…

计算机视觉与模式识别 · 计算机科学 2022-06-01 Aitor Alvarez-Gila , Joost van de Weijer , Yaxing Wang , Estibaliz Garrote

Diffusion models have demonstrated impressive performance in text-to-image generation. They utilize a text encoder and cross-attention blocks to infuse textual information into images at a pixel level. However, their capability to generate…

计算机视觉与模式识别 · 计算机科学 2023-06-06 Luping Liu , Zijian Zhang , Yi Ren , Rongjie Huang , Xiang Yin , Zhou Zhao

Three-dimensional segmentation in magnetic resonance images (MRI), which reflects the true shape of the objects, is challenging since high-resolution isotropic MRIs are rare and typical MRIs are anisotropic, with the out-of-plane dimension…

图像与视频处理 · 电气工程与系统科学 2023-03-15 Hanxue Gu , Hongyu He , Roy Colglazier , Jordan Axelrod , Robert French , Maciej A Mazurowski

The application of Vision-language foundation models (VLFMs) to remote sensing (RS) imagery has garnered significant attention due to their superior capability in various downstream tasks. A key challenge lies in the scarcity of…

计算机视觉与模式识别 · 计算机科学 2025-07-23 Yiguo He , Junjie Zhu , Yiying Li , Xiaoyu Zhang , Chunping Qiu , Jun Wang , Qiangjuan Huang , Ke Yang

The success of deep reinforcement learning (RL) and imitation learning (IL) in vision-based robotic manipulation typically hinges on the expense of large scale data collection. With simulation, data to train a policy can be collected…

机器人学 · 计算机科学 2021-07-06 Daniel Ho , Kanishka Rao , Zhuo Xu , Eric Jang , Mohi Khansari , Yunfei Bai

Although existing image caption models can produce promising results using recurrent neural networks (RNNs), it is difficult to guarantee that an object we care about is contained in generated descriptions, for example in the case that the…

计算机视觉与模式识别 · 计算机科学 2019-04-12 Yue Zheng , Yali Li , Shengjin Wang

Diffusion models (DMs) have become dominant in visual generation but suffer performance drop when tested on resolutions that differ from the training scale, whether lower or higher. In fact, the key challenge in generating variable-scale…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Guohui Zhang , Jiangtong Tan , Linjiang Huang , Zhonghang Yuan , Mingde Yao , Jie Huang , Feng Zhao

We introduce the Multi-Instance Generation (MIG) task, which focuses on generating multiple instances within a single image, each accurately placed at predefined positions with attributes such as category, color, and shape, strictly…

计算机视觉与模式识别 · 计算机科学 2025-05-06 Dewei Zhou , You Li , Fan Ma , Zongxin Yang , Yi Yang

Multimodal fusion of remote sensing images serves as a core technology for overcoming the limitations of single-source data and improving the accuracy of surface information extraction, which exhibits significant application value in fields…

计算机视觉与模式识别 · 计算机科学 2026-01-12 Siyu Zhang , Lianlei Shan , Runhe Qiu

In this paper, we introduce a new method for generating an object image from text attributes on a desired location, when the base image is given. One step further to the existing studies on text-to-image generation mainly focusing on the…

计算机视觉与模式识别 · 计算机科学 2018-08-16 Hyojin Park , YoungJoon Yoo , Nojun Kwak

Metric Cross-View Geo-Localization (MCVGL) aims to estimate the 3-DoF camera pose (position and heading) by matching ground and satellite images. In this work, instead of pinhole and satellite images, we study robust MCVGL using holistic…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Junwei Zheng , Ruize Dai , Ruiping Liu , Zichao Zeng , Yufan Chen , Fangjinhua Wang , Kunyu Peng , Kailun Yang , Jiaming Zhang , Rainer Stiefelhagen

Compared with natural images, remote sensing images (RSIs) have the unique characteristic. i.e., larger intraclass variance, which makes semantic segmentation for remote sensing images more challenging. Moreover, existing semantic…

计算机视觉与模式识别 · 计算机科学 2025-03-24 Wei Zhang , Mengting Ma , Yizhen Jiang , Rongrong Lian , Zhenkai Wu , Kangning Cui , Xiaowen Ma

Most text-to-3D generators build upon off-the-shelf text-to-image models trained on billions of images. They use variants of Score Distillation Sampling (SDS), which is slow, somewhat unstable, and prone to artifacts. A mitigation is to…

计算机视觉与模式识别 · 计算机科学 2024-02-14 Luke Melas-Kyriazi , Iro Laina , Christian Rupprecht , Natalia Neverova , Andrea Vedaldi , Oran Gafni , Filippos Kokkinos

We present DIPO, a novel framework for the controllable generation of articulated 3D objects from a pair of images: one depicting the object in a resting state and the other in an articulated state. Compared to the single-image approach,…

计算机视觉与模式识别 · 计算机科学 2025-05-29 Ruiqi Wu , Xinjie Wang , Liu Liu , Chunle Guo , Jiaxiong Qiu , Chongyi Li , Lichao Huang , Zhizhong Su , Ming-Ming Cheng

In this work, we study the problem of generating novel images from complex multimodal prompt sequences. While existing methods achieve promising results for text-to-image generation, they often struggle to capture fine-grained details from…

计算机视觉与模式识别 · 计算机科学 2024-05-29 Amandeep Kumar , Muzammal Naseer , Sanath Narayan , Rao Muhammad Anwer , Salman Khan , Hisham Cholakkal

The emergence of generative models has revolutionized the field of remote sensing (RS) image generation. Despite generating high-quality images, existing methods are limited in relying mainly on text control conditions, and thus do not…

计算机视觉与模式识别 · 计算机科学 2024-09-04 Datao Tang , Xiangyong Cao , Xingsong Hou , Zhongyuan Jiang , Junmin Liu , Deyu Meng