中文
相关论文

相关论文: TV-3DG: Mastering Text-to-3D Customized Generation…

200 篇论文

Text-to-3D generation has recently seen significant progress. To enhance its practicality in real-world applications, it is crucial to generate multiple independent objects with interactions, similar to layer-compositing in 2D image…

计算机视觉与模式识别 · 计算机科学 2024-07-24 Zizheng Yan , Jiapeng Zhou , Fanpeng Meng , Yushuang Wu , Lingteng Qiu , Zisheng Ye , Shuguang Cui , Guanying Chen , Xiaoguang Han

The standard definition generation task requires to automatically produce mono-lingual definitions (e.g., English definitions for English words), but ignores that the generated definitions may also consist of unfamiliar words for language…

计算与语言 · 计算机科学 2023-06-12 Hengyuan Zhang , Dawei Li , Yanran Li , Chenming Shang , Chufan Shi , Yong Jiang

Efficiently training accurate deep models for weakly supervised semantic segmentation (WSSS) with image-level labels is challenging and important. Recently, end-to-end WSSS methods have become the focus of research due to their high…

计算机视觉与模式识别 · 计算机科学 2023-02-28 Rongtao Xu , Changwei Wang , Jiaxi Sun , Shibiao Xu , Weiliang Meng , Xiaopeng Zhang

We introduce Delta Denoising Score (DDS), a novel scoring function for text-based image editing that guides minimal modifications of an input image towards the content described in a target prompt. DDS leverages the rich generative prior of…

计算机视觉与模式识别 · 计算机科学 2023-04-17 Amir Hertz , Kfir Aberman , Daniel Cohen-Or

Vision foundation models (VFMs) such as DINO have led to a paradigm shift in 2D camera-based perception towards extracting generalized features to support many downstream tasks. Recent works introduce self-supervised cross-modal knowledge…

计算机视觉与模式识别 · 计算机科学 2025-11-24 Hariprasath Govindarajan , Maciej K. Wozniak , Marvin Klingner , Camille Maurice , B Ravi Kiran , Senthil Yogamani

While promptable segmentation (\textit{e.g.}, SAM) has shown promise for various segmentation tasks, it still requires manual visual prompts for each object to be segmented. In contrast, task-generic promptable segmentation aims to reduce…

计算机视觉与模式识别 · 计算机科学 2025-08-15 Chao Yin , Hao Li , Kequan Yang , Jide Li , Pinpin Zhu , Xiaoqiang Li

Current diffusion-based super-resolution (SR) approaches achieve commendable performance at the cost of high inference overhead. Therefore, distillation techniques are utilized to accelerate the multi-step teacher model into one-step…

计算机视觉与模式识别 · 计算机科学 2025-11-10 Weiyi You , Mingyang Zhang , Leheng Zhang , Xingyu Zhou , Kexuan Shi , Shuhang Gu

Score-based generative models (SGMs) have recently emerged as a promising class of generative models. The key idea is to produce high-quality images by recurrently adding Gaussian noises and gradients to a Gaussian sample until converging…

计算机视觉与模式识别 · 计算机科学 2022-06-13 Hengyuan Ma , Li Zhang , Xiatian Zhu , Jingfeng Zhang , Jianfeng Feng

The controllability of 3D object generation methods is achieved through input text. Existing text-to-3D object generation methods primarily focus on generating a single object based on a single object description. However, these methods…

计算机视觉与模式识别 · 计算机科学 2024-09-04 Shaorong Sun , Shuchao Pang , Yazhou Yao , Xiaoshui Huang

We propose Noise Conditional Variational Score Distillation (NCVSD), a novel method for distilling pretrained diffusion models into generative denoisers. We achieve this by revealing that the unconditional score function implicitly…

计算机视觉与模式识别 · 计算机科学 2025-06-12 Xinyu Peng , Ziyang Zheng , Yaoming Wang , Han Li , Nuowen Kan , Wenrui Dai , Chenglin Li , Junni Zou , Hongkai Xiong

Open-world object counting leverages the robust text-image alignment of pre-trained vision-language models (VLMs) to enable counting of arbitrary categories in images specified by textual queries. However, widely adopted naive fine-tuning…

计算机视觉与模式识别 · 计算机科学 2025-04-25 Yiming Zhao , Guorong Li , Laiyun Qing , Amin Beheshti , Jian Yang , Michael Sheng , Yuankai Qi , Qingming Huang

Recent 3D generative models have achieved remarkable performance in synthesizing high resolution photorealistic images with view consistency and detailed 3D shapes, but training them for diverse domains is challenging since it requires…

计算机视觉与模式识别 · 计算机科学 2023-04-03 Gwanghyun Kim , Se Young Chun

Convolutional neural networks (CNNs) are highly successful for super-resolution (SR) but often require sophisticated architectures with heavy memory cost and computational overhead, significantly restricts their practical deployments on…

计算机视觉与模式识别 · 计算机科学 2021-05-26 Yanbo Wang , Shaohui Lin , Yanyun Qu , Haiyan Wu , Zhizhong Zhang , Yuan Xie , Angela Yao

In recent years, Denoising Diffusion Probabilistic Models (DDPMs) have demonstrated exceptional performance in various 2D generative tasks. Following this success, DDPMs have been extended to 3D shape generation, surpassing previous…

计算机视觉与模式识别 · 计算机科学 2023-09-14 Cristian Sbrolli , Paolo Cudrano , Matteo Frosi , Matteo Matteucci

Although the vision-and-language pretraining (VLP) equipped cross-modal image-text retrieval (ITR) has achieved remarkable progress in the past two years, it suffers from a major drawback: the ever-increasing size of VLP models restricts…

多媒体 · 计算机科学 2022-07-05 Jun Rao , Liang Ding , Shuhan Qi , Meng Fang , Yang Liu , Li Shen , Dacheng Tao

Large Language Models (LLMs) have demonstrated remarkable capabilities in code generation. However, the quality of the generated code is heavily dependent on the structure and composition of the prompts used. Crafting high-quality prompts…

软件工程 · 计算机科学 2025-04-08 Jinyang Li , Sangwon Hyun , M. Ali Babar

Single image-to-3D generation is pivotal for crafting controllable 3D assets. Given its under-constrained nature, we attempt to leverage 3D geometric priors from a novel view diffusion model and 2D appearance priors from an image generation…

计算机视觉与模式识别 · 计算机科学 2024-10-10 Shuzhou Yang , Yu Wang , Haijie Li , Jiarui Meng , Yanmin Wu , Xiandong Meng , Jian Zhang

To address the data scarcity associated with 3D assets, 2D-lifting techniques such as Score Distillation Sampling (SDS) have become a widely adopted practice in text-to-3D generation pipelines. However, the diffusion models used in these…

计算机视觉与模式识别 · 计算机科学 2025-01-23 Utkarsh Nath , Rajeev Goel , Eun Som Jeon , Changhoon Kim , Kyle Min , Yezhou Yang , Yingzhen Yang , Pavan Turaga

Synthetic Data Generation (SDG), leveraging Large Language Models (LLMs), has recently been recognized and broadly adopted as an effective approach to improve the performance of smaller but more resource and compute efficient LLMs through…

机器学习 · 计算机科学 2026-03-25 Srideepika Jayaraman , Achille Fokoue , Dhaval Patel , Jayant Kalagnanam

Distilling video generation models to extremely low inference budgets (e.g., 2--4 NFEs) is crucial for real-time deployment, yet remains challenging. Trajectory-style consistency distillation often becomes conservative under complex video…

计算机视觉与模式识别 · 计算机科学 2026-04-06 Xingtong Ge , Yi Zhang , Yushi Huang , Dailan He , Xiahong Wang , Bingqi Ma , Guanglu Song , Yu Liu , Jun Zhang
‹ 上一页 1 8 9 10 下一页 ›