中文
相关论文

相关论文: RockGPT: Reconstructing three-dimensional digital …

200 篇论文

We present VideoGPT: a conceptually simple architecture for scaling likelihood based generative modeling to natural videos. VideoGPT uses VQ-VAE that learns downsampled discrete latent representations of a raw video by employing 3D…

计算机视觉与模式识别 · 计算机科学 2021-09-16 Wilson Yan , Yunzhi Zhang , Pieter Abbeel , Aravind Srinivas

This manuscript presents PoreDiT, a novel generative model designed for high-efficiency digital rock reconstruction at gigavoxel scales. Addressing the significant challenges in digital rock physics (DRP), particularly the trade-off between…

人工智能 · 计算机科学 2026-04-14 Yizhuo Huang , Baoquan Sun , Haibo Huang

In this study, we introduce T2M-HiFiGPT, a novel conditional generative framework for synthesizing human motion from textual descriptions. This framework is underpinned by a Residual Vector Quantized Variational AutoEncoder (RVQ-VAE) and a…

计算机视觉与模式识别 · 计算机科学 2023-12-27 Congyi Wang

Pore-scale modeling of rock images based on information in 3D micro-computed tomography data is crucial for studying complex subsurface processes such as CO2 and brine multiphase flow during Geologic Carbon Storage (GCS). While deep…

计算机视觉与模式识别 · 计算机科学 2024-10-30 Zihan Ren , Sanjay Srinivasan , Dustin Crandall

Learning-based 3D reconstruction models, represented by Visual Geometry Grounded Transformers (VGGTs), have made remarkable progress with the use of large-scale transformers. Their prohibitive computational and memory costs severely hinder…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Weilun Feng , Haotong Qin , Mingqiang Wu , Chuanguang Yang , Yuqi Li , Xiangqi Li , Zhulin An , Libo Huang , Yulun Zhang , Michele Magno , Yongjun Xu

In this work, we investigate a simple and must-known conditional generative framework based on Vector Quantised-Variational AutoEncoder (VQ-VAE) and Generative Pre-trained Transformer (GPT) for human motion generation from textural…

计算机视觉与模式识别 · 计算机科学 2023-09-26 Jianrong Zhang , Yangsong Zhang , Xiaodong Cun , Shaoli Huang , Yong Zhang , Hongwei Zhao , Hongtao Lu , Xi Shen

We present TAR3D, a novel framework that consists of a 3D-aware Vector Quantized-Variational AutoEncoder (VQ-VAE) and a Generative Pre-trained Transformer (GPT) to generate high-quality 3D assets. The core insight of this work is to migrate…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Xuying Zhang , Yutong Liu , Yangguang Li , Renrui Zhang , Yufei Liu , Kai Wang , Wanli Ouyang , Zhiwei Xiong , Peng Gao , Qibin Hou , Ming-Ming Cheng

Micro-CT scanning of rocks significantly enhances our understanding of pore-scale physics in porous media. With advancements in pore-scale simulation methods, such as pore network models, it is now possible to accurately simulate multiphase…

图像与视频处理 · 电气工程与系统科学 2024-09-19 Zihan Ren , Sanjay Srinivasan

Reconstructing topologically consistent facial geometry is crucial for the digital avatar creation pipelines. Existing methods either require tedious manual efforts, lack generalization to in-the-wild data, or are constrained by the limited…

计算机视觉与模式识别 · 计算机科学 2025-11-27 Xin Ming , Yuxuan Han , Tianyu Huang , Feng Xu

Fragment-Based Drug Discovery (FBDD) is a popular approach in early drug development, but designing effective linkers to combine disconnected molecular fragments into chemically and pharmacologically viable candidates remains challenging.…

机器学习 · 计算机科学 2025-09-24 Xuefeng Liu , Songhao Jiang , Qinan Huang , Tinson Xu , Ian Foster , Mengdi Wang , Hening Lin , Rick Stevens

Each grid block in a 3D geological model requires a rock type that represents all physical and chemical properties of that block. The properties that classify rock types are lithology, permeability, and capillary pressure. Scientists and…

机器学习 · 计算机科学 2022-01-06 Omar Alfarisi , Djamel Ouzzane , Mohamed Sassi , Tiejun Zhang

The de novo generation of molecules with desirable properties is a critical challenge, where diffusion models are computationally intensive and autoregressive models struggle with error propagation. In this work, we introduce the Graph…

机器学习 · 计算机科学 2025-12-03 Haozhuo Zheng , Cheng Wang , Yang Liu

The reconstruction of three-dimensional dynamic scenes is a well-established yet challenging task within the domain of computer vision. In this paper, we propose a novel approach that combines the domains of 3D geometry reconstruction and…

计算机视觉与模式识别 · 计算机科学 2025-09-11 David Stotko , Reinhard Klein

Note: The final version of this article was published in Computers and Geosciences, Volume 206, January 2026, 106038. DOI: 10.1016/j.cageo.2025.106038. Readers should refer to the published version for the most up-to-date content.…

Reconstructing dynamic 4D scenes is challenging, as it requires robust disentanglement of dynamic objects from the static background. While 3D foundation models like VGGT provide accurate 3D geometry, their performance drops markedly when…

计算机视觉与模式识别 · 计算机科学 2025-11-26 Yu Hu , Chong Cheng , Sicheng Yu , Xiaoyang Guo , Hao Wang

We introduce BrickGPT, the first approach for generating physically stable interconnecting brick assembly models from text prompts. To achieve this, we construct a large-scale, physically stable dataset of brick structures, along with their…

计算机视觉与模式识别 · 计算机科学 2025-11-14 Ava Pun , Kangle Deng , Ruixuan Liu , Deva Ramanan , Changliu Liu , Jun-Yan Zhu

Inverse design of solid-state materials with desired properties represents a formidable challenge in materials science. Although recent generative models have demonstrated potential, their adoption has been hindered by limitations such as…

材料科学 · 物理学 2024-08-15 Yan Chen , Xueru Wang , Xiaobin Deng , Yilun Liu , Xi Chen , Yunwei Zhang , Lei Wang , Hang Xiao

We are witnessing a proliferation of textured 3D models captured from the real world with automatic photo-reconstruction tools. Digital 3D models of this class come with a unique set of characteristics and defects -- especially concerning…

图形学 · 计算机科学 2020-12-29 Andrea Maggiordomo , Federico Ponchio , Paolo Cignoni , Marco Tarini

With the advance of diffusion models, today's video generation has achieved impressive quality. But generating temporal consistent long videos is still challenging. A majority of video diffusion models (VDMs) generate long videos in an…

计算机视觉与模式识别 · 计算机科学 2024-06-18 Kaifeng Gao , Jiaxin Shi , Hanwang Zhang , Chunping Wang , Jun Xiao

Architecture embodies aesthetic, cultural, and historical values, standing as a tangible testament to human civilization. Researchers have long leveraged virtual reality (VR), mixed reality (MR), and augmented reality (AR) to enable…

图形学 · 计算机科学 2025-09-26 Yuze Wang , Luo Yang , Junyi Wang , Yue Qi
‹ 上一页 1 2 3 10 下一页 ›