中文
相关论文

相关论文: SceneAR: Scene-based Micro Narratives for Sharing …

200 篇论文

The techniques for 3D indoor scene capturing are widely used, but the meshes produced leave much to be desired. In this paper, we propose "RoomDreamer", which leverages powerful natural language to synthesize a new room with a different…

计算机视觉与模式识别 · 计算机科学 2023-05-22 Liangchen Song , Liangliang Cao , Hongyu Xu , Kai Kang , Feng Tang , Junsong Yuan , Yang Zhao

Storyboarding is widely used for creating 3D animations. Animators use the 2D sketches in storyboards as references to craft the desired 3D animations through a trial-and-error process. The traditional approach requires exceptional…

图形学 · 计算机科学 2025-04-29 Lei Zhong , Chuan Guo , Yiming Xie , Jiawei Wang , Changjian Li

Recent work on image content manipulation based on vision-language pre-training models has been effectively extended to text-driven 3D scene editing. However, existing schemes for 3D scene editing still exhibit certain shortcomings,…

计算机视觉与模式识别 · 计算机科学 2024-07-11 Shuangkang Fang , Yufeng Wang , Yi-Hsuan Tsai , Yi Yang , Wenrui Ding , Shuchang Zhou , Ming-Hsuan Yang

Site-specific outdoor AR experiences are typically authored using static 3D models, but are deployed in physical environments that change over time. As a result, virtual content may become misaligned with its intended real-world referents,…

人机交互 · 计算机科学 2025-08-12 Nels Numan , Jessica Van Brummelen , Ziwen Lu , Anthony Steed

Promising performance has been achieved for visual perception on the point cloud. However, the current methods typically rely on labour-extensive annotations on the scene scans. In this paper, we explore how synthetic models alleviate the…

计算机视觉与模式识别 · 计算机科学 2022-03-22 Runnan Chen , Xinge Zhu , Nenglun Chen , Dawei Wang , Wei Li , Yuexin Ma , Ruigang Yang , Wenping Wang

We introduce a new task called Multimodal Named Entity Recognition (MNER) for noisy user-generated data such as tweets or Snapchat captions, which comprise short text with accompanying images. These social media posts often come in…

计算与语言 · 计算机科学 2018-02-23 Seungwhan Moon , Leonardo Neves , Vitor Carvalho

Computation capabilities of recent mobile devices enable natural feature processing for Augmented Reality (AR). However, mobile AR applications are still faced with scalability and performance challenges. In this paper, we propose CloudAR,…

多媒体 · 计算机科学 2018-05-09 Wenxiao Zhang , Sikun Lin , Farshid Hassani Bijarbooneh , Hao Fei Cheng , And Pan Hui

In this paper, we propose Text2Scene, a model that generates various forms of compositional scene representations from natural language descriptions. Unlike recent works, our method does NOT use Generative Adversarial Networks (GANs).…

计算机视觉与模式识别 · 计算机科学 2019-06-11 Fuwen Tan , Song Feng , Vicente Ordonez

Bridging the physical and digital world through interaction remains a core challenge in augmented reality (AR). Existing systems target single objects, limiting support for planning, comparison, and assembly tasks that depend on…

Tuning-free personalized image generation methods have achieved significant success in maintaining facial consistency, i.e., identities, even with multiple characters. However, the lack of holistic consistency in scenes with multiple…

计算机视觉与模式识别 · 计算机科学 2024-09-20 Zhengguang Zhou , Jing Li , Huaxia Li , Nemo Chen , Xu Tang

The mainstream image captioning models rely on Convolutional Neural Network (CNN) image features to generate captions via recurrent models. Recently, image scene graphs have been used to augment captioning models so as to leverage their…

计算机视觉与模式识别 · 计算机科学 2021-08-18 Kien Nguyen , Subarna Tripathi , Bang Du , Tanaya Guha , Truong Q. Nguyen

This paper introduces M2AR, a new web-based, two- and three-dimensional modeling environment that enables the modeling and execution of augmented reality applications without requiring programming knowledge. The platform is based on a 3D…

人机交互 · 计算机科学 2024-10-08 Fabian Muff , Hans-Georg Fill

Augmented reality (AR) requires the seamless integration of visual, auditory, and linguistic channels for optimized human-computer interaction. While auditory and visual inputs facilitate real-time and contextual user guidance, the…

计算与语言 · 计算机科学 2023-10-19 Jing Bi , Nguyen Manh Nguyen , Ali Vosoughi , Chenliang Xu

Connecting current observations with prior experiences helps robots adapt and plan in new, unseen 3D environments. Recently, 3D scene analogies have been proposed to connect two 3D scenes, which are smooth maps that align scene regions with…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Junho Kim , Young Min Kim

We present MS2Mesh-XR, a novel multi-modal sketch-to-mesh generation pipeline that enables users to create realistic 3D objects in extended reality (XR) environments using hand-drawn sketches assisted by voice inputs. In specific, users can…

计算机视觉与模式识别 · 计算机科学 2024-12-13 Yuqi Tong , Yue Qiu , Ruiyang Li , Shi Qiu , Pheng-Ann Heng

Scene, as the crucial unit of storytelling in movies, contains complex activities of actors and their interactions in a physical environment. Identifying the composition of scenes serves as a critical step towards semantic understanding of…

计算机视觉与模式识别 · 计算机科学 2020-04-29 Anyi Rao , Linning Xu , Yu Xiong , Guodong Xu , Qingqiu Huang , Bolei Zhou , Dahua Lin

Visually exploring in a real-world 4D spatiotemporal space freely in VR has been a long-term quest. The task is especially appealing when only a few or even single RGB cameras are used for capturing the dynamic scene. To this end, we…

计算机视觉与模式识别 · 计算机科学 2023-02-21 Liangchen Song , Anpei Chen , Zhong Li , Zhang Chen , Lele Chen , Junsong Yuan , Yi Xu , Andreas Geiger

Text-to-image models are becoming increasingly popular, revolutionizing the landscape of digital art creation by enabling highly detailed and creative visual content generation. These models have been widely employed across various domains,…

计算机视觉与模式识别 · 计算机科学 2025-01-09 Matthew Zheng , Enis Simsar , Hidir Yesiltepe , Federico Tombari , Joel Simon , Pinar Yanardag

We address the task of indoor scene generation by generating a sequence of objects, along with their locations and orientations conditioned on a room layout. Large-scale indoor scene datasets allow us to extract patterns from user-designed…

计算机视觉与模式识别 · 计算机科学 2021-04-05 Xinpeng Wang , Chandan Yeshwanth , Matthias Nießner

A method for generating narratives by analyzing single images or image sequences is presented, inspired by the time immemorial tradition of Narrative Art. The proposed method explores the multimodal capabilities of GPT-4o to interpret…

计算与语言 · 计算机科学 2024-08-22 Edirlei Soares de Lima , Marco A. Casanova , Antonio L. Furtado