English
Related papers

Related papers: WorldAct: Activating Monolithic 3D Worlds into Int…

200 papers

Despite the success achieved by existing image generation and editing methods, current models still struggle with complex problems including intricate text prompts, and the absence of verification and self-correction mechanisms makes the…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Zhenyu Wang , Aoxue Li , Zhenguo Li , Xihui Liu

Synthetic data has been a critical tool for training scene text detection and recognition models. On the one hand, synthetic word images have proven to be a successful substitute for real images in training scene text recognizers. On the…

Computer Vision and Pattern Recognition · Computer Science 2020-08-19 Shangbang Long , Cong Yao

We present Dynamic ReAct, a novel approach for enabling ReAct agents to efficiently operate with extensive Model Control Protocol (MCP) tool sets that exceed the contextual memory limitations of large language models. Our approach addresses…

Software Engineering · Computer Science 2025-09-29 Nishant Gaurav , Adit Akarsh , Ankit Ranjan , Manoj Bajaj

Generative models trained on internet data have revolutionized how text, image, and video content can be created. Perhaps the next milestone for generative models is to simulate realistic experience in response to actions taken by humans,…

Artificial Intelligence · Computer Science 2024-09-27 Sherry Yang , Yilun Du , Kamyar Ghasemipour , Jonathan Tompson , Leslie Kaelbling , Dale Schuurmans , Pieter Abbeel

Common knowledge indicates that the process of constructing image datasets usually depends on the time-intensive and inefficient method of manual collection and annotation. Large models offer a solution via data generation. Nonetheless,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-14 Haoran Sun , Haoyu Bian , Shaoning Zeng , Yunbo Rao , Xu Xu , Lin Mei , Jianping Gou

Real-world data collection for embodied agents remains costly and unsafe, calling for scalable, realistic, and simulator-ready 3D environments. However, existing scene-generation systems often rely on rule-based or task-specific pipelines,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Hongchi Xia , Xuan Li , Zhaoshuo Li , Qianli Ma , Jiashu Xu , Ming-Yu Liu , Yin Cui , Tsung-Yi Lin , Wei-Chiu Ma , Shenlong Wang , Shuran Song , Fangyin Wei

Despite the remarkable progress of Multimodal Large Language Models (MLLMs) in 2D vision-language tasks, their application to complex 3D scene manipulation remains underexplored. In this paper, we bridge this critical gap by tackling three…

Computer Vision and Pattern Recognition · Computer Science 2026-01-08 Zhengfei Kuang , Rui Lin , Long Zhao , Gordon Wetzstein , Saining Xie , Sanghyun Woo

Understanding 3D scenes in open-world settings poses fundamental challenges for vision and robotics, particularly due to the limitations of closed-vocabulary supervision and static annotations. To address this, we propose a unified…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Fei Yu , Quan Deng , Shengeng Tang , Yuehua Li , Lechao Cheng

We introduce Drag4D, an interactive framework that integrates object motion control within text-driven 3D scene generation. This framework enables users to define 3D trajectories for the 3D objects generated from a single image, seamlessly…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Minjun Kang , Inkyu Shin , Taeyeop Lee , In So Kweon , Kuk-Jin Yoon

We introduce ReplaceAnything3D model (RAM3D), a novel text-guided 3D scene editing method that enables the replacement of specific objects within a scene. Given multi-view images of a scene, a text prompt describing the object to replace,…

Computer Vision and Pattern Recognition · Computer Science 2024-02-01 Edward Bartrum , Thu Nguyen-Phuoc , Chris Xie , Zhengqin Li , Numair Khan , Armen Avetisyan , Douglas Lanman , Lei Xiao

With the development of deep neural networks, the demand for a significant amount of annotated training data becomes the performance bottlenecks in many fields of research and applications. Image synthesis can generate annotated images…

Computer Vision and Pattern Recognition · Computer Science 2020-12-24 Minghui Liao , Boyu Song , Shangbang Long , Minghang He , Cong Yao , Xiang Bai

Interactions play a key role in understanding objects and scenes, for both virtual and real world agents. We introduce a new general representation for proximal interactions among physical objects that is agnostic to the type of objects or…

Recent work in 3D scene understanding is moving beyond purely spatial analysis toward functional scene understanding. However, existing methods often consider functional relationships between object pairs in isolation, failing to capture…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Zhengyu Fu , René Zurbrügg , Kaixian Qu , Marc Pollefeys , Marco Hutter , Hermann Blum , Zuria Bauer

Digital storytelling, essential in entertainment, education, and marketing, faces challenges in production scalability and flexibility. The StoryAgent framework, introduced in this paper, utilizes Large Language Models and generative tools…

Computation and Language · Computer Science 2024-06-24 Samuel S. Sohn , Danrui Li , Sen Zhang , Che-Jui Chang , Mubbasir Kapadia

Recent advances in video diffusion transformers have enabled interactive gaming world models that allow users to explore generated environments over extended horizons. However, existing approaches struggle with precise action control and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Jisu Nam , Yicong Hong , Chun-Hao Paul Huang , Feng Liu , JoungBin Lee , Jiyoung Kim , Siyoon Jin , Yunsung Lee , Jaeyoon Jung , Suhwan Choi , Seungryong Kim , Yang Zhou

Recent advancements in object-centric text-to-3D generation have shown impressive results. However, generating complex 3D scenes remains an open challenge due to the intricate relations between objects. Moreover, existing methods are…

Computer Vision and Pattern Recognition · Computer Science 2024-12-31 Yu-Hsiang Huang , Wei Wang , Sheng-Yu Huang , Yu-Chiang Frank Wang

Recent advances in vision-language models have enabled mobile GUI agents to perceive visual interfaces and execute user instructions, but reliable prediction of action consequences remains critical for long-horizon and high-risk…

Artificial Intelligence · Computer Science 2026-05-25 Weikai Xu , Kun Huang , Yunren Feng , Jiaxing Li , Yuhan Chen , Yuxuan Liu , Zhizheng Jiang , Heng Qu , Pengzhi Gao , Wei Liu , Jian Luan , Xiaolin Hu , Bo An

This paper presents GaussEdit, a framework for adaptive 3D scene editing guided by text and image prompts. GaussEdit leverages 3D Gaussian Splatting as its backbone for scene representation, enabling convenient Region of Interest selection…

Graphics · Computer Science 2025-10-01 Zhenyu Shu , Junlong Yu , Kai Chao , Shiqing Xin , Ligang Liu

Semantics has enabled 3D scene understanding and affordance-driven object interaction. However, robots operating in real-world environments face a critical limitation: they cannot anticipate how objects move. Long-horizon mobile…

Video generation has achieved impressive quality, but it still suffers from artifacts such as temporal inconsistency and violation of physical laws. Leveraging 3D scenes can fundamentally resolve these issues by providing precise control…

Computer Vision and Pattern Recognition · Computer Science 2024-12-02 Zhaofang Qian , Abolfazl Sharifi , Tucker Carroll , Ser-Nam Lim