English
Related papers

Related papers: Unify-Agent: A Unified Multimodal Agent for World-…

200 papers

Achieving general-purpose robotic manipulation requires robots to seamlessly bridge high-level semantic intent with low-level physical interaction in unstructured environments. However, existing approaches falter in zero-shot…

Robotics · Computer Science 2026-02-16 Haichao Liu , Yuanjiang Xue , Yuheng Zhou , Haoyuan Deng , Yinan Liang , Lihua Xie , Ziwei Wang

Advancing complex reasoning in large language models relies on high-quality, verifiable datasets, yet human annotation remains cost-prohibitive and difficult to scale. Current synthesis paradigms often face a recurring trade-off:…

Artificial Intelligence · Computer Science 2026-02-04 Zhengbo Jiao , Shaobo Wang , Zifan Zhang , Xuan Ren , Wei Wang , Bing Zhao , Hu Wei , Linfeng Zhang

In recent years, significant progress has been made in both image generation and generated image detection. Despite their rapid, yet largely independent, development, these two fields have evolved distinct architectural paradigms: the…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Yanran Zhang , Wenzhao Zheng , Yifei Li , Bingyao Yu , Yu Zheng , Lei Chen , Jiwen Lu , Jie Zhou

In this paper, we investigate the problem of embodied multi-agent cooperation, where decentralized agents must cooperate given only egocentric views of the world. To effectively plan in this setting, in contrast to learning world dynamics…

Computer Vision and Pattern Recognition · Computer Science 2025-04-17 Hongxin Zhang , Zeyuan Wang , Qiushi Lyu , Zheyuan Zhang , Sunli Chen , Tianmin Shu , Behzad Dariush , Kwonjoon Lee , Yilun Du , Chuang Gan

The rapid progress in embodied artificial intelligence has highlighted the necessity for more advanced and integrated models that can perceive, interpret, and predict environmental dynamics. In this context, World Models (WMs) have been…

This paper describes our research on AI agents embodied in visual, virtual or physical forms, enabling them to interact with both users and their environments. These agents, which include virtual avatars, wearable devices, and robots, are…

Recent advances in artificial intelligence have been driven by the presence of increasingly realistic and complex simulated environments. However, many of the existing environments provide either unrealistic visuals, inaccurate physics, low…

Recent advances in human preference alignment have significantly improved multimodal generation and understanding. A key approach is to train reward models that provide supervision signals for preference optimization. However, existing…

Computer Vision and Pattern Recognition · Computer Science 2026-02-26 Yibin Wang , Yuhang Zang , Hao Li , Cheng Jin , Jiaqi Wang

The fashion domain encompasses a variety of real-world multimodal tasks, including multimodal retrieval and multimodal generation. The rapid advancements in artificial intelligence generated content, particularly in technologies like large…

Computer Vision and Pattern Recognition · Computer Science 2024-10-15 Xiangyu Zhao , Yuehan Zhang , Wenlong Zhang , Xiao-Ming Wu

Unified multimodal models (UMMs) have emerged as a powerful paradigm in fundamental cross-modality research, demonstrating significant potential in both image understanding and generation. However, existing research in the face domain…

Computer Vision and Pattern Recognition · Computer Science 2026-01-14 Junzhe Li , Sifan Zhou , Liya Guo , Xuerui Qiu , Linrui Xu , Delin Qu , Tingting Long , Chun Fan , Ming Li , Hehe Fan , Jun Liu , Shuicheng Yan

Enabling embodied agents to imagine future states is essential for robust and generalizable visual navigation. Yet, state-of-the-art systems typically rely on modular designs that decouple navigation planning from visual world modeling,…

Artificial Intelligence · Computer Science 2026-03-24 Yifei Dong , Fengyi Wu , Guangyu Chen , Lingdong Kong , Xu Zhu , Qiyu Hu , Yuxuan Zhou , Jingdong Sun , Jun-Yan He , Qi Dai , Alexander G. Hauptmann , Zhi-Qi Cheng

With the recent fast development of generative models, instruction-based image editing has shown great potential in generating high-quality images. However, the quality of editing highly depends on carefully designed instructions, placing…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Mingde Yao , Zhiyuan You , King-Man Tam , Menglu Wang , Tianfan Xue

Multimodal fake news detection typically demands complex architectures and substantial computational resources, posing deployment challenges in real-world settings. We introduce UNITE-FND, a novel framework that reframes multimodal fake…

Machine Learning · Computer Science 2025-02-18 Arka Mukherjee , Shreya Ghosh

Real-world multimodal agents solve multi-step workflows grounded in visual evidence. For example, an agent can troubleshoot a device by linking a wiring photo to a schematic and validating the fix with online documentation, or plan a trip…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Zhaochen Su , Jincheng Gao , Hangyu Guo , Zhenhua Liu , Lueyang Zhang , Xinyu Geng , Shijue Huang , Peng Xia , Guanyu Jiang , Cheng Wang , Yue Zhang , Yi R. Fung , Junxian He

Despite recent advances in diffusion models, AI generated images still often contain visual artifacts that compromise realism. Although more thorough pre-training and bigger models might reduce artifacts, there is no assurance that they can…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Jaehyun Park , Minyoung Ahn , Minkyu Kim , Jonghyun Lee , Jae-Gil Lee , Dongmin Park

Visual compliance verification is a critical yet underexplored problem in computer vision, especially in domains such as media, entertainment, and advertising where content must adhere to complex and evolving policy rules. Existing methods…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Rahul Ghosh , Baishali Chaudhury , Hari Prasanna Das , Meghana Ashok , Ryan Razkenari , Long Chen , Sungmin Hong , Chun-Hao Liu

We introduce a multicrossmodal LLM-agent framework motivated by the growing volume and diversity of materials-science data ranging from high-resolution microscopy and dynamic simulation videos to tabular experiment logs and sprawling…

Materials Science · Physics 2025-05-22 Adib Bazgir , Rama chandra Praneeth Madugula , Yuwen Zhang

Deep learning techniques have become widely utilized in histopathology image classification due to their superior performance. However, this success heavily relies on the availability of substantial labeled data, which necessitates…

Image and Video Processing · Electrical Eng. & Systems 2024-10-15 Meng Li , Chaoyi Li , Can Peng , Brian C. Lovell

Recent advances in large language model (LLM) have empowered autonomous agents to perform multi-turn interactions with tools and environments. However, scaling such agent training is limited by the lack of diverse and reliable environments.…

Artificial Intelligence · Computer Science 2026-05-26 Zhaoyang Wang , Canwen Xu , Boyi Liu , Yite Wang , Siwei Han , Zhewei Yao , Huaxiu Yao , Yuxiong He

Image reconstruction and image synthesis are important for handling incomplete multimodal imaging data, but existing methods require various task-specific models, complicating training and deployment workflows. We introduce Any2all, a…

Image and Video Processing · Electrical Eng. & Systems 2026-02-10 Weijie Gan , Xucheng Wang , Tongyao Wang , Wenshang Wang , Chunwei Ying , Yuyang Hu , Yasheng Chen , Hongyu An , Ulugbek S. Kamilov
‹ Prev 1 4 5 6 7 8 10 Next ›