中文
相关论文

相关论文: DoraemonGPT: Toward Understanding Dynamic Scenes w…

200 篇论文

Conversation agents fueled by Large Language Models (LLMs) are providing a new way to interact with visual data. While there have been initial attempts for image-based conversation models, this work addresses the under-explored field of…

计算机视觉与模式识别 · 计算机科学 2024-06-11 Muhammad Maaz , Hanoona Rasheed , Salman Khan , Fahad Shahbaz Khan

In the pursuit of efficient automated content creation, procedural generation, leveraging modifiable parameters and rule-based systems, emerges as a promising approach. Nonetheless, it could be a demanding endeavor, given its intricate…

计算机视觉与模式识别 · 计算机科学 2024-05-30 Chunyi Sun , Junlin Han , Weijian Deng , Xinlong Wang , Zishan Qin , Stephen Gould

Large language models (LLMs) have revolutionized AI, but are constrained by limited context windows, hindering their utility in tasks like extended conversations and document analysis. To enable using context beyond limited context windows,…

人工智能 · 计算机科学 2024-02-13 Charles Packer , Sarah Wooders , Kevin Lin , Vivian Fang , Shishir G. Patil , Ion Stoica , Joseph E. Gonzalez

Building on the advances of language models, Large Multimodal Models (LMMs) have contributed significant improvements in video understanding. While the current video LMMs utilize advanced Large Language Models (LLMs), they rely on either…

计算机视觉与模式识别 · 计算机科学 2024-06-14 Muhammad Maaz , Hanoona Rasheed , Salman Khan , Fahad Khan

A 3D scene graph represents a compact scene model by capturing both the objects present and the semantic relationships between them, making it a promising structure for robotic applications. To effectively interact with users, an embodied…

计算机视觉与模式识别 · 计算机科学 2025-08-07 Tatiana Zemskova , Dmitry Yudin

Vision language models (VLMs) have experienced rapid advancements through the integration of large language models (LLMs) with image-text pairs, yet they struggle with detailed regional visual understanding due to limited spatial awareness…

计算机视觉与模式识别 · 计算机科学 2024-03-05 Qiushan Guo , Shalini De Mello , Hongxu Yin , Wonmin Byeon , Ka Chun Cheung , Yizhou Yu , Ping Luo , Sifei Liu

Recent research on Large Language Models (LLMs) has led to remarkable advancements in general NLP AI assistants. Some studies have further explored the use of LLMs for planning and invoking models or APIs to address more general multi-modal…

计算机视觉与模式识别 · 计算机科学 2023-06-29 Difei Gao , Lei Ji , Luowei Zhou , Kevin Qinghong Lin , Joya Chen , Zihan Fan , Mike Zheng Shou

Existing MLLMs encounter significant challenges in modeling the temporal context within long videos. Currently, mainstream Agent-based methods use external tools to assist a single MLLM in answering long video questions. Despite such…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Boyu Chen , Zhengrong Yue , Siran Chen , Zikang Wang , Yang Liu , Peng Li , Yali Wang

Large language models (LLMs), such as ChatGPT/GPT-4, have proven to be powerful tools in promoting the user experience as an AI assistant. The continuous works are proposing multi-modal large language models (MLLM), empowering LLMs with the…

计算与语言 · 计算机科学 2023-10-23 Ziqiang Zheng , Jipeng Zhang , Tuan-Anh Vu , Shizhe Diao , Yue Him Wong Tim , Sai-Kit Yeung

World models are progressively being employed across diverse fields, extending from basic environment simulation to complex scenario construction. However, existing models are mainly trained on domain-specific states and actions, and…

人工智能 · 计算机科学 2024-10-01 Zhiqi Ge , Hongzhe Huang , Mingze Zhou , Juncheng Li , Guoming Wang , Siliang Tang , Yueting Zhuang

Trained with an unprecedented scale of data, large language models (LLMs) like ChatGPT and GPT-4 exhibit the emergence of significant reasoning abilities from model scaling. Such a trend underscored the potential of training LLMs with…

计算机视觉与模式识别 · 计算机科学 2023-10-20 Gengze Zhou , Yicong Hong , Qi Wu

Dynamic scenes contain intricate spatio-temporal information, crucial for mobile robots, UAVs, and autonomous driving systems to make informed decisions. Parsing these scenes into semantic triplets <Subject-Predicate-Object> for accurate…

计算机视觉与模式识别 · 计算机科学 2025-05-08 Hang Zhang , Zhuoling Li , Jun Liu

Event cameras record visual information as asynchronous pixel change streams, excelling at scene perception under unsatisfactory lighting or high-dynamic conditions. Existing multimodal large language models (MLLMs) concentrate on natural…

计算机视觉与模式识别 · 计算机科学 2024-12-03 Shaoyu Liu , Jianing Li , Guanghui Zhao , Yunjian Zhang , Xin Meng , Fei Richard Yu , Xiangyang Ji , Ming Li

Existing change detection methods often lack the versatility to handle diverse real-world queries and the intelligence for comprehensive analysis. This paper presents a general agent framework, integrating Large Language Models (LLM) with…

人工智能 · 计算机科学 2026-01-08 Zixuan Xiao , Jun Ma

The evolution of text to visual components facilitates people's daily lives, such as generating image, videos from text and identifying the desired elements within the images. Computer vision models involving the multimodal abilities in the…

计算机视觉与模式识别 · 计算机科学 2024-03-25 Chris Kelly , Luhui Hu , Jiayin Hu , Yu Tian , Deshun Yang , Bang Yang , Cindy Yang , Zihao Li , Zaoshan Huang , Yuexian Zou

World models empower model-based agents to interactively explore, reason, and plan within imagined environments for real-world decision-making. However, the high demand for interactivity poses challenges in harnessing recent advancements in…

计算机视觉与模式识别 · 计算机科学 2024-11-01 Jialong Wu , Shaofeng Yin , Ningya Feng , Xu He , Dong Li , Jianye Hao , Mingsheng Long

Robotic agents must master common sense and long-term sequential decisions to solve daily tasks through natural language instruction. The developments in Large Language Models (LLMs) in natural language processing have inspired efforts to…

机器人学 · 计算机科学 2024-09-16 Yaran Chen , Wenbo Cui , Yuanwen Chen , Mining Tan , Xinyao Zhang , Dongbin Zhao , He Wang

Building models that can understand and reason about 3D scenes is difficult owing to the lack of data sources for 3D supervised training and large-scale training regimes. In this work we ask - How can the knowledge in a pre-trained language…

计算机视觉与模式识别 · 计算机科学 2024-08-14 Shivam Chandhok

The task of estimating the world model describing the dynamics of a real world process assumes immense importance for anticipating and preparing for future outcomes. For applications such as video surveillance, robotics applications,…

计算机视觉与模式识别 · 计算机科学 2025-05-22 Hao Tang , Kevin Ellis , Suhas Lohit , Michael J. Jones , Moitreya Chatterjee

We are interested in enabling visual planning for complex long-horizon tasks in the space of generated videos and language, leveraging recent advances in large generative models pretrained on Internet-scale data. To this end, we present…

‹ 上一页 1 2 3 10 下一页 ›