中文
相关论文

相关论文: ImageChain: Advancing Sequential Image-to-Text Rea…

200 篇论文

A picture is worth a thousand words, thus, it is crucial for conversational agents to understand, perceive, and effectively respond with pictures. However, we find that directly employing conventional image generation techniques is…

计算与语言 · 计算机科学 2024-02-09 Xiaowen Sun , Jiazhan Feng , Yuxuan Wang , Yuxuan Lai , Xingyu Shen , Dongyan Zhao

Visual storytelling aims to generate human-level narrative language (i.e., a natural paragraph with multiple sentences) from a photo streams. A typical photo story consists of a global timeline with multi-thread local storylines, where each…

计算机视觉与模式识别 · 计算机科学 2016-06-03 Yu Liu , Jianlong Fu , Tao Mei , Chang Wen Chen

In this paper, we advance the study of AI-augmented reasoning in the context of Human-Computer Interaction (HCI), psychology and cognitive science, focusing on the critical task of visual perception. Specifically, we investigate the…

人机交互 · 计算机科学 2025-04-18 Shravan Chaudhari , Trilokya Akula , Yoon Kim , Tom Blake

Large language models (LLMs) are increasingly applied to sequential decision-making through in-context learning (ICL), yet their effectiveness is highly sensitive to prompt quality. Effective prompts should meet three principles: focus on…

人工智能 · 计算机科学 2025-11-19 Ruomeng Ding , Wei Cheng , Minglai Shao , Chen Zhao

We introduce OmChat, a model designed to excel in handling long contexts and video understanding tasks. OmChat's new architecture standardizes how different visual inputs are processed, making it more efficient and adaptable. It uses a…

计算机视觉与模式识别 · 计算机科学 2024-07-09 Tiancheng Zhao , Qianqian Zhang , Kyusong Lee , Peng Liu , Lu Zhang , Chunxin Fang , Jiajia Liao , Kelei Jiang , Yibo Ma , Ruochen Xu

We propose a novel cognitively-inspired method to improve and interpret physical simulation in vision-language models. Our ``Chain of Time" method involves generating a series of intermediate images during a simulation, and it is motivated…

计算机视觉与模式识别 · 计算机科学 2025-11-04 YingQiao Wang , Eric Bigelow , Boyi Li , Tomer Ullman

Text-to-image models are powerful for producing high-quality images based on given text prompts, but crafting these prompts often requires specialized vocabulary. To address this, existing methods train rewriting models with supervision…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Hongji Yang , Yucheng Zhou , Wencheng Han , Jianbing Shen

This paper delves into the capabilities of large language models (LLMs), specifically focusing on advancing the theoretical comprehension of chain-of-thought prompting. We investigate how LLMs can be effectively induced to generate a…

计算与语言 · 计算机科学 2024-06-07 Rasul Tutunov , Antoine Grosnit , Juliusz Ziomek , Jun Wang , Haitham Bou-Ammar

Compressing long chains of thought (CoT) into compact latent tokens is crucial for efficient reasoning with large language models (LLMs). Recent studies employ autoencoders to achieve this by reconstructing textual CoT from latent tokens,…

计算机视觉与模式识别 · 计算机科学 2026-02-02 Xiaoshu Chen , Sihang Zhou , Ke Liang , Taichun Zhou , Xinwang Liu

End-to-end text-image machine translation (TIMT), which directly translates textual content in images across languages, is crucial for real-world multilingual scene understanding. Despite advances in vision-language large models (VLLMs),…

计算机视觉与模式识别 · 计算机科学 2026-03-26 Gengluo Li , Chengquan Zhang , Yupu Liang , Huawen Shen , Yaping Zhang , Pengyuan Lyu , Weinong Wang , Xingyu Wan , Gangyan Zeng , Han Hu , Can Ma , Yu Zhou

Recent advances in multimodal reasoning models have demonstrated impressive capabilities across text and vision. However, even leading models exhibit redundant self-reflection when generating lengthy reasoning chains. While training-free…

计算机视觉与模式识别 · 计算机科学 2025-11-24 Yuan Zhang , Ming Lu , Junwen Pan , Tao Huang , Kuan Cheng , Qi She , Shanghang Zhang

Although large language models (LLMs) have demonstrated impressive potential on simple tasks, their breadth of scope, lack of transparency, and insufficient controllability can make them less effective when assisting humans on more complex…

人机交互 · 计算机科学 2022-03-21 Tongshuang Wu , Michael Terry , Carrie J. Cai

While sequential reasoning enhances the capability of Vision-Language Models (VLMs) to execute complex multimodal tasks, their reliability in grounding these reasoning chains within actual visual evidence remains insufficiently explored. We…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Rory Driscoll , Alexandros Christoforos , Chadbourne Davis

While multimodal large language models (MLLMs) exhibit strong performance on single-video tasks (e.g., video question answering), their capability for spatiotemporal pattern reasoning across multiple videos remains a critical gap in pattern…

计算机视觉与模式识别 · 计算机科学 2026-01-07 Nannan Zhu , Yonghao Dong , Teng Wang , Xueqian Li , Shengjun Deng , Yijia Wang , Zheng Hong , Tiantian Geng , Guo Niu , Hanyan Huang , Xiongfei Yao , Shuaiwei Jiao

We explore the use of large language models (LLMs) for next-utterance anticipation in human dialogue. Despite recent advances in LLMs demonstrating their ability to engage in natural conversations with users, we show that even leading…

人工智能 · 计算机科学 2026-05-12 Yueyi Yang , Haotian Liu , Fang Kang , Mengqi Zhang , Zheng Lian , Hao Tang , Haoyu Chen

In an era defined by the explosive growth of data and rapid technological advancements, Multimodal Large Language Models (MLLMs) stand at the forefront of artificial intelligence (AI) systems. Designed to seamlessly integrate diverse data…

Multimodal large language models (MLLMs) can process text presented as images, yet they often perform worse than when the same content is provided as textual tokens. We systematically diagnose this "modality gap" by evaluating seven MLLMs…

计算与语言 · 计算机科学 2026-05-26 Kaiser Sun , Xiaochuang Yuan , Hongjun Liu , Chen Zhao , Cheng Zhang , Mark Dredze , Fan Bai

Video captioning is a critical task in the field of multimodal machine learning, aiming to generate descriptive and coherent textual narratives for video content. While large vision-language models (LVLMs) have shown significant progress,…

计算机视觉与模式识别 · 计算机科学 2024-12-17 Ji-jun Park , Soo-joon Choi

Recently, Multimodal Large Language Models (MLLMs) have achieved exceptional performance across diverse tasks, continually surpassing previous expectations regarding their capabilities. Nevertheless, their proficiency in perceiving emotions…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Daiqing Wu , Dongbao Yang , Sicheng Zhao , Can Ma , Yu Zhou

Multimodal Large Language Models (MLLMs) strive to achieve a profound, human-like understanding of and interaction with the physical world, but often exhibit a shallow and incoherent integration when acquiring information (Perception) and…

‹ 上一页 1 8 9 10 下一页 ›