English
Related papers

Related papers: MultiSurf-GPT: Facilitating Context-Aware Reasonin…

200 papers

In-context learning (ICL) involves reasoning from given contextual examples. As more modalities comes, this procedure is becoming more challenging as the interleaved input modalities convolutes the understanding process. This is exemplified…

Computer Vision and Pattern Recognition · Computer Science 2023-10-10 Yixin Chen , Shuai Zhang , Boran Han , Jiaya Jia

Accurate point tracking in surgical environments remains challenging due to complex visual conditions, including smoke occlusion, specular reflections, and tissue deformation. While existing surgical tracking datasets provide coordinate…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Rulin Zhou , Wenlong He , An Wang , Jianhang Zhang , Xuanhui Zeng , Xi Zhang , Chaowei Zhu , Haijun Hu , Hongliang Ren

Multi-spectral imagery plays a crucial role in diverse Remote Sensing applications including land-use classification, environmental monitoring and urban planning. These images are widely adopted because their additional spectral bands…

Computer Vision and Pattern Recognition · Computer Science 2025-09-24 Ganesh Mallya , Yotam Gigi , Dahun Kim , Maxim Neumann , Genady Beryozkin , Tomer Shekel , Anelia Angelova

Few-shot molecular property prediction (FSMPP) is essential in drug discovery and materials design, where high-quality labeled data are often scarce and expensive to obtain. Despite the promising performance of existing methods, especially…

Computational Engineering, Finance, and Science · Computer Science 2026-05-14 Zeyu Wang , Xin Zheng , Yao Lu , Shanqing Yu , Qi Xuan , Shirui Pan

With the rapid advancement of large language models (LLMs), intelligent conversational assistants have demonstrated remarkable capabilities across various domains. However, they still mainly rely on explicit textual input and do not know…

Human-Computer Interaction · Computer Science 2025-12-29 Ziyan Zhang , Nan Gao , Zhiqiang Nie , Shantanu Pal , Haining Zhang

The demand of applying semantic segmentation model on mobile devices has been increasing rapidly. Current state-of-the-art networks have enormous amount of parameters hence unsuitable for mobile devices, while other small memory footprint…

Computer Vision and Pattern Recognition · Computer Science 2019-04-15 Tianyi Wu , Sheng Tang , Rui Zhang , Yongdong Zhang

Efficiently capturing consistent and complementary semantic features in a multimodal conversation context is crucial for Multimodal Emotion Recognition in Conversation (MERC). Existing methods mainly use graph structures to model dialogue…

Computation and Language · Computer Science 2024-05-06 Tao Meng , Fuchen Zhang , Yuntao Shou , Wei Ai , Nan Yin , Keqin Li

Maritime Multi-Scene Recognition is crucial for enhancing the capabilities of intelligent marine robotics, particularly in applications such as marine conservation, environmental monitoring, and disaster response. However, this task…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Xinyu Xi , Hua Yang , Shentai Zhang , Yijie Liu , Sijin Sun , Xiuju Fu

Multimodal Large Language Models (MLLMs) exhibit impressive capabilities across a variety of tasks, especially when equipped with carefully designed visual prompts. However, existing studies primarily focus on logical reasoning and visual…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Dingning Liu , Cheng Wang , Peng Gao , Renrui Zhang , Xinzhu Ma , Yuan Meng , Zhihui Wang

Predicting the future location of mobile objects reinforces location-aware services with proactive intelligence and helps businesses and decision-makers with better planning and near real-time scheduling in different applications such as…

Physical environment understanding is vital in delivering immersive and interactive mobile augmented reality (AR) user experiences. Recently, we have witnessed a transition in the design of environment understanding systems, from visual…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-10-18 Yiqin Zhao , Ashkan Ganj , Tian Guo

Because of the growing interest for mobile device and pervasive applications deployed on cloud computing, the providing of intelligent and ubiquitous context-aware applications that take into account the user's context is one of the main…

Software Engineering · Computer Science 2021-04-05 Asmae Benali , Bouchra El Asri , Houda Kriouile

Multimodal few-shot learning is challenging due to the large domain gap between vision and language modalities. Existing methods are trying to communicate visual concepts as prompts to frozen language models, but rely on hand-engineered…

Computer Vision and Pattern Recognition · Computer Science 2023-03-01 Ivona Najdenkoska , Xiantong Zhen , Marcel Worring

Spatial synchronization in roadside scenarios is essential for integrating data from multiple sensors at different locations. Current methods using cascading spatial transformation (CST) often lead to cumulative errors in large-scale…

Signal Processing · Electrical Eng. & Systems 2023-11-09 Yong Li , Zhiguo Zhao , Yunli Chen , Rui Tian

Language-driven grasp detection has the potential to revolutionize human-robot interaction by allowing robots to understand and execute grasping tasks based on natural language commands. However, existing approaches face two key challenges.…

Robotics · Computer Science 2025-07-22 Quang Nguyen , Tri Le , Huy Nguyen , Thieu Vo , Tung D. Ta , Baoru Huang , Minh N. Vu , Anh Nguyen

Context information brings new opportunities for efficient and effective applications and services on mobile devices. A wide range of research has exploited context dependency, i.e., the relations between context(s) and the outcome, to…

Human-Computer Interaction · Computer Science 2012-09-11 Ahmad Rahmati , Clayton Shepard , Chad Tossell , Lin Zhong , Philip Kortum

Imagine interconnected objects with embedded artificial intelligence (AI), empowered to sense the environment, see it, hear it, touch it, interact with it, and move. As future networks of intelligent objects come to life, tremendous new…

In the context of Synthetic Aperture Radar (SAR) image recognition, traditional methods often struggle with the intrinsic limitations of SAR data, such as weak texture, high noise, and ambiguous object boundaries. This work explores a novel…

Signal Processing · Electrical Eng. & Systems 2025-07-15 Chaoran Li , Xingguo Xu , Siyuan Mu

With the development of foundation models such as large language models, zero-shot transfer learning has become increasingly significant. This is highlighted by the generative capabilities of NLP models like GPT-4, and the retrieval-based…

Machine Learning · Computer Science 2024-06-25 Yuhan Li , Peisong Wang , Zhixun Li , Jeffrey Xu Yu , Jia Li

This study investigates the potential of a multimodal large language model (LLM), specifically ChatGPT-4o, to perform human-like interpretations of traffic scenes using static dashcam images. Herein, we focus on three judgment tasks…

Computer Vision and Pattern Recognition · Computer Science 2025-07-14 Yuki Yoshihara , Linjing Jiang , Nihan Karatas , Hitoshi Kanamori , Asuka Harada , Takahiro Tanaka
‹ Prev 1 4 5 6 7 8 10 Next ›