English
Related papers

Related papers: Multi-modal user interface control detection using…

200 papers

Multi-modal 3D object detection has been an active research topic in autonomous driving. Nevertheless, it is non-trivial to explore the cross-modal feature fusion between sparse 3D points and dense 2D pixels. Recent approaches either fuse…

Computer Vision and Pattern Recognition · Computer Science 2022-10-19 Xin Li , Botian Shi , Yuenan Hou , Xingjiao Wu , Tianlong Ma , Yikang Li , Liang He

Object detection is a critical problem for the safe interaction between autonomous vehicles and road users. Deep-learning methodologies allowed the development of object detection approaches with better performance. However, there is still…

Computer Vision and Pattern Recognition · Computer Science 2021-07-28 Andrés Gómez , Thomas Genevois , Jerome Lussereau , Christian Laugier

Graphical User Interface (GUI) provides visual bridges between software apps and end users. However, due to the compatibility of software or hardware, UI display issues such as text overlap, blurred screen, image missing always occur during…

Software Engineering · Computer Science 2021-07-07 Yuhui Su , Zhe Liu , Chunyang Chen , Junjie Wang , Qing Wang

Combining LiDAR and camera data has shown potential in enhancing short-distance object detection in autonomous driving systems. Yet, the fusion encounters difficulties with extended distance detection due to the contrast between LiDAR's…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Tanmoy Dam , Sanjay Bhargav Dharavath , Sameer Alam , Nimrod Lilith , Supriyo Chakraborty , Mir Feroskhan

Visual In-Context Learning (VICL) has emerged as a powerful paradigm, enabling models to perform novel visual tasks by learning from in-context examples. The dominant "retrieve-then-prompt" approach typically relies on selecting the single…

Computer Vision and Pattern Recognition · Computer Science 2026-01-16 Wenwen Liao , Jianbo Yu , Yuansong Wang , Qingchao Jiang , Xiaofeng Yang

Humans usually perceive the world in a multimodal way that vision, touch, sound are utilised to understand surroundings from various dimensions. These senses are combined together to achieve a synergistic effect where the learning is more…

Computer Vision and Pattern Recognition · Computer Science 2021-12-30 Guanqun Cao , Shan Luo

Multimodal camera-LiDAR fusion technology has found extensive application in 3D object detection, demonstrating encouraging performance. However, existing methods exhibit significant performance degradation in challenging scenarios…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Sixian Liu , Chen Xu , Qiang Wang , Donghai Shi , Yiwen Li

This paper presents the design, implementation, and evolution of a comprehensive multimodal room-monitoring system that integrates synchronized video and audio processing for real-time activity recognition and anomaly detection. We describe…

Sound · Computer Science 2025-11-25 Aman Verma , Keshav Samdani , Mohd. Samiuddin Shafi

As mobile computing technology rapidly evolves, deploying efficient object detection algorithms on mobile devices emerges as a pivotal research area in computer vision. This study zeroes in on optimizing the YOLOv7 algorithm to boost its…

Computer Vision and Pattern Recognition · Computer Science 2024-03-05 Wenkai Gong

The existing state-of-the-art point descriptor relies on structure information only, which omit the texture information. However, texture information is crucial for our humans to distinguish a scene part. Moreover, the current…

Computer Vision and Pattern Recognition · Computer Science 2021-11-19 Xiaoshui Huang , Wentao Qu , Yifan Zuo , Yuming Fang , Xiaowei Zhao

Multi-modal salient object detection (MSOD) aims to boost saliency detection performance by integrating visible sources with depth or thermal infrared ones. Existing methods generally design different fusion schemes to handle certain issues…

Computer Vision and Pattern Recognition · Computer Science 2024-06-04 Kunpeng Wang , Zhengzheng Tu , Chenglong Li , Cheng Zhang , Bin Luo

Autonomous graphical user interface (GUI) agents aim to facilitate task automation by interacting with the user interface without manual intervention. Recent studies have investigated eliciting the capabilities of large language models…

Computation and Language · Computer Science 2024-06-10 Zhuosheng Zhang , Aston Zhang

Feedback is crucial for every design process, such as user interface (UI) design, and automating design critiques can significantly improve the efficiency of the design workflow. Although existing multimodal large language models (LLMs)…

Artificial Intelligence · Computer Science 2025-05-26 Peitong Duan , Chin-Yi Cheng , Bjoern Hartmann , Yang Li

Multimodal sentiment analysis is a key technology in the fields of human-computer interaction and affective computing. Accurately recognizing human emotional states is crucial for facilitating smooth communication between humans and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-07 Wangyuan Zhu , Jun Yu

Scene understanding using multi-modal data is necessary in many applications, e.g., autonomous navigation. To achieve this in a variety of situations, existing models must be able to adapt to shifting data distributions without arduous data…

Computer Vision and Pattern Recognition · Computer Science 2023-08-24 Cody Simons , Dripta S. Raychaudhuri , Sk Miraj Ahmed , Suya You , Konstantinos Karydis , Amit K. Roy-Chowdhury

With the recently increasing capabilities of modern vehicles, novel approaches for interaction emerged that go beyond traditional touch-based and voice command approaches. Therefore, hand gestures, head pose, eye gaze, and speech have been…

Human-Computer Interaction · Computer Science 2022-11-08 Amr Gomaa

Accurate, scalable traffic monitoring is critical for real-time and long-term transportation management, particularly during disruptions such as natural disasters, large construction projects, or major policy changes like New York City's…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Fan Zuo , Donglin Zhou , Jingqin Gao , Kaan Ozbay

Multimodal surface material classification plays a critical role in advancing tactile perception for robotic manipulation and interaction. In this paper, we present Surformer v2, an enhanced multi-modal classification architecture designed…

Robotics · Computer Science 2025-09-08 Manish Kansana , Sindhuja Penchala , Shahram Rahimi , Noorbakhsh Amiri Golilarz

The human visual perception system has very strong robustness and contextual awareness in a variety of image processing tasks. This robustness and the perception ability of contextual awareness is closely related to the characteristics of…

Computer Vision and Pattern Recognition · Computer Science 2019-12-24 Aiqing Fang , Xinbo Zhao , Yanning Zhang

Single-modal object detection tasks often experience performance degradation when encountering diverse scenarios. In contrast, multimodal object detection tasks can offer more comprehensive information about object features by integrating…

Computer Vision and Pattern Recognition · Computer Science 2024-12-25 Chang Liu , Xin Ma , Xiaochen Yang , Yuxiang Zhang , Yanni Dong