English
Related papers

Related papers: OLMD: Orientation-aware Long-term Motion Decouplin…

200 papers

Kalman filter (KF) based methods for multi-object tracking (MOT) make an assumption that objects move linearly. While this assumption is acceptable for very short periods of occlusion, linear estimates of motion for prolonged time can be…

Computer Vision and Pattern Recognition · Computer Science 2023-03-17 Jinkun Cao , Jiangmiao Pang , Xinshuo Weng , Rawal Khirodkar , Kris Kitani

Distributed machine learning (ML) over wireless networks hinges on accurate channel state information (CSI) and efficient exchange of high-dimensional model updates. These demands are governed by channel coherence time and bandwidth, which…

Information Theory · Computer Science 2026-03-10 Mehdi Karbalayghareh , David J. Love , Christopher G. Brinton

Large language models (LLMs) are effective at capturing complex, valuable conceptual representations from textual data for a wide range of real-world applications. However, in fields like Intelligent Fault Diagnosis (IFD), incorporating…

Artificial Intelligence · Computer Science 2024-12-03 Hamzah A. A. M. Qaid , Bo Zhang , Dan Li , See-Kiong Ng , Wei Li

Multi-agent motion prediction is challenging because it aims to foresee the future trajectories of multiple agents (\textit{e.g.} pedestrians) simultaneously in a complicated scene. Existing work addressed this challenge by either learning…

Computer Vision and Pattern Recognition · Computer Science 2020-08-04 Chaofan Tao , Qinhong Jiang , Lixin Duan , Ping Luo

SLAM (Simultaneous Localisation and Mapping) is a crucial component for robotic systems, providing a map of an environment, the current location and previous trajectory of a robot. While 3D LiDAR SLAM has received notable improvements in…

Robotics · Computer Science 2025-04-29 Leon Davies , Baihua Li , Mohamad Saada , Simon Sølvsten , Qinggang Meng

Vision-Language Navigation in Continuous Environments (VLN-CE) requires agents to learn complex reasoning from long-horizon human interactions. While Multi-modal Large Language Models (MLLMs) have driven recent progress, current training…

Robotics · Computer Science 2026-03-11 Haoyuan Li , Rui Liu , Hehe Fan , Yi Yang

Infrared object detection focuses on identifying and locating objects in complex environments (\eg, dark, snow, and rain) where visible imaging cameras are disabled by poor illumination. However, due to low contrast and weak edge…

Computer Vision and Pattern Recognition · Computer Science 2026-01-15 Fan Liu , Ting Wu , Chuanyi Zhang , Liang Yao , Xing Ma , Yuhui Zheng

Large Language Models (LLMs) have made substantial advancements in the field of robotic and autonomous driving. This study presents the first Occupancy-based Large Language Model (Occ-LLM), which represents a pioneering effort to integrate…

Robotics · Computer Science 2025-02-11 Tianshuo Xu , Hao Lu , Xu Yan , Yingjie Cai , Bingbing Liu , Yingcong Chen

Multimodal recommendation enhances accuracy by leveraging visual and textual signals, and its success largely depends on learning high-quality cross-modal representations. Recent advances in Large Vision-Language Models (LVLMs) offer…

Information Retrieval · Computer Science 2026-04-28 Zhongtao Rao , Peilin Zhou , Dading Chong , Zhiwei Chen , Shoujin Wang , Nan Tang

The ultimate goal of continuous sign language recognition(CSLR) is to facilitate the communication between special people and normal people, which requires a certain degree of real-time and deploy-ability of the model. However, in the…

Computer Vision and Pattern Recognition · Computer Science 2023-04-04 Qidan Zhu , Jing Li , Fei Yuan , Quan Gan

In this paper, we design the joint decoding (JD) of non-orthogonal multiple access (NOMA) systems employing short block length codes. We first proposed a low-complexity soft-output ordered-statistics decoding (LC-SOSD) based on a decoding…

Sign language is commonly used by deaf or mute people to communicate but requires extensive effort to master. It is usually performed with the fast yet delicate movement of hand gestures, body posture, and even facial expressions. Current…

Computer Vision and Pattern Recognition · Computer Science 2021-10-13 Songyao Jiang , Bin Sun , Lichen Wang , Yue Bai , Kunpeng Li , Yun Fu

Reliable, drift-free global localization presents significant challenges yet remains crucial for autonomous navigation in large-scale dynamic environments. In this paper, we introduce a tightly-coupled Semantic-LiDAR-Inertial-Wheel Odometry…

Robotics · Computer Science 2025-09-19 Haoxuan Jiang , Peicong Qian , Yusen Xie , Linwei Zheng , Xiaocong Li , Ming Liu , Jun Ma

The vast number of parameters in large language models (LLMs) endows them with remarkable capabilities, allowing them to excel in a variety of natural language processing tasks. However, this complexity also presents challenges, making LLMs…

Computation and Language · Computer Science 2023-10-24 Mingzhe Du , Anh Tuan Luu , Bin Ji , See-kiong Ng

Accurate traffic congestion classification is essential for intelligent transportation systems and real-time urban traffic management. This paper presents a multimodal framework combining open-vocabulary visual-language reasoning (CLIP),…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Yu-Hsuan Lin

Instruction tuning is widely used to improve a pre-trained Multimodal Large Language Model (MLLM) by training it on curated task-specific datasets, enabling better comprehension of human instructions. However, it is infeasible to collect…

Computation and Language · Computer Science 2025-05-30 Haiyang Guo , Fanhu Zeng , Ziwei Xiang , Fei Zhu , Da-Han Wang , Xu-Yao Zhang , Cheng-Lin Liu

In this survey, we analyze the newest machine learning (ML) techniques for optical orthogonal frequency division multiplexing (O-OFDM)-based optical communications. ML has been proposed to mitigate channel and transceiver imperfections. For…

Machine Learning · Computer Science 2021-05-10 Hichem Mrabet , Elias Giaccoumidis , Iyad Dayoub

Out-of-distribution (OOD) detection is crucial for model reliability, as it identifies samples from unknown classes and reduces errors due to unexpected inputs. Vision-Language Models (VLMs) such as CLIP are emerging as powerful tools for…

Computer Vision and Pattern Recognition · Computer Science 2024-07-15 Yabin Zhang , Wenjie Zhu , Chenhang He , Lei Zhang

Stance detection is an active task in natural language processing (NLP) that aims to identify the author's stance towards a particular target within a text. Given the remarkable language understanding capabilities and encyclopedic prior…

Computation and Language · Computer Science 2024-08-12 Junxia Ma , Changjiang Wang , Hanwen Xing , Dongming Zhao , Yazhou Zhang

Multimodal Large Language Models (MLLMs) have shown strong performance in document image tasks, especially Optical Character Recognition (OCR). However, they struggle with Document Image Machine Translation (DIMT), which requires handling…

Computation and Language · Computer Science 2025-07-14 Yupu Liang , Yaping Zhang , Zhiyang Zhang , Zhiyuan Chen , Yang Zhao , Lu Xiang , Chengqing Zong , Yu Zhou
‹ Prev 1 8 9 10 Next ›