English
Related papers

Related papers: Process signature-driven high spatio-temporal reso…

200 papers

We study the problem of aligning a video that captures a local portion of an environment to the 2D LiDAR scan of the entire environment. We introduce a method (VioLA) that starts with building a semantic map of the local scene from the…

Computer Vision and Pattern Recognition · Computer Science 2023-11-09 Jun-Jee Chao , Selim Engin , Nikhil Chavan-Dafle , Bhoram Lee , Volkan Isler

Automated surgical gesture recognition is of great importance in robot-assisted minimally invasive surgery. However, existing methods assume that training and testing data are from the same domain, which suffers from severe performance…

Computer Vision and Pattern Recognition · Computer Science 2021-07-20 Xueying Shi , Yueming Jin , Qi Dou , Jing Qin , Pheng-Ann Heng

Visual grounding, which aims to ground a visual region via natural language, is a task that heavily relies on cross-modal alignment. Existing works utilized uni-modal pre-trained models to transfer visual or linguistic knowledge separately…

Computer Vision and Pattern Recognition · Computer Science 2024-09-06 Linhui Xiao , Xiaoshan Yang , Fang Peng , Yaowei Wang , Changsheng Xu

Vision Language Models (VLMs) have undergone significant advancements, particularly with the emergence of mobile-oriented VLMs, which offer a wide range of application scenarios. However, the substantial computational requirements for…

Machine Learning · Computer Science 2025-12-25 Yuanhao Xi , Xiaohuan Bing , Ramin Yahyapour

Optical Coherence Tomography Angiography (OCTA) is a crucial tool in the clinical screening of retinal diseases, allowing for accurate 3D imaging of blood vessels through non-invasive scanning. However, the hardware-based approach for…

Image and Video Processing · Electrical Eng. & Systems 2024-08-22 Shuhan Li , Dong Zhang , Xiaomeng Li , Chubin Ou , Lin An , Yanwu Xu , Kwang-Ting Cheng

The multi-resolution approximation (MRA) of Gaussian processes was recently proposed to conduct likelihood-based inference for massive spatial data sets. An advantage of the methodology is that it can be parallelized. We implemented the MRA…

Computation · Statistics 2019-05-07 Huang Huang , Lewis R. Blake , Dorit M. Hammerling

Establishing pixel/voxel-level or region-level correspondences is the core challenge in image registration. The latter, also known as region-based correspondence representation, leverages paired regions of interest (ROIs) to enable regional…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Shiqi Huang , Tingfa Xu , Wen Yan , Dean Barratt , Yipeng Hu

The Internet of Things (IoT) and mobile technology have significantly transformed healthcare by enabling real-time monitoring and diagnosis of patients. Recognizing medical-related human activities (MRHA) is pivotal for healthcare systems,…

Computer Vision and Pattern Recognition · Computer Science 2025-01-14 Subrata Kumer Paul , Abu Saleh Musa Miah , Rakhi Rani Paul , Md. Ekramul Hamid , Jungpil Shin , Md Abdur Rahim

Multimodal sentiment analysis (MSA) integrates various modalities, such as text, image, and audio, to provide a more comprehensive understanding of sentiment. However, effective MSA is challenged by alignment and fusion issues. Alignment…

Computer Vision and Pattern Recognition · Computer Science 2025-12-08 Yuhua Wen , Qifei Li , Yingying Zhou , Yingming Gao , Zhengqi Wen , Jianhua Tao , Ya Li

High dynamic range (HDR) imaging is of fundamental importance in modern digital photography pipelines and used to produce a high-quality photograph with well exposed regions despite varying illumination across the image. This is typically…

Image and Video Processing · Electrical Eng. & Systems 2024-07-24 Sibi Catley-Chandar , Thomas Tanay , Lucas Vandroux , Aleš Leonardis , Gregory Slabaugh , Eduardo Pérez-Pellitero

Vision Language Action (VLA) models derive their generalization capability from diverse training data, yet collecting embodied robot interaction data remains prohibitively expensive. In contrast, human demonstration videos are far more…

Robotic-assisted surgery (RAS) is established in clinical practice, and automated surgical skill assessment utilizing multimodal data offers transformative potential for surgical analytics and education. However, developing effective…

Multi-modal unsupervised domain adaptation (MM-UDA) for 3D semantic segmentation is a practical solution to embed semantic understanding in autonomous systems without expensive point-wise annotations. While previous MM-UDA methods can…

Computer Vision and Pattern Recognition · Computer Science 2023-09-22 Haozhi Cao , Yuecong Xu , Jianfei Yang , Pengyu Yin , Shenghai Yuan , Lihua Xie

Vision-Language-Action (VLA) models are increasingly expected to not only complete robot tasks, but also follow human instructions about how those tasks should be executed. However, existing robot datasets usually pair trajectories with…

Decision-makers often encounter uncertainty, and the distribution of uncertain parameters plays a crucial role in making reliable decisions. However, complete information is rarely available. The sample average approximation (SAA) approach…

Optimization and Control · Mathematics 2025-08-27 Ziliang Jin , Jianqiang Cheng , Daniel Zhuoyu Long , Kai Pan

Reconstructing articulated objects into high-fidelity digital twins is crucial for applications such as robotic manipulation and interactive simulation. Recent self-supervised methods using differentiable rendering frameworks like 3D…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Xuelu Li , Zhaonan Wang , Xiaogang Wang , Lei Wu , Manyi Li , Changhe Tu

Large language models are first pre-trained on trillions of tokens and then instruction-tuned or aligned to specific preferences. While pre-training remains out of reach for most researchers due to the compute required, fine-tuning has…

Computation and Language · Computer Science 2024-06-10 Megh Thakkar , Quentin Fournier , Matthew D Riemer , Pin-Yu Chen , Amal Zouaq , Payel Das , Sarath Chandar

Continuous space-time video super-resolution (C-STVSR) aims to simultaneously enhance video resolution and frame rate at an arbitrary scale. Recently, implicit neural representation (INR) has been applied to video restoration, representing…

Computer Vision and Pattern Recognition · Computer Science 2025-10-02 Yunfan Lu , Yusheng Wang , Zipeng Wang , Pengteng Li , Bin Yang , Hui Xiong

Event-based vision has been rapidly growing in recent years justified by the unique characteristics it presents such as its high temporal resolutions (~1us), high dynamic range (>120dB), and output latency of only a few microseconds. This…

Computer Vision and Pattern Recognition · Computer Science 2023-01-03 Zaid El-Shair , Samir Rawashdeh

The convergence of robotics and virtual reality (VR) has enabled safer and more efficient workflows in high-risk laboratory settings, particularly virology labs. As biohazard complexity increases, minimizing direct human exposure while…

Robotics · Computer Science 2025-06-18 Farha Abdul Wasay , Mohammed Abdul Rahman , Hania Ghouse
‹ Prev 1 8 9 10 Next ›