English
Related papers

Related papers: CoPESD: A Multi-Level Surgical Motion Dataset for …

200 papers

Within the domain of medical analysis, extensive research has explored the potential of mutual learning between Masked Autoencoders(MAEs) and multimodal data. However, the impact of MAEs on intermodality remains a key challenge. We…

Image and Video Processing · Electrical Eng. & Systems 2024-06-03 Lei Li , Tianfang Zhang , Xinglin Zhang , Jiaqi Liu , Bingqi Ma , Yan Luo , Tao Chen

Ultrasound video segmentation is clinically valuable yet difficult due to speckle noise, weak boundaries, and rapid anatomical deformation. Recent promptable foundation models enable point-guided segmentation, but their direct deployment in…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Ruiqiang Xiao , Zhaohu Xing , Yijun Yang , Zhenyan Han , Weiming Wang , Kaishun Wu , Lei Zhu

Segmentation-based, two-stage neural network has shown excellent results in the surface defect detection, enabling the network to learn from a relatively small number of samples. In this work, we introduce end-to-end training of the…

Computer Vision and Pattern Recognition · Computer Science 2020-07-16 Jakob Božič , Domen Tabernik , Danijel Skočaj

Due to the difficulty of collecting real paired data, most existing desmoking methods train the models by synthesizing smoke, generalizing poorly to real surgical scenarios. Although a few works have explored single-image real-world…

Computer Vision and Pattern Recognition · Computer Science 2024-08-16 Renlong Wu , Zhilu Zhang , Shuohao Zhang , Longfei Gou , Haobin Chen , Lei Zhang , Hao Chen , Wangmeng Zuo

This paper introduces a SSSUMO, semi-supervised deep learning approach for submovement decomposition that achieves state-of-the-art accuracy and speed. While submovement analysis offers valuable insights into motor control, existing methods…

Human-Computer Interaction · Computer Science 2025-07-14 Evgenii Rudakov , Jonathan Shock , Otto Lappi , Benjamin Ultan Cowley

We introduce a novel sequential modeling approach which enables learning a Large Vision Model (LVM) without making use of any linguistic data. To do this, we define a common format, "visual sentences", in which we can represent raw images…

Computer Vision and Pattern Recognition · Computer Science 2023-12-04 Yutong Bai , Xinyang Geng , Karttikeya Mangalam , Amir Bar , Alan Yuille , Trevor Darrell , Jitendra Malik , Alexei A Efros

In the evolving landscape of transportation systems, integrating Large Language Models (LLMs) offers a promising frontier for advancing intelligent decision-making across various applications. This paper introduces a novel 3-dimensional…

Machine Learning · Computer Science 2024-12-17 Dexter Le , Aybars Yunusoglu , Karn Tiwari , Murat Isik , I. Can Dikmen

Large Language Models (LLMs) and Vision-Language Models (VLMs) have emerged as promising candidates for end-to-end autonomous driving. However, these models typically face challenges in inference latency, action precision, and…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Jiaru Zhang , Manav Gagvani , Can Cui , Juntong Peng , Ruqi Zhang , Ziran Wang

3D medical image analysis is essential for modern healthcare, yet traditional task-specific models are inadequate due to limited generalizability across diverse clinical scenarios. Multimodal large language models (MLLMs) offer a promising…

Computer Vision and Pattern Recognition · Computer Science 2025-10-22 Yiming Shi , Xun Zhu , Kaiwen Wang , Ying Hu , Chenyi Guo , Miao Li , Ji Wu

Deep learning can predict depth maps and capsule ego-motion from capsule endoscopy videos, aiding in 3D scene reconstruction and lesion localization. However, the collisions of the capsule endoscopies within the gastrointestinal tract cause…

Computer Vision and Pattern Recognition · Computer Science 2024-12-24 Long Bai , Beilei Cui , Liangyu Wang , Yanheng Li , Shilong Yao , Sishen Yuan , Yanan Wu , Yang Zhang , Max Q. -H. Meng , Zhen Li , Weiping Ding , Hongliang Ren

Large language models (LLMs) have demonstrated impressive capabilities in a wide range of downstream natural language processing tasks. Nevertheless, their considerable sizes and memory demands hinder practical deployment, underscoring the…

Computation and Language · Computer Science 2026-03-17 Haolei Bai , Siyong Jian , Tuo Liang , Yu Yin , Huan Wang

Reliable control of myoelectric prostheses is often hindered by high inter-subject variability and the clinical impracticality of high-density sensor arrays. This study proposes a deep learning framework for accurate gesture recognition…

We present OpenSeeD, a simple Open-vocabulary Segmentation and Detection framework that jointly learns from different segmentation and detection datasets. To bridge the gap of vocabulary and annotation granularity, we first introduce a…

Computer Vision and Pattern Recognition · Computer Science 2023-03-21 Hao Zhang , Feng Li , Xueyan Zou , Shilong Liu , Chunyuan Li , Jianfeng Gao , Jianwei Yang , Lei Zhang

Dynamic magnetic resonance imaging (MRI) plays an indispensable role in cardiac diagnosis. To enable fast imaging, the k-space data can be undersampled but the image reconstruction poses a great challenge of high-dimensional processing.…

Image and Video Processing · Electrical Eng. & Systems 2024-10-03 Zi Wang , Min Xiao , Yirong Zhou , Chengyan Wang , Naiming Wu , Yi Li , Yiwen Gong , Shufu Chang , Yinyin Chen , Liuhong Zhu , Jianjun Zhou , Congbo Cai , He Wang , Di Guo , Guang Yang , Xiaobo Qu

Large Vision-Language Models (LVLMs) process multimodal inputs consisting of text tokens and vision tokens extracted from images or videos. Due to the rich visual information, a single image can generate thousands of vision tokens, leading…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Zicong Tang , Ziyang Ma , Suqing Wang , Zuchao Li , Lefei Zhang , Hai Zhao , Yun Li , Qianren Wang

Intraoperative navigation relies heavily on precise 3D reconstruction to ensure accuracy and safety during surgical procedures. However, endoscopic scenarios present unique challenges, including sparse features and inconsistent lighting,…

Graphics · Computer Science 2025-07-17 Yuchao Zheng , Jianing Zhang , Guochen Ning , Hongen Liao

Confocal laser endomicroscopy (CLE) is a non-invasive, real-time imaging modality that can be used for in-situ, in-vivo imaging and the microstructural analysis of mucous structures. The diagnosis using CLE is, however, complicated by…

Computer Vision and Pattern Recognition · Computer Science 2025-11-17 Nils Porsche , Flurin Müller-Diesing , Sweta Banerjee , Miguel Goncalves , Marc Aubreville

Pulmonary nodule detection plays an important role in lung cancer screening with low-dose computed tomography (CT) scans. It remains challenging to build nodule detection deep learning models with good generalization performance due to…

Computer Vision and Pattern Recognition · Computer Science 2020-02-10 Yuemeng Li , Yong Fan

Downsampling images and labels, often necessitated by limited resources or to expedite network training, leads to the loss of small objects and thin boundaries. This undermines the segmentation network's capacity to interpret images…

Image and Video Processing · Electrical Eng. & Systems 2024-10-27 Shahzad Ali , Yu Rim Lee , Soo Young Park , Won Young Tak , Soon Ki Jung

Segmentation of axon and myelin from microscopy images of the nervous system provides useful quantitative information about the tissue microstructure, such as axon density and myelin thickness. This could be used for instance to document…

Computer Vision and Pattern Recognition · Computer Science 2018-07-12 Aldo Zaimi , Maxime Wabartha , Victor Herman , Pierre-Louis Antonsanti , Christian Samuel Perone , Julien Cohen-Adad
‹ Prev 1 8 9 10 Next ›