English
Related papers

Related papers: ReCAP: Recursive Cross Attention Network for Pseud…

200 papers

We present RASO, a foundation model designed to Recognize Any Surgical Object, offering robust open-set recognition capabilities across a broad range of surgical procedures and object classes, in both surgical images and videos. RASO…

Computer Vision and Pattern Recognition · Computer Science 2025-05-07 Jiajie Li , Brian R Quaranto , Chenhui Xu , Ishan Mishra , Ruiyang Qin , Dancheng Liu , Peter C W Kim , Jinjun Xiong

We propose a novel multi-modal and multi-task architecture for simultaneous low level gesture and surgical task classification in Robot Assisted Surgery (RAS) videos.Our end-to-end architecture is based on the principles of a long…

Computer Vision and Pattern Recognition · Computer Science 2018-05-03 Duygu Sarikaya , Khurshid A. Guru , Jason J. Corso

We study an $\ell_{1}$-regularized generalized least-squares (GLS) estimator for high-dimensional regressions with autocorrelated errors. Specifically, we consider the case where errors are assumed to follow an autoregressive process,…

Methodology · Statistics 2025-10-17 Kaveh S. Nobari , Alex Gibberd

Purpose: Basic surgical skills of suturing and knot tying are an essential part of medical training. Having an automated system for surgical skills assessment could help save experts time and improve training efficiency. There have been…

Computer Vision and Pattern Recognition · Computer Science 2017-02-28 Aneeq Zia , Yachna Sharma , Vinay Bettadapura , Eric L. Sarin , Irfan Essa

Neural ranking models (NRMs) have demonstrated effective performance in several information retrieval (IR) tasks. However, training NRMs often requires large-scale training data, which is difficult and expensive to obtain. To address this…

Information Retrieval · Computer Science 2023-04-19 Yen-Chieh Lien , Hamed Zamani , W. Bruce Croft

Open-Vocabulary Remote Sensing Image Segmentation (OVRSIS), an emerging task that adapts Open-Vocabulary Segmentation (OVS) to the remote sensing (RS) domain, remains underexplored due to the absence of a unified evaluation benchmark and…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Bingyu Li , Haocheng Dong , Da Zhang , Zhiyuan Zhao , Junyu Gao , Xuelong Li

Recently, graph-based semi-supervised learning and pseudo-labeling have gained attention due to their effectiveness in reducing the need for extensive data annotations. Pseudo-labeling uses predictions from unlabeled data to improve model…

Computer Vision and Pattern Recognition · Computer Science 2025-10-30 Jingjun Bi , Fadi Dornaika

Evaluating surgeon skill has predominantly been a subjective task. Development of objective methods for surgical skill assessment are of increased interest. Recently, with technological advances such as robotic-assisted minimally invasive…

Machine Learning · Computer Science 2016-11-18 Mahtab J. Fard , Sattar Ameri , Ratna B. Chinnam , Abhilash K. Pandya , Michael D. Klein , R. Darin Ellis

Vision-based surgical skill assessment (SSA) enables objective and scalable evaluation of operative performance. Progress in this field is constrained by the high cost and time demands for manual annotation of quantitative skill scores, as…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Dimitrios Anastasiou , Razvan Caramalau , Jialang Xu , Runlong He , Freweini Tesfai , Matthew Boal , Nader Francis , Danail Stoyanov , Evangelos B. Mazomenos

The recurrent neural network-transducer (RNNT) is a promising approach for automatic speech recognition (ASR) with the introduction of a prediction network that autoregressively considers linguistic aspects. To train the autoregressive…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-26 Takafumi Moriya , Takanori Ashihara , Hiroshi Sato , Kohei Matsuura , Tomohiro Tanaka , Ryo Masumura

Open-vocabulary 3D Scene Graph (3DSG) can enhance various downstream tasks in robotics by leveraging structured semantic representations, yet current 3DSG construction methods suffer from semantic inconsistencies caused by noisy cross-image…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Yue Chang , Rufeng Chen , Zhaofan Zhang , Yi Chen , Yifan Tian , Sihong Xie

While recent advances in deep learning for surgical scene segmentation have demonstrated promising results on single-centre and single-imaging modality data, these methods usually do not generalise to unseen distribution (i.e., from other…

Computer Vision and Pattern Recognition · Computer Science 2025-12-03 Mansoor Ali , Maksim Richards , Gilberto Ochoa-Ruiz , Sharib Ali

Endoscopic Sinus and Skull Base Surgeries (ESSBSs) is a challenging and potentially dangerous surgical procedure, and objective skill assessment is the key components to improve the effectiveness of surgical training, to re-validate…

Machine Learning · Computer Science 2021-12-07 Yangming Li , Randall Bly , Sarah Akkina , Rajeev C. Saxena , Ian Humphreys , Mark Whipple , Kris Moe , Blake Hannaford

The core objective of image captioning is to achieve lossless semantic compression from visual signals into textual modalities. However, the reliance on manually curated reference texts for evaluation essentially forces models to mimic…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Ziyun Chen , Fan Liu , Liang Yao , Chuanyi Zhang , Yuye Ma , Wei Zhou

Generic group-based RL assumes that sampled rollout groups are already usable learning signals. We show that this assumption breaks down in sparse-hit generative recommendation, where many sampled groups never become learnable at all. We…

Machine Learning · Computer Science 2026-04-27 Peiyan Zhang , Hanmo Liu , Chengxuan Tong , Yuxia Wu , Wei Guo , Yong Liu

Semi-supervised learning has received considerable attention for its potential to leverage abundant unlabeled data to enhance model robustness. Pseudo labeling is a widely used strategy in semi supervised learning. However, existing methods…

Computer Vision and Pattern Recognition · Computer Science 2024-11-14 Tao Wang , Xinlin Zhang , Yuanbin Chen , Yuanbo Zhou , Longxuan Zhao , Tao Tan , Tong Tong

Most 3D human mesh regressors are fully supervised with 3D pseudo-GT human model parameters and weakly supervised with GT 2D/3D joint coordinates as the 3D pseudo-GTs bring great performance gain. The 3D pseudo-GTs are obtained by…

Computer Vision and Pattern Recognition · Computer Science 2022-04-20 Gyeongsik Moon , Hongsuk Choi , Kyoung Mu Lee

Open-set recognition (OSR) aims to simultaneously detect unknown-class samples and classify known-class samples. Most of the existing OSR methods are inductive methods, which generally suffer from the domain shift problem that the learned…

Computer Vision and Pattern Recognition · Computer Science 2022-07-14 Jiayin Sun , Qiulei Dong

Large language models exhibit strong reasoning capabilities, yet often rely on shortcuts such as surface pattern matching and answer memorization rather than genuine logical inference. We propose Shortcut-Aware Reasoning Training (SART), a…

Computation and Language · Computer Science 2026-03-24 Hongyu Cao , Kunpeng Liu , Dongjie Wang , Yanjie Fu

In natural language processing tasks, pure reinforcement learning (RL) fine-tuning methods often suffer from inefficient exploration and slow convergence; while supervised fine-tuning (SFT) methods, although efficient in training, have…

Computation and Language · Computer Science 2025-09-17 Min Zeng , Jingfei Sun , Xueyou Luo , Caiquan Liu , Shiqi Zhang , Li Xie , Xiaoxin Chen