English
Related papers

Related papers: The NGT200 Dataset: Geometric Multi-View Isolated …

200 papers

Self-supervised learning (SSL) has emerged as a crucial technique in image processing, encoding, and understanding, especially for developing today's vision foundation models that utilize large-scale datasets without annotations to enhance…

Computer Vision and Pattern Recognition · Computer Science 2025-01-08 Chuang Niu , Wenjun Xia , Hongming Shan , Ge Wang

Building ASR systems robust to foreign-accented speech is an important challenge in today's globalized world. A prior study explored the way to enhance the performance of phonetic token-based ASR on accented speech by reproducing the…

Sound · Computer Science 2026-01-28 Kentaro Onda , Satoru Fukayama , Daisuke Saito , Nobuaki Minematsu

Sign language recognition and translation technologies have the potential to increase access and inclusion of deaf signing communities, but research progress is bottlenecked by a lack of representative data. We introduce a new resource for…

Computation and Language · Computer Science 2023-10-03 Lee Kezar , Elana Pontecorvo , Adele Daniels , Connor Baer , Ruth Ferster , Lauren Berger , Jesse Thomason , Zed Sevcikova Sehyr , Naomi Caselli

This paper explores sentence-level multilingual Visual Speech Recognition (VSR) that can recognize different languages with a single trained model. As the massive multilingual modeling of visual data requires huge computational costs, we…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-19 Minsu Kim , Jeong Hun Yeo , Se Jin Park , Hyeongseop Rha , Yong Man Ro

Self-supervised learning is emerging in fine-grained visual recognition with promising results. However, existing self-supervised learning methods are often susceptible to irrelevant patterns in self-supervised tasks and lack the capability…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 ShuaiHeng Li , Qing Cai , Fan Zhang , Menghuan Zhang , Yangyang Shu , Zhi Liu , Huafeng Li , Lingqiao Liu

Surgical image segmentation is essential for robot-assisted surgery and intraoperative guidance. However, existing methods are constrained to predefined categories, produce one-shot predictions without adaptive refinement, and lack…

Computer Vision and Pattern Recognition · Computer Science 2026-02-11 Ange Lou , Yamin Li , Qi Chang , Nan Xi , Luyuan Xie , Zichao Li , Tianyu Luan

Existing multilingual vision-language (VL) benchmarks often only cover a handful of languages. Consequently, evaluations of large vision-language models (LVLMs) predominantly target high-resource languages, underscoring the need for…

Computation and Language · Computer Science 2025-02-19 Fabian David Schmidt , Florian Schneider , Chris Biemann , Goran Glavaš

Glass surfaces create complex interactions of reflected and transmitted light, making single-image reflection removal (SIRR) challenging. Existing datasets suffer from limited physical realism in synthetic data or insufficient scale in real…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Yu Guo , Zhiqiang Lao , Xiyun Song , Yubin Zhou , Heather Yu

Pre-trained Vision-Language Models (VLMs) utilizing extensive image-text paired data have demonstrated unprecedented image-text association capabilities, achieving remarkable results across various downstream tasks. A critical challenge is…

Computer Vision and Pattern Recognition · Computer Science 2024-09-10 Zilun Zhang , Tiancheng Zhao , Yulong Guo , Jianwei Yin

We introduce the problem of zero-shot sign language recognition (ZSSLR), where the goal is to leverage models learned over the seen sign class examples to recognize the instances of unseen signs. To this end, we propose to utilize the…

Computer Vision and Pattern Recognition · Computer Science 2019-07-25 Yunus Can Bilge , Nazli Ikizler-Cinbis , Ramazan Gokberk Cinbis

Signed Language Processing (SLP) concerns the automated processing of signed languages, the main means of communication of Deaf and hearing impaired individuals. SLP features many different tasks, ranging from sign recognition to…

Computation and Language · Computer Science 2024-07-25 Federico Tavella , Viktor Schlegel , Marta Romeo , Aphrodite Galata , Angelo Cangelosi

Accurate biomedical image classification under low-resource conditions remains challenging due to limited annotations, subtle inter-class visual differences, and complex disease semantics. While vision--language models offer a promising…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Xiaoliu Luo , Minxue Xiao , Ting Xie , Mengzhu Wang , Huiqing Qi , Joey Tianyi Zhou , Taiping Zhang , Xu Wang

Sign languages serve as essential communication systems for individuals with hearing and speech impairments. However, digital linguistic dataset resources for underrepresented sign languages, such as Nepali Sign Language (NSL), remain…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Birat Poudel , Satyam Ghimire , Sijan Bhattarai , Saurav Bhandari , Suramya Sharma Dahal

Large Vision--Language Models (LVLMs) hold great promise for advancing optical remote sensing (RS) analysis, yet existing reasoning segmentation frameworks couple linguistic reasoning and pixel prediction through end-to-end supervised…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Xu Zhang , Junyao Ge , Yang Zheng , Kaitai Guo , Jimin Liang

Implicit neural representation (INR) has recently emerged as a promising paradigm for signal representations. Typically, INR is parameterized by a multiplayer perceptron (MLP) which takes the coordinates as the inputs and generates…

Computer Vision and Pattern Recognition · Computer Science 2024-06-07 Zhicheng Cai

Here we present Symplectically Integrated Symbolic Regression (SISR), a novel technique for learning physical governing equations from data. SISR employs a deep symbolic regression approach, using a multi-layer LSTM-RNN with mutation to…

Machine Learning · Computer Science 2022-09-07 Daniel M. DiPietro , Bo Zhu

Sign Language Recognition (SLR) systems aim to be embedded in video stream platforms to recognize the sign performed in front of a camera. SLR research has taken advantage of recent advances in pose estimation models to use skeleton…

Computer Vision and Pattern Recognition · Computer Science 2023-04-13 David Laines , Gissella Bejarano , Miguel Gonzalez-Mendoza , Gilberto Ochoa-Ruiz

We present GLIPv2, a grounded VL understanding model, that serves both localization tasks (e.g., object detection, instance segmentation) and Vision-Language (VL) understanding tasks (e.g., VQA, image captioning). GLIPv2 elegantly unifies…

Computer Vision and Pattern Recognition · Computer Science 2022-10-13 Haotian Zhang , Pengchuan Zhang , Xiaowei Hu , Yen-Chun Chen , Liunian Harold Li , Xiyang Dai , Lijuan Wang , Lu Yuan , Jenq-Neng Hwang , Jianfeng Gao

The effective extraction of spatial-angular features plays a crucial role in light field image super-resolution (LFSR) tasks, and the introduction of convolution and Transformers leads to significant improvement in this area. Nevertheless,…

Image and Video Processing · Electrical Eng. & Systems 2025-12-02 Zeke Zexi Hu , Xiaoming Chen , Vera Yuk Ying Chung , Yiran Shen

Multi-modal large language models (MLLMs) have demonstrated remarkable success in vision and visual-language tasks within the natural image domain. Owing to the significant diversities between the natural and remote sensing (RS) images, the…

Computer Vision and Pattern Recognition · Computer Science 2024-03-11 Wei Zhang , Miaoxin Cai , Tong Zhang , Yin Zhuang , Xuerui Mao