English
Related papers

Related papers: Adaptive Texture-aware Masking for Self-Supervised…

200 papers

Inspired by the masked language modeling (MLM) in natural language processing tasks, the masked image modeling (MIM) has been recognized as a strong self-supervised pre-training method in computer vision. However, the high random mask ratio…

Computer Vision and Pattern Recognition · Computer Science 2024-06-04 Zhaowen Li , Yousong Zhu , Zhiyang Chen , Wei Li , Chaoyang Zhao , Rui Zhao , Ming Tang , Jinqiao Wang

Learning semantically meaningful representations from unstructured 3D point clouds remains a central challenge in computer vision, especially in the absence of large-scale labeled datasets. While masked point modeling (MPM) is widely used…

Computer Vision and Pattern Recognition · Computer Science 2025-06-30 Remco F. Leijenaar , Hamidreza Kasaei

Deep learning models for medical image classification usually achieve promising results but typically rely on large, annotated datasets or standard transfer learning from ImageNet. Self-Supervised Learning (SSL) has emerged as a powerful…

Computer Vision and Pattern Recognition · Computer Science 2026-05-07 Joao Batista Florindo , Amanda Pontes de Oliveira Ornelas

Predicting attributes in the landmark free facial images is itself a challenging task which gets further complicated when the face gets occluded due to the usage of masks. Smart access control gates which utilize identity verification or…

Computer Vision and Pattern Recognition · Computer Science 2022-01-12 Prerana Mukherjee , Vinay Kaushik , Ronak Gupta , Ritika Jha , Daneshwari Kankanwadi , Brejesh Lall

The recent progress in self-supervised learning has successfully combined Masked Image Modeling (MIM) with Siamese Networks, harnessing the strengths of both methodologies. Nonetheless, certain challenges persist when integrating…

Computer Vision and Pattern Recognition · Computer Science 2024-11-12 Kirill Vishniakov , Eric Xing , Zhiqiang Shen

Low-dose dental cone beam computed tomography (CBCT) has been increasingly used for maxillofacial modeling. However, the presence of metallic inserts, such as implants, crowns, and dental filling, causes severe streaking and shading…

Image and Video Processing · Electrical Eng. & Systems 2022-02-09 Chang Min Hyun , Taigyntuya Bayaraa , Hye Sun Yun , Tae Jun Jang , Hyoung Suk Park , Jin Keun Seo

Recently impressive performance has been achieved in Concept Bottleneck Models (CBM) by utilizing the image-text alignment learned by a large pre-trained vision-language model (i.e. CLIP). However, there exist two key limitations in concept…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Minghong Zhong , Guoshuai Zou , Kanghao Chen , Dexia Chen , Ruixuan Wang

The computer-assisted radiologic informative report is currently emerging in dental practice to facilitate dental care and reduce time consumption in manual panoramic radiographic interpretation. However, the amount of dental radiographs…

Computer Vision and Pattern Recognition · Computer Science 2022-10-21 Amani Almalki , Longin Jan Latecki

Cone-beam computed tomography (CBCT) is an important tool facilitating computer aided interventions, despite often suffering from artifacts that pose challenges for accurate interpretation. While the degraded image quality can affect…

Image and Video Processing · Electrical Eng. & Systems 2024-07-02 Maximilian E. Tschuchnig , Philipp Steininger , Michael Gadermayr

Self-Supervised Learning (SSL) has emerged as a powerful paradigm to mitigate the reliance on large, annotated datasets, a common bottleneck in medical image analysis. However, standard SSL methods, which rely on simple geometric and color…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Joao Batista Florindo

Understanding the internal representations of large language models is crucial for ensuring their reliability and safety, with sparse autoencoders (SAEs) emerging as a promising interpretability approach. However, current SAE training…

Machine Learning · Computer Science 2025-10-13 T. Ed Li , Junyu Ren

Continual Test-Time Adaptation (CTTA) is proposed to migrate a source pre-trained model to continually changing target distributions, addressing real-world dynamism. Existing CTTA methods mainly rely on entropy minimization or…

Computer Vision and Pattern Recognition · Computer Science 2024-03-28 Jiaming Liu , Ran Xu , Senqiao Yang , Renrui Zhang , Qizhe Zhang , Zehui Chen , Yandong Guo , Shanghang Zhang

Cone-beam computed tomography (CBCT) is an imaging modality widely used in head and neck diagnostics due to its accessibility and lower radiation dose. However, its relatively long acquisition times make it susceptible to patient motion,…

Recently, masked image modeling (MIM), which learns visual representations by reconstructing the masked patches of an image, has dominated self-supervised learning in computer vision. However, the pre-training of MIM always takes massive…

Computer Vision and Pattern Recognition · Computer Science 2024-10-01 Jie Gui , Tuo Chen , Minjing Dong , Zhengqi Liu , Hao Luo , James Tin-Yau Kwok , Yuan Yan Tang

Cone-beam computed tomography (CBCT) is routinely collected during image-guided radiation therapy (IGRT) to provide updated patient anatomy information for cancer treatments. However, CBCT images often suffer from streaking artifacts and…

Computer Vision and Pattern Recognition · Computer Science 2023-11-02 Jiarui Zhu , Werxing Chen , Hongfei Sun , Shaohua Zhi , Jing Qin , Jing Cai , Ge Ren

Accurate segmentation of teeth and pulp in Cone-Beam Computed Tomography (CBCT) is vital for clinical applications like treatment planning and diagnosis. However, this process requires extensive expertise and is exceptionally…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Zhi Qin Tan , Xiatian Zhu , Owen Addison , Yunpeng Li

Recent advances in skeleton-based person re-identification (re-ID) obtain impressive performance via either hand-crafted skeleton descriptors or skeleton representation learning with deep learning paradigms. However, they typically require…

Computer Vision and Pattern Recognition · Computer Science 2022-06-22 Haocong Rao , Chunyan Miao

3D-aware portrait editing has a wide range of applications in multiple fields. However, current approaches are limited due that they can only perform mask-guided or text-based editing. Even by fusing the two procedures into a model, the…

Computer Vision and Pattern Recognition · Computer Science 2024-07-08 Kangneng Zhou , Daiheng Gao , Xuan Wang , Jie Zhang , Peng Zhang , Xusen Sun , Longhao Zhang , Shiqi Yang , Bang Zhang , Liefeng Bo , Yaxing Wang , Ming-Ming Cheng

Text images are unique in their dual nature, encompassing both visual and linguistic information. The visual component encompasses structural and appearance-based features, while the linguistic dimension incorporates contextual and semantic…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Yifei Zhang , Chang Liu , Jin Wei , Xiaomeng Yang , Yu Zhou , Can Ma , Xiangyang Ji

We propose ADIOS, a masked image model (MIM) framework for self-supervised learning, which simultaneously learns a masking function and an image encoder using an adversarial objective. The image encoder is trained to minimise the distance…

Computer Vision and Pattern Recognition · Computer Science 2022-07-07 Yuge Shi , N. Siddharth , Philip H. S. Torr , Adam R. Kosiorek