English
Related papers

Related papers: AI Challenger : A Large-scale Dataset for Going De…

200 papers

Dataset bias in vision-language tasks is becoming one of the main problems which hinders the progress of our community. Existing solutions lack a principled analysis about why modern image captioners easily collapse into dataset bias. In…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Xu Yang , Hanwang Zhang , Jianfei Cai

Fully automatic semantic segmentation of highly specific semantic classes and complex shapes may not meet the accuracy standards demanded by scientists. In such cases, human-centered AI solutions, able to assist operators while preserving…

Computer Vision and Pattern Recognition · Computer Science 2021-12-24 Gaia Pavoni , Massimiliano Corsini , Federico Ponchio , Alessandro Muntoni , Paolo Cignoni

Empowered by large datasets, e.g., ImageNet, unsupervised learning on large-scale data has enabled significant advances for classification tasks. However, whether the large-scale unsupervised semantic segmentation can be achieved remains…

Computer Vision and Pattern Recognition · Computer Science 2022-11-04 Shanghua Gao , Zhong-Yu Li , Ming-Hsuan Yang , Ming-Ming Cheng , Junwei Han , Philip Torr

We introduce the novel problem of identifying the photographer behind a photograph. To explore the feasibility of current computer vision techniques to address this problem, we created a new dataset of over 180,000 images taken by 41…

Computer Vision and Pattern Recognition · Computer Science 2016-06-02 Christopher Thomas , Adriana Kovashka

Datasets (semi-)automatically collected from the web can easily scale to millions of entries, but a dataset's usefulness is directly related to how clean and high-quality its examples are. In this paper, we describe and publicly release an…

Computer Vision and Pattern Recognition · Computer Science 2020-08-24 Houda Alberts , Iacer Calixto

Recent years have witnessed increasing attention in cartoon media, powered by the strong demands of industrial applications. As the first step to understand this media, cartoon face recognition is a crucial but less-explored task with few…

Computer Vision and Pattern Recognition · Computer Science 2020-06-30 Yi Zheng , Yifan Zhao , Mengyuan Ren , He Yan , Xiangju Lu , Junhui Liu , Jia Li

Anatomical landmark detection (ALD) from a medical image is crucial for a wide array of clinical applications. While existing methods achieve quite some success in ALD, they often struggle to balance global context with computational…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Xiaoqian Zhou , Zhen Huang , Heqin Zhu , Qingsong Yao , S. Kevin Zhou

Current captioning datasets focus on object-centric captions, describing the visible objects in the image, e.g. "people eating food in a park". Although these datasets are useful to evaluate the ability of Vision & Language models to…

Computation and Language · Computer Science 2023-09-26 Michele Cafagna , Kees van Deemter , Albert Gatt

Continual learning with vision-language models like CLIP offers a pathway toward scalable machine learning systems by leveraging its transferable representations. Existing CLIP-based methods adapt the pre-trained image encoder by adding…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Mao-Lin Luo , Zi-Hao Zhou , Tong Wei , Min-Ling Zhang

Data-driven computational approaches have evolved to enable extraction of information from medical images with a reliability, accuracy and speed which is already transforming their interpretation and exploitation in clinical practice. While…

Image and Video Processing · Electrical Eng. & Systems 2019-10-22 Tom Vercauteren , Mathias Unberath , Nicolas Padoy , Nassir Navab

In this study, we introduce a new problem raised by social media and photojournalism, named Image Address Localization (IAL), which aims to predict the readable textual address where an image was taken. Existing two-stage approaches involve…

Computer Vision and Pattern Recognition · Computer Science 2024-07-12 Shixiong Xu , Chenghao Zhang , Lubin Fan , Gaofeng Meng , Shiming Xiang , Jieping Ye

Vision transformers in vision-language models typically use the same amount of compute for every image, regardless of whether it is simple or complex. We propose ICAR (Image Complexity-Aware Retrieval), an adaptive computation approach that…

Information Retrieval · Computer Science 2026-01-16 Mikel Williams-Lekuona , Georgina Cosma

In this work, we announce a comprehensive well curated and opensource dataset with millions of samples for pre-college and college level problems in mathematicsand science. A preliminary set of results using transformer architecture with…

Numerical Analysis · Mathematics 2021-10-01 Neeraj Kollepara , Snehith Kumar Chatakonda , Pawan Kumar

Generating a description of an image is called image captioning. Image captioning requires to recognize the important objects, their attributes and their relationships in an image. It also needs to generate syntactically and semantically…

Computer Vision and Pattern Recognition · Computer Science 2018-10-16 Md. Zakir Hossain , Ferdous Sohel , Mohd Fairuz Shiratuddin , Hamid Laga

Recent studies of the applications of conversational AI tools, such as chatbots powered by large language models, to complex real-world knowledge work have shown limitations related to reasoning and multi-step problem solving. Specifically,…

Artificial Intelligence · Computer Science 2024-03-06 Nova Spivack , Sam Douglas , Michelle Crames , Tim Connors

Research in 3D mapping is crucial for smart city applications, yet the cost of acquiring 3D data often hinders progress. Visual localization, particularly monocular camera position estimation, offers a solution by determining the camera's…

Real-world image classification tasks tend to be complex, where expert labellers are sometimes unsure about the classes present in the images, leading to the issue of learning with noisy labels (LNL). The ill-posedness of the LNL task…

Computer Vision and Pattern Recognition · Computer Science 2024-05-02 Zheng Zhang , Cuong Nguyen , Kevin Wells , Thanh-Toan Do , Gustavo Carneiro

Interpreting a large number of neurons in deep learning is difficult. Our proposed `CLAssifier-DECoder' architecture (ClaDec) facilitates the understanding of the output of an arbitrary layer of neurons or subsets thereof. It uses a decoder…

Computer Vision and Pattern Recognition · Computer Science 2022-03-09 Johannes Schneider , Michail Vlachos

Recent successes in learning-based image classification, however, heavily rely on the large number of annotated training samples, which may require considerable human efforts. In this paper, we propose a novel active learning framework,…

Computer Vision and Pattern Recognition · Computer Science 2017-01-16 Keze Wang , Dongyu Zhang , Ya Li , Ruimao Zhang , Liang Lin

With the rapid advancement of generative models, highly realistic image synthesis has posed new challenges to digital security and media credibility. Although AI-generated image detection methods have partially addressed these concerns, a…

Computer Vision and Pattern Recognition · Computer Science 2025-09-12 Chunxiao Li , Xiaoxiao Wang , Meiling Li , Boming Miao , Peng Sun , Yunjian Zhang , Xiangyang Ji , Yao Zhu