中文
相关论文

相关论文: Njobvu-AI: An open-source tool for collaborative i…

200 篇论文

MirBot is a collaborative application for smartphones that allows users to perform object recognition. This app can be used to take a photograph of an object, select the region of interest and obtain the most likely class (dog, chair, etc.)…

计算机视觉与模式识别 · 计算机科学 2020-06-05 Antonio Pertusa , Antonio-Javier Gallego , Marisa Bernabeu

Camera traps offer enormous new opportunities in ecological studies, but current automated image analysis methods often lack the contextual richness needed to support impactful conservation outcomes. Here we present an integrated approach…

计算机视觉与模式识别 · 计算机科学 2024-11-22 Paul Fergus , Carl Chalmers , Naomi Matthews , Stuart Nixon , Andre Burger , Oliver Hartley , Chris Sutherland , Xavier Lambin , Steven Longmore , Serge Wich

In this work, we propose an open-vocabulary object detection method that, based on image-caption pairs, learns to detect novel object classes along with a given set of known classes. It is a two-stage training approach that first uses a…

计算机视觉与模式识别 · 计算机科学 2022-07-29 Maria A. Bravo , Sudhanshu Mittal , Thomas Brox

Automated wildlife surveys based on drone imagery and object detection technology are a powerful and increasingly popular tool in conservation biology. Most detectors require training images with annotated bounding boxes, which are tedious,…

计算机视觉与模式识别 · 计算机科学 2025-04-11 Giacomo May , Emanuele Dalsasso , Benjamin Kellenberger , Devis Tuia

Autonomous agents that navigate Graphical User Interfaces (GUIs) to automate tasks like document editing and file management can greatly enhance computer workflows. While existing research focuses on online settings, desktop environments,…

One of the long-term challenges of robotics is to enable robots to interact with humans in the visual world via natural language, as humans are visual animals that communicate through language. Overcoming this challenge requires the ability…

计算机视觉与模式识别 · 计算机科学 2020-01-07 Yuankai Qi , Qi Wu , Peter Anderson , Xin Wang , William Yang Wang , Chunhua Shen , Anton van den Hengel

We present a novel data-efficient semi-supervised framework to improve the generalization of image captioning models. Constructing a large-scale labeled image captioning dataset is an expensive task in terms of labor, time, and cost. In…

计算机视觉与模式识别 · 计算机科学 2023-01-27 Dong-Jin Kim , Tae-Hyun Oh , Jinsoo Choi , In So Kweon

Machine learning for remote sensing imaging relies on up-to-date and accurate labels for model training and testing. Labelling remote sensing imagery is time and cost intensive, requiring expert analysis. Previous labelling tools rely on…

计算机视觉与模式识别 · 计算机科学 2026-01-28 Tulsi Patel , Mark W. Jones , Thomas Redfern

Labeling data is an important step in the supervised machine learning lifecycle. It is a laborious human activity comprised of repeated decision making: the human labeler decides which of several potential labels to apply to each example.…

We consider the problem of finding spatial configurations of multiple objects in images, e.g., a mobile inspection robot is tasked to localize abandoned tools on the floor. We define the spatial configuration of objects by first-order logic…

Facial recognition is a key enabling component for emerging Internet of Things (IoT) services such as smart homes or responsive offices. Through the use of deep neural networks, facial recognition has achieved excellent performance.…

计算机视觉与模式识别 · 计算机科学 2019-08-27 Chris Xiaoxuan Lu , Xuan Kan , Bowen Du , Changhao Chen , Hongkai Wen , Andrew Markham , Niki Trigoni , John Stankovic

The lack of labeled training data has limited the development of natural language processing tools, such as named entity recognition, for many languages spoken in developing countries. Techniques such as distant and weak supervision can be…

计算与语言 · 计算机科学 2020-04-01 David Ifeoluwa Adelani , Michael A. Hedderich , Dawei Zhu , Esther van den Berg , Dietrich Klakow

Curating high-quality, domain-specific datasets is a major bottleneck for deploying robust vision systems, requiring complex trade-offs between data quality, diversity, and cost when researching vast, unlabeled data lakes. We introduce…

Multimodal large language models (MLLMs) that think with images can interactively use tools to reason about visual inputs, but current approaches often rely on a narrow set of tools with limited real-world necessity and scalability. In this…

计算机视觉与模式识别 · 计算机科学 2025-12-04 Zirun Guo , Minjie Hong , Feng Zhang , Kai Jia , Tao Jin

The recent rapid and tremendous success of deep convolutional neural networks (CNN) on many challenging computer vision tasks largely derives from the accessibility of the well-annotated ImageNet and PASCAL VOC datasets. Nevertheless,…

计算机视觉与模式识别 · 计算机科学 2017-12-29 Xiaosong Wang , Le Lu , Hoo-chang Shin , Lauren Kim , Mohammadhadi Bagheri , Isabella Nogues , Jianhua Yao , Ronald M. Summers

Attributes act as intermediate representations that enable parameter sharing between classes, a must when training data is scarce. We propose to view attribute-based image classification as a label-embedding problem: each class is embedded…

计算机视觉与模式识别 · 计算机科学 2016-10-05 Zeynep Akata , Florent Perronnin , Zaid Harchaoui , Cordelia Schmid

Natural Language Image Editing (NLIE) aims to use natural language instructions to edit images. Since novices are inexperienced with image editing techniques, their instructions are often ambiguous and contain high-level abstractions that…

计算与语言 · 计算机科学 2020-02-13 Tzu-Hsiang Lin , Alexander Rudnicky , Trung Bui , Doo Soon Kim , Jean Oh

With the renaissance of deep learning, neural networks have achieved promising results on many natural language understanding (NLU) tasks. Even though the source codes of many neural network models are publicly available, there is still a…

计算与语言 · 计算机科学 2020-11-30 Nham Le , Tuan Lai , Trung Bui , Doo Soon Kim

There are many web-based visualization systems available to date, each having its strengths and limitations. The goals these systems set out to accomplish influence design decisions and determine how reusable and scalable they are. Weave is…

软件工程 · 计算机科学 2017-09-01 Andrew Dufilie , Georges Grinstein

The ability to interpret a scene is an important capability for a robot that is supposed to interact with its environment. The knowledge of what is in front of the robot is, for example, relevant for navigation, manipulation, or planning.…

机器人学 · 计算机科学 2019-02-04 Andres Milioto , Cyrill Stachniss