English
Related papers

Related papers: The AAU Multimodal Annotation Toolboxes: Annotatin…

200 papers

Image captioning is a multimodal problem that has drawn extensive attention in both the natural language processing and computer vision community. In this paper, we present a novel image captioning architecture to better explore semantics…

Computer Vision and Pattern Recognition · Computer Science 2020-06-23 Zhan Shi , Xu Zhou , Xipeng Qiu , Xiaodan Zhu

Video-Based Design (VBD) uses video as a primary medium for analyzing user interactions, prototyping, and generating design insights. However, current VBD workflows are constrained by labor-intensive, inconsistent manual annotations that…

Human-Computer Interaction · Computer Science 2025-11-20 Tianhao He , Evangelos Niforatos , Gerd Kortuem

Object-level spatial-temporal understanding is essential for video question answering, yet existing multimodal large language models (MLLMs) encode frames holistically and lack explicit mechanisms for fine-grained object grounding. Recent…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Zekun Qian , Ruize Han , Wei Feng

We present ASSIST, an object-wise neural radiance field as a panoptic representation for compositional and realistic simulation. Central to our approach is a novel scene node data structure that stores the information of each object in a…

Computer Vision and Pattern Recognition · Computer Science 2023-11-13 Zhide Zhong , Jiakai Cao , Songen Gu , Sirui Xie , Weibo Gao , Liyi Luo , Zike Yan , Hao Zhao , Guyue Zhou

In recent years, narrative visualization has gained much attention. Researchers have proposed different design spaces for various narrative visualization genres and scenarios to facilitate the creation process. As users' needs grow and…

Human-Computer Interaction · Computer Science 2023-03-23 Qing Chen , Shixiong Cao , Jiazhe Wang , Nan Cao

Automatic generation of video captions is a fundamental challenge in computer vision. Recent techniques typically employ a combination of Convolutional Neural Networks (CNNs) and Recursive Neural Networks (RNNs) for video captioning. These…

Computer Vision and Pattern Recognition · Computer Science 2019-04-30 Nayyer Aafaq , Naveed Akhtar , Wei Liu , Syed Zulqarnain Gilani , Ajmal Mian

The rapid growth and availability of event sequence data across domains requires effective analysis and exploration methods to facilitate decision-making. Visual analytics combines computational techniques with interactive visualizations,…

In this manuscript, we introduce a semi-automatic scene graph annotation tool for images, the GeneAnnotator. This software allows human annotators to describe the existing relationships between participators in the visual scene in the form…

Computer Vision and Pattern Recognition · Computer Science 2021-09-07 Zhixuan Zhang , Chi Zhang , Zhenning Niu , Le Wang , Yuehu Liu

Web pages contain a large variety of information, but are largely designed for use by graphical web browsers. Mobile access to web-based information often requires presenting HTML web pages using channels that are limited in their graphical…

Human-Computer Interaction · Computer Science 2007-05-23 Zhiyan Shao , Robert Capra , Manuel A. Perez-Quinones

For art investigations of paintings, multiple imaging technologies, such as visual light photography, infrared reflectography, ultraviolet fluorescence photography, and x-radiography are often used. For a pixel-wise comparison, the…

Computer Vision and Pattern Recognition · Computer Science 2022-08-19 Aline Sindel , Andreas Maier , Vincent Christlein

Pretraining general-purpose visual features has become a crucial part of tackling many computer vision tasks. While one can learn such features on the extensively-annotated ImageNet dataset, recent approaches have looked at ways to allow…

Computer Vision and Pattern Recognition · Computer Science 2020-08-05 Mert Bulent Sariyildiz , Julien Perez , Diane Larlus

An annotation consists of a portion of information that is associated with a piece of content in order to explain something about the content or to add more information. The use of annotations as a tool in the educational field has positive…

Computation and Language · Computer Science 2025-01-28 Joaquín Gayoso-Cabada , Antonio Sarasa-Cabezuelo , José-Luis Sierra

Inferring detailed 3D geometry of the scene is crucial for robotics applications, simulation, and 3D content creation. However, such information is hard to obtain, and thus very few datasets support it. In this paper, we propose an…

Computer Vision and Pattern Recognition · Computer Science 2020-10-27 Tianchang Shen , Jun Gao , Amlan Kar , Sanja Fidler

In the field of affective computing, where research continually advances at a rapid pace, the demand for user-friendly tools has become increasingly apparent. In this paper, we present the AffectToolbox, a novel software system that aims to…

Human-Computer Interaction · Computer Science 2024-02-26 Silvan Mertes , Dominik Schiller , Michael Dietz , Elisabeth André , Florian Lingenfelser

We present an architecture for integrating real-time, multimodal input into a computational agent's contextual model. Using a human-avatar interaction in a virtual world, we treat aligned gesture and speech as an ensemble where content may…

Human-Computer Interaction · Computer Science 2019-09-19 Nikhil Krishnaswamy , James Pustejovsky

Current state-of-the-art segmentation techniques for ocular images are critically dependent on large-scale annotated datasets, which are labor-intensive to gather and often raise privacy concerns. In this paper, we present a novel…

Computer Vision and Pattern Recognition · Computer Science 2022-12-09 Darian Tomašević , Peter Peer , Vitomir Štruc

In this work we propose a new automatic image annotation model, dubbed {\bf diverse and distinct image annotation} (D2IA). The generative model D2IA is inspired by the ensemble of human annotations, which create semantically relevant, yet…

Computer Vision and Pattern Recognition · Computer Science 2018-04-03 Baoyuan Wu , Weidong Chen , Peng Sun , Wei Liu , Bernard Ghanem , Siwei Lyu

The production of 2D animation follows an industry-standard workflow, encompassing four essential stages: character design, keyframe animation, in-betweening, and coloring. Our research focuses on reducing the labor costs in the above…

Computer Vision and Pattern Recognition · Computer Science 2025-01-31 Yihao Meng , Hao Ouyang , Hanlin Wang , Qiuyu Wang , Wen Wang , Ka Leong Cheng , Zhiheng Liu , Yujun Shen , Huamin Qu

This paper argues in favor of the adoption of annotation practices for multimodal datasets that recognize and represent the inherently perspectivized nature of multimodal communication. To support our claim, we present a set of annotation…

The size of an individual cell type, such as a red blood cell, does not vary much among humans. We use this knowledge as a prior for classifying and detecting cells in images with only a few ground truth bounding box annotations, while most…

Computer Vision and Pattern Recognition · Computer Science 2022-11-14 Hari Om Aggrawal , Dipam Goswami , Vinti Agarwal
‹ Prev 1 8 9 10 Next ›