English
Related papers

Related papers: GAVIN: Gaze-Assisted Voice-Based Implicit Note-tak…

200 papers

Visual voice activity detection (V-VAD) uses visual features to predict whether a person is speaking or not. V-VAD is useful whenever audio VAD (A-VAD) is inefficient either because the acoustic signal is difficult to analyze or because it…

Computer Vision and Pattern Recognition · Computer Science 2020-10-19 Sylvain Guy , Stéphane Lathuilière , Pablo Mesejo , Radu Horaud

Despite a significant improvement in the educational aids in terms of effective teaching-learning process, most of the educational content available to the students is less than optimal in the context of being up-to-date, exhaustive and…

Computers and Society · Computer Science 2015-08-12 Anamika Chhabra , S. R. S. Iyengar , Poonam Saini , Rajesh Shreedhar Bhat

People are producing more written material then anytime in the history. The increase is so high that professionals from the various fields are no more able to cope with this amount of publications. Text mining tools can offer tools to help…

Artificial Intelligence · Computer Science 2016-02-03 Nikola Milosevic

The laborious and costly nature of affect annotation is a key detrimental factor for obtaining large scale corpora with valid and reliable affect labels. Motivated by the lack of tools that can effectively determine an annotator's…

In this paper, we present an acoustic database, designed to drive and support research on voiced enabled technologies inside moving vehicles. The recording process involves (i) recordings of acoustic impulse responses, acquired under static…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-28 Nikolaos Stefanakis , Marinos Kalaitzakis , Andreas Symiakakis , Stefanos Papadakis , Despoina Pavlidi

In human dialogue, nonverbal information such as nodding and facial expressions is as crucial as verbal information, and spoken dialogue systems are also expected to express such nonverbal behaviors. We focus on nodding, which is critical…

Human-Computer Interaction · Computer Science 2025-08-05 Kazushi Kato , Koji Inoue , Divesh Lala , Keiko Ochi , Tatsuya Kawahara

Audio-Visual Embodied Navigation aims to enable agents to autonomously navigate to sound sources in unknown 3D environments using auditory cues. While current AVN methods excel on in-distribution sound sources, they exhibit poor…

Sound · Computer Science 2025-10-15 Yi Wang , Yinfeng Yu , Fuchun Sun , Liejun Wang , Wendong Zheng

Visual gaze estimation, with its wide-ranging application scenarios, has garnered increasing attention within the research community. Although existing approaches infer gaze solely from image signals, recent advances in visual-language…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Jun Wang , Hao Ruan , Liangjian Wen , Yong Dai , Mingjie Wang

We introduce the Guideline-Centered Annotation Methodology (GCAM), a novel data annotation methodology designed to report the annotation guidelines associated with each data sample. Our approach addresses three key limitations of the…

Computation and Language · Computer Science 2024-12-11 Federico Ruggeri , Eleonora Misino , Arianna Muti , Katerina Korre , Paolo Torroni , Alberto Barrón-Cedeño

Generative AI (GenAI) image tools are increasingly used in design practice, enabling rapid ideation but offering limited support for refinement tasks such as adjusting layout, scale, or visual attributes. While text prompts and inpainting…

Human-Computer Interaction · Computer Science 2026-02-10 Hyerim Park , Phuong Thao Tran , Andre Luckow , Ceenu George , Michael Sedlmair , Malin Eiband

Vision and voice are two vital keys for agents' interaction and learning. In this paper, we present a novel indoor navigation model called Memory Vision-Voice Indoor Navigation (MVV-IN), which receives voice commands and analyzes multimodal…

Computer Vision and Pattern Recognition · Computer Science 2020-09-02 Liqi Yan , Dongfang Liu , Yaoxian Song , Changbin Yu

Visualization literacy assessments typically rely on correctness to classify performance, providing little evidence about how readers arrive at their answers. We argue that gaze can address this gap as an implicit process signal that…

Human-Computer Interaction · Computer Science 2026-03-25 Kathrin Schnizer

Natural Language Understanding (NLU) models are typically trained in a supervised learning framework. In the case of intent classification, the predicted labels are predefined and based on the designed annotation schema while the labelling…

AI models rely on annotated data to learn pattern and perform prediction. Annotation is usually a labor-intensive step that require associating labels ranging from a simple classification label to more complex tasks such as object…

Computer Vision and Pattern Recognition · Computer Science 2025-09-05 Safouane El Ghazouali , Umberto Michelucci

Annotations in Visual Analytics (VA) have become a common means to support the analysis by integrating additional information into the VA system. That additional information often depends on the current process step in the visual analysis.…

Human-Computer Interaction · Computer Science 2020-08-21 Christoph Schmidt , Paul Rosenthal , Heidrun Schumann

Successful learning depends on learners' ability to sustain attention, which is particularly challenging in online education due to limited teacher interaction. A potential indicator for attention is gaze synchrony, demonstrating predictive…

Human-Computer Interaction · Computer Science 2024-04-02 Babette Bühler , Efe Bozkir , Hannah Deininger , Peter Gerjets , Ulrich Trautwein , Enkelejda Kasneci

The Geneva Affective Picture Database WordNet Annotation Tool (GWAT) is a user-friendly web application for manual annotation of pictures in Geneva Affective Picture Database (GAPED) with WordNet. The annotation tool has an intuitive…

Human-Computer Interaction · Computer Science 2017-07-03 Marko Horvat , Dujo Duvnjak , Davor Jug

This paper is on active learning where the goal is to reduce the data annotation burden by interacting with a (human) oracle during training. Standard active learning methods ask the oracle to annotate data samples. Instead, we take a…

Computer Vision and Pattern Recognition · Computer Science 2017-08-03 Miriam W. Huijser , Jan C. van Gemert

This paper has two goals. First, we present the turn-taking annotation layers created for 95 minutes of conversational speech of the Graz Corpus of Read and Spontaneous Speech (GRASS), available to the scientific community. Second, we…

Computation and Language · Computer Science 2025-04-15 Anneliese Kelterer , Barbara Schuppler

Modern personalized recommendation services often rely on user feedback, either explicit or implicit, to improve the quality of services. Explicit feedback refers to behaviors like ratings, while implicit feedback refers to behaviors like…

Information Retrieval · Computer Science 2023-11-22 Guanyu Lin , Chen Gao , Yu Zheng , Yinfeng Li , Jianxin Chang , Yanan Niu , Yang Song , Kun Gai , Zhiheng Li , Depeng Jin , Yong Li