English
Related papers

Related papers: GAVIN: Gaze-Assisted Voice-Based Implicit Note-tak…

200 papers

With the rapid adoption of multimodal large language models (MLLMs) across diverse applications, there is a pressing need for task-centered, high-quality training data. A key limitation of current training datasets is their reliance on…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Xiaoyu Lin , Aniket Ghorpade , Hansheng Zhu , Justin Qiu , Dea Rrozhani , Monica Lama , Mick Yang , Zixuan Bian , Ruohan Ren , Alan B. Hong , Jiatao Gu , Chris Callison-Burch

Gaze following estimates gaze targets of in-scene person by understanding human behavior and scene information. Existing methods usually analyze scene images for gaze following. However, compared with visual images, audio also provides…

Computer Vision and Pattern Recognition · Computer Science 2024-01-17 Yuqi Hou , Zhongqun Zhang , Nora Horanyi , Jaewon Moon , Yihua Cheng , Hyung Jin Chang

An annotation consists of a portion of information that is associated with a piece of content in order to explain something about the content or to add more information. The use of annotations as a tool in the educational field has positive…

Computation and Language · Computer Science 2025-01-28 Joaquín Gayoso-Cabada , Antonio Sarasa-Cabezuelo , José-Luis Sierra

The potential of using gaze as an input modality in the mobile context is growing. While users often encumber themselves by carrying objects and using mobile devices while walking, the impact of encumbrance on gaze input performance remains…

Human-Computer Interaction · Computer Science 2025-12-19 Omar Namnakani , Yasmeen Abdrabou , John H. Williamson , Mohamed Khamis

Visual-audio navigation (VAN) is attracting more and more attention from the robotic community due to its broad applications, \emph{e.g.}, household robots and rescue robots. In this task, an embodied agent must search for and navigate to…

Robotics · Computer Science 2023-06-22 Hongcheng Wang , Yuxuan Wang , Fangwei Zhong , Mingdong Wu , Jianwei Zhang , Yizhou Wang , Hao Dong

Audio captioning aims to generate text descriptions of audio clips. In the real world, many objects produce similar sounds. How to accurately recognize ambiguous sounds is a major challenge for audio captioning. In this work, inspired by…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-30 Xubo Liu , Qiushi Huang , Xinhao Mei , Haohe Liu , Qiuqiang Kong , Jianyuan Sun , Shengchen Li , Tom Ko , Yu Zhang , Lilian H. Tang , Mark D. Plumbley , Volkan Kılıç , Wenwu Wang

PanGEA, the Panoramic Graph Environment Annotation toolkit, is a lightweight toolkit for collecting speech and text annotations in photo-realistic 3D environments. PanGEA immerses annotators in a web-based simulation and allows them to move…

Computer Vision and Pattern Recognition · Computer Science 2021-03-24 Alexander Ku , Peter Anderson , Jordi Pont-Tuset , Jason Baldridge

Rationales in the form of manually annotated input spans usually serve as ground truth when evaluating explainability methods in NLP. They are, however, time-consuming and often biased by the annotation process. In this paper, we debate…

Computation and Language · Computer Science 2024-03-01 Stephanie Brandl , Oliver Eberle , Tiago Ribeiro , Anders Søgaard , Nora Hollenstein

Taking notes quickly while effectively capturing key information can be challenging, especially when watching videos that present simultaneous visual and auditory streams. Manually taken notes often miss crucial details due to the…

Human-Computer Interaction · Computer Science 2025-01-31 Faria Huq , Abdus Samee , David Chuan-en Lin , Xiaodi Alice Tang , Jeffrey P. Bigham

Active Learning aims to minimize annotation effort by selecting the most useful instances from a pool of unlabeled data. However, typical active learning methods overlook the presence of distinct example groups within a class, whose…

Machine Learning · Computer Science 2024-10-14 Michalis Korakakis , Andreas Vlachos , Adrian Weller

Instance segmentation is a computer vision task where separate objects in an image are detected and segmented. State-of-the-art deep neural network models require large amounts of labeled data in order to perform well in this task. Making…

Computer Vision and Pattern Recognition · Computer Science 2022-02-21 Tuomas Sormunen , Arttu Lämsä , Miguel Bordallo Lopez

This handbook is a hands-on guide on how to approach text annotation tasks. It provides a gentle introduction to the topic, an overview of theoretical concepts as well as practical advice. The topics covered are mostly technical, but…

Turn-taking prediction is crucial for seamless interactions. This study introduces a novel, lightweight framework for accurate turn-taking prediction in triadic conversations without relying on computationally intensive methods. Unlike…

Human-Computer Interaction · Computer Science 2025-05-30 Seongsil Heo , Calvin Murdock , Michael Proulx , Christi Miller

Computer vision is widely used in the fields of driverless, face recognition and 3D reconstruction as a technology to help or replace human eye perception images or multidimensional data through computers. Nowadays, with the development and…

Computer Vision and Pattern Recognition · Computer Science 2021-10-11 Ming Li , ChenHao Guo

In recent years we have witnessed an increasing number of interactive systems on handheld mobile devices which utilise gaze as a single or complementary interaction modality. This trend is driven by the enhanced computational power of these…

Human-Computer Interaction · Computer Science 2023-07-04 Yaxiong Lei , Shijing He , Mohamed Khamis , Juan Ye

Robust gaze estimation is a challenging task, even for deep CNNs, due to the non-availability of large-scale labeled data. Moreover, gaze annotation is a time-consuming process and requires specialized hardware setups. We propose MTGLS: a…

Computer Vision and Pattern Recognition · Computer Science 2021-12-14 Shreya Ghosh , Munawar Hayat , Abhinav Dhall , Jarrod Knibbe

While walking meetings offer a healthy alternative to sit-down meetings, they also pose practical challenges. Taking notes is difficult while walking, which limits the potential of walking meetings. To address this, we designed the Walking…

Human-Computer Interaction · Computer Science 2023-04-06 Luke Haliburton , Natalia Bartłomiejczyk , Paweł W. Woźniak , Albrecht Schmidt , Jasmin Niess

Active vision -- where a policy controls its own gaze during manipulation -- has emerged as a key capability for imitation learning, with multiple independent systems demonstrating its benefits in the past year. Yet there is no shared…

Robotics · Computer Science 2026-05-11 Giacomo Spigler

Attention level estimation systems have a high potential in many use cases, such as human-robot interaction, driver modeling and smart home systems, since being able to measure a person's attention level opens the possibility to natural…

Computer Vision and Pattern Recognition · Computer Science 2019-01-25 Andrea Coifman , Péter Rohoska , Miklas S. Kristoffersen , Sven E. Shepstone , Zheng-Hua Tan

In challenging real-life conditions such as extreme head-pose, occlusions, and low-resolution images where the visual information fails to estimate visual attention/gaze direction, audio signals could provide important and complementary…

Computer Vision and Pattern Recognition · Computer Science 2022-08-15 Shreya Ghosh , Abhinav Dhall , Munawar Hayat , Jarrod Knibbe
‹ Prev 1 4 5 6 7 8 10 Next ›