English
Related papers

Related papers: Tools for online tutorials: comparing capture devi…

200 papers

Remote collaboration tools for conferencing and presentation are gaining significant popularity during the COVID-19 pandemic period. Most prior work has issues, such as a) limited support for media types, b) lack of interactivity, for…

Multimedia · Computer Science 2022-04-14 Chunxu Tang , Beinan Wang , C. Y. Roger Chen , Huijun Wu

Recent advances in hardware sophistication related to graphics display, audio and video devices made available a large number of multimedia and hypermedia applications. These multimedia applications need to store and retrieve the different…

Information Retrieval · Computer Science 2010-01-05 Rajkumar Kannan , Frederic Andres , Balakrishnan Ramadoss

We present an open-source library for seamless robot control through motion capture using smartphones and smartwatches. Our library features three modes: Watch Only Mode, enabling control with a single smartwatch; Upper Arm Mode, offering…

Robotics · Computer Science 2024-06-04 Fabian C Weigend , Neelesh Kumar , Oya Aran , Heni Ben Amor

Generating interaction-centric videos, such as those depicting humans or robots interacting with objects, is crucial for embodied intelligence, as they provide rich and diverse visual priors for robot learning, manipulation policy training,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-24 Gen Li , Bo Zhao , Jianfei Yang , Laura Sevilla-Lara

This paper presents the AToMiC (Authoring Tools for Multimedia Content) dataset, designed to advance research in image/text cross-modal retrieval. While vision-language pretrained transformers have led to significant improvements in…

This paper addresses the challenge of robotic grasping of general objects. Similar to prior research, the task reads a single-view 3D observation (i.e., point clouds) captured by a depth camera as input. Crucially, the success of object…

Robotics · Computer Science 2024-07-23 Kangqi Ma , Hao Dong , Yadong Mu

Multimodal scene search of conversations is essential for unlocking valuable insights into social dynamics and enhancing our communication. While experts in conversational analysis have their own knowledge and skills to find key scenes, a…

Human-Computer Interaction · Computer Science 2024-02-20 Riku Arakawa , Kiyosu Maeda , Hiromu Yakura

This article methodologically reflects on how social media scholars can effectively engage with speech-based data in their analyses. While contemporary media studies have embraced textual, visual, and relational data, the aural dimension…

Social and Information Networks · Computer Science 2024-12-18 Hongrui Jin

Online social platforms centered around content creators often allow comments on content, where creators moderate the comments they receive. As creators can face overwhelming numbers of comments, with some of them harassing or hateful,…

Human-Computer Interaction · Computer Science 2022-02-18 Shagun Jhaver , Quan Ze Chen , Detlef Knauss , Amy Zhang

The task of describing video content in natural language is commonly referred to as video captioning. Unlike conventional video captions, which are typically brief and widely available, long-form paragraph descriptions in natural language…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Mihai Masala , Marius Leordeanu

This paper explores the usage of multimodal image-to-text models to enhance text-based item retrieval. We propose utilizing pre-trained image captioning and tagging models, such as instructBLIP and CLIP, to generate text-based product…

Information Retrieval · Computer Science 2024-02-14 Jason Tang , Garrin McGoldrick , Marie Al-Ghossein , Ching-Wei Chen

Tutorial videos of mobile apps have become a popular and compelling way for users to learn unfamiliar app features. To make the video accessible to the users, video creators always need to annotate the actions in the video, including what…

Human-Computer Interaction · Computer Science 2023-08-08 Sidong Feng , Chunyang Chen , Zhenchang Xing

Context: Online collaborative creation of models is becoming commonplace. Collaborative modeling using chatbots and natural language may lower the barriers to modeling for users from different domains. Objective: We compare the perceived…

Human-Computer Interaction · Computer Science 2024-08-27 Ranci Ren , John W. Castro , Santiago R. Acuña , Oscar Dieste , Silvia T. Acuña

Video is a powerful medium for communication and storytelling, yet reauthoring existing footage remains challenging. Even simple edits often demand expertise, time, and careful planning, constraining how creators envision and shape their…

Human-Computer Interaction · Computer Science 2026-04-07 Sitong Wang , Anh Truong , Lydia B. Chilton , Dingzeyu Li

This work presents a next-generation human-robot interface that can infer and realize the user's manipulation intention via sight only. Specifically, we develop a system that integrates near-eye-tracking and robotic manipulation to enable…

Robotics · Computer Science 2023-05-16 Shaochen Wang , Wei Zhang , Zhangli Zhou , Jiaxi Cao , Ziyang Chen , Kang Chen , Bin Li , Zhen Kan

Recently, the rise of large-scale vision-language pretrained models like CLIP, coupled with the technology of Parameter-Efficient FineTuning (PEFT), has captured substantial attraction in video action recognition. Nevertheless, prevailing…

Computer Vision and Pattern Recognition · Computer Science 2024-01-23 Mengmeng Wang , Jiazheng Xing , Boyuan Jiang , Jun Chen , Jianbiao Mei , Xingxing Zuo , Guang Dai , Jingdong Wang , Yong Liu

Modern knowledge workers typically need to use multiple resources, such as documents, web pages, and applications, at the same time. This complexity in their computing environments forces workers to restore various resources in the course…

Human-Computer Interaction · Computer Science 2022-09-27 Donghan Hu , Sang Won Lee

Despite a plethora of research dedicated to designing HITs for non-workstations, there is a lack of research looking specifically into workers' perceptions of the suitability of these devices for managing and completing work. In this work,…

Human-Computer Interaction · Computer Science 2024-09-10 Senjuti Dutta , Scott Ruoti , Rhema Linder , Alex C. Williams , Anastasia Kuzminykh

The real-world is inherently multi-modal at its core. Our tools observe and take snapshots of it, in digital form, such as videos or sounds, however much of it is lost. Similarly for actions and information passing between humans, languages…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Mihai-Cristian Pîrvu , Marius Leordeanu

Dense video captioning is a task of localizing interesting events from an untrimmed video and producing textual description (captions) for each localized event. Most of the previous works in dense video captioning are solely based on visual…

Computer Vision and Pattern Recognition · Computer Science 2020-05-07 Vladimir Iashin , Esa Rahtu