中文
相关论文

相关论文: Tools for online tutorials: comparing capture devi…

200 篇论文

Remote collaboration tools for conferencing and presentation are gaining significant popularity during the COVID-19 pandemic period. Most prior work has issues, such as a) limited support for media types, b) lack of interactivity, for…

多媒体 · 计算机科学 2022-04-14 Chunxu Tang , Beinan Wang , C. Y. Roger Chen , Huijun Wu

Recent advances in hardware sophistication related to graphics display, audio and video devices made available a large number of multimedia and hypermedia applications. These multimedia applications need to store and retrieve the different…

信息检索 · 计算机科学 2010-01-05 Rajkumar Kannan , Frederic Andres , Balakrishnan Ramadoss

We present an open-source library for seamless robot control through motion capture using smartphones and smartwatches. Our library features three modes: Watch Only Mode, enabling control with a single smartwatch; Upper Arm Mode, offering…

机器人学 · 计算机科学 2024-06-04 Fabian C Weigend , Neelesh Kumar , Oya Aran , Heni Ben Amor

Generating interaction-centric videos, such as those depicting humans or robots interacting with objects, is crucial for embodied intelligence, as they provide rich and diverse visual priors for robot learning, manipulation policy training,…

计算机视觉与模式识别 · 计算机科学 2025-11-24 Gen Li , Bo Zhao , Jianfei Yang , Laura Sevilla-Lara

This paper presents the AToMiC (Authoring Tools for Multimedia Content) dataset, designed to advance research in image/text cross-modal retrieval. While vision-language pretrained transformers have led to significant improvements in…

This paper addresses the challenge of robotic grasping of general objects. Similar to prior research, the task reads a single-view 3D observation (i.e., point clouds) captured by a depth camera as input. Crucially, the success of object…

机器人学 · 计算机科学 2024-07-23 Kangqi Ma , Hao Dong , Yadong Mu

Multimodal scene search of conversations is essential for unlocking valuable insights into social dynamics and enhancing our communication. While experts in conversational analysis have their own knowledge and skills to find key scenes, a…

人机交互 · 计算机科学 2024-02-20 Riku Arakawa , Kiyosu Maeda , Hiromu Yakura

This article methodologically reflects on how social media scholars can effectively engage with speech-based data in their analyses. While contemporary media studies have embraced textual, visual, and relational data, the aural dimension…

社会与信息网络 · 计算机科学 2024-12-18 Hongrui Jin

Online social platforms centered around content creators often allow comments on content, where creators moderate the comments they receive. As creators can face overwhelming numbers of comments, with some of them harassing or hateful,…

人机交互 · 计算机科学 2022-02-18 Shagun Jhaver , Quan Ze Chen , Detlef Knauss , Amy Zhang

The task of describing video content in natural language is commonly referred to as video captioning. Unlike conventional video captions, which are typically brief and widely available, long-form paragraph descriptions in natural language…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Mihai Masala , Marius Leordeanu

This paper explores the usage of multimodal image-to-text models to enhance text-based item retrieval. We propose utilizing pre-trained image captioning and tagging models, such as instructBLIP and CLIP, to generate text-based product…

信息检索 · 计算机科学 2024-02-14 Jason Tang , Garrin McGoldrick , Marie Al-Ghossein , Ching-Wei Chen

Tutorial videos of mobile apps have become a popular and compelling way for users to learn unfamiliar app features. To make the video accessible to the users, video creators always need to annotate the actions in the video, including what…

人机交互 · 计算机科学 2023-08-08 Sidong Feng , Chunyang Chen , Zhenchang Xing

Context: Online collaborative creation of models is becoming commonplace. Collaborative modeling using chatbots and natural language may lower the barriers to modeling for users from different domains. Objective: We compare the perceived…

人机交互 · 计算机科学 2024-08-27 Ranci Ren , John W. Castro , Santiago R. Acuña , Oscar Dieste , Silvia T. Acuña

Video is a powerful medium for communication and storytelling, yet reauthoring existing footage remains challenging. Even simple edits often demand expertise, time, and careful planning, constraining how creators envision and shape their…

人机交互 · 计算机科学 2026-04-07 Sitong Wang , Anh Truong , Lydia B. Chilton , Dingzeyu Li

This work presents a next-generation human-robot interface that can infer and realize the user's manipulation intention via sight only. Specifically, we develop a system that integrates near-eye-tracking and robotic manipulation to enable…

机器人学 · 计算机科学 2023-05-16 Shaochen Wang , Wei Zhang , Zhangli Zhou , Jiaxi Cao , Ziyang Chen , Kang Chen , Bin Li , Zhen Kan

Recently, the rise of large-scale vision-language pretrained models like CLIP, coupled with the technology of Parameter-Efficient FineTuning (PEFT), has captured substantial attraction in video action recognition. Nevertheless, prevailing…

计算机视觉与模式识别 · 计算机科学 2024-01-23 Mengmeng Wang , Jiazheng Xing , Boyuan Jiang , Jun Chen , Jianbiao Mei , Xingxing Zuo , Guang Dai , Jingdong Wang , Yong Liu

Modern knowledge workers typically need to use multiple resources, such as documents, web pages, and applications, at the same time. This complexity in their computing environments forces workers to restore various resources in the course…

人机交互 · 计算机科学 2022-09-27 Donghan Hu , Sang Won Lee

Despite a plethora of research dedicated to designing HITs for non-workstations, there is a lack of research looking specifically into workers' perceptions of the suitability of these devices for managing and completing work. In this work,…

人机交互 · 计算机科学 2024-09-10 Senjuti Dutta , Scott Ruoti , Rhema Linder , Alex C. Williams , Anastasia Kuzminykh

The real-world is inherently multi-modal at its core. Our tools observe and take snapshots of it, in digital form, such as videos or sounds, however much of it is lost. Similarly for actions and information passing between humans, languages…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Mihai-Cristian Pîrvu , Marius Leordeanu

Dense video captioning is a task of localizing interesting events from an untrimmed video and producing textual description (captions) for each localized event. Most of the previous works in dense video captioning are solely based on visual…

计算机视觉与模式识别 · 计算机科学 2020-05-07 Vladimir Iashin , Esa Rahtu