English
Related papers

Related papers: Multi-Modal End-User Programming of Web-Based Virt…

200 papers

Grounding natural language instructions on the web to perform previously unseen tasks enables accessibility and automation. We introduce a task and dataset to train AI agents from open-domain, step-by-step instructions originally written…

Computation and Language · Computer Science 2021-04-06 Nancy Xu , Sam Masling , Michael Du , Giovanni Campagna , Larry Heck , James Landay , Monica S Lam

The improved competence of generative models can help building multi-modal virtual assistants that leverage modalities beyond language. By observing humans performing multi-step tasks, one can build assistants that have situational…

Computer Vision and Pattern Recognition · Computer Science 2025-01-22 Pha Nguyen , Sailik Sengupta , Girik Malik , Arshit Gupta , Bonan Min

Despite huge gains in performance in natural language understanding via large language models in recent years, voice assistants still often fail to meet user expectations. In this study, we conducted a mixed-methods analysis of how voice…

Human-Computer Interaction · Computer Science 2023-03-06 Amanda Baughan , Allison Mercurio , Ariel Liu , Xuezhi Wang , Jilin Chen , Xiao Ma

People use videos to learn new recipes, exercises, and crafts. Such videos remain difficult for blind and low vision (BLV) people to follow as they rely on visual comparison. Our observations of visual rehabilitation therapists (VRTs)…

Human-Computer Interaction · Computer Science 2025-07-28 Mina Huh , Zihui Xue , Ujjaini Das , Kumar Ashutosh , Kristen Grauman , Amy Pavel

Our research investigates the capability of modern multimodal reasoning models, powered by Large Language Models (LLMs), to facilitate vision-powered assistants for multi-step daily activities. Such assistants must be able to 1) encode…

Computer Vision and Pattern Recognition · Computer Science 2024-08-14 Mrinal Verghese , Brian Chen , Hamid Eghbalzadeh , Tushar Nagarajan , Ruta Desai

Language agents, built on top of language models (LMs), are systems that can interact with complex environments, such as the open web. In this work, we examine whether such agents can perform realistic and time-consuming tasks on the web,…

Computation and Language · Computer Science 2024-10-22 Ori Yoran , Samuel Joseph Amouyal , Chaitanya Malaviya , Ben Bogin , Ofir Press , Jonathan Berant

While intelligent virtual assistants like Siri, Alexa, and Google Assistant have become ubiquitous in modern life, they still face limitations in their ability to follow multi-step instructions and accomplish complex goals articulated in…

Machine Learning · Computer Science 2023-12-13 Yanchu Guan , Dong Wang , Zhixuan Chu , Shiyu Wang , Feiyue Ni , Ruihua Song , Longfei Li , Jinjie Gu , Chenyi Zhuang

Advances in multimodal large language models enable automatic video narration and question answering (VQA), offering scalable alternatives to labor-intensive, human-authored audio descriptions (ADs) for blind and low vision (BLV) viewers.…

Human-Computer Interaction · Computer Science 2026-03-17 Maryam Cheema , Sina Elahimanesh , Pooyan Fazli , Hasti Seifi

Visual Language Action (VLA) models are a multi-modal class of Artificial Intelligence (AI) systems that integrate visual perception, natural language understanding, and action planning to enable agents to interpret their environment,…

Software Engineering · Computer Science 2025-08-04 Pablo Valle , Chengjie Lu , Shaukat Ali , Aitor Arrieta

Intelligent conversational agents and virtual assistants, such as chatbots and voice assistants, have the potential of augmenting health service capacity to screen symptoms and deliver healthcare interventions. In this paper, we developed…

Human-Computer Interaction · Computer Science 2022-02-07 Abdalsalam Almzayyen , Angel Vela de la Garza Evia , Nick Coronato , Mehdi Boukhechba

This paper presents the design of the machine learning architecture that underlies the Alexa Skills Kit (ASK) a large scale Spoken Language Understanding (SLU) Software Development Kit (SDK) that enables developers to extend the…

Interactions with virtual assistants typically start with a trigger phrase followed by a command. In this work, we explore the possibility of making these interactions more natural by eliminating the need for a trigger phrase. Our goal is…

Today, intelligent voice assistant (VA) software like Amazon's Alexa, Google's Voice Assistant (GVA) and Apple's Siri have millions of users. These VAs often collect and analyze huge user data for improving their functionality. However,…

Cryptography and Security · Computer Science 2021-10-08 Vandit Sharma , Mainack Mondal

Virtual assistants (VAs) have seen increased use in recent years due to their ease of use for daily tasks. Despite their growing prevalence, their security and privacy implications are still not well understood. To address this gap, we…

Cryptography and Security · Computer Science 2023-12-25 Borna Kalhor , Sanchari Das

Vision-Language-Action (VLA) models demonstrate promising generalization in robotic manipulation, driven by advances in large-scale vision and language pre-training. This progress can be misleading. Despite the zero-shot perception and…

A considerable part of the success experienced by Voice-controlled virtual assistants (VVA) is due to the emotional and personalized experience they deliver, with humor being a key component in providing an engaging interaction. In this…

Machine Learning · Computer Science 2019-12-09 Alejandro Mottini , Amber Roy Chowdhury

Virtual try-on has made significant progress in recent years. This paper addresses how to achieve multifunctional virtual try-on guided solely by text instructions, including full outfit change and local editing. Previous methods primarily…

Computer Vision and Pattern Recognition · Computer Science 2025-07-09 Yujie Hu , Xuanyu Zhang , Weiqi Li , Jian Zhang

In this work, we present and evaluate SELMA, a Speech-Enabled Language Model for virtual Assistant interactions that integrates audio and text as inputs to a Large Language Model (LLM). SELMA is designed to handle three primary and two…

Sound · Computer Science 2025-02-04 Dominik Wagner , Alexander Churchill , Siddharth Sigtia , Erik Marchi

Procedural activity assistants potentially support humans in a variety of settings, from our daily lives, e.g., cooking or assembling flat-pack furniture, to professional situations, e.g., manufacturing or biological experiments. Despite…

Computation and Language · Computer Science 2025-10-02 Kimihiro Hasegawa , Wiradee Imrattanatrai , Masaki Asada , Ken Fukuda , Teruko Mitamura

This project successfully developed, evaluated and integrated a Voice User Interface (VUI) into a web application that we are developing for immersive molecular graphics. Said app provides augmented and virtual reality (AR and VR)…

Human-Computer Interaction · Computer Science 2026-03-04 Fabio Cortes Rodriguez , Luciano Abriata