中文
相关论文

相关论文: VoCopilot: Voice-Activated Tracking of Everyday In…

200 篇论文

Large Multimodal Models (LMMs) have shown strong potential for assisting users in tasks, such as programming, content creation, and information access, yet their interaction remains largely limited to traditional interfaces such as desktops…

人机交互 · 计算机科学 2026-02-12 Liuchuan Yu , Yongqi Zhang , Lap-Fai Yu

In this work, we introduce Speech-Copilot, a modular framework for instruction-oriented speech-processing tasks that minimizes human effort in toolset construction. Unlike end-to-end methods using large audio-language models, Speech-Copilot…

音频与语音处理 · 电气工程与系统科学 2024-09-24 Chun-Yi Kuan , Chih-Kai Yang , Wei-Ping Huang , Ke-Han Lu , Hung-yi Lee

Creating music is iterative, requiring varied methods at each stage. However, existing AI music systems fall short in orchestrating multiple subsystems for diverse needs. To address this gap, we introduce Loop Copilot, a novel system that…

声音 · 计算机科学 2024-09-02 Yixiao Zhang , Akira Maezawa , Gus Xia , Kazuhiko Yamamoto , Simon Dixon

In the rapidly evolving landscape of human-robot collaboration, effective communication between humans and robots is crucial for complex task execution. Traditional request-response systems often lack naturalness and may hinder efficiency.…

机器人学 · 计算机科学 2024-09-12 Davide Ferrari , Filippo Alberi , Cristian Secchi

Understanding and classifying user personas is critical for delivering effective personalization. While persona information offers valuable insights, its full potential is realized only when contextualized, linking user characteristics with…

人机交互 · 计算机科学 2026-02-05 Saleh Afzoon , Amin Beheshti , Usman Naseem

Think-Aloud Computing, a method for capturing users' verbalized thoughts during software tasks, allows eliciting rich contextual insights into evolving intentions, struggles, and decision-making processes of users in real-time. However,…

人机交互 · 计算机科学 2026-02-11 Frederic Gmeiner , John Thompson , George Fitzmaurice , Justin Matejka

In this paper, we introduce a novel Face-to-Face spoken dialogue model. It processes audio-visual speech from user input and generates audio-visual speech as the response, marking the initial step towards creating an avatar chatbot system…

计算机视觉与模式识别 · 计算机科学 2024-08-05 Se Jin Park , Chae Won Kim , Hyeongseop Rha , Minsu Kim , Joanna Hong , Jeong Hun Yeo , Yong Man Ro

Voice assistants have become quite popular lately while in parallel they are an important part of smarthome systems. Through their voice assistants, users can perform various tasks, control other devices and enjoy third party services. The…

密码学与安全 · 计算机科学 2021-09-06 Georgios Germanos , Dimitris Kavallieros , Nicholas Kolokotronis , Nikolaos Georgiou

With the development of smart devices, such as the Amazon Echo and Apple's HomePod, speech data have become a new dimension of big data. However, privacy and security concerns may hinder the collection and sharing of real-world speech data,…

密码学与安全 · 计算机科学 2020-04-17 Yaowei Han , Sheng Li , Yang Cao , Qiang Ma , Masatoshi Yoshikawa

Using our voices to access, and interact with, online services raises concerns about the trade-offs between convenience, privacy, and security. The conflict between maintaining privacy and ensuring input authenticity has often been hindered…

计算机与社会 · 计算机科学 2023-02-28 Ranya Aloufi , Hamed Haddadi , David Boyle

There have been significant innovations in media technologies in the recent years. While these developments have improved experiences for individual users, design of multi-user interfaces still remains a challenge. A relatively unexplored…

人机交互 · 计算机科学 2018-11-20 Sumit Shekhar , Aditya Siddhant , Anindya Shankar Bhandari , Nishant Yadav

Silent and whispered speech offer promise for always-available voice interaction with AI, yet existing methods struggle to balance vocabulary size, wearability, silence, and noise robustness. We present NasoVoce, a nose-bridge-mounted…

人机交互 · 计算机科学 2026-03-12 Jun Rekimoto , Yu Nishimura , Bojian Yang

Healthcare professionals work in complex, high-stakes environments where effective communication is critical for care delivery, team coordination, and individual well-being. However, communication activity in everyday clinical settings…

Since the 1990s, the telephone has been the primary mode of communication. However, Voice over Internet Protocol (VoIP), which is a highly straightforward and affordable form of data transfer, is now becoming an important part of daily…

密码学与安全 · 计算机科学 2025-07-30 M. Matsive Ali

Dialogue models falter in noisy, multi-speaker environments, often producing irrelevant responses and awkward turn-taking. We present AV-Dialog, the first multimodal dialog framework that uses both audio and visual cues to track the target…

计算与语言 · 计算机科学 2025-11-17 Tuochao Chen , Bandhav Veluri , Hongyu Gong , Shyamnath Gollakota

Audio has become an increasingly crucial biometric modality due to its ability to provide an intuitive way for humans to interact with machines. It is currently being used for a range of applications, including person authentication to…

声音 · 计算机科学 2023-07-14 Rishabh Ranjan , Mayank Vatsa , Richa Singh

Visual Question-Answering, a technology that generates textual responses from an image and natural language question, has progressed significantly. Notably, it can aid in tracking and inquiring about daily activities, crucial in healthcare…

机器学习 · 计算机科学 2024-10-29 Wenqiang Chen , Jiaxuan Cheng , Leyao Wang , Wei Zhao , Wojciech Matusik

This work focuses on reliable detection of bird sound emissions as recorded in the open field. Acoustic detection of avian sounds can be used for the automatized monitoring of multiple bird taxa and querying in long-term recordings for…

声音 · 计算机科学 2016-09-28 Ilyas Potamitis

User identification plays a pivotal role in how we interact with our mobile devices. Many existing authentication approaches require active input from the user or specialized sensing hardware, and studies on mobile device usage show…

人机交互 · 计算机科学 2020-04-13 Yilin Yang , Chen Wang , Yingying Chen , Yan Wang

Recent advancements in zero-shot text-to-speech (TTS) modeling have led to significant strides in generating high-fidelity and diverse speech. However, dialogue generation, along with achieving human-like naturalness in speech, continues to…

音频与语音处理 · 电气工程与系统科学 2024-12-17 Leying Zhang , Yao Qian , Long Zhou , Shujie Liu , Dongmei Wang , Xiaofei Wang , Midia Yousefi , Yanmin Qian , Jinyu Li , Lei He , Sheng Zhao , Michael Zeng
‹ 上一页 1 2 3 10 下一页 ›