中文
相关论文

相关论文: VoiceAlign: A Shimming Layer for Enhancing the Usa…

200 篇论文

Large Language Models (LLMs) have revolutionized natural language processing, but their application to speech-based tasks remains challenging due to the complexities of integrating audio and text modalities. This paper introduces Ichigo, a…

计算与语言 · 计算机科学 2025-04-07 Alan Dao , Dinh Bach Vu , Huy Hoang Ha

Audio-Visual Large Language Models (AVLLMs) are emerging as unified interfaces to multimodal perception. We present the first mechanistic interpretability study of AVLLMs, analyzing how audio and visual features evolve and fuse through…

人工智能 · 计算机科学 2026-04-06 Ramaneswaran Selvakumar , Kaousheik Jayakumar , S Sakshi , Sreyan Ghosh , Ruohan Gao , Dinesh Manocha

Deploying conversational voice agents with large language models faces a critical challenge: cloud-based foundation models provide deep reasoning and domain knowledge but introduce latency that disrupts natural conversation, while on-device…

计算与语言 · 计算机科学 2025-11-11 Vidya Srinivas , Zachary Englhardt , Maximus Powers , Shwetak Patel , Vikram Iyer

Generating responsive listener head dynamics with nuanced emotions and expressive reactions is crucial for practical dialogue modeling in various virtual avatar animations. Previous studies mainly focus on the direct short-term production…

计算机视觉与模式识别 · 计算机科学 2025-05-08 Shiying Li , Xingqun Qi , Bingkun Yang , Chen Weile , Zezhao Tian , Muyi Sun , Qifeng Liu , Man Zhang , Zhenan Sun

Deploying modern Speech Language Models (SpeechLMs) in streaming settings requires systems that provide low latency, high throughput, and strong guarantees of streamability. Existing systems fall short of supporting diverse models flexibly…

While voice user interfaces offer increased accessibility due to hands-free and eyes-free interactions, older adults often have challenges such as constructing structured requests and perceiving how such devices operate. Voice-first user…

AI alignment is about ensuring AI systems only pursue goals and activities that are beneficial to humans. Most of the current approach to AI alignment is to learn what humans value from their behavioural data. This paper proposes a…

Large Language Model (LLM) agents deployed for real-world tasks face a fundamental dilemma: user requests are underspecified, yet agents must decide whether to act on incomplete information or interrupt users for clarification. Existing…

计算与语言 · 计算机科学 2026-01-13 Yijiang River Dong , Tiancheng Hu , Zheng Hui , Caiqi Zhang , Ivan Vulić , Andreea Bobu , Nigel Collier

Vision-language-action models (VLAs) have become increasingly popular in robot manipulation for their end-to-end design and remarkable performance. However, existing VLAs rely heavily on vision-language models (VLMs) that only support…

机器人学 · 计算机科学 2025-02-24 Wei Zhao , Pengxiang Ding , Min Zhang , Zhefei Gong , Shuanghao Bai , Han Zhao , Donglin Wang

The burgeoning field of on-device AI communication, where devices exchange information directly through embedded foundation models, such as language models (LMs), requires robust, efficient, and generalizable communication frameworks.…

信息论 · 计算机科学 2024-07-02 Ju-Hyung Lee , Dong-Ho Lee , Joohan Lee , Jay Pujara

Adaptive interfaces can help users perform sequential decision-making tasks like robotic teleoperation given noisy, high-dimensional command signals (e.g., from a brain-computer interface). Recent advances in human-in-the-loop machine…

机器人学 · 计算机科学 2023-09-08 Jensen Gao , Siddharth Reddy , Glen Berseth , Anca D. Dragan , Sergey Levine

Recent advances in diffusion-based lip-syncing generative models have demonstrated their ability to produce highly synchronized talking face videos for visual dubbing. Although these models excel at lip synchronization, they often struggle…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Yanyu Zhu , Lichen Bai , Jintao Xu , Hai-tao Zheng

Embodied agents designed to assist users with tasks must engage in natural language interactions, interpret instructions, execute actions, and communicate effectively to resolve issues. However, collecting large-scale, diverse datasets of…

计算与语言 · 计算机科学 2024-11-01 Daniel Philipov , Vardhan Dongre , Gokhan Tur , Dilek Hakkani-Tür

This paper analyses Conversational AI multi-agent interoperability frameworks and describes the novel architecture proposed by the Open Voice Interoperability initiative (Linux Foundation AI and DATA), also known briefly as OVON (Open Voice…

人工智能 · 计算机科学 2024-07-30 Diego Gosmar , Deborah A. Dahl , Emmett Coin

With the rapid advancement and adoption of Audio Large Language Models (ALLMs), voice agents are now being deployed in high-stakes domains such as banking, customer service, and IT support. However, their vulnerabilities to adversarial…

密码学与安全 · 计算机科学 2026-02-10 Xiang Li , Pin-Yu Chen , Wenqi Wei

As Speech Language Models (SLMs) transition from personal devices to shared, multi-user environments such as smart homes, a new challenge emerges: the model is expected to distinguish between users to manage information flow appropriately.…

音频与语音处理 · 电气工程与系统科学 2026-01-29 Yuxiang Wang , Hongyu Liu , Dekun Chen , Xueyao Zhang , Zhizheng Wu

The combination of Large Language Models (LLM) and Automatic Speech Recognition (ASR), when deployed on edge devices (called edge ASR-LLM), can serve as a powerful personalized assistant to enable audio-based interaction for users. Compared…

Speech synthesis, voice cloning, and voice conversion techniques present severe privacy and security threats to users of voice user interfaces (VUIs). These techniques transform one or more elements of a speech signal, e.g., identity and…

密码学与安全 · 计算机科学 2021-07-23 Ranya Aloufi , Hamed Haddadi , David Boyle

The increasing deployment of autonomous AI agents on the web is hampered by a fundamental misalignment: agents must infer affordances from human-oriented user interfaces, leading to brittle, inefficient, and insecure interactions. To…

人机交互 · 计算机科学 2025-11-17 Sven Schultze , Meike Verena Kietzmann , Nils-Lucas Schönfeld , Ruth Stock-Homburg

Recent advancements in vision-language models have achieved remarkable results in making language models understand vision inputs. However, a unified approach to align these models across diverse tasks such as image captioning and visual…

计算机视觉与模式识别 · 计算机科学 2025-03-26 Kartik Jangra , Aman Kumar Singh , Yashwani Mann , Geetanjali Rathee