English
Related papers

Related papers: End-to-end Listen, Look, Speak and Act

200 papers

Recent advances in spoken dialogue language models have shifted from turn-based to full-duplex designs, where the model continuously listens to the user while generating responses. However, existing duplex backbones still lack a native…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-21 Haoyang Zhang , Jun Chen , Donghang Wu , Yuxin Li , Yuxin Zhang , Xiangyu Tony Zhang , Che Liu , Qingjian Lin , Yizhou Peng , Hexin Liu , Eng Siong Chng , Chao Yan , Boyong Wu , Yechang Huang , Xuerui Yang , Fei Tian

Current Vision-Language-Action (VLA) models are often constrained by a rigid, static interaction paradigm, which lacks the ability to see, hear, speak, and act concurrently as well as handle real-time user interruptions dynamically. This…

Vision Language Action (VLA) models promise an open-vocabulary interface that can translate perceptual ambiguity into semantically grounded driving decisions, yet they still treat language as a static prior fixed at inference time. As a…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-30 Ziang Guo , Feng Yang , Xuefeng Zhang , Jiaqi Guo , Kun Zhao , Yixiao Zhou , Peng Lu , Sifa Zheng , Zufeng Zhang

Vision-language-action models (VLAs) have shown generalization capabilities in robotic manipulation tasks by inheriting from vision-language models (VLMs) and learning action generation. Most VLA models focus on interpreting vision and…

We introduce Ella, an embodied social agent capable of lifelong learning within a community in a 3D open world, where agents accumulate experiences and acquire knowledge through everyday visual observations and social interactions. At the…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Hongxin Zhang , Zheyuan Zhang , Zeyuan Wang , Zunzhe Zhang , Lixing Fang , Qinhong Zhou , Chuang Gan

Interactive virtual humanoid agent is a crucial interface with the physical world. A relatively complete humanoid agent first needs to have face and body, then possess both verbal and non-verbal (such as eye contact, facial expression, lip…

Computer Vision and Pattern Recognition · Computer Science 2024-08-07 Tenglong Ao

Spoken language understanding, which extracts intents and/or semantic concepts in utterances, is conventionally formulated as a post-processing of automatic speech recognition. It is usually trained with oracle transcripts, but needs to…

Sound · Computer Science 2020-07-30 Viet-Trung Dang , Tianyu Zhao , Sei Ueno , Hirofumi Inaguma , Tatsuya Kawahara

In this work, we extend the instruction-tuned Llama-2 model with end-to-end general-purpose speech processing and reasoning abilities while maintaining the wide range of original LLM capabilities, without using any carefully curated paired…

Computation and Language · Computer Science 2024-04-16 Yassir Fathullah , Chunyang Wu , Egor Lakomkin , Ke Li , Junteng Jia , Yuan Shangguan , Jay Mahadeokar , Ozlem Kalinli , Christian Fuegen , Mike Seltzer

The rapid development of Vision-Language models (VLMs) and Multimodal Language Models (MLLMs) in autonomous driving research has significantly reshaped the landscape by enabling richer scene understanding, context-aware reasoning, and more…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Karthik Mohan , Sonam Singh , Amit Arvind Kale

Emergency Medical Services (EMS) responders often operate under time-sensitive conditions, facing cognitive overload and inherent risks, requiring essential skills in critical thinking and rapid decision-making. This paper presents…

Artificial Intelligence · Computer Science 2024-10-27 Keshara Weerasinghe , Saahith Janapati , Xueren Ge , Sion Kim , Sneha Iyer , John A. Stankovic , Homa Alemzadeh

Humans possess a unified cognitive ability to perceive, comprehend, and interact with the physical world. Why can't large language models replicate this holistic understanding? Through a systematic analysis of existing training paradigms in…

Synthesized speech from articulatory movements can have real-world use for patients with vocal cord disorders, situations requiring silent speech, or in high-noise environments. In this work, we present EMA2S, an end-to-end multimodal…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-10 Yu-Wen Chen , Kuo-Hsuan Hung , Shang-Yi Chuang , Jonathan Sherman , Wen-Chin Huang , Xugang Lu , Yu Tsao

Large Language Models (LLMs) suffer severe catastrophic forgetting when adapted sequentially to new tasks in a continual learning (CL) setting. Existing approaches are fundamentally limited: replay-based methods are impractical and…

Machine Learning · Computer Science 2026-01-08 Shristi Das Biswas , Yue Zhang , Anwesan Pal , Radhika Bhargava , Kaushik Roy

In this paper, we propose GTA-VLA(Guide, Think, Act), an interactive Vision-Language-Action (VLA) framework that enables spatially steerable embodied reasoning by allowing users to guide robot policies with explicit visual cues. Existing…

Robotics · Computer Science 2026-05-14 Yiran Ling , Qing Lian , Jinghang Li , Qing Jiang , Tianming Zhang , Xiaoke Jiang , Chuanxiu Liu , Jie Liu , Lei Zhang

Generating lifelike conversational avatars requires modeling not just isolated speakers, but the dynamic, reciprocal interaction of speaking and listening. However, modeling the listener is exceptionally challenging: direct audio-driven…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Xuangeng Chu , Ruicong Liu , Yifei Huang , Yun Liu , Yichen Peng , Bo Zheng

Large Language Models (LLMs) have recently shown remarkable ability to process not only text but also multimodal inputs such as speech and audio. However, most existing models primarily focus on analyzing input signals using text…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-20 Junyi Ao , Dekun Chen , Xiaohai Tian , Wenjie Feng , Jun Zhang , Lu Lu , Yuxuan Wang , Haizhou Li , Zhizheng Wu

Previous research on empathetic dialogue systems has mostly focused on generating responses given certain emotions. However, being empathetic not only requires the ability of generating emotional responses, but more importantly, requires…

Computation and Language · Computer Science 2019-08-22 Zhaojiang Lin , Andrea Madotto , Jamin Shin , Peng Xu , Pascale Fung

Voice Assistants such as Alexa, Siri, and Google Assistant typically use a two-stage Spoken Language Understanding pipeline; first, an Automatic Speech Recognition (ASR) component to process customer speech and generate text transcriptions,…

Computation and Language · Computer Science 2020-12-17 Subendhu Rongali , Beiye Liu , Liwei Cai , Konstantine Arkoudas , Chengwei Su , Wael Hamza

To operate effectively in the real world, robots should integrate multimodal reasoning with precise action generation. However, existing vision-language-action (VLA) models often sacrifice one for the other, narrow their abilities to…

Robotics · Computer Science 2026-03-04 Shuai Yang , Hao Li , Bin Wang , Yilun Chen , Yang Tian , Tai Wang , Hanqing Wang , Feng Zhao , Yiyi Liao , Jiangmiao Pang

In spoken question answering, the systems are designed to answer questions from contiguous text spans within the related speech transcripts. However, the most natural way that human seek or test their knowledge is via human conversations.…

Computation and Language · Computer Science 2022-05-02 Chenyu You , Nuo Chen , Fenglin Liu , Shen Ge , Xian Wu , Yuexian Zou
‹ Prev 1 2 3 10 Next ›