中文
相关论文

相关论文: Phoenix-VAD: Streaming Semantic Endpoint Detection…

200 篇论文

Dialogue models falter in noisy, multi-speaker environments, often producing irrelevant responses and awkward turn-taking. We present AV-Dialog, the first multimodal dialog framework that uses both audio and visual cues to track the target…

计算与语言 · 计算机科学 2025-11-17 Tuochao Chen , Bandhav Veluri , Hongyu Gong , Shyamnath Gollakota

Recent advances in multimodal large language models (MLLMs) have enabled impressive progress in vision-language understanding, yet their high computational cost limits deployment in resource-constrained scenarios such as robotic…

计算机视觉与模式识别 · 计算机科学 2025-11-26 Quoc-Huy Trinh

Spoken language models (SLMs) have advanced rapidly in recent years, accompanied by a growing number of evaluation benchmarks. However, most existing benchmarks emphasize task completion and capability scaling, while remaining poorly…

计算与语言 · 计算机科学 2026-01-13 Zehan Li , Hongjie Chen , Qing Wang , Yuxin Zhang , Jing Zhou , Hang Lv , Mengjie Du , Yaodong Song , Jie Lian , Jian Kang , Jie Li , Yongxiang Li , Xuelong Li

Active perception enables robots to dynamically gather information by adjusting their viewpoints, a crucial capability for interacting with complex, partially observable environments. In this paper, we present AP-VLM, a novel framework that…

机器人学 · 计算机科学 2025-06-10 Venkatesh Sripada , Samuel Carter , Frank Guerin , Amir Ghalamzan

While Vision-Language Models (VLMs) and Multimodal Large Language Models (MLLMs) have shown strong generalisation in detecting image and video deepfakes, their use for audio deepfake detection remains largely unexplored. In this work, we…

声音 · 计算机科学 2026-01-05 Akanksha Chuchra , Shukesh Reddy , Sudeepta Mishra , Abhijit Das , Abhinav Dhall

Journalists face mounting challenges in monitoring ever-expanding digital information streams to identify newsworthy content. While traditional automation tools gather information at scale, they struggle with the editorial judgment needed…

人机交互 · 计算机科学 2025-10-01 Nick Hagar , Ethan Silver , Clare Spencer , Nicholas Diakopoulos

Data discovery in data lakes with ever increasing datasets has long been recognized as a big challenge in the realm of data management, especially for semantic search of and hierarchical global catalog generation of tables. While large…

数据库 · 计算机科学 2025-02-24 Qi An , Chihua Ying , Yuqing Zhu , Yihao Xu , Manwei Zhang , Jianmin Wang

Active speaker detection plays a vital role in human-machine interaction. Recently, a few end-to-end audiovisual frameworks emerged. However, these models' inference time was not explored and are not applicable for real-time applications…

声音 · 计算机科学 2022-11-24 Fiseha B. Tesema , Zheyuan Lin , Shiqiang Zhu , Wei Song , Jason Gu , Hong Wu

Vision-language models (VLMs) have recently emerged as powerful representation learning systems that align visual observations with natural language concepts, offering new opportunities for semantic reasoning in safety-critical autonomous…

计算机视觉与模式识别 · 计算机科学 2026-02-19 Ross Greer , Maitrayee Keskar , Angel Martinez-Sanchez , Parthib Roy , Shashank Shriram , Mohan Trivedi

Autonomous driving systems require real-time environmental perception to ensure user safety and experience. Streaming perception is a task of reporting the current state of the world, which is used to evaluate the delay and accuracy of…

计算机视觉与模式识别 · 计算机科学 2023-09-14 Yihui Huang , Ningjiang Chen

Pretrained vision-language models (VLMs), such as CLIP, achieve remarkable zero-shot performance, yet their downstream potential hinges on effective fine-tuning. Most adaptation methods typically focus on refining representation from…

计算机视觉与模式识别 · 计算机科学 2025-07-29 Liang Chen , Ghazi Shazan Ahmad , Tianjun Yao , Lingqiao Liu , Zhiqiang Shen

Multimodal large language models (MLLMs) have advanced zero-shot end-to-end Vision-Language Navigation (VLN), yet robust navigation requires not only semantic understanding but also predictive modeling of environment dynamics and spatial…

计算机视觉与模式识别 · 计算机科学 2026-03-05 Zehua Fan , Wenqi Lyu , Wenxuan Song , Linge Zhao , Yifei Yang , Xi Wang , Junjie He , Lida Huang , Haiyan Liu , Bingchuan Sun , Guangjun Bao , Xuanyao Mao , Liang Xu , Yan Wang , Feng Gao

Large Language Model (LLM)-enhanced agents become increasingly prevalent in Human-AI communication, offering vast potential from entertainment to professional domains. However, current multi-modal dialogue systems overlook the acoustic…

计算与语言 · 计算机科学 2024-06-19 Haoqiu Yan , Yongxin Zhu , Kai Zheng , Bing Liu , Haoyu Cao , Deqiang Jiang , Linli Xu

Despite their advanced reasoning capabilities, state-of-the-art Multimodal Large Language Models (MLLMs) demonstrably lack a core component of human intelligence: the ability to `read the room' and assess deception in complex social…

计算机视觉与模式识别 · 计算机科学 2025-11-21 Caixin Kang , Yifei Huang , Liangyang Ouyang , Mingfang Zhang , Ruicong Liu , Yoichi Sato

Vision-language-action (VLA) models have demonstrated exceptional performance in natural language-driven perception and control. However, the high computational cost of VLA models poses significant efficiency challenges, particularly for…

机器人学 · 计算机科学 2026-03-31 Yiran Shi , Dongqi Guo , Tianchen Zhao , Feng Gao , Liangzhi Shi , Chao Yu , ZhiJian Mo , Qihua Xiao , XiaoShuai Peng , Qingmin Liao , Yu Wang

Deep learning (DL)-based Semantic Communications (SemCom) is becoming critical to maximize overall efficiency of communication networks. Nevertheless, SemCom is sensitive to wireless channel uncertainties, source outliers, and suffer from…

机器学习 · 计算机科学 2025-02-18 Jianhua Pei , Cheng Feng , Ping Wang , Hina Tabassum , Dongyuan Shi

Recent multi-modal Large Language Models (LLMs) such as GPT-4o have demonstrated strong capabilities of direct speech interaction. However, the lack of specialized and comprehensive benchmarks for end-to-end speech LLM evaluation hinders…

计算与语言 · 计算机科学 2025-09-29 Linhao Zhang , Jian Zhang , Bokai Lei , Chuhan Wu , Aiwei Liu , Wei Jia , Xiao Zhou

Writing mathematical notation requires substantial effort, diverting cognitive resources from conceptual understanding to documentation mechanics, significantly impacting individuals with fine motor disabilities (FMDs). Current limits of…

人机交互 · 计算机科学 2025-08-12 Kenneth Ge , Ryan Paul , Priscilla Zhang , JooYoung Seo

By implicitly recognizing a user based on his/her speech input, speaker identification enables many downstream applications, such as personalized system behavior and expedited shopping checkouts. Based on whether the speech content is…

机器学习 · 计算机科学 2021-06-21 Ruirui Li , Chelsea J. -T. Ju , Zeya Chen , Hongda Mao , Oguz Elibol , Andreas Stolcke

How to leverage dynamic contextual information in end-to-end speech recognition has remained an active research area. Previous solutions to this problem were either designed for specialized use cases that did not generalize well to…

‹ 上一页 1 8 9 10 下一页 ›