中文
相关论文

相关论文: ROMA: Real-time Omni-Multimodal Assistant with Int…

200 篇论文

In this article, we provide both analytical and numerical performance analysis of multi-service oriented multiple access (MOMA), a recently proposed non-orthogonal multiple-access scheme for scenarios with a massive number of concurrent…

信息论 · 计算机科学 2017-10-31 Nassar Ksairi , Mérouane Debbah

Training multimodal large language models (MLLMs) for video understanding requires large-scale annotated data spanning diverse tasks such as object counting, question answering, and segmentation. However, collecting and annotating…

计算机视觉与模式识别 · 计算机科学 2026-04-15 Tanzila Rahman , Renjie Liao , Leonid Sigal

We introduce Omni-MMSI, a new task that requires comprehensive social interaction understanding from raw audio, vision, and speech input. The task involves perceiving identity-attributed social cues (e.g., who is speaking what) and…

计算机视觉与模式识别 · 计算机科学 2026-04-02 Xinpeng Li , Bolin Lai , Hardy Chen , Shijian Deng , Cihang Xie , Yuyin Zhou , James Matthew Rehg , Yapeng Tian

GPT-4o, an all-encompassing model, represents a milestone in the development of large multi-modal language models. It can understand visual, auditory, and textual modalities, directly output audio, and support flexible duplex interaction.…

音频与语音处理 · 电气工程与系统科学 2024-11-06 Zhifei Xie , Changqiao Wu

Streaming video understanding demands more than watching longer videos: assistants must decide when to speak in real time, balancing responsiveness against verbosity. Yet most video-language models (VideoLLMs) are trained for offline…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Zichen Wen , Boxue Yang , Junlong Ke , Jiajie Huang , Chenfei Liao , Junxi Wang , Xuyang Liu , Linfeng Zhang

Recent advances in Large Language Models (LLMs) have enabled the development of Video-LLMs, advancing multimodal learning by bridging video data with language tasks. However, current video understanding models struggle with processing long…

计算机视觉与模式识别 · 计算机科学 2025-01-24 Haomiao Xiong , Zongxin Yang , Jiazuo Yu , Yunzhi Zhuge , Lu Zhang , Jiawen Zhu , Huchuan Lu

What makes good representations for video understanding, such as anticipating future activities, or answering video-conditioned questions? While earlier approaches focus on end-to-end learning directly from video pixels, we propose to…

计算机视觉与模式识别 · 计算机科学 2024-07-16 Shijie Wang , Qi Zhao , Minh Quan Do , Nakul Agarwal , Kwonjoon Lee , Chen Sun

Chatbots via large language models (LLMs) generate fluent responses but often struggle with when to speak, especially for brief, timely listener reactions during ongoing dialogue. We present a multimodal strategy for LLMs, which leverages…

计算机视觉与模式识别 · 计算机科学 2026-05-21 Zikai Liao , Yi Ouyang , Yi-Lun Lee , Chen-Ping Yu , Yi-Hsuan Tsai , Zhaozheng Yin

Procedural tasks with multiple ordered steps are ubiquitous in daily life. Recent advances in multimodal large language models (MLLMs) have enabled personal assistants that support daily activities. However, existing systems primarily…

人工智能 · 计算机科学 2026-05-07 Lilin Xu , Bufang Yang , Siyang Jiang , Kaiwei Liu , Kaiyuan Hou , Yuang Fan , Hongkai Chen , Zhenyu Yan , Xiaofan Jiang

The ability to comprehend audio--which includes speech, non-speech sounds, and music--is crucial for AI agents to interact effectively with the world. We present MMAU, a novel benchmark designed to evaluate multimodal audio understanding…

音频与语音处理 · 电气工程与系统科学 2024-10-28 S Sakshi , Utkarsh Tyagi , Sonal Kumar , Ashish Seth , Ramaneswaran Selvakumar , Oriol Nieto , Ramani Duraiswami , Sreyan Ghosh , Dinesh Manocha

Recent advances in diffusion-based video generation have substantially improved visual fidelity and temporal coherence. However, most existing approaches remain task-specific and rely primarily on textual instructions, limiting their…

Extracting real-time insights from multi-modal data streams from various domains such as healthcare, intelligent transportation, and satellite remote sensing remains a challenge. High computational demands and limited knowledge scope…

计算机视觉与模式识别 · 计算机科学 2025-01-27 Murugan Sankaradas , Ravi K. Rajendran , Srimat T. Chakradhar

Recommendation systems have become popular and effective tools to help users discover their interesting items by modeling the user preference and item property based on implicit interactions (e.g., purchasing and clicking). Humans perceive…

信息检索 · 计算机科学 2023-02-10 Hongyu Zhou , Xin Zhou , Zhiwei Zeng , Lingzi Zhang , Zhiqi Shen

Recently, semantic communication (SC) has garnered increasing attention for its efficiency, yet it remains vulnerable to semantic jamming attacks. These attacks entail introducing crafted perturbation signals to legitimate signals over the…

信号处理 · 电气工程与系统科学 2025-01-03 Kequan Zhou , Guangyi Zhang , Yunlong Cai , Qiyu Hu , Guanding Yu

Recent multimodal large language models (MLLMs) achieve strong performance on reactive question answering, but real-world streaming assistants require proactive reasoning over continuous visual inputs. Existing benchmarks mainly study…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Jinzhao Li , Yinuo Chen , Wenxuan Song , Yijia Lei , Yichi Zhang , Honglei Yan , Panwang Pan , Miao Liu

Referring audio-visual segmentation (RAVS) has recently seen significant advancements, yet challenges remain in integrating multimodal information and deeply understanding and reasoning about audiovisual content. To extend the boundaries of…

计算机视觉与模式识别 · 计算机科学 2025-08-01 Kaining Ying , Henghui Ding , Guangquan Jie , Yu-Gang Jiang

Real-world live retrieval-augmented generation (RAG) systems face significant challenges when processing user queries that are often noisy, ambiguous, and contain multiple intents. While RAG enhances large language models (LLMs) with…

计算与语言 · 计算机科学 2025-06-27 Guanting Dong , Xiaoxi Li , Yuyao Zhang , Mengjie Deng

We introduce Vinci, a real-time embodied smart assistant built upon an egocentric vision-language model. Designed for deployment on portable devices such as smartphones and wearable cameras, Vinci operates in an "always on" mode,…

Most popular goal-oriented dialogue agents are capable of understanding the conversational context. However, with the surge of virtual assistants with screen, the next generation of agents are required to also understand screen context in…

机器学习 · 计算机科学 2021-11-26 Sanchit Agarwal , Jan Jezabek , Arijit Biswas , Emre Barut , Shuyang Gao , Tagyoung Chung

Long-horizon omnimodal question answering answers questions by reasoning over text, images, audio, and video. Despite recent progress on OmniLLMs, low-resource long audio-video QA still suffers from costly dense encoding, weak fine-grained…

计算与语言 · 计算机科学 2026-03-31 Yifan Zhu , Xinyu Mu , Tao Feng , Zhonghong Ou , Yuning Gong , Haoran Luo