中文
相关论文

相关论文: Mobile-O: Unified Multimodal Understanding and Gen…

200 篇论文

Vision-and-Language Navigation (VLN) requires agents to autonomously navigate complex environments via visual images and natural language instructions--remains highly challenging. Recent research on enhancing language-guided navigation…

人工智能 · 计算机科学 2026-02-10 Changxin Huang , Lv Tang , Zhaohuan Zhan , Lisha Yu , Runhao Zeng , Zun Liu , Zhengjie Wang , Jianqiang Li

This paper investigates robust semantic communications over multiple-input multiple-output (MIMO) fading channels. Current semantic communications over MIMO channels mainly focus on channel adaptive encoding and decoding, which lacks…

信息论 · 计算机科学 2024-07-09 Yiheng Duan , Tong Wu , Zhiyong Chen , Meixia Tao

Interpreting the decisions of deep learning models has been actively studied since the explosion of deep neural networks. One of the most convincing interpretation approaches is salience-based visual interpretation, such as Grad-CAM, where…

计算机视觉与模式识别 · 计算机科学 2023-10-17 Yiming Lei , Zilong Li , Yangyang Li , Junping Zhang , Hongming Shan

End-to-end human animation, such as audio-driven talking human generation, has undergone notable advancements in the recent few years. However, existing methods still struggle to scale up as large general video generation models, limiting…

计算机视觉与模式识别 · 计算机科学 2025-07-01 Gaojie Lin , Jianwen Jiang , Jiaqi Yang , Zerong Zheng , Chao Liang

Recent advances in Large Multi-modal Models (LMMs) have demonstrated their remarkable success as general-purpose multi-modal assistants, with particular focuses on holistic image- and video-language understanding. Conversely, less attention…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Ye Liu , Zongyang Ma , Junfu Pu , Zhongang Qi , Yang Wu , Ying Shan , Chang Wen Chen

Recent advances in vision-language pre-training have enabled machines to perform better in multimodal object discrimination (e.g., image-text semantic alignment) and image synthesis (e.g., text-to-image generation). On the other hand,…

计算机视觉与模式识别 · 计算机科学 2023-06-02 Xiao Dong , Runhui Huang , Xiaoyong Wei , Zequn Jie , Jianxing Yu , Jian Yin , Xiaodan Liang

MiniMind-O is an open 0.1B-scale omni model built on the MiniMind language model. It accepts text, speech, and image inputs, and returns both text and streaming speech. The release includes model code, checkpoints, and the main Parquet…

声音 · 计算机科学 2026-05-06 Jingyao Gong

Cross-view geo-localization (CVGL) plays a vital role in drone-based multimedia applications, enabling precise localization by matching drone-captured aerial images against geo-tagged satellite databases in GNSS-denied environments.…

计算机视觉与模式识别 · 计算机科学 2026-01-09 Jian Sun , Kangdao Liu , Chi Zhang , Chuangquan Chen , Junge Shen , C. L. Philip Chen , Chi-Man Vong

In emergencies, the ability to quickly and accurately gather environmental data and command information, and to make timely decisions, is particularly critical. Traditional semantic communication frameworks, primarily based on a single…

计算机视觉与模式识别 · 计算机科学 2024-08-13 Weiqi Fu , Lianming Xu , Xin Wu , Haoyang Wei , Li Wang

Speech-driven gestures and facial animations are fundamental to expressive digital avatars in games, virtual production, and interactive media. However, existing methods are either limited to a single modality for audio motion alignment,…

Current vision-language models have been explored for multi-modal embedding tasks like information retrieval. However, they face significant challenges in real-world queries and targets involving diverse modality combinations, as existing…

计算机视觉与模式识别 · 计算机科学 2026-05-07 Jiajun Qin , Yuan Pu , Zhuolun He , Seunggeun Kim , David Z. Pan , Bei Yu

In the current landscape of artificial intelligence, foundation models serve as the bedrock for advancements in both language and vision domains. OpenAI GPT-4 has emerged as the pinnacle in large language models (LLMs), while the computer…

计算机视觉与模式识别 · 计算机科学 2023-11-20 Chris Kelly , Luhui Hu , Cindy Yang , Yu Tian , Deshun Yang , Bang Yang , Zaoshan Huang , Zihao Li , Yuexian Zou

Unified models aim to support both understanding and generation by encoding images into discrete tokens and processing them alongside text within a single autoregressive framework. This unified design offers architectural simplicity and…

计算机视觉与模式识别 · 计算机科学 2026-03-13 Ziyao Wang , Chen Chen , Jingtao Li , Weiming Zhuang , Jiabo Huang , Ang Li , Lingjuan Lyu

Unified vision large language models (VLLMs) have recently achieved impressive advancements in both multimodal understanding and generation, powering applications such as visual question answering and text-guided image synthesis. However,…

计算与语言 · 计算机科学 2025-09-19 Pengyu Wang , Shaojun Zhou , Chenkun Tan , Xinghao Wang , Wei Huang , Zhen Ye , Zhaowei Li , Botian Jiang , Dong Zhang , Xipeng Qiu

Multi-modal generative AI (Artificial Intelligence) has attracted increasing attention from both academia and industry. Particularly, two dominant families of techniques have emerged: i) Multi-modal large language models (LLMs) demonstrate…

人工智能 · 计算机科学 2025-11-26 Xin Wang , Yuwei Zhou , Bin Huang , Hong Chen , Wenwu Zhu

Precise modeling of channel multipath is essential for understanding wireless propagation environments and optimizing communication systems. In particular, sixth-generation (6G) artificial intelligence (AI)-native communication systems…

信号处理 · 电气工程与系统科学 2025-11-20 Zengrui Han , Lu Bai , Xuesong Cai , Xiang Cheng

We propose the first joint audio-video generation framework that brings engaging watching and listening experiences simultaneously, towards high-quality realistic videos. To generate joint audio-video pairs, we propose a novel Multi-Modal…

计算机视觉与模式识别 · 计算机科学 2023-03-27 Ludan Ruan , Yiyang Ma , Huan Yang , Huiguo He , Bei Liu , Jianlong Fu , Nicholas Jing Yuan , Qin Jin , Baining Guo

Ensuring model explainability and robustness is essential for reliable deployment of deep vision systems. Current methods for evaluating robustness rely on collecting and annotating extensive test sets. While this is common practice, the…

计算机视觉与模式识别 · 计算机科学 2024-10-10 Yinong Oliver Wang , Eileen Li , Jinqi Luo , Zhaoning Wang , Fernando De la Torre

Recent motion-aware large language models have demonstrated promising potential in unifying motion comprehension and generation. However, existing approaches primarily focus on coarse-grained motion-text modeling, where text describes the…

计算机视觉与模式识别 · 计算机科学 2025-04-04 Bizhu Wu , Jinheng Xie , Keming Shen , Zhe Kong , Jianfeng Ren , Ruibin Bai , Rong Qu , Linlin Shen

Multimodal learning has rapidly advanced visual understanding, largely via multimodal large language models (MLLMs) that use powerful LLMs as cognitive cores. In visual generation, however, these powerful core models are typically reduced…

计算机视觉与模式识别 · 计算机科学 2025-12-15 Han Lin , Xichen Pan , Ziqi Huang , Ji Hou , Jialiang Wang , Weifeng Chen , Zecheng He , Felix Juefei-Xu , Junzhe Sun , Zhipeng Fan , Ali Thabet , Mohit Bansal , Chu Wang