中文
相关论文

相关论文: OmniDataComposer: A Unified Data Structure for Mul…

200 篇论文

Real-world problems are often dependent on multiple data modalities, making multimodal fusion essential for leveraging diverse information sources. In high-stakes domains, such as in healthcare, understanding how each modality contributes…

神经与进化计算 · 计算机科学 2025-05-19 Mafalda Malafaia , Thalea Schlender , Tanja Alderliesten , Peter A. N. Bosman

Recently, the AI community has made significant strides in developing powerful foundation models, driven by large-scale multimodal datasets. However, for audio representation learning, existing datasets suffer from limitations in the…

声音 · 计算机科学 2024-09-10 Luoyi Sun , Xuenan Xu , Mengyue Wu , Weidi Xie

The ability of Large Language Models (LLMs) to generate structured outputs that follow arbitrary schemas is crucial to a wide range of downstream tasks that require diverse structured representations of results such as information…

计算与语言 · 计算机科学 2025-11-25 James Y. Huang , Wenxuan Zhou , Nan Xu , Fei Wang , Qin Liu , Sheng Zhang , Hoifung Poon , Muhao Chen

Recent advancements in leveraging pre-trained 2D diffusion models achieve the generation of high-quality novel views from a single in-the-wild image. However, existing works face challenges in producing controllable novel views due to the…

计算机视觉与模式识别 · 计算机科学 2025-07-04 Yunhan Yang , Shuo Chen , Yukun Huang , Xiaoyang Wu , Yuan-Chen Guo , Edmund Y. Lam , Hengshuang Zhao , Tong He , Xihui Liu

Urban scene synthesis with video generation models has recently shown great potential for autonomous driving. Existing video generation approaches to autonomous driving primarily focus on RGB video generation and lack the ability to support…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Guile Wu , David Huang , Dongfeng Bai , Bingbing Liu

We introduce OmnixR, an evaluation suite designed to benchmark SoTA Omni-modality Language Models, such as GPT-4o and Gemini. Evaluating OLMs, which integrate multiple modalities such as text, vision, and audio, presents unique challenges.…

Multimodal sentiment analysis, a pivotal task in affective computing, seeks to understand human emotions by integrating cues from language, audio, and visual signals. While many recent approaches leverage complex attention mechanisms and…

计算与语言 · 计算机科学 2025-05-09 Nischal Mandal , Yang Li

Big Data are rapidly produced from various heterogeneous data sources. They are of different types (text, image, video or audio) and have different levels of reliability and completeness. One of the most interesting architectures that deal…

人工智能 · 计算机科学 2021-08-11 Siham Yousfi , Maryem Rhanoui , Dalila Chiadmi

Pedestrian Attribute Recognition is a foundational computer vision task that provides essential support for downstream applications, including person retrieval in video surveillance and intelligent retail analytics. However, existing…

计算机视觉与模式识别 · 计算机科学 2026-03-06 Minghe Xu , Rouying Wu , Jiarui Xu , Minhao Sun , Zikang Yan , Xiao Wang , ChiaWei Chu , Yu Li

Autonomous vehicles (AVs) are poised to redefine transportation by enhancing road safety, minimizing human error, and optimizing traffic efficiency. The success of AVs depends on their ability to interpret complex, dynamic environments…

多媒体 · 计算机科学 2025-07-11 Abolfazl Zarghani , Amirhossein Ebrahimi , Amir Malekesfandiari

The convergence of text, visual, and audio data is a key step towards human-like artificial intelligence, however the current Vision-Language-Speech landscape is dominated by encoder-only models which lack generative abilities. We propose…

With the advancement of generative models, the synthesis of different sensory elements such as music, visuals, and speech has achieved significant realism. However, the approach to generate multi-sensory outputs has not been fully explored,…

计算机视觉与模式识别 · 计算机科学 2024-08-22 Minheng Ni , Chenfei Wu , Huaying Yuan , Zhengyuan Yang , Ming Gong , Lijuan Wang , Zicheng Liu , Wangmeng Zuo , Nan Duan

The widespread adoption of mobile devices and data collection technologies has led to an exponential increase in trajectory data, presenting significant challenges in spatio-temporal data mining, particularly for efficient and accurate…

计算机视觉与模式识别 · 计算机科学 2025-05-26 Yuanshao Zhu , James Jianqiao Yu , Xiangyu Zhao , Xiao Han , Qidong Liu , Xuetao Wei , Yuxuan Liang

Multi-sensor modal fusion has demonstrated strong advantages in 3D object detection tasks. However, existing methods that fuse multi-modal features require transforming features into the bird's eye view space and may lose certain…

计算机视觉与模式识别 · 计算机科学 2023-10-10 Chunyong Hu , Hang Zheng , Kun Li , Jianyun Xu , Weibo Mao , Maochun Luo , Lingxuan Wang , Mingxia Chen , Qihao Peng , Kaixuan Liu , Yiru Zhao , Peihan Hao , Minzhe Liu , Kaicheng Yu

We introduce JamendoMaxCaps, a large-scale music-caption dataset featuring over 362,000 freely licensed instrumental tracks from the renowned Jamendo platform. The dataset includes captions generated by a state-of-the-art captioning model,…

声音 · 计算机科学 2025-05-19 Abhinaba Roy , Renhang Liu , Tongyu Lu , Dorien Herremans

Given a video and a set of input object masks, an omnimatte method aims to decompose the video into semantically meaningful layers containing individual objects along with their associated effects, such as shadows and reflections. Existing…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Yao-Chih Lee , Erika Lu , Sarah Rumbley , Michal Geyer , Jia-Bin Huang , Tali Dekel , Forrester Cole

Utilizing multi-modal data enhances scene understanding by providing complementary semantic and geometric information. Existing methods fuse features or distill knowledge from multiple modalities into a unified representation, improving…

计算机视觉与模式识别 · 计算机科学 2025-06-05 Jialei Chen , Xu Zheng , Danda Pani Paudel , Luc Van Gool , Hiroshi Murase , Daisuke Deguchi

Whole-body multi-modal human motion generation poses two primary challenges: creating an effective motion generation mechanism and integrating various modalities, such as text, speech, and music, into a cohesive framework. Unlike previous…

计算机视觉与模式识别 · 计算机科学 2025-10-17 Zhe Li , Weihao Yuan , Weichao Shen , Siyu Zhu , Zilong Dong , Chang Xu

Omni-modal Large Language Models (OLLMs) that process text, images, videos, and audio introduce new challenges for safety and value guardrails in human-AI interaction. Prior guardrail research largely targets unimodal settings and typically…

人工智能 · 计算机科学 2025-12-03 Boyu Zhu , Xiaofei Wen , Wenjie Jacky Mo , Tinghui Zhu , Yanan Xie , Peng Qi , Muhao Chen

We present PresentAgent, a multimodal agent that transforms long-form documents into narrated presentation videos. While existing approaches are limited to generating static slides or text summaries, our method advances beyond these…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Jingwei Shi , Zeyu Zhang , Biao Wu , Yanjie Liang , Meng Fang , Ling Chen , Yang Zhao