中文
相关论文

相关论文: Multi-Modal Scene Graph with Kolmogorov-Arnold Exp…

200 篇论文

Kolmogorov-Arnold Networks(KANs), as a theoretically efficient neural network architecture, have garnered attention for their potential in capturing complex patterns. However, their application in computer vision remains relatively…

计算机视觉与模式识别 · 计算机科学 2024-11-15 Yueyang Cang , Yu hang liu , Li Shi

Cross-modal retrieval between videos and texts has attracted growing attentions due to the rapid emergence of videos on the web. The current dominant approach for this problem is to learn a joint embedding space to measure cross-modal…

计算机视觉与模式识别 · 计算机科学 2020-03-03 Shizhe Chen , Yida Zhao , Qin Jin , Qi Wu

Video Question Answering (VQA) inherently relies on multimodal reasoning, integrating visual, temporal, and linguistic cues to achieve a deeper understanding of video content. However, many existing methods rely on feeding frame-level…

Textual graph-based retrieval-augmented generation (GraphRAG) has emerged as a powerful paradigm for enhancing large language models (LLMs) in domain-specific question answering. While existing approaches primarily focus on zero-shot…

信息检索 · 计算机科学 2026-03-26 Yukun Wu , Lihui Liu

Due to its complexity, graph learning-based multi-modal integration and classification is one of the most challenging obstacles for disease prediction. To effectively offset the negative impact between modalities in the process of…

机器学习 · 计算机科学 2025-02-14 Jin Liu , Junbin Mao , Hanhe Lin , Hulin Kuang , Shirui Pan , Xusheng Wu , Shan Xie , Fei Liu , Yi Pan

Multimodal Summarization (MMS) aims to generate concise textual summaries by understanding and integrating information across videos, transcripts, and images. However, existing approaches still suffer from three main challenges: (1)…

计算机视觉与模式识别 · 计算机科学 2026-03-09 Xiaoxing You , Qiang Huang , Lingyu Li , Xiaojun Chang , Jun Yu

Systems that can find correspondences between multiple modalities, such as between speech and images, have great potential to solve different recognition and data analysis tasks in an unsupervised manner. This work studies multimodal…

计算机视觉与模式识别 · 计算机科学 2024-03-08 Khazar Khorrami , Okko Räsänen

We present a continuation to our previous work, in which we developed the MR-CKR framework to reason with knowledge overriding across contexts organized in multi-relational hierarchies. Reasoning is realized via ASP with algebraic measures,…

人工智能 · 计算机科学 2023-05-04 Loris Bozzato , Thomas Eiter , Rafael Kiesel , Daria Stepanova

Multimodal Sarcasm Explanation (MuSE) is a new yet challenging task, which aims to generate a natural language sentence for a multimodal social post (an image as well as its caption) to explain why it contains sarcasm. Although the existing…

计算与语言 · 计算机科学 2023-06-30 Liqiang Jing , Xuemeng Song , Kun Ouyang , Mengzhao Jia , Liqiang Nie

In recent years, anomaly events detection in crowd scenes attracts many researchers' attention, because of its importance to public safety. Existing methods usually exploit visual information to analyze whether any abnormal events have…

计算机视觉与模式识别 · 计算机科学 2021-10-29 Junyu Gao , Maoguo Gong , Xuelong Li

Audio-Visual Video Parsing (AVVP) task aims to parse the event categories and occurrence times from audio and visual modalities in a given video. Existing methods usually focus on implicitly modeling audio and visual features through weak…

多媒体 · 计算机科学 2025-05-06 Yaru Chen , Peiliang Zhang , Fei Li , Faegheh Sardari , Ruohao Guo , Zhenbo Li , Wenwu Wang

Emotion Recognition in Conversations (ERC) presents unique challenges, requiring models to capture the temporal flow of multi-turn dialogues and to effectively integrate cues from multiple modalities. We propose Mixture of Speech-Text…

计算与语言 · 计算机科学 2026-02-27 Soumya Dutta , Smruthi Balaji , Sriram Ganapathy

Multimodal recommendation aims to integrate collaborative signals with heterogeneous content such as visual and textual information, but remains challenged by modality-specific noise, semantic inconsistency, and unstable propagation over…

信息检索 · 计算机科学 2026-02-02 Wei Yang , Rui Zhong , Yiqun Chen , Chi Lu , Peng Jiang

Audio-visual video parsing is the task of categorizing a video at the segment level with weak labels, and predicting them as audible or visible events. Recent methods for this task leverage the attention mechanism to capture the semantic…

计算机视觉与模式识别 · 计算机科学 2023-10-12 Yaru Chen , Ruohao Guo , Xubo Liu , Peipei Wu , Guangyao Li , Zhenbo Li , Wenwu Wang

Audio-Visual scene understanding is a challenging problem due to the unstructured spatial-temporal relations that exist in the audio signals and spatial layouts of different objects and various texture patterns in the visual images.…

计算机视觉与模式识别 · 计算机科学 2023-01-03 Liguang Zhou , Yuhongze Zhou , Xiaonan Qi , Junjie Hu , Tin Lun Lam , Yangsheng Xu

This paper focuses on the Audio-Visual Question Answering (AVQA) task that aims to answer questions derived from untrimmed audible videos. To generate accurate answers, an AVQA model is expected to find the most informative audio-visual…

计算机视觉与模式识别 · 计算机科学 2023-12-21 Zhangbin Li , Dan Guo , Jinxing Zhou , Jing Zhang , Meng Wang

Recent advancements in deep learning for image classification predominantly rely on convolutional neural networks (CNNs) or Transformer-based architectures. However, these models face notable challenges in medical imaging, particularly in…

计算机视觉与模式识别 · 计算机科学 2025-02-26 Zhuoqin Yang , Jiansong Zhang , Xiaoling Luo , Zheng Lu , Linlin Shen

Visual dialogue is a challenging task that needs to extract implicit information from both visual (image) and textual (dialogue history) contexts. Classical approaches pay more attention to the integration of the current question, vision…

计算机视觉与模式识别 · 计算机科学 2020-08-31 Xiaoze Jiang , Siyi Du , Zengchang Qin , Yajing Sun , Jing Yu

In this report, our approach to tackling the task of ActivityNet 2018 Kinetics-600 challenge is described in detail. Though spatial-temporal modelling methods, which adopt either such end-to-end framework as I3D \cite{i3d} or two-stage…

计算机视觉与模式识别 · 计算机科学 2018-06-28 Dongliang He , Fu Li , Qijie Zhao , Xiang Long , Yi Fu , Shilei Wen

This report presents SceneNet and KnowledgeNet, our approaches developed for the HD-EPIC VQA Challenge 2025. SceneNet leverages scene graphs generated with a multi-modal large language model (MLLM) to capture fine-grained object…

计算机视觉与模式识别 · 计算机科学 2025-06-11 Agnese Taluzzi , Davide Gesualdi , Riccardo Santambrogio , Chiara Plizzari , Francesca Palermo , Simone Mentasti , Matteo Matteucci