English
Related papers

Related papers: The Topology of Multimodal Fusion: Why Current Arc…

200 papers

Contemporary ML separates the static structure of parameters from the dynamic flow of inference, yielding systems that lack the sample efficiency and thermodynamic frugality of biological cognition. In this theoretical work, we propose…

Machine Learning · Computer Science 2025-12-09 Xin Li

Grassmannian manifold offers a powerful carrier for geometric representation learning by modelling high-dimensional data as low-dimensional subspaces. However, existing approaches predominantly rely on static single-subspace…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Xuan Yu , Tianyang Xu

Unified multimodal models for image generation and understanding represent a significant step toward AGI and have attracted widespread attention from researchers. The main challenge of this task lies in the difficulty in establishing an…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Dian Zheng , Manyuan Zhang , Hongyu Li , Kai Zou , Hongbo Liu , Ziyu Guo , Kaituo Feng , Yexin Liu , Ying Luo , Hongsheng Li

Recent years have seen remarkable progress in both multimodal understanding models and image generation models. Despite their respective successes, these two domains have evolved independently, leading to distinct architectural paradigms:…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Shanshan Zhao , Xinjie Zhang , Jintao Guo , Jiakui Hu , Lunhao Duan , Minghao Fu , Yong Xien Chng , Guo-Hua Wang , Qing-Guo Chen , Zhao Xu , Weihua Luo , Kaifu Zhang

Generative AI has democratized content creation, but popular chatbot-based interfaces often prioritize execution, generating fully rendered artifacts right away. This issue can lead to premature convergence and design fixation, where users…

Human-Computer Interaction · Computer Science 2026-04-07 Chao Wen , Tung Phung , Pronita Mehrotra , Sumit Gulwani , Roger E. Beaty , Tomohiro Nagashima , Adish Singla

The long-standing goal of multimodal AI is to build unified models in which visual understanding and visual generation mutually enhance one another. Despite recent works such as BAGEL, BLIP3o achieves remarkable progress; In practice,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Yujun Tong , Dongliang Chang , Zijin Yin , Xintong Liu , Yuanchen Fang , Zhanyu Ma

Turn-taking, aiming to decide when the next speaker can start talking, is an essential component in building human-robot spoken dialogue systems. Previous studies indicate that multimodal cues can facilitate this challenging task. However,…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-22 Jiudong Yang , Peiying Wang , Yi Zhu , Mingchao Feng , Meng Chen , Xiaodong He

Recently, the rapid advancement of multimodal domains has driven a data-centric paradigm shift in graph ML, transitioning from text-attributed to multimodal-attributed graphs. This advancement significantly enhances data representation and…

Artificial Intelligence · Computer Science 2026-01-30 Xunkai Li , Zhengyu Wu , Zekai Chen , Henan Sun , Daohan Su , Guang Zeng , Hongchao Qin , Rong-Hua Li , Guoren Wang

Talking head generation is to synthesize a lip-synchronized talking head video by inputting an arbitrary face image and corresponding audio clips. Existing methods ignore not only the interaction and relationship of cross-modal information,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 Sen Chen , Zhilei Liu , Jiaxing Liu , Longbiao Wang

We address the challenging task of text-driven 3D human-object interaction (HOI) motion generation. Existing methods primarily rely on a direct text-to-HOI mapping, which suffers from three key limitations due to the significant…

Computer Vision and Pattern Recognition · Computer Science 2026-02-12 Yin Wang , Ziyao Zhang , Zhiying Leng , Haitian Liu , Frederick W. B. Li , Mu Li , Xiaohui Liang

The study of irreducible higher-order interactions has become a core topic of study in complex systems. Two of the most well-developed frameworks, topological data analysis and multivariate information theory, aim to provide formal tools…

Information Theory · Computer Science 2025-04-15 Thomas F. Varley , Pedro A. M. Mediano , Alice Patania , Josh Bongard

Deep neural networks for 3D point cloud understanding have achieved remarkable success in object classification and recognition, yet recent work shows that these models remain highly vulnerable to adversarial perturbations. Existing 3D…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Gayathry Chandramana Krishnan Nampoothiry , Raghuram Venkatapuram , Anirban Ghosh , Ayan Dutta

Brain connectomics is still largely dominated by pairwise-based models, such as graphs, which cannot represent circulatory or higher-order functional interactions. In this paper, we propose a multimodal framework based on Topological Signal…

Neurons and Cognition · Quantitative Biology 2026-04-01 Breno C. Bispo , Stefania Sardellitti , Juliano B. Lima , Fernando A. N. Santos

Despite extensive investment in artificial intelligence, 95% of enterprises report no measurable profit impact from AI deployments (MIT, 2025). In this theoretical paper, we argue that this gap reflects paradigmatic lock-in that channels AI…

Computers and Society · Computer Science 2025-09-15 Diana A. Wolfe , Alice Choe , Fergus Kidd

Topology optimization(TO) is widely used in engineering because of its ability to save material and optimize structural performance. Although prior work has explored 2D human-centered design tool for TO, the results are often limited in…

Human-Computer Interaction · Computer Science 2026-04-24 Shuyue Feng , Cedric Caremel , Yoshihiro Kawahara

Fusing multiple modalities has proven effective for multimodal information processing. However, the incongruity between modalities poses a challenge for multimodal fusion, especially in affect recognition. In this study, we first analyze…

Computation and Language · Computer Science 2023-11-14 Yaoting Wang , Yuanchao Li , Paul Pu Liang , Louis-Philippe Morency , Peter Bell , Catherine Lai

Multimodal learning faces a fundamental tension between deep, fine-grained fusion and computational scalability. While cross-attention models achieve strong performance through exhaustive pairwise fusion, their quadratic complexity is…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Yusuf Shihata

Vision-Language Models (VLMs) such as CLIP learn a shared embedding space for images and text, yet their representations remain geometrically separated, a phenomenon known as the modality gap. This gap limits tasks requiring cross-modal…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Hongyuan Liu , Qinli Yang , Wen Li , Zhong Zhang , Jiaming Liu , Wei Han , Zhili Qin , Jinxia Guo , Junming Shao

Natural human interactions for Mixed Reality Applications are overwhelmingly multimodal: humans communicate intent and instructions via a combination of visual, aural and gestural cues. However, supporting low-latency and accurate…

Human-Computer Interaction · Computer Science 2020-12-21 Darshana Rathnayake , Ashen de Silva , Dasun Puwakdandawa , Lakmal Meegahapola , Archan Misra , Indika Perera

People perceive the world with different senses, such as sight, hearing, smell, and touch. Processing and fusing information from multiple modalities enables Artificial Intelligence to understand the world around us more easily. However,…

Computer Vision and Pattern Recognition · Computer Science 2023-07-06 Zecheng Liu , Jia Wei , Rui Li , Jianlong Zhou
‹ Prev 1 2 3 10 Next ›