English
Related papers

Related papers: OTCR: Optimal Transmission, Compression and Repres…

200 papers

Diffusion policies have shown promise in learning complex behaviors from demonstrations, particularly for tasks requiring precise control and long-term planning. However, they face challenges in robustness when encountering distribution…

Machine Learning · Computer Science 2025-02-24 Mingyang Sun , Pengxiang Ding , Weinan Zhang , Donglin Wang

Vehicle-Infrastructure Collaborative Perception (VICP) is pivotal for resolving occlusion in autonomous driving, yet the trade-off between communication bandwidth and feature redundancy remains a critical bottleneck. While intermediate…

Computer Vision and Pattern Recognition · Computer Science 2026-01-07 Li Wang , Boqi Li , Hang Chen , Xingjian Wu , Yichen Wang , Jiewen Tan , Xinyu Zhang , Huaping Liu

Visual Document Retrieval (VDR) requires representations that capture both fine-grained visual details and global document structure to ensure retrieval efficacy while maintaining computational efficiency. Existing VDR models struggle to…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Fengbin Zhu , Zijing Cai , Yuzhe Wang , Pengyang Shao , Wenjie Wang , Fuli Feng , Richang Hong , Tat-Seng Chua

Semi-supervised learning has made remarkable strides by effectively utilizing a limited amount of labeled data while capitalizing on the abundant information present in unlabeled data. However, current algorithms often prioritize aligning…

Computer Vision and Pattern Recognition · Computer Science 2024-05-31 Zhiquan Tan , Kaipeng Zheng , Weiran Huang

Multimodality Representation Learning, as a technique of learning to embed information from different modalities and their correlations, has achieved remarkable success on a variety of applications, such as Visual Question Answering (VQA),…

Artificial Intelligence · Computer Science 2024-03-04 Muhammad Arslan Manzoor , Sarah Albarri , Ziting Xian , Zaiqiao Meng , Preslav Nakov , Shangsong Liang

Multimodal brain decoding aims to reconstruct semantic information that is consistent with visual stimuli from brain activity signals such as fMRI, and then generate readable natural language descriptions. However, multimodal brain decoding…

Machine Learning · Computer Science 2026-04-21 Xuanyu Hu

Multimodal Entity Linking (MEL) aims to link ambiguous mentions in multimodal contexts to entities in a multimodal knowledge graph. A pivotal challenge is to fully leverage multi-element correlations between mentions and entities to bridge…

Computation and Language · Computer Science 2024-06-06 Zefeng Zhang , Jiawei Sheng , Chuang Zhang , Yunzhi Liang , Wenyuan Zhang , Siqi Wang , Tingwen Liu

Encoding only the task-related information from the raw data, \ie, disentangled representation learning, can greatly contribute to the robustness and generalizability of models. Although significant advances have been made by regularizing…

Computer Vision and Pattern Recognition · Computer Science 2024-08-15 Zhuohang Dang , Minnan Luo , Chengyou Jia , Guang Dai , Jihong Wang , Xiaojun Chang , Jingdong Wang

The autonomous driving community has shown significant interest in 3D occupancy prediction, driven by its exceptional geometric perception and general object recognition capabilities. To achieve this, current works try to construct a…

Computer Vision and Pattern Recognition · Computer Science 2024-04-12 Qihang Ma , Xin Tan , Yanyun Qu , Lizhuang Ma , Zhizhong Zhang , Yuan Xie

Recent grid-based document representations like BERTgrid allow the simultaneous encoding of the textual and layout information of a document in a 2D feature map so that state-of-the-art image segmentation and/or object detection models can…

Computation and Language · Computer Science 2021-05-26 Weihong Lin , Qifang Gao , Lei Sun , Zhuoyao Zhong , Kai Hu , Qin Ren , Qiang Huo

Out-of-distribution (OOD) detection plays a crucial role in ensuring the safety and reliability of deep neural networks in various applications. While there has been a growing focus on OOD detection in visual data, the field of textual OOD…

Computation and Language · Computer Science 2024-04-10 Li-Ming Zhan , Bo Liu , Xiao-Ming Wu

The Mixture of Experts (MoE) has emerged as a highly successful technique in deep learning, based on the principle of divide-and-conquer to maximize model capacity without significant additional computational cost. Even in the era of…

Computation and Language · Computer Science 2024-09-02 Boan Liu , Liang Ding , Li Shen , Keqin Peng , Yu Cao , Dazhao Cheng , Dacheng Tao

Visual commonsense understanding requires Vision Language (VL) models to not only understand image and text but also cross-reference in-between to fully integrate and achieve comprehension of the visual scene described. Recently, various…

Computer Vision and Pattern Recognition · Computer Science 2023-10-24 Zhecan Wang , Haoxuan You , Yicheng He , Wenhao Li , Kai-Wei Chang , Shih-Fu Chang

Vision Language models (VLMs) have demonstrated strong performance across a wide range of benchmarks, yet they often suffer from modality dominance, where predictions rely disproportionately on a single modality. Prior approaches primarily…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Seulgi Kim , Mohit Prabhushankar , Ghassan AlRegib

We introduce ABot-OCR, an end-to-end vision-language model that transcribes a page image directly into clean Markdown in a single forward pass. By doing so, our approach completely eliminates the need for brittle modular orchestration. To…

Computer Vision and Pattern Recognition · Computer Science 2026-05-28 Kaitao Jiang , Ruiyan Gong , Xiaolong Cheng , Kangning Niu , Tianlun Li , Mu Xu

DeepSeek-OCR utilizes an optical 2D mapping approach to achieve high-ratio vision-text compression, claiming to decode text tokens exceeding ten times the input visual tokens. While this suggests a promising solution for the LLM…

Computation and Language · Computer Science 2026-01-09 Yunhao Liang , Ruixuan Ying , Bo Li , Hong Li , Kai Yan , Qingwen Li , Min Yang , Okamoto Satoshi , Zhe Cui , Shiwen Ni

Learning effective visual representations that generalize well without human supervision is a fundamental problem in order to apply Machine Learning to a wide variety of tasks. Recently, two families of self-supervised methods, contrastive…

Machine Learning · Computer Science 2021-12-07 Kuang-Huei Lee , Anurag Arnab , Sergio Guadarrama , John Canny , Ian Fischer

The key challenge in unaligned multimodal language sequences lies in effectively integrating information from various modalities to obtain a refined multimodal joint representation. Recently, the disentangle and fuse methods have achieved…

Computation and Language · Computer Science 2024-09-20 Fan Qian , Jiqing Han , Jianchen Li , Yongjun He , Tieran Zheng , Guibin Zheng

Open Information Extraction (OpenIE) aims to extract relational tuples from open-domain sentences. Traditional rule-based or statistical models have been developed based on syntactic structures of sentences, identified by syntactic parsers.…

Computation and Language · Computer Science 2022-12-06 Kuicai Dong , Aixin Sun , Jung-Jae Kim , Xiaoli Li

The safe operation of autonomous vehicles (AVs) is highly dependent on their understanding of the surroundings. For this, the task of 3D semantic occupancy prediction divides the space around the sensors into voxels, and labels each voxel…

Computer Vision and Pattern Recognition · Computer Science 2025-05-07 Zhenxing Ming , Julie Stephany Berrio , Mao Shan , Yaoqi Huang , Hongyu Lyu , Nguyen Hoang Khoi Tran , Tzu-Yun Tseng , Stewart Worrall