English
Related papers

Related papers: Hi-SAM: A Hierarchical Structure-Aware Multi-modal…

200 papers

The recent Segment Anything Model (SAM) represents a big leap in scaling up segmentation models, allowing for powerful zero-shot capabilities and flexible prompting. Despite being trained with 1.1 billion masks, SAM's mask prediction…

Computer Vision and Pattern Recognition · Computer Science 2023-10-24 Lei Ke , Mingqiao Ye , Martin Danelljan , Yifan Liu , Yu-Wing Tai , Chi-Keung Tang , Fisher Yu

User queries in real-world recommendation systems often combine structured constraints (e.g., category, attributes) with unstructured preferences (e.g., product descriptions or reviews). We introduce HyST (Hybrid retrieval over…

Information Retrieval · Computer Science 2025-08-26 Jiyoon Myung , Jihyeon Park , Joohyung Han

In recent years, the research community has shown a lot of interest to panoramic images that offer a 360-degree directional perspective. Multiple data modalities can be fed, and complimentary characteristics can be utilized for more robust…

Computer Vision and Pattern Recognition · Computer Science 2023-08-21 Suresh Guttikonda , Jason Rambach

While vision-language models (VLMs) have made significant progress in multimodal perception (e.g., open-vocabulary object detection) with simple language queries, state-of-the-art VLMs still show limited ability to perceive complex queries…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Sojung An , Kwanyong Park , Yong Jae Lee , Donghyun Kim

Recently segment anything model (SAM) has shown powerful segmentation capability and has drawn great attention in computer vision fields. Massive following works have developed various applications based on the pre-trained SAM and achieved…

Computer Vision and Pattern Recognition · Computer Science 2025-01-09 Han Shu , Wenshuo Li , Yehui Tang , Yiman Zhang , Yihao Chen , Houqiang Li , Yunhe Wang , Xinghao Chen

Major progress on language models (LMs) in recent years has largely resulted from moving away from specialized models designed for specific tasks, to general models based on powerful architectures (e.g. the Transformer) that learn…

Machine Learning · Computer Science 2025-07-16 Sukjun Hwang , Brandon Wang , Albert Gu

Person re-identification (ReID) aims to retrieve target pedestrian images given either visual queries (image-to-image, I2I) or textual descriptions (text-to-image, T2I). Although both tasks share a common retrieval objective, they pose…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Linhan Zhou , Shuang Li , Neng Dong , Yonghang Tai , Yafei Zhang , Huafeng Li

Foundation models in language and vision benefit from a unified discrete token interface that converts raw inputs into sequences for scalable pre-training and inference. For graphs, an effective tokenizer should yield reusable discrete…

Information Retrieval · Computer Science 2026-05-28 Yang Xiang , Li Fan , Chenke Yin , Lutz Oettershagen , Chengtao Ji

Recent token merging techniques for Vision Transformers (ViTs) provide substantial speedups by reducing the number of tokens processed by self-attention, often without retraining. However, their direct application to the Segment Anything…

Multimodal sentiment analysis (MSA) integrates various modalities, such as text, image, and audio, to provide a more comprehensive understanding of sentiment. However, effective MSA is challenged by alignment and fusion issues. Alignment…

Computer Vision and Pattern Recognition · Computer Science 2025-12-08 Yuhua Wen , Qifei Li , Yingying Zhou , Yingming Gao , Zhengqi Wen , Jianhua Tao , Ya Li

Vehicle recognition is a fundamental problem in SAR image interpretation. However, robustly recognizing vehicle targets is a challenging task in SAR due to the large intraclass variations and small interclass variations. Additionally, the…

Computer Vision and Pattern Recognition · Computer Science 2024-09-24 Weijie Li , Wei Yang , Wenpeng Zhang , Tianpeng Liu , Yongxiang Liu , Li Liu

Current multimodal recommendation models have extensively explored the effective utilization of multimodal information; however, their reliance on ID embeddings remains a performance bottleneck. Even with the assistance of multimodal…

Information Retrieval · Computer Science 2024-10-28 Kangning Zhang , Jiarui Jin , Yingjie Qin , Ruilong Su , Jianghao Lin , Yong Yu , Weinan Zhang

Multi-modal recommender system focuses on utilizing rich modal information ( i.e., images and textual descriptions) of items to improve recommendation performance. The current methods have achieved remarkable success with the powerful…

Information Retrieval · Computer Science 2025-08-20 Shouxing Ma , Yawen Zeng , Shiqing Wu , Guandong Xu

Multimodal image registration is a fundamental task and a prerequisite for downstream cross-modal analysis. Despite recent progress in shared feature extraction and multi-scale architectures, two key limitations remain. First, some methods…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Chunlei Zhang , Jiahao Xia , Yun Xiao , Bo Jiang , Jian Zhang

Theory of mind (ToM) enables AI systems to infer agents' hidden goals and mental states, but existing approaches focus mainly on small human understandable gridworld spaces. We introduce HiVAE, a hierarchical variational architecture that…

Machine Learning · Computer Science 2026-02-20 Nigel Doering , Rahath Malladi , Arshia Sangwan , David Danks , Tauhidur Rahman

Transformer-based LLMs achieve strong results on many language tasks; however, long inputs remain challenging because context windows are finite, and prefill latency and memory grow rapidly with prompt length. Flat token-stream processing…

Computation and Language · Computer Science 2026-05-26 Maryam Haghifam , Zifan He , Jason Cong , Yizhou Sun

Reliable learning of multimodal data (e.g., multi-omics) is a widely concerning issue, especially in safety-critical applications such as medical diagnosis. However, low-quality data induced by multimodal noise poses a major challenge in…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Shu Shen , C. L. Philip Chen , Tong Zhang

The performance of deep learning models in remote sensing (RS) strongly depends on the availability of high-quality labeled data. However, collecting large-scale annotations is costly and time-consuming, while vast amounts of unlabeled…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Wei Huang , Zhitong Xiong , Chenying Liu , Xiao Xiang Zhu

Being a popular mode of text-based communication in multilingual communities, code-mixing in online social media has became an important subject to study. Learning the semantics and morphology of code-mixed language remains a key challenge,…

Computation and Language · Computer Science 2022-04-28 Ayan Sengupta , Tharun Suresh , Md Shad Akhtar , Tanmoy Chakraborty

Implicit feedback, such as user clicks, serves as the primary data source for modern recommender systems. However, click interactions inherently contain substantial noise, including accidental clicks, clickbait-induced interactions, and…

Information Retrieval · Computer Science 2026-02-18 Xikai Yang , Yang Wang , Yilin Li , Sebastian Sun