中文
相关论文

相关论文: MVB: A Large-Scale Dataset for Baggage Re-Identifi…

200 篇论文

Connected Vision Systems (CVS) are transforming a variety of applications, including autonomous vehicles, smart cities, surveillance, and human-robot interaction. These systems harness multi-view multi-camera (MVMC) data to provide enhanced…

计算机视觉与模式识别 · 计算机科学 2025-10-14 Muhammad Munsif , Waqas Ahmad , Amjid Ali , Mohib Ullah , Adnan Hussain , Sung Wook Baik

State-of-the-art video object detection methods maintain a memory structure, either a sliding window or a memory queue, to enhance the current frame using attention mechanisms. However, we argue that these memory structures are not…

计算机视觉与模式识别 · 计算机科学 2024-02-02 Guanxiong Sun , Yang Hua , Guosheng Hu , Neil Robertson

Existing 4D human datasets fall short for fashion-specific research, lacking either realistic garment dynamics or task-specific annotations. Synthetic datasets suffer from a realism gap, whereas real-world captures lack the detailed…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Hunor Laczkó , Libang Jia , Loc-Phat Truong , Diego Hernández , Sergio Escalera , Jordi Gonzalez , Meysam Madadi

The recent progress in self-supervised learning has successfully combined Masked Image Modeling (MIM) with Siamese Networks, harnessing the strengths of both methodologies. Nonetheless, certain challenges persist when integrating…

计算机视觉与模式识别 · 计算机科学 2024-11-12 Kirill Vishniakov , Eric Xing , Zhiqiang Shen

Multi-frequency Electrical Impedance Tomography (mfEIT) is an emerging biomedical imaging modality to reveal frequency-dependent conductivity distributions in biomedical applications. Conventional model-based image reconstruction methods…

图像与视频处理 · 电气工程与系统科学 2021-05-27 Zhou Chen , Jinxi Xiang , Pierre Bagnaninchi , Yunjie Yang

3D object detection is critical for autonomous driving, yet it remains fundamentally challenging to simultaneously maximize computational efficiency and capture long-range spatial dependencies. We observed that Mamba-based models, with…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Longhui Zheng , Qiming Xia , Xiaolu Chen , Zhaoliang Liu , Chenglu Wen

Existing multimodal machine translation (MMT) datasets consist of images and video captions or general subtitles, which rarely contain linguistic ambiguity, making visual information not so effective to generate appropriate translations. We…

计算与语言 · 计算机科学 2022-05-27 Yihang Li , Shuichiro Shimizu , Weiqi Gu , Chenhui Chu , Sadao Kurohashi

Bird's-eye view (BEV) perception has garnered significant attention in autonomous driving in recent years, in part because BEV representation facilitates multi-modal sensor fusion. BEV representation enables a variety of perception tasks…

计算机视觉与模式识别 · 计算机科学 2025-03-28 Goodarz Mehr , Azim Eskandarian

Conventional person re-identification (ReID) research is often limited to single-modality sensor data from static cameras, which fails to address the complexities of real-world scenarios where multi-modal signals are increasingly prevalent.…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Ruiyang Ha , Songyi Jiang , Bin Li , Bikang Pan , Yihang Zhu , Junjie Zhang , Xiatian Zhu , Shaogang Gong , Jingya Wang

Multimodal embedding models have been crucial in enabling various downstream tasks such as semantic similarity, information retrieval, and clustering over different modalities. However, existing multimodal embeddings like VLM2Vec, E5-V, GME…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Rui Meng , Ziyan Jiang , Ye Liu , Mingyi Su , Xinyi Yang , Yuepeng Fu , Can Qin , Zeyuan Chen , Ran Xu , Caiming Xiong , Yingbo Zhou , Wenhu Chen , Semih Yavuz

Fusion of 2D images and 3D point clouds is important because information from dense images can enhance sparse point clouds. However, fusion is challenging because 2D and 3D data live in different spaces. In this work, we propose MVPNet…

计算机视觉与模式识别 · 计算机科学 2019-10-01 Maximilian Jaritz , Jiayuan Gu , Hao Su

Semantic segmentation has been one of the leading research interests in computer vision recently. It serves as a perception foundation for many fields, such as robotics and autonomous driving. The fast development of semantic segmentation…

计算机视觉与模式识别 · 计算机科学 2020-05-19 Ye Lyu , George Vosselman , Guisong Xia , Alper Yilmaz , Michael Ying Yang

Visible-infrared person re-identification (V-I ReID) seeks to match images of individuals captured over a distributed network of RGB and IR cameras. The task is challenging due to the significant differences between V and I modalities,…

计算机视觉与模式识别 · 计算机科学 2023-05-02 Arthur Josi , Mahdi Alehdaghi , Rafael M. O. Cruz , Eric Granger

In this work, we explore a deep learning based automated visual inspection and verification algorithm, based on the Siamese Neural Network architecture. Consideration is also given to how the input pairs of images can affect the performance…

计算机视觉与模式识别 · 计算机科学 2024-09-04 John Oyekan , Liam Quantrill , Christopher Turner , Ashutosh Tiwari

Recent open-vocabulary 3D scene understanding approaches mainly focus on training 3D networks through contrastive learning with point-text pairs or by distilling 2D features into 3D models via point-pixel alignment. While these methods show…

计算机视觉与模式识别 · 计算机科学 2025-07-01 Xingyilang Yin , Jiale Wang , Xi Yang , Mutian Xu , Xu Gu , Nannan Wang

Human head detection, keypoint estimation, and 3D head model fitting are essential tasks with many applications. However, traditional real-world datasets often suffer from bias, privacy, and ethical concerns, and they have been recorded in…

计算机视觉与模式识别 · 计算机科学 2024-12-06 Orest Kupyn , Eugene Khvedchenia , Christian Rupprecht

We introduce MVGenMaster, a multi-view diffusion model enhanced with 3D priors to address versatile Novel View Synthesis (NVS) tasks. MVGenMaster leverages 3D priors that are warped using metric depth and camera poses, significantly…

计算机视觉与模式识别 · 计算机科学 2025-03-07 Chenjie Cao , Chaohui Yu , Shang Liu , Fan Wang , Xiangyang Xue , Yanwei Fu

Multi-view product image queries can improve retrieval performance over single view queries significantly. In this paper, we investigated the performance of deep convolutional neural networks (ConvNets) on multi-view product image search.…

计算机视觉与模式识别 · 计算机科学 2017-05-02 Muhammet Bastan , Ozgur Yilmaz

The rapid advancement of AI-generated multimodal video-audio content has raised significant concerns regarding information security and content authenticity. Existing synthetic video datasets predominantly focus on the visual modality…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Mengxue Hu , Yunfeng Diao , Changtao Miao , Zhiqing Guo , Jianshu Li , Zhe Li , Joey Tianyi Zhou

To tackle the challeging problem of multi-person 3D pose estimation from a single image, we propose a multi-view matching (MVM) method in this work. The MVM method generates reliable 3D human poses from a large-scale video dataset, called…

计算机视觉与模式识别 · 计算机科学 2020-04-14 Yeji Shen , C. -C. Jay Kuo