English
Related papers

Related papers: ViP: Video Platform for PyTorch

200 papers

Partially Relevant Video Retrieval (PRVR) is a practical yet challenging task that involves retrieving videos based on queries relevant to only specific segments. While existing works follow the paradigm of developing models to process…

Computer Vision and Pattern Recognition · Computer Science 2025-08-27 Yi Pan , Yujia Zhang , Michael Kampffmeyer , Xiaoguang Zhao

Multi-modal Large Language Models (MLLMs) for Visual Question Answering (VQA) often suffer from dual limitations: knowledge hallucination and insufficient fine-grained visual perception. Crucially, we identify that commonsense graphs and…

Computer Vision and Pattern Recognition · Computer Science 2026-05-28 Zhiyang Li , Ao Ke , Yukun Cao , Xike Xie

While significant progress has been made on understanding hand-object interactions in computer vision, it is still very challenging for robots to perform complex dexterous manipulation. In this paper, we propose a new platform and pipeline…

Machine Learning · Computer Science 2022-07-07 Yuzhe Qin , Yueh-Hua Wu , Shaowei Liu , Hanwen Jiang , Ruihan Yang , Yang Fu , Xiaolong Wang

We present VLMEvalKit: an open-source toolkit for evaluating large multi-modality models based on PyTorch. The toolkit aims to provide a user-friendly and comprehensive framework for researchers and developers to evaluate existing…

Human vision is able to capture the part-whole hierarchical information from the entire scene. This paper presents the Visual Parser (ViP) that explicitly constructs such a hierarchy with transformers. ViP divides visual representations…

Computer Vision and Pattern Recognition · Computer Science 2022-01-11 Shuyang Sun , Xiaoyu Yue , Song Bai , Philip Torr

Large Vision-Language Models (LVLMs) have made significant strides in the field of video understanding in recent times. Nevertheless, existing video benchmarks predominantly rely on text prompts for evaluation, which often require complex…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Yiming Zhao , Yu Zeng , Yukun Qi , YaoYang Liu , Xikun Bao , Lin Chen , Zehui Chen , Qing Miao , Chenxi Liu , Jie Zhao , Feng Zhao

Extracting structured information from videos is critical for numerous downstream applications in the industry. In this paper, we define a significant task of extracting hierarchical key information from visual texts on videos. To fulfill…

Information Retrieval · Computer Science 2024-01-10 Siyu An , Ye Liu , Haoyuan Peng , Di Yin

Multimodal learning, which involves integrating information from various modalities such as text, images, audio, and video, is pivotal for numerous complex tasks like visual question answering, cross-modal retrieval, and caption generation.…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 G. Thomas Hudson , Dean Slack , Thomas Winterbottom , Jamie Sterling , Chenghao Xiao , Junjie Shentu , Noura Al Moubayed

Many believe that the successes of deep learning on image understanding problems can be replicated in the realm of video understanding. However, due to the scale and temporal nature of video, the span of video understanding problems and the…

Computer Vision and Pattern Recognition · Computer Science 2021-10-05 Matthew Hutchinson , Vijay Gadepally

We propose DeepV2D, an end-to-end deep learning architecture for predicting depth from video. DeepV2D combines the representation ability of neural networks with the geometric principles governing image formation. We compose a collection of…

Computer Vision and Pattern Recognition · Computer Science 2020-04-29 Zachary Teed , Jia Deng

Machine learning is an important part of the data science field. In petrophysics, machine learning algorithms and applications have been widely approached. In this context, Vietnam Petroleum Institute (VPI) has researched and deployed…

Machine Learning · Computer Science 2024-10-10 Anh Tuan Nguyen

In this paper we report on a multimedia communication system including a VCoIP (Video Conferencing over IP) software with a distributed architecture and its applications for teaching scenarios. It is a simple, ready-to-use scheme for…

Multimedia · Computer Science 2007-05-23 Hans L. Cycon , Thomas C. Schmidt , Matthias Waehlisch , Mark Palkow , Henrik Regensburg

We present VPNeXt, a new and simple model for the Plain Vision Transformer (ViT). Unlike the many related studies that share the same homogeneous paradigms, VPNeXt offers a fresh perspective on dense representation based on ViT. In more…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Xikai Tang , Ye Huang , Guangqiang Yin , Lixin Duan

This is the preprint version of our paper on ICONIP2015. The proposed platform supports the integrated VRGIS functions including 3D spatial analysis functions, 3D visualization for spatial process and serves for 3D globe and digital city.…

Human-Computer Interaction · Computer Science 2015-08-11 Weixi Wang , Zhihan Lv , Xiaoming Li , Weiping Xu , Baoyun Zhang

Diffusion Transformers (DiTs) can generate short photorealistic videos, yet directly training and sampling longer videos with full attention across the video remains computationally challenging. Alternative methods break long videos down…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Bhishma Dedhia , David Bourgin , Krishna Kumar Singh , Yuheng Li , Yan Kang , Zhan Xu , Niraj K. Jha , Yuchen Liu

Visual texts embedded in videos carry rich semantic information, which is crucial for both holistic video understanding and fine-grained reasoning about local human actions. However, existing video understanding benchmarks largely overlook…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Zhoufaran Yang , Yan Shu , Jing Wang , Zhifei Yang , Yan Zhang , Yu Li , Keyang Lu , Gangyan Zeng , Shaohui Liu , Yu Zhou , Nicu Sebe

The rapid advancement of multimodal large language models has demonstrated impressive capabilities, yet nearly all operate in an offline paradigm, hindering real-time interactivity. Addressing this gap, we introduce the Real-tIme Video…

Computer Vision and Pattern Recognition · Computer Science 2026-03-05 Yansong Shi , Qingsong Zhao , Tianxiang Jiang , Xiangyu Zeng , Yi Wang , Limin Wang

Deep learning-based vision is characterized by intricate frameworks that often necessitate a profound understanding, presenting a barrier to newcomers and limiting broad adoption. With many researchers grappling with the constraints of…

Computer Vision and Pattern Recognition · Computer Science 2023-11-13 Fabi Prezja

Robotic manipulation requires understanding both the 3D spatial structure of the environment and its temporal evolution, yet most existing policies overlook one or both. They typically rely on 2D visual observations and backbones pretrained…

Recent years have witnessed the significant development of learning-based video compression methods, which aim at optimizing objective or perceptual quality and bit rates. In this paper, we introduce deep video compression with perceptual…

Image and Video Processing · Electrical Eng. & Systems 2021-10-11 Saiping Zhang , Marta Mrak , Luis Herranz , Marc Górriz , Shuai Wan , Fuzheng Yang