English
Related papers

Related papers: Omnidirectional MediA Format (OMAF): Toolbox for V…

200 papers

With the rapid growth in deepfake video content, we require improved and generalizable methods to detect them. Most existing detection methods either use uni-modal cues or rely on supervised training to capture the dissonance between the…

Computer Vision and Pattern Recognition · Computer Science 2024-06-06 Trevine Oorloff , Surya Koppisetti , Nicolò Bonettini , Divyaraj Solanki , Ben Colman , Yaser Yacoob , Ali Shahriyari , Gaurav Bharaj

Seamless integration of virtual and physical worlds in augmented reality benefits from the system semantically "understanding" the physical environment. AR research has long focused on the potential of context awareness, demonstrating novel…

Human-Computer Interaction · Computer Science 2024-10-08 Chengyuan Xu , Radha Kumaran , Noah Stier , Kangyou Yu , Tobias Höllerer

Multimodal embedding models aim to map heterogeneous inputs, such as text, images, videos, and audio, into a shared semantic space. However, existing methods and benchmarks remain largely limited to partial modality coverage, making it…

Information Retrieval · Computer Science 2026-04-28 Haohang Huang , Xuan Lu , Mingyi Su , Xuan Zhang , Ziyan Jiang , Ping Nie , Kai Zou , Tomas Pfister , Wenhu Chen , Wei Zhang , Xiaoyu Shen , Rui Meng

This paper considers the design of beamforming for orthogonal time frequency space modulation assisted non-orthogonal multiple access (OTFS-NOMA) networks, in which a high-mobility user is sharing the spectrum with multiple low-mobility…

Information Theory · Computer Science 2019-11-01 Zhiguo Ding

In visual computing, 3D geometry is represented in many different forms including meshes, point clouds, voxel grids, level sets, and depth images. Each representation is suited for different tasks thus making the transformation of one…

Computer Vision and Pattern Recognition · Computer Science 2022-09-02 Trevor Houchens , Cheng-You Lu , Shivam Duggal , Rao Fu , Srinath Sridhar

Multiview images have flexible field of view (FoV) but inflexible depth of field (DoF). To overcome the limitation of multiview images on visual tasks, in this paper, we present varifocal multiview (VFMV) images with flexible DoF. VFMV…

Image and Video Processing · Electrical Eng. & Systems 2021-11-22 Kejun Wu , Qiong Liu , Guoan Li , Gangyi Jiang , You Yang

Multi-camera dynamic Augmented Reality (AR) applications require a camera pose estimation to leverage individual information from each camera in one common system. This can be achieved by combining contextual information, such as markers or…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Shiyu Li , Hannah Schieber , Kristoffer Waldow , Benjamin Busam , Julian Kreimeier , Daniel Roth

This paper explores the effectiveness of Multimodal Large Language models (MLLMs) as assistive technologies for visually impaired individuals. We conduct a user survey to identify adoption patterns and key challenges users face with such…

Video understanding requires reasoning at multiple spatiotemporal resolutions -- from short fine-grained motions to events taking place over longer durations. Although transformer architectures have recently advanced the state-of-the-art,…

Computer Vision and Pattern Recognition · Computer Science 2022-06-01 Shen Yan , Xuehan Xiong , Anurag Arnab , Zhichao Lu , Mi Zhang , Chen Sun , Cordelia Schmid

The overview-detail design pattern, characterized by an overview of multiple items and a detailed view of a selected item, is ubiquitously implemented across software interfaces. Designers often try to account for all users, but ultimately…

Human-Computer Interaction · Computer Science 2025-03-12 Bryan Min , Allen Chen , Yining Cao , Haijun Xia

We present an Object-aware Feature Aggregation (OFA) module for video object detection (VID). Our approach is motivated by the intriguing property that video-level object-aware knowledge can be employed as a powerful semantic prior to help…

Computer Vision and Pattern Recognition · Computer Science 2020-10-26 Qichuan Geng , Hong Zhang , Na Jiang , Xiaojuan Qi , Liangjun Zhang , Zhong Zhou

The challenge of open-vocabulary recognition lies in the model has no clue of new categories it is applied to. Existing works have proposed different methods to embed category cues into the model, \eg, through few-shot fine-tuning,…

Computer Vision and Pattern Recognition · Computer Science 2024-06-10 Zehong Ma , Shiliang Zhang , Longhui Wei , Qi Tian

This survey provides a comprehensive overview of recent advances in multimodal alignment and fusion within the field of machine learning, driven by the increasing availability and diversity of data modalities such as text, images, audio,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Songtao Li , Hao Tang

Accurate 3D perception is essential for understanding the environment in autonomous driving. Recent advancements in 3D semantic occupancy prediction have leveraged camera-LiDAR fusion to improve robustness and accuracy. However, current…

Computer Vision and Pattern Recognition · Computer Science 2024-12-31 Minjae Seong , Jisong Kim , Geonho Bang , Hawook Jeong , Jun Won Choi

A wide choice of cinematic lenses enables motion-picture creators to adapt image visual-appearance to their creative vision. Such choice does not exist in the realm of real-time computer graphics, where only one type of perspective…

Graphics · Computer Science 2024-02-09 Jakub Maksymilian Fober

Accurate 3D volumetric mapping is critical for autonomous underwater vehicles operating in obstacle-rich environments. Vision-based perception provides high-resolution data but fails in turbid conditions, while sonar is robust to lighting…

Robotics · Computer Science 2026-03-17 Ivana Collado-Gonzalez , John McConnell , Brendan Englot

Multiple-input multiple-output (MIMO) techniques have recently demonstrated significant potentials in visible light communications (VLC), as they can overcome the modulation bandwidth limitation and provide substantial improvement in terms…

Information Theory · Computer Science 2017-02-07 Hanaa Marshoud , Paschalis C. Sofotasios , Sami Muhaidat , Bayan S. Sharif , George K. Karagiannidis

Despite the growing number of unmanned aerial vehicles (UAVs) and aerial videos, there is a paucity of studies focusing on the aesthetics of aerial videos that can provide valuable information for improving the aesthetic quality of aerial…

Computer Vision and Pattern Recognition · Computer Science 2020-11-05 Qi Kuang , Xin Jin , Qinping Zhao , Bin Zhou

While recent image warping approaches achieved remarkable success on existing benchmarks, they still require training separate models for each specific task and cannot generalize well to different camera models or customized manipulations.…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Kang Liao , Zongsheng Yue , Zhonghua Wu , Chen Change Loy

Most instruction-driven 3D editing methods rely on 2D models to guide the explicit and iterative optimization of 3D representations. This paradigm, however, suffers from two primary drawbacks. First, it lacks a universal design of different…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Chen Liyi , Wang Pengfei , Zhang Guowen , Ma Zhiyuan , Zhang Lei
‹ Prev 1 8 9 10 Next ›