中文
相关论文

相关论文: Mono-to-stereo through parametric stereo generatio…

200 篇论文

Stereo video generation has been gaining increasing attention with recent advancements in video diffusion models. However, most existing methods focus on generating 3D stereoscopic videos from monocular 2D videos. These approaches typically…

This paper addresses the problem of remote sensing image pan-sharpening from the perspective of generative adversarial learning. We propose a novel deep neural network based method named PSGAN. To the best of our knowledge, this is one of…

计算机视觉与模式识别 · 计算机科学 2020-12-22 Qingjie Liu , Huanyu Zhou , Qizhi Xu , Xiangyu Liu , Yunhong Wang

While many text-to-audio systems produce monophonic or fixed-stereo outputs, generating audio with user-defined spatial properties remains a challenge. Existing deep learning-based spatialization methods often rely on latent-space…

声音 · 计算机科学 2025-09-16 Tutti Chi , Letian Gao , Yixiao Zhang

In this paper, we propose a new deep learning-based method for estimating room layout given a pair of 360 panoramas. Our system, called Position-aware Stereo Merging Network or PSMNet, is an end-to-end joint layout-pose estimator. PSMNet…

计算机视觉与模式识别 · 计算机科学 2022-03-31 Haiyan Wang , Will Hutchcroft , Yuguang Li , Zhiqiang Wan , Ivaylo Boyadzhiev , Yingli Tian , Sing Bing Kang

Reconstructing the 3D shape of an object using several images under different light sources is a very challenging task, especially when realistic assumptions such as light propagation and attenuation, perspective viewing geometry and…

计算机视觉与模式识别 · 计算机科学 2020-09-15 Fotios Logothetis , Ignas Budvytis , Roberto Mecca , Roberto Cipolla

This paper presents a probabilistic approach for online dense reconstruction using a single monocular camera moving through the environment. Compared to spatial stereo, depth estimation from motion stereo is challenging due to insufficient…

机器人学 · 计算机科学 2019-03-27 Yonggen Ling , Kaixuan Wang , Shaojie Shen

Recently, diffusion models have achieved great success in mono-channel audio generation. However, when it comes to stereo audio generation, the soundscapes often have a complex scene of multiple objects and directions. Controlling stereo…

声音 · 计算机科学 2025-02-26 Peiwen Sun , Sitong Cheng , Xiangtai Li , Zhen Ye , Huadai Liu , Honggang Zhang , Wei Xue , Yike Guo

We present StereoWorld, a camera-conditioned stereo world model that jointly learns appearance and binocular geometry for end-to-end stereo video generation.Unlike monocular RGB or RGBD approaches, StereoWorld operates exclusively within…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Yang-Tian Sun , Zehuan Huang , Yifan Niu , Lin Ma , Yan-Pei Cao , Yuewen Ma , Xiaojuan Qi

Monocular depth estimation aims at estimating a pixelwise depth map for a single image, which has wide applications in scene understanding and autonomous driving. Existing supervised and unsupervised methods face great challenges.…

计算机视觉与模式识别 · 计算机科学 2018-08-21 Xiaoyang Guo , Hongsheng Li , Shuai Yi , Jimmy Ren , Xiaogang Wang

Numerous models have shown great success in the fields of speech recognition as well as speech synthesis, but models for speech to speech processing have not been heavily explored. We propose Speech to Speech Synthesis Network (STSSN), a…

声音 · 计算机科学 2026-02-20 Bjorn Johnson , Jared Levy

Stereo matching is one of the widely used techniques for inferring depth from stereo images owing to its robustness and speed. It has become one of the major topics of research since it finds its applications in autonomous driving, robotic…

计算机视觉与模式识别 · 计算机科学 2021-09-22 Viny Saajan Victor , Peter Neigel

We conduct a thorough study of photometric stereo under nearby point light source illumination, from modeling to numerical solution, through calibration. In the classical formulation of photometric stereo, the luminous fluxes are assumed to…

计算机视觉与模式识别 · 计算机科学 2017-09-05 Yvain Quéau , Bastien Durix , Tao Wu , Daniel Cremers , François Lauze , Jean-Denis Durou

Predicting accurate normal maps of objects from two-dimensional images in regions of complex structure and spatial material variations is challenging using photometric stereo methods due to the influence of surface reflection properties…

计算机视觉与模式识别 · 计算机科学 2024-04-16 Kai Luo , Yakun Ju , Lin Qi , Kaixuan Wang , Junyu Dong

We introduce MonSter++, a geometric foundation model for multi-view depth estimation, unifying rectified stereo matching and unrectified multi-view stereo. Both tasks fundamentally recover metric depth from correspondence search and…

计算机视觉与模式识别 · 计算机科学 2025-09-26 Junda Cheng , Wenjing Liao , Zhipeng Cai , Longliang Liu , Gangwei Xu , Xianqi Wang , Yuzhou Wang , Zikang Yuan , Yong Deng , Jinliang Zang , Yangyang Shi , Jinhui Tang , Xin Yang

Human auditory perception is shaped by moving sound sources in 3D space, yet prior work in generative sound modelling has largely been restricted to mono signals or static spatial audio. In this work, we introduce a framework for generating…

声音 · 计算机科学 2025-09-29 Yunyi Liu , Shaofan Yang , Kai Li , Xu Li

Stereo image and video generation, stereo geometry estimation, and condition-controlled view synthesis require paired data in which the variables that determine binocular geometry -- camera baseline, intrinsics, scene depth, and camera…

计算机视觉与模式识别 · 计算机科学 2026-05-25 Yangzhi Cui , Feng Qiao , Nathan Jacobs

Purpose: Stereo matching methods that enable depth estimation are crucial for visualization enhancement applications in computer-assisted surgery (CAS). Learning-based stereo matching methods are promising to predict accurate results on…

计算机视觉与模式识别 · 计算机科学 2023-02-07 Zixin Yang , Richard Simon , Cristian A. Linte

Self-supervised stereo matching holds great promise by eliminating the reliance on expensive ground-truth data. Its dominant paradigm, based on photometric consistency, is however fundamentally hindered by the occlusion challenge -- an…

计算机视觉与模式识别 · 计算机科学 2025-11-07 Ruizhi Yang , Xingqiang Li , Jiajun Bai , Jinsong Du

We study the problem of stereo singing voice cancellation, a subtask of music source separation, whose goal is to estimate an instrumental background from a stereo mix. We explore how to achieve performance similar to large state-of-the-art…

声音 · 计算机科学 2024-01-23 Clara Borrelli , James Rae , Dogac Basaran , Matt McVicar , Mehrez Souden , Matthias Mauch

In the field of remote sensing, the scarcity of stereo-matched and particularly lack of accurate ground truth data often hinders the training of deep neural networks. The use of synthetically generated images as an alternative, alleviates…

计算机视觉与模式识别 · 计算机科学 2024-04-16 Vasudha Venkatesan , Daniel Panangian , Mario Fuentes Reyes , Ksenia Bittner