中文
相关论文

相关论文: Linton Stereo Illusion: Response on Johnston (1991…

200 篇论文

We present another explanation for the moon illusion, the phenomenon in which the moon looks larger near the horizon than near the zenith. In our model of the moon illusion, the sky is considered a spatially-contiguous and…

计算机视觉与模式识别 · 计算机科学 2016-09-29 Joseph Antonides , Toshiro Kubota

Blind people face a lot of problems in their daily routines. They have to struggle a lot just to do their day-to-day chores. In this paper, we have proposed a system with the objective to help the visually impaired by providing audio aid…

计算机视觉与模式识别 · 计算机科学 2019-11-21 Nikhil Thakurdesai , Anupam Tripathi , Dheeraj Butani , Smita Sankhe

Warning: this paper contains material which may be offensive or upsetting. While much of recent work has focused on the detection of hate speech and overtly offensive content, very little research has explored the more subtle but equally…

计算与语言 · 计算机科学 2021-12-03 Teyun Kwon , Anandha Gopalan

Transforming a large language model (LLM) into a Vision-Language Model (VLM) can be achieved by mapping the visual tokens from a vision encoder into the embedding space of an LLM. Intriguingly, this mapping can be as simple as a shallow MLP…

计算机视觉与模式识别 · 计算机科学 2026-02-26 Benno Krojer , Shravan Nayak , Oscar Mañas , Vaibhav Adlakha , Desmond Elliott , Siva Reddy , Marius Mosbach

Adjusting transparency is a common method of mitigating occlusion but is often detrimental for understanding the relative depth relationships between objects as well as removes potentially important information from the occluding object. We…

人机交互 · 计算机科学 2025-07-01 George Bell , Alma Cantu

Interpretability and explainability have gained more and more attention in the field of machine learning as they are crucial when it comes to high-stakes decisions and troubleshooting. Since both provide information about predictors and…

机器学习 · 计算机科学 2024-04-26 Benjamin Leblanc , Pascal Germain

Conventional neural network models (CNN), loosely inspired by the primate visual system, have been shown to predict neural responses in the visual cortex. However, the relationship between CNNs and the visual system is incomplete due to…

计算机视觉与模式识别 · 计算机科学 2021-08-24 Reem Abdel-Salam

It is difficult for people to interpret the decision-making in the inference process of deep neural networks. Visual explanation is one method for interpreting the decision-making of deep learning. It analyzes the decision-making of 2D CNNs…

计算机视觉与模式识别 · 计算机科学 2021-11-01 Masahiro Mitsuhara , Tsubasa Hirakawa , Takayoshi Yamashita , Hironobu Fujiyoshi

Stereo cameras are a popular choice for obstacle avoidance for outdoor lighweight, low-cost robotics applications. However, they are unable to sense thin and reflective objects well. Currently, many algorithms are tuned to perform well on…

计算机视觉与模式识别 · 计算机科学 2019-10-14 John Keller , Sebastian Scherer

This work analyzes the difficulties in learning and teaching Einstein's theory of special relativity. An extensive bibliographic review has been performed, considering articles published in the most relevant journals on science education,…

Large vision-language models (LVLMs), which integrate a vision encoder (VE) with a large language model, have achieved remarkable success across various tasks. However, there are still crucial challenges in LVLMs such as object…

计算机视觉与模式识别 · 计算机科学 2025-10-13 Hoigi Seo , Dong Un Kang , Hyunjin Cho , Joohoon Lee , Se Young Chun

Humans can approach complex visual problems by mentally simulating intermediate visual steps, rather than reasoning through language alone. Inspired by this, several works on Vision-Language Models have recently explored chain-of-thought…

计算机视觉与模式识别 · 计算机科学 2026-05-20 André G. Viveiros , Nuno Gonçalves , André F. T. Martins , Matthias Lindemann

Dynamic stereo matching is the task of estimating consistent disparities from stereo videos with dynamic objects. Recent learning-based methods prioritize optimal performance on a single stereo pair, resulting in temporal inconsistencies.…

计算机视觉与模式识别 · 计算机科学 2024-03-19 Junpeng Jing , Ye Mao , Krystian Mikolajczyk

When viewing stereoscopic displays, people may not always be able to stay exactly in front of the display. It is known that viewing stereoscopic display from different vertical angles lead to different visual discomfort. However, the…

人机交互 · 计算机科学 2018-11-22 Yaohua Xie , Danli Wang , Fang Sun

Large language models (LLMs) are known to generate plausible but false information across a wide range of contexts, yet the real-world magnitude and consequences of this hallucination problem remain poorly understood. Here we leverage a…

数字图书馆 · 计算机科学 2026-05-11 Zhenyue Zhao , Yihe Wang , Toby Stuart , Mathijs De Vaan , Paul Ginsparg , Yian Yin

As the Universe expands, the redshift of distant sources changes with time. Here we discuss gravitational lensing phenomena that are consequence of the redshift drift between lensed source, gravitational lens, and observer. When the source…

宇宙学与河外天体物理 · 物理学 2022-06-01 Giovanni Covone , Mauro Sereno

The Large Synoptic Survey Telescope (LSST) is a wide-field imaging system of unprecedented etendue. The initial goal of the project is to carry out a ten year imaging survey in six broad passbands (ugrizy) that cover $350 nm < \lambda < 1.1…

天体物理仪器与方法 · 物理学 2019-05-14 Christopher W. Stubbs , Katrin Heitmann

Large Vision-Language Models (LVLMs) demonstrate significant progress in multimodal understanding and reasoning, yet object hallucination remains a critical challenge. While existing research focuses on mitigating language priors or…

计算机视觉与模式识别 · 计算机科学 2026-01-08 Yuxuan Xia , Siheng Wang , Peng Li

Vision-and-language navigation (VLN) is a multimodal task where an agent follows natural language instructions and navigates in visual environments. Multiple setups have been proposed, and researchers apply new model architectures or…

计算机视觉与模式识别 · 计算机科学 2022-05-05 Wanrong Zhu , Yuankai Qi , Pradyumna Narayana , Kazoo Sone , Sugato Basu , Xin Eric Wang , Qi Wu , Miguel Eckstein , William Yang Wang

Reinforcement Fine-Tuning (RFT) on flow-based models is crucial for preference alignment. However, they often introduce visual hallucinations like over-optimized details and semantic misalignment. This work preliminarily explores why visual…

计算机视觉与模式识别 · 计算机科学 2026-02-04 Xiaofeng Tan , Jun Liu , Yuanting Fan , Bin-Bin Gao , Xi Jiang , Xiaochen Chen , Jinlong Peng , Chengjie Wang , Hongsong Wang , Feng Zheng