中文
相关论文

相关论文: Context encoding enables machine learning-based qu…

200 篇论文

Image quality assessment (IQA) is an active research area in the field of image processing. Most prior works focus on visual quality of natural images captured by cameras. In this paper, we explore visual quality of scanned documents,…

图像与视频处理 · 电气工程与系统科学 2023-07-26 Justin Yang , Peter Bauer , Todd Harris , Changhyung Lee , Hyeon Seok Seo , Jan P Allebach , Fengqing Zhu

Visual Question Answering (VQA) becomes one of the most active research problems in the medical imaging domain. A well-known VQA challenge is the intrinsic diversity between the image and text modalities, and in the medical VQA task, there…

计算机视觉与模式识别 · 计算机科学 2023-02-28 Yuan Zhou , Jing Mei , Yiqin Yu , Tanveer Syeda-Mahmood

Masked Autoencoder (MAE) has recently been shown to be effective in pre-training Vision Transformers (ViT) for natural image analysis. By reconstructing full images from partially masked inputs, a ViT encoder aggregates contextual…

图像与视频处理 · 电气工程与系统科学 2023-04-24 Lei Zhou , Huidong Liu , Joseph Bae , Junjun He , Dimitris Samaras , Prateek Prasanna

We seek to semantically describe a set of images, capturing both the attributes of single images and the variations within the set. Our procedure is analogous to Principle Component Analysis, in which the role of projection vectors is…

计算机视觉与模式识别 · 计算机科学 2022-10-24 Oded Hupert , Idan Schwartz , Lior Wolf

In photoacoustic computed tomography (PACT) with short-pulsed laser excitation, wideband acoustic signals are generated in biological tissues with frequencies related to the effective shapes and sizes of the optically absorbing targets.…

Optical scattering presents a major obstacle to high resolution imaging in biological tissue and other turbid media. Conventional photoacoustic imaging can partially overcome this obstacle, enabling imaging of optical absorption in the…

Interpretability is significant in computational pathology, leading to the development of multimodal information integration from histopathological image and corresponding text data.However, existing multimodal methods have limited…

计算机视觉与模式识别 · 计算机科学 2026-01-22 Kangcheng Zhou , Jun Jiang , Qing Zhang , Shuang Zheng , Qingli Li , Shugong Xu

We show that it is possible to estimate the shape of an object by measuring only the fluctuations of a probing field, allowing us to expose the object to a minimal light intensity. This scheme, based on noise measurements through homodyne…

量子物理 · 物理学 2020-01-30 Jeremy B. Clark , Zhifan Zhou , Quentin Glorieux , Alberto M. Marino , Paul D. Lett

This report introduces a novel optofluidic platform based on piezo-MEMS technology, capable of identifying subtle variations in the fluid concentration. The system utilizes piezoelectric micromachined ultrasound transducers (PMUTs) as…

Images encode both the state of the world and its content. The former is useful for tasks such as planning and control, and the latter for classification. The automatic extraction of this information is challenging because of the…

人工智能 · 计算机科学 2020-12-09 Christine Allen-Blanchette , Kostas Daniilidis

The Visual Question Answering (VQA) task combines challenges for processing data with both Visual and Linguistic processing, to answer basic `common sense' questions about given images. Given an image and a question in natural language, the…

计算机视觉与模式识别 · 计算机科学 2020-12-24 Yash Srivastava , Vaishnav Murali , Shiv Ram Dubey , Snehasis Mukherjee

Diffusion probabilistic models (DPMs) have achieved remarkable quality in image generation that rivals GANs'. But unlike GANs, DPMs use a set of latent variables that lack semantic meaning and cannot serve as a useful representation for…

计算机视觉与模式识别 · 计算机科学 2022-03-14 Konpat Preechakul , Nattanat Chatthee , Suttisak Wizadwongsa , Supasorn Suwajanakorn

Existing approaches to image reconstruction in photoacoustic computed tomography (PACT) with acoustically heterogeneous media are limited to weakly varying media, are computationally burdensome, and/or cannot effectively mitigate the…

医学物理 · 物理学 2013-03-25 Chao Huang , Kun Wang , Liming Nie , Lihong V. Wang , Mark A. Anastasio

Photoacoustic imaging (PAI) is a promising medical imaging modality providing the spatial resolution of ultrasound (US) imaging and the contrast of pure optical imaging. For linear-array PAI, a beamformer has to be used as the…

医学物理 · 物理学 2018-02-15 Moein Mozaffarzadeh , Yan Yan , Mohammad Mehrmohammadi , Bahador Makkiabadi

Non-invasively focusing light into strongly scattering media, such as biological tissue, is highly desirable but challenging. Recently, wavefront shaping technologies guided by ultrasonic encoding or photoacoustic sensing have been…

光学 · 物理学 2014-02-05 Puxiang Lai , Lidai Wang , Jian Wei Tay , Lihong V. Wang

Significance: Compressed sensing (CS) uses special measurement designs combined with powerful mathematical algorithms to reduce the amount of data to be collected while maintaining image quality. This is relevant to almost any imaging…

计算机视觉与模式识别 · 计算机科学 2024-02-27 Markus Haltmeier , Matthias Ye , Karoline Felbermayer , Florian Hinterleitner , Peter Burgholzer

In this paper our objectives are, first, networks that can embed audio and visual inputs into a common space that is suitable for cross-modal retrieval; and second, a network that can localize the object that sounds in an image, given the…

计算机视觉与模式识别 · 计算机科学 2018-07-27 Relja Arandjelović , Andrew Zisserman

Visual understanding is inherently contextual -- what we focus on in an image depends on the task at hand. For instance, given an image of a person holding a bouquet of flowers, we may focus on either the person such as their clothing, or…

Deep neural networks have shown striking progress and obtained state-of-the-art results in many AI research fields in the recent years. However, it is often unsatisfying to not know why they predict what they do. In this paper, we address…

计算机视觉与模式识别 · 计算机科学 2016-09-12 Yash Goyal , Akrit Mohapatra , Devi Parikh , Dhruv Batra

The attention mechanism is a critical component of Large Language Models (LLMs) that allows tokens in a sequence to interact with each other, but is order-invariant. Incorporating position encoding (PE) makes it possible to address by…

计算与语言 · 计算机科学 2024-05-31 Olga Golovneva , Tianlu Wang , Jason Weston , Sainbayar Sukhbaatar
‹ 上一页 1 8 9 10 下一页 ›