Related papers: Photonic Modes Prediction via Multi-Modal Diffusio…
Large pre-trained models have had a significant impact on computer vision by enabling multi-modal learning, where the CLIP model has achieved impressive results in image classification, object detection, and semantic segmentation. However,…
Photonic computation started to shape the future of fast, efficient and accessible computation. The advantages brought by light based Diffractive Deep Neural Networks (D2NN), are shown to be overwhelmingly advantageous especially in…
Multi-plane light converter (MPLC) designs supporting hundreds of modes are attractive in high-throughput optical communications. These photonic structures typically comprise >10 phase masks in free space, with millions of independent…
The resolution of optical imaging is classically limited by the width of the point-spread function, which in turn is determined by the Rayleigh length. Recently, spatial-mode demultiplexing (SPADE) has been proposed as a method to achieve…
Generating consistent human images with controllable pose and appearance is essential for applications in virtual try on, image editing, and digital human creation. Current methods often suffer from occlusions, garment style drift, and pose…
Text-to-image diffusion models have made significant progress in generating naturalistic images from textual inputs, and demonstrate the capacity to learn and represent complex visual-semantic relationships. While these diffusion models…
Numerical results are presented for single-mode guidance, which is based on photonic band gap (PBG) effect, in one-dimensional planar all-dielectric light-guiding systems. In such systems there may be two kinds of light-speed point (the…
According to a recent proposal [S. Takayama et al., Appl. Phys. Lett. 87, 061107 (2005)], the triangular lattice of triangular air holes may allow to achieve a complete photonic band gap in two-dimensional photonic crystal slabs. In this…
Diffusion models have found phenomenal success as expressive priors for solving inverse problems, but their extension beyond natural images to more structured scientific domains remains limited. Motivated by applications in materials…
Photonic crystals (PhCs) are periodic dielectric structures that exhibit unique electromagnetic properties, such as the creation of band gaps where electromagnetic wave propagation is inhibited. Accurately predicting dispersion relations,…
Recent advances in vision language models (VLM) have been driven by contrastive models such as CLIP, which learn to associate visual information with their corresponding text descriptions. However, these models have limitations in…
Plasmonic hyperbolic metasurfaces have emerged as an effective platform for manipulating the propagation of light. Here, confined modes on arrays of silver nanoridges that exhibit hyperbolic dispersion are used to demonstrate and model a…
In this work, a numerical modal decomposition approach is applied to model the optical field of laser light after propagating through a highly multi-mode fiber. The algorithm for the decomposition is based on the reconstruction of measured…
Diffusion models, such as Stable Diffusion, have shown incredible performance on text-to-image generation. Since text-to-image generation often requires models to generate visual concepts with fine-grained details and attributes specified…
Subwavelength photonic structures and metamaterials provide revolutionary approaches for controlling light. The inverse design methods proposed for these subwavelength structures are vital to the development of new photonic devices.…
The Skinned Multi-Person Linear (SMPL) model plays a crucial role in 3D human pose estimation, providing a streamlined yet effective representation of the human body. However, ensuring the validity of SMPL configurations during tasks such…
We introduce Transfusion, a recipe for training a multi-modal model over discrete and continuous data. Transfusion combines the language modeling loss function (next token prediction) with diffusion to train a single transformer over…
Multimodal medical images play a crucial role in the precise and comprehensive clinical diagnosis. Diffusion model is a powerful strategy to synthesize the required medical images. However, existing approaches still suffer from the problem…
Feature representation plays a crucial role in visual correspondence, and recent methods for image matching resort to deeply stacked convolutional layers. These models, however, are both monolithic and static in the sense that they…
In our work we investigate the propagation of optical modes in nanoscale hybrid plasmonic waveguides. Frequency domain Maxwell equations based simulations are implemented to study properties of mixed modes in 3D. The results of our analysis…