English
Related papers

Related papers: A Conceptual Model of Intelligent Multimedia Data …

200 papers

The ability to read, understand and find important information from written text is a critical skill in our daily lives for our independence, comfort and safety. However, a significant part of our society is affected by partial vision…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Wiktor Mucha , Florin Cuconasu , Naome A. Etori , Valia Kalokyri , Giovanni Trappolini

Continual learning is essential for medical image classification systems to adapt to dynamically evolving clinical environments. The integration of multimodal information can significantly enhance continual learning of image classes.…

Computer Vision and Pattern Recognition · Computer Science 2025-08-06 Jiantao Tan , Peixian Ma , Kanghao Chen , Zhiming Dai , Ruixuan Wang

Ideation is a critical component of video-based design (VBD), where videos serve as the primary medium for design exploration and inspiration. The emergence of generative AI offers considerable potential to enhance this process by…

Human-Computer Interaction · Computer Science 2024-11-07 Tianhao He , Andrija Stankovic , Evangelos Niforatos , Gerd Kortuem

Multimodal large language models (MLLMs) have achieved remarkable progress, yet the object hallucination remains a critical challenge for reliable deployment. In this paper, we present an in-depth analysis of instruction token embeddings…

Machine Learning · Computer Science 2026-05-13 Runhe Lai , Xinhua Lu , Yanqi Wu , Jinlun Ye , Weijiang Yu , Ruixuan Wang

Large language models (LLMs) can capture rich representations of concepts that are useful for real-world tasks. However, language alone is limited. While existing LLMs excel at text-based inferences, health applications require that models…

Computation and Language · Computer Science 2023-05-26 Xin Liu , Daniel McDuff , Geza Kovacs , Isaac Galatzer-Levy , Jacob Sunshine , Jiening Zhan , Ming-Zher Poh , Shun Liao , Paolo Di Achille , Shwetak Patel

We propose LENS, a modular approach for tackling computer vision problems by leveraging the power of large language models (LLMs). Our system uses a language model to reason over outputs from a set of independent and highly descriptive…

Computation and Language · Computer Science 2023-06-29 William Berrios , Gautam Mittal , Tristan Thrush , Douwe Kiela , Amanpreet Singh

Stroboscopic light stimulation (SLS) on closed eyes typically induces simple visual hallucinations (VHs), characterised by vivid, geometric and colourful patterns. A dataset of 862 sentences, extracted from 422 open subjective reports, was…

Computation and Language · Computer Science 2025-02-26 Romy Beauté , David J. Schwartzman , Guillaume Dumas , Jennifer Crook , Fiona Macpherson , Adam B. Barrett , Anil K. Seth

For the last decade, there has been a push to use multi-dimensional (latent) spaces to represent concepts; and yet how to manipulate these concepts or reason with them remains largely unclear. Some recent methods exploit multiple latent…

Computer Vision and Pattern Recognition · Computer Science 2024-09-24 Lorenzo Olearo , Giorgio Longari , Simone Melzi , Alessandro Raganato , Rafael Peñaloza

Opportunistic photo capture (e.g., slides, exhibits, or artifacts) is a common strategy for preserving information encountered in information-rich environments for later revisitation. While fast and minimally disruptive, such photo…

Human-Computer Interaction · Computer Science 2026-04-13 Ashwin Ram , Aeneas Leon Sommer , Martin Schmitz , Jürgen Steimle

We present Firefly, a new browser-based interactive tool for visualizing 3D particle data sets. On a typical personal computer, Firefly can simultaneously render and enable real-time interactions with > ~10 million particles, and can…

Instrumentation and Methods for Astrophysics · Physics 2023-07-24 Alexander B. Gurvich , Aaron M. Geller

Imaging, scattering, and spectroscopy are fundamental in understanding and discovering new functional materials. Contemporary innovations in automation and experimental techniques have led to these measurements being performed much faster…

Machine Learning · Computer Science 2022-01-12 Tatiana Konstantinova , Phillip M. Maffettone , Bruce Ravel , Stuart I. Campbell , Andi M. Barbour , Daniel Olds

Embeddings have become a pivotal means to represent complex, multi-faceted information about entities, concepts, and relationships in a condensed and useful format. Nevertheless, they often preclude direct interpretation. While downstream…

Entropy-based inference methods have gained traction for improving the reliability of Large Language Models (LLMs). However, many existing approaches, such as entropy minimization techniques, suffer from high computational overhead and fail…

Machine Learning · Computer Science 2026-01-27 Jin Li , Zhebo Wang , Tianliang Lu , Mohan Li , Wenpeng Xing , Meng Han

Large language models (LLMs) often map semantically related prompts to similar internal representations at specific layers, even when their surface forms differ widely. We show that this behavior can be explained through Iterated Function…

Computation and Language · Computer Science 2026-01-21 Sotirios Panagiotis Chytas , Vikas Singh

Multispectral object detection is critical for safety-sensitive applications such as autonomous driving and surveillance, where robust perception under diverse illumination conditions is essential. However, the limited availability of…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Manuel Nkegoum , Minh-Tan Pham , Élisa Fromont , Bruno Avignon , Sébastien Lefèvre

A body of recent work in modeling neural activity focuses on recovering low-dimensional latent features that capture the statistical structure of large-scale neural populations. Most such approaches have focused on linear generative models,…

Neurons and Cognition · Quantitative Biology 2016-10-26 Yuanjun Gao , Evan Archer , Liam Paninski , John P. Cunningham

This paper introduces a new method of data-driven microscope design for virtual fluorescence microscopy. Our results show that by including a model of illumination within the first layers of a deep convolutional neural network, it is…

Image and Video Processing · Electrical Eng. & Systems 2020-04-23 Colin L. Cooke , Fanjie Kong , Amey Chaware , Kevin C. Zhou , Kanghyun Kim , Rong Xu , D. Michael Ando , Samuel J. Yang , Pavan Chandra Konda , Roarke Horstmeyer

Open-vocabulary 3D scene understanding enables users to segment novel objects in complex 3D environments through natural language. However, existing approaches remain slow, memory-intensive, and overly complex due to iterative optimization…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Jaehun Bang , Jinhyeok Kim , Minji Kim , Seungheon Jeong , Kyungdon Joo

We introduce Lumos, the first end-to-end multimodal question-answering system with text understanding capabilities. At the core of Lumos is a Scene Text Recognition (STR) component that extracts text from first person point-of-view images,…

Event-based multimodal large language models (MLLMs) enable robust perception in high-speed and low-light scenarios, addressing key limitations of frame-based MLLMs. However, current event-based MLLMs often rely on dense image-like…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Shaoyu Liu , Jianing Li , Guanghui Zhao , Yunjian Zhang , Wen Jiang , Ming Li , Xiangyang Ji