English
Related papers

Related papers: A Conceptual Model of Intelligent Multimedia Data …

200 papers

Multimodal Large Language Models (MLLMs) have achieved significant advances in integrating visual and linguistic information, yet their ability to reason about complex and real-world scenarios remains limited. The existing benchmarks are…

Phase-only spatial light modulators (SLMs) are used in optical systems for several purposes. In this article, the main landmarks of SLM-based imaging systems are surveyed. In addition to conventional two-dimensional imaging, these systems…

Optics · Physics 2026-03-10 Joseph Rosen

Multimodal health sensing offers rich behavioral signals for assessing mental health, yet translating these numerical time-series measurements into natural language remains challenging. Current LLMs cannot natively ingest long-duration…

The quest for deeper understanding of biological systems has driven the acquisition of increasingly larger multidimensional image datasets. Inspecting and manipulating data of this complexity is very challenging in traditional visualization…

Graphics · Computer Science 2018-08-23 Stanislav Pidhorskyi , Michael Morehead , Quinn Jones , George Spirou , Gianfranco Doretto

Domain-specific knowledge can significantly contribute to addressing a wide variety of vision tasks. However, the generation of such knowledge entails considerable human labor and time costs. This study investigates the potential of Large…

We present new state-of-the-art lens models for strong gravitational lensing systems from the Sloan Lens ACS (SLACS) survey, developed within a Bayesian framework that employs high-dimensional (pixellated), data-driven priors for the…

Instruction tuning has become the de facto method to equip large language models (LLMs) with the ability of following user instructions. Usually, hundreds of thousands or millions of instruction-following pairs are employed to fine-tune the…

Computation and Language · Computer Science 2023-11-28 Qianlong Du , Chengqing Zong , Jiajun Zhang

Phishing and disinformation are popular social engineering attacks with attackers invariably applying influence cues in texts to make them more appealing to users. We introduce Lumen, a learning-based framework that exposes influence cues…

Computation and Language · Computer Science 2021-07-23 Hanyu Shi , Mirela Silva , Daniel Capecci , Luiz Giovanini , Lauren Czech , Juliana Fernandes , Daniela Oliveira

Advancements in foundation models have made it possible to conduct applications in various downstream tasks. Especially, the new era has witnessed a remarkable capability to extend Large Language Models (LLMs) for tackling tasks of 3D scene…

Computer Vision and Pattern Recognition · Computer Science 2025-07-18 Yifan Xu , Chao Zhang , Hanqi Jiang , Xiaoyan Wang , Ruifei Ma , Yiwei Li , Zihao Wu , Zeju Li , Xiangde Liu

New era has unlocked exciting possibilities for extending Large Language Models (LLMs) to tackle 3D vision-language tasks. However, most existing 3D multimodal LLMs (MLLMs) rely on compressing holistic 3D scene information or segmenting…

Computer Vision and Pattern Recognition · Computer Science 2025-07-23 Xiaoyan Wang , Zeju Li , Yifan Xu , Jiaxing Qi , Zhifei Yang , Ruifei Ma , Xiangde Liu , Chao Zhang

We introduce DRESS, a novel approach for generating stylized large language model (LLM) responses through representation editing. Existing methods like prompting and fine-tuning are either insufficient for complex style adaptation or…

Computation and Language · Computer Science 2025-01-27 Xinyu Ma , Yifeng Xu , Yang Lin , Tianlong Wang , Xu Chu , Xin Gao , Junfeng Zhao , Yasha Wang

Autonomous 3D scanning of open-world target structures via drones remains challenging despite broad applications. Existing paradigms rely on restrictive assumptions or effortful human priors, limiting practicality, efficiency, and…

Few-shot learning (FSL) has attracted considerable attention recently. Among existing approaches, the metric-based method aims to train an embedding network that can make similar samples close while dissimilar samples as far as possible and…

Computer Vision and Pattern Recognition · Computer Science 2022-10-13 Bin Xiao , Chien-Liang Liu , Wen-Hoar Hsaio

Flying Triangulation sensors enable a free-hand and motion-robust 3D data acquisition of complex shaped objects. The measurement principle is based on a multi-line light-sectioning approach and uses sophisticated algorithms for real-time…

Computer Vision and Pattern Recognition · Computer Science 2013-05-20 Florian Willomitzer , Svenja Ettl , Christian Faber , Gerd Häusler

Maintaining high data quality is crucial for reliable data analysis and machine learning (ML). However, existing data quality management tools often lack automation, interactivity, and integration with ML workflows. This demonstration paper…

Databases · Computer Science 2025-01-29 Mohamed Abdelaal , Samuel Lokadjaja , Arne Kreuz , Harald Schöning

In this article we introduce PINGSoft, a set of IDL routines designed to visualise and manipulate, in an interactive and friendly way, Integral Field Spectroscopic data. The package is optimised for large databases and a fast visualisation…

Instrumentation and Methods for Astrophysics · Physics 2015-05-19 F. F. Rosales-Ortega

Large Vision-Language Models (LVLMs) have demonstrated strong multimodal reasoning capabilities on long and complex documents. However, their high memory footprint makes them impractical for deployment on resource-constrained edge devices.…

Computer Vision and Pattern Recognition · Computer Science 2025-11-24 Tanveer Hannan , Dimitrios Mallios , Parth Pathak , Faegheh Sardari , Thomas Seidl , Gedas Bertasius , Mohsen Fayyaz , Sunando Sengupta

Dynamic speckle imaging, typically used in laser-illuminated surface diagnostics, has proven valuable for assessing biological activity. In this work, we demonstrate its feasibility in an endoscopic context using a disposable bronchoscope.…

Instrumentation and Detectors · Physics 2025-05-01 Aurélien plyer , Elise Colin , Enrique Garcia-Caurel

Large Language Models (LLMs) have been successfully used in many natural-language tasks and applications including text generation and AI chatbots. They also are a promising new technology for concept-oriented deep learning (CODL). However,…

Machine Learning · Computer Science 2023-09-21 Daniel T. Chang

Recent works have shown that Multimodal Large Language Models (MLLMs) are highly vulnerable to hidden-pattern visual illusions, where the hidden content is imperceptible to models but obvious to humans. This deficiency highlights a…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Jinzhe Tu , Ruilei Guo , Zihan Guo , Junxiao Yang , Shiyao Cui , Minlie Huang
‹ Prev 1 4 5 6 7 8 10 Next ›