Related papers: Multi-Interactive-Modality based Modeling for Myop…
Developmental Dyslexia (DD) is a learning disability related to the acquisition of reading skills that affects about 5% of the population. DD can have an enormous impact on the intellectual and personal development of affected children, so…
Light pollution is one of the most rapidly increasing types of environmental degradation. To limit this pollution several effective practices have been defined: shields on lighting fixtures to prevent direct upward light; no over lighting,…
This research is comparing learning environments to students dropout intentions. While using statistics I looked at data and the correlations between two articles to see how the two studies looked side to side. Learning environments and…
The hallucination of large multimodal models (LMMs), providing responses that appear correct but are actually incorrect, limits their reliability and applicability. This paper aims to study the hallucination problem of LMMs in video…
Attention is a key factor for successful learning, with research indicating strong associations between (in)attention and learning outcomes. This dissertation advanced the field by focusing on the automated detection of attention-related…
This work presents mEBAL, a multimodal database for eye blink detection and attention level estimation. The eye blink frequency is related to the cognitive activity and automatic detectors of eye blinks have been proposed for many tasks…
We ask where, and under what conditions, dyslexic reading costs arise in a large-scale naturalistic reading dataset. Using eye-tracking aligned to word-level features (word length, frequency, and predictability), we model how each feature…
This paper addresses the challenging problem of estimating the general visual attention of people in images. Our proposed method is designed to work across multiple naturalistic social scenarios and provides a full picture of the subject's…
Recent Alzheimer's disease (AD) patient studies have focused on retinal analysis, as the retina is the only part of the central nervous system which can be imaged non-invasively by optical methods. However as this is a relatively new…
Postural stability is linked to vision in everyone, since when the eyes are closed stability decreases by a factor of 2 or more. However, in persons with dyslexia postural stability is often deficient even when the eyes are open, since they…
Retinitis pigmentosa (RP), one of the leading causes of vision loss and blindness globally, is a progressive retinal disease involving the degradation of photoreceptors (7) and/or retinal pigment epithelial cells (14). Affecting…
Age-related macular degeneration (AMD) is the leading cause of visual impairment among elderly in the world. Early detection of AMD is of great importance, as the vision loss caused by this disease is irreversible and permanent. Color…
This paper investigates the causal impact of the parental environment on the student's academic performance in mathematics, literature and English (as a foreign language), using a new database covering all children aged 8 to 15 of the…
Large-scale vision-language pre-trained (VLP) models are prone to hallucinate non-existent visual objects when generating text based on visual information. In this paper, we systematically study the object hallucination problem from three…
While Multimodal Large Language Models (MLLMs) have experienced significant advancement in visual understanding and reasoning, their potential to serve as powerful, flexible, interpretable, and text-driven models for Image Quality…
Vision-language model (VLM) hallucination is commonly linked to imbalanced allocation of attention across input modalities: system, image and text. However, existing mitigation strategies tend towards an image-centric interpretation of…
Engaging with natural environments and representations of nature has been shown to improve mood states and reduce cognitive decline in older adults. The current study evaluated the use of virtual reality (VR) for presenting immersive 360…
Multi-modal Large Langue Models (MLLMs) often process thousands of visual tokens, which consume a significant portion of the context window and impose a substantial computational burden. Prior work has empirically explored visual token…
Despite the remarkable performance of foundation vision-language models, the shared representation space for text and vision can also encode harmful label associations detrimental to fairness. While prior work has uncovered bias in…
Eye-based information channels include the pupils, gaze, saccades, fixational movements, and numerous forms of eye opening and closure. Pupil size variation indicates cognitive load and emotion, while a person's gaze direction is said to be…