English
Related papers

Related papers: Large Language Models and Non-Negative Matrix Fact…

200 papers

Audio Large Language Models (AudioLLMs) have achieved strong results in semantic tasks like speech recognition and translation, but remain limited in modeling paralinguistic cues such as emotion. Existing approaches often treat emotion…

Computation and Language · Computer Science 2025-09-30 Wenyu Zhang , Yingxu He , Geyu Lin , Zhuohan Liu , Shuo Sun , Bin Wang , Xunlong Zou , Jeremy H. M. Wong , Qiongqiong Wang , Hardik B. Sailor , Nancy F. Chen , Ai Ti Aw

We present a structured matrix factorization approach to analyzing calcium imaging recordings of large neuronal ensembles. Our goal is to simultaneously identify the locations of the neurons, demix spatially overlapping components, and…

Neurons and Cognition · Quantitative Biology 2014-09-11 Eftychios A. Pnevmatikakis , Yuanjun Gao , Daniel Soudry , David Pfau , Clay Lacefield , Kira Poskanzer , Randy Bruno , Rafael Yuste , Liam Paninski

While multimodal large language models offer a promising solution to the "black box" nature of health AI by generating interpretable reasoning traces, verifying the validity of these traces remains a critical challenge. Existing evaluation…

Biological multimodal large language models (MLLMs) have emerged as powerful foundation models for scientific discovery. However, existing models are specialized to a single modality, limiting their ability to solve inherently cross-modal…

Machine Learning · Computer Science 2026-03-17 Wonbin Lee , Dongki Kim , Sung Ju Hwang

While models in audio and speech processing are becoming deeper and more end-to-end, they as a consequence need expensive training on large data, and are often brittle. We build on a classical model of human hearing and make it…

Sound · Computer Science 2024-09-16 Ruolan Leslie Famularo , Dmitry N. Zotkin , Shihab A. Shamma , Ramani Duraiswami

Passive Acoustic Monitoring (PAM) analysis is often hindered by the intensive manual effort needed to create labelled training data. This study introduces a synthetic data framework to generate large volumes of richly labelled training data…

Sound · Computer Science 2025-07-23 Kaspar Soltero , Tadeu Siqueira , Stefanie Gutschmidt

This paper presents a scalable method for integrating compositional morphological representations into a vector-based probabilistic language model. Our approach is evaluated in the context of log-bilinear language models, rendered suitably…

Computation and Language · Computer Science 2014-05-19 Jan A. Botha , Phil Blunsom

Ultrasound interpretation requires both precise lesion localization and holistic clinical reasoning, yet existing methods typically excel at only one of these capabilities: specialized detectors offer strong localization but limited…

Computer Vision and Pattern Recognition · Computer Science 2026-05-01 Jing Zhang , Wentao Jiang , Tao Huang , Zhiwei Wang , Jianxin Liu , Jian Chen , Ping Ye , Gang Wang , Zengmao Wang , Bo Du , Dacheng Tao

This work presents a novel approach for speaker diarization to leverage lexical information provided by automatic speech recognition. We propose a speaker diarization system that can incorporate word-level speaker turn probabilities with…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-16 Tae Jin Park , Kyu J. Han , Jing Huang , Xiaodong He , Bowen Zhou , Panayiotis Georgiou , Shrikanth Narayanan

This paper describes several improvements to a new method for signal decomposition that we recently formulated under the name of Differentiable Dictionary Search (DDS). The fundamental idea of DDS is to exploit a class of powerful deep…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-29 Lukáš Samuel Marták , Rainer Kelz , Gerhard Widmer

Audio-Language Models (ALMs), trained on paired audio-text data, are designed to process, understand, and reason about audio-centric multimodal content. Unlike traditional supervised approaches that use predefined labels, ALMs leverage…

Sound · Computer Science 2026-03-13 Yi Su , Jisheng Bai , Qisheng Xu , Kele Xu , Yong Dou

The growing demand for home healthcare calls for tools that can support care delivery. In this study, we explore automatic health assessment from voice using real-world home care visit data, leveraging the diverse patient information it…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-22 Yu-Wen Chen , William Ho , Sasha M. Vergez , Grace Flaherty , Pallavi Gupta , Zhihong Zhang , Maryam Zolnoori , Margaret V. McDonald , Maxim Topaz , Zoran Kostic , Julia Hirschberg

Systems that can associate images with their spoken audio captions are an important step towards visually grounded language learning. We describe a scalable method to automatically generate diverse audio for image captioning datasets. This…

Computer Vision and Pattern Recognition · Computer Science 2019-09-20 Gabriel Ilharco , Yuan Zhang , Jason Baldridge

Acoustic and linguistic analysis for elderly emotion recognition is an under-studied and challenging research direction, but essential for the creation of digital assistants for the elderly, as well as unobtrusive telemonitoring of elderly…

Computation and Language · Computer Science 2020-09-09 Gizem Soğancıoğlu , Oxana Verkholyak , Heysem Kaya , Dmitrii Fedotov , Tobias Cadèe , Albert Ali Salah , Alexey Karpov

Depression is one of the most prevalent mental health disorders globally. In recent years, multi-modal data, such as speech, video, and transcripts, has been increasingly used to develop AI-assisted depression assessment systems. Large…

The validity of medical studies based on real-world clinical data, such as observational studies, depends on critical assumptions necessary for drawing causal conclusions about medical interventions. Many published studies are flawed…

Artificial Intelligence · Computer Science 2024-07-30 Ahmed Alaa , Rachael V. Phillips , Emre Kıcıman , Laura B. Balzer , Mark van der Laan , Maya Petersen

Large language models reveal deep comprehension and fluent generation in the field of multi-modality. Although significant advancements have been achieved in audio multi-modality, existing methods are rarely leverage language model for…

Sound · Computer Science 2024-08-06 Hualei Wang , Jianguo Mao , Zhifang Guo , Jiarui Wan , Hong Liu , Xiangdong Wang

Large language models (LLMs) have emerged as transformative tools in medicine, with strong capabilities in language understanding, reasoning, and structured information extraction. Radiation oncology is particularly well suited for LLM…

With the recent success of representation learning methods, which includes deep learning as a special case, there has been considerable interest in developing techniques that incorporate known physical constraints into the learned…

Machine Learning · Computer Science 2024-01-02 Harsha Vardhan Tetali , Joel B. Harley , Benjamin D. Haeffele

Verification of biomedical claims is critical for healthcare decision-making, public health policy and scientific research. We present an interactive biomedical claim verification system by integrating LLMs, transparent model explanations,…

Human-Computer Interaction · Computer Science 2025-03-03 Siting Liang , Daniel Sonntag
‹ Prev 1 4 5 6 7 8 10 Next ›