English
Related papers

Related papers: BMFM-RNA: whole-cell expression decoding improves …

200 papers

Physiological signals such as electrocardiograms (ECG) and electroencephalograms (EEG) provide complementary insights into human health and cognition, yet multi-modal integration is challenging due to limited multi-modal labeled data, and…

This paper presents Masked ELMo, a new RNN-based model for language model pre-training, evolved from the ELMo language model. Contrary to ELMo which only uses independent left-to-right and right-to-left contexts, Masked ELMo learns fully…

Computation and Language · Computer Science 2020-10-12 Gregory Senay , Emmanuelle Salin

The synergistic interpretation of anatomical information from computed tomography (CT) and metabolic information from positron emission tomography (PET) is important to oncologic imaging. However, existing deep learning methods for PET/CT…

Image and Video Processing · Electrical Eng. & Systems 2026-05-22 Xiaofeng Liu , Qianru Zhang , Thibault Marin , Menghua Xia , Chi Liu , Georges El Fakhri , Jinsong Ouyang

Understanding cell identity and function through single-cell level sequencing data remains a key challenge in computational biology. We present a novel framework that leverages gene-specific textual annotations from the NCBI Gene database…

Genomics · Quantitative Biology 2025-05-14 Douglas Jiang , Zilin Dai , Luxuan Zhang , Qiyi Yu , Haoqi Sun , Feng Tian

In this work, we significantly enhance masked particle modeling (MPM), a self-supervised learning scheme for constructing highly expressive representations of unordered sets relevant to developing foundation models for high-energy physics.…

High Energy Physics - Phenomenology · Physics 2024-10-02 Matthew Leigh , Samuel Klein , François Charton , Tobias Golling , Lukas Heinrich , Michael Kagan , Inês Ochoa , Margarita Osadchy

Encoder-decoder transformer models have achieved great success on various vision-language (VL) tasks, but they suffer from high inference latency. Typically, the decoder takes up most of the latency because of the auto-regressive decoding.…

Computer Vision and Pattern Recognition · Computer Science 2023-11-16 Peng Tang , Pengkai Zhu , Tian Li , Srikar Appalaraju , Vijay Mahadevan , R. Manmatha

Multimodal Large Language Models have demonstrated powerful cross-modal understanding and reasoning capabilities in general domains. However, in the electromagnetic (EM) domain, they still face challenges such as data scarcity and…

Foundational models, pretrained on a large scale, have demonstrated substantial success across non-medical domains. However, training these models typically requires large, comprehensive datasets, which contrasts with the smaller and more…

Computer Vision and Pattern Recognition · Computer Science 2023-11-17 Raphael Schäfer , Till Nicke , Henning Höfener , Annkristin Lange , Dorit Merhof , Friedrich Feuerhake , Volkmar Schulz , Johannes Lotz , Fabian Kiessling

Commonly-used transformer language models depend on a tokenization schema which sets an unchangeable subword vocabulary prior to pre-training, destined to be applied to all downstream tasks regardless of domain shift, novel word formations,…

Computation and Language · Computer Science 2021-08-03 Yuval Pinter , Amanda Stent , Mark Dredze , Jacob Eisenstein

Large language models (LLMs) excel across diverse tasks but face significant deployment challenges due to high inference costs. LLM inference comprises prefill (compute-bound) and decode (memory-bound) stages, with decode dominating latency…

Artificial Intelligence · Computer Science 2025-08-13 Woojeong Kim , Junxiong Wang , Jing Nathan Yan , Mohamed Abdelfattah , Alexander M. Rush

Foundational deep learning (DL) models are general models, trained on large, diverse, and unlabelled datasets, typically using self-supervised learning techniques have led to significant advancements especially in natural language…

Signal Processing · Electrical Eng. & Systems 2024-11-18 Ahmed Aboulfotouh , Ashkan Eshaghbeigi , Dimitrios Karslidis , Hatem Abou-Zeid

Encoder-decoder models have made great progress on handwritten mathematical expression recognition recently. However, it is still a challenge for existing methods to assign attention to image features accurately. Moreover, those…

Computer Vision and Pattern Recognition · Computer Science 2021-05-18 Wenqi Zhao , Liangcai Gao , Zuoyu Yan , Shuai Peng , Lin Du , Ziyin Zhang

Brain computer interface (BCI) research, as well as increasing portions of the field of neuroscience, have found success deploying large-scale artificial intelligence (AI) pre-training methods in conjunction with vast public repositories of…

Neurons and Cognition · Quantitative Biology 2025-06-03 Mattson Ogg , Rahul Hingorani , Diego Luna , Griffin W. Milsap , William G. Coon , Clara A. Scholl

Foundation models such as the recently introduced Segment Anything Model (SAM) have achieved remarkable results in image segmentation tasks. However, these models typically require user interaction through handcrafted prompts such as…

Computer Vision and Pattern Recognition · Computer Science 2024-10-01 Mélanie Gaillochet , Christian Desrosiers , Hervé Lombaert

Since the emergence of the ImageNet dataset, the pretraining and fine-tuning approach has become widely adopted in computer vision due to the ability of ImageNet-pretrained models to learn a wide variety of visual features. However, a…

Computer Vision and Pattern Recognition · Computer Science 2024-10-22 Pablo Meseguer , Rocío del Amor , Adrian Colomer , Valery Naranjo

During the diagnostic process, clinicians leverage multimodal information, such as chief complaints, medical images, and laboratory-test results. Deep-learning models for aiding diagnosis have yet to meet this requirement. Here we report a…

Computer Vision and Pattern Recognition · Computer Science 2023-06-02 Hong-Yu Zhou , Yizhou Yu , Chengdi Wang , Shu Zhang , Yuanxu Gao , Jia Pan , Jun Shao , Guangming Lu , Kang Zhang , Weimin Li

In oncology, Positron Emission Tomography-Computed Tomography (PET/CT) is widely used in cancer diagnosis, staging, and treatment monitoring, as it combines anatomical details from CT with functional metabolic activity and molecular marker…

Motivation: Predictive modelling of gene expression is a powerful framework for the in silico exploration of transcriptional regulatory interactions through the integration of high-throughput -omics data. A major limitation of previous…

Genomics · Quantitative Biology 2018-08-14 David M Budden , Daniel G Hurley , Edmund J Crampin

Multimodal survival methods combining gigapixel histology whole-slide images (WSIs) and transcriptomic profiles are particularly promising for patient prognostication and stratification. Current approaches involve tokenizing the WSIs into…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Andrew H. Song , Richard J. Chen , Guillaume Jaume , Anurag J. Vaidya , Alexander S. Baras , Faisal Mahmood

In this paper, we propose a new progressive pre-training method for image understanding tasks which leverages RGB-D datasets. The method utilizes Multi-Modal Contrastive Masked Autoencoder and Denoising techniques. Our proposed approach…

Computer Vision and Pattern Recognition · Computer Science 2024-09-17 Muhammad Abdullah Jamal , Omid Mohareri
‹ Prev 1 4 5 6 7 8 10 Next ›