English
Related papers

Related papers: Probing 3D Chromatin Structure Awareness in Evo2 D…

200 papers

Current Large Language Models (LLMs) for understanding proteins primarily treats amino acid sequences as a text modality. Meanwhile, Protein Language Models (PLMs), such as ESM-2, have learned massive sequential evolutionary knowledge from…

Machine Learning · Computer Science 2024-12-17 Nuowei Liu , Changzhi Sun , Tao Ji , Junfeng Tian , Jianxin Tang , Yuanbin Wu , Man Lan

Do multilingual embedding models encode a language-general representation of proficiency? We investigate this by training linear and non-linear probes on hidden-state activations from Qwen3-Embedding (0.6B, 4B, 8B) to predict CEFR…

Computation and Language · Computer Science 2026-04-09 Laurits Lyngbaek , Ross Deans Kristensen-McLachlan

Motivation: Protein folding is a dynamic process during which a protein's amino acid sequence undergoes a series of 3-dimensional (3D) conformational changes en route to reaching a native 3D structure; the resulting 3D structural…

Biomolecules · Quantitative Biology 2026-04-09 Aydin Wells , Khalique Newaz , Jennifer Morones , Jianlin Cheng , Tijana Milenković

Tokens serve as the basic units of representation in DNA language models (DNALMs), yet their design remains underexplored. Unlike natural language, DNA lacks inherent token boundaries or predefined compositional rules, making tokenization a…

Machine Learning · Computer Science 2026-04-13 Nan Huang , Xiaoxiao Zhou , Junxia Cui , Mario Tapia-Pacheco , Tiffany Amariuta , Yang Li , Jingbo Shang

Equivariant graph neural network (GNN) methods for antibody complementarity-determining region (CDR) design achieve the highest sequence recovery but suffer from severe vocabulary collapse. The current best GNN methods over-predict very few…

Machine Learning · Computer Science 2026-05-21 Mansoor Ahmed , Sujin Lee , Umar Khayaz , Murray Patterson

Self-supervised word embedding algorithms such as word2vec provide a minimal setting for studying representation learning in language modeling. We examine the quartic Taylor approximation of the word2vec loss around the origin, and we show…

Machine Learning · Computer Science 2025-10-20 Dhruva Karkada , James B. Simon , Yasaman Bahri , Michael R. DeWeese

Contact interface properties are important in determining the performances of devices based on atomically thin two-dimensional (2D) materials, especially those with short channels. Understanding the contact interface is therefore quite…

Materials Science · Physics 2020-05-01 Bo Han , Chen Yang , Xiaolong Xu , Yuehui Li , Ruochen Shi , Kaihui Liu , Haicheng Wang , Yu Ye , Jing Lu , Dapeng Yu , Peng Gao

Epigenetics plays a key role in cellular differentiation and maintaining cell identity, enabling cells to regulate their genetic activity without altering the DNA sequence. Epigenetic regulation occurs within the context of hierarchically…

Quantitative Methods · Quantitative Biology 2024-09-11 Daria Stepanova , Meritxell Brunet Guasch , Helen M. Byrne , Tomás Alarcón

There is substantial interest in developing artificial intelligence systems to support radiologists across tasks ranging from segmentation to report generation. Existing computed tomography (CT) foundation models have largely focused on…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Rubén Moreno-Aguado , Alba Magallón , Victor Moreno , Yingying Fang , Guang Yang

The advent of foundation models has revolutionized the fields of natural language processing and computer vision, paving the way for their application in autonomous driving (AD). This survey presents a comprehensive review of more than 40…

Machine Learning · Computer Science 2024-09-06 Haoxiang Gao , Zhongruo Wang , Yaqian Li , Kaiwen Long , Ming Yang , Yiqing Shen

Egocentric video-language pretraining has significantly advanced video representation learning. Humans perceive and interact with a fully 3D world, developing spatial awareness that extends beyond text-based understanding. However, most…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Boshen Xu , Yuting Mei , Xinbi Liu , Sipeng Zheng , Qin Jin

This paper investigates the application of the transformer architecture in protein folding, as exemplified by DeepMind's AlphaFold project, and its implications for the understanding of so-called large language models. The prevailing…

Computers and Society · Computer Science 2024-12-10 Fabian Offert , Paul Kim , Qiaoyu Cai

Dense retrieval calls for discriminative embeddings to represent the semantic relationship between query and document. It may benefit from the using of large language models (LLMs), given LLMs' strong capability on semantic understanding.…

Computation and Language · Computer Science 2025-11-25 Zheng Liu , Chaofan Li , Shitao Xiao , Yingxia Shao , Defu Lian

Despite their black-box nature, deep learning models are extensively used in image-based drug discovery to extract feature vectors from single cells in microscopy images. To better understand how these networks perform representation…

Image and Video Processing · Electrical Eng. & Systems 2024-03-27 Vivek Gopalakrishnan , Jingzhe Ma , Zhiyong Xie

Recent advances in multimodal large language models (MLLMs) have shown impressive reasoning capabilities. However, further enhancing existing MLLMs necessitates high-quality vision-language datasets with carefully curated task complexities,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-30 Xiuwei Chen , Wentao Hu , Hanhui Li , Jun Zhou , Zisheng Chen , Meng Cao , Yihan Zeng , Kui Zhang , Yu-Jie Yuan , Jianhua Han , Hang Xu , Xiaodan Liang

Recent advances in generative models, particularly diffusion and auto-regressive models, have revolutionized fields like computer vision and natural language processing. However, their application to structure-based drug design (SBDD)…

Machine Learning · Computer Science 2025-07-29 Yi He , Ailun Wang , Zhi Wang , Yu Liu , Xingyuan Xu , Wen Yan

Embodied scene understanding requires not only comprehending visual-spatial information that has been observed but also determining where to explore next in the 3D physical world. Existing 3D Vision-Language (3D-VL) models primarily focus…

Computer Vision and Pattern Recognition · Computer Science 2025-07-31 Ziyu Zhu , Xilin Wang , Yixuan Li , Zhuofan Zhang , Xiaojian Ma , Yixin Chen , Baoxiong Jia , Wei Liang , Qian Yu , Zhidong Deng , Siyuan Huang , Qing Li

Essential life processes take place across multiple space and time scales in living organisms but understanding their mechanistic interactions remains an ongoing challenge. Advanced multiscale modeling techniques are providing new…

Molecular Networks · Quantitative Biology 2025-04-08 Achal Mahajan , Erik J. Navarro , William Poole , Carlos F Lopez

The 3D organisation of the genome in interphase cells is not a randomly folded polymer. Rather, experiments show that chromosomes arrange into a network of 3D compartments that correlate with biological processes, such as transcription,…

Biological Physics · Physics 2018-11-06 Kumar Rajendra , Ludvig Lizana , Per Stenberg

Self-supervised visual foundation models produce powerful embeddings that achieve remarkable performance on a wide range of downstream tasks. However, unlike vision-language models such as CLIP, self-supervised visual features are not…