中文
相关论文

相关论文: Protein Circuit Tracing via Cross-layer Transcoder…

200 篇论文

A plethora of protein language models have been released in recent years. Yet comparatively little work has addressed how to best sample from them to optimize desired biological properties. We fill this gap by proposing a flexible,…

机器学习 · 计算机科学 2026-05-08 Calvin McCarter , Nick Bhattacharya , Sebastian W. Ober , Hunter Elliott

Large Language Models (LLMs) have shown remarkable ability in solving complex tasks, making them a promising tool for enhancing tabular learning. However, existing LLM-based methods suffer from high resource requirements, suboptimal…

机器学习 · 计算机科学 2025-05-12 Ruxue Shi , Hengrui Gu , Xu Shen , Xin Wang

Protein language models (PLMs) have demonstrated remarkable capabilities in learning relationships between protein sequences and functions. However, finetuning these large models requires substantial computational resources, often with…

机器学习 · 计算机科学 2025-12-09 Shuo Zhang , Jian K. Liu

Comprehending the long-timescale dynamics of protein-ligand complexes is very important for drug discovery and structural biology, but it continues to be computationally challenging for large biomolecular systems. We introduce…

生物大分子 · 定量生物学 2025-08-26 Rakesh Thakur , Riya Gupta

Mechanistic interpretability seeks to understand how Large Language Models (LLMs) represent and process information. Recent approaches based on dictionary learning and transcoders enable representing model computation in terms of sparse,…

Transformer-based vision-language models (VLMs) contain substantial depth redundancy, yet the effect of removing specific decoder layers remains poorly understood, especially for domains that require tight coupling between perception and…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Saeed Khaki , Nima Safaei , Kamal Ginotra

Probabilistic Circuits (PCs) are deep generative models that support exact and efficient probabilistic inference. Yet in autoregressive language modeling, PCs still lag behind Transformer-based large language models (LLMs), suggesting an…

机器学习 · 计算机科学 2026-05-14 Zhiyu Zhao , Xuejie Liu , Muhan Zhang , Anji Liu

We present Pre-trained Machine Reader (PMR), a novel method for retrofitting pre-trained masked language models (MLMs) to pre-trained machine reading comprehension (MRC) models without acquiring labeled data. PMR can resolve the discrepancy…

计算与语言 · 计算机科学 2023-10-17 Weiwen Xu , Xin Li , Wenxuan Zhang , Meng Zhou , Wai Lam , Luo Si , Lidong Bing

Though the pre-trained contextualized language model (PrLM) has made a significant impact on NLP, training PrLMs in languages other than English can be impractical for two reasons: other languages often lack corpora sufficient for training…

计算与语言 · 计算机科学 2021-07-28 Zuchao Li , Kevin Parnow , Hai Zhao , Zhuosheng Zhang , Rui Wang , Masao Utiyama , Eiichiro Sumita

Protein language models are a powerful tool for learning protein representations through pre-training on vast protein sequence datasets. However, traditional protein language models lack explicit structural supervision, despite its…

生物大分子 · 定量生物学 2024-02-09 Zuobai Zhang , Jiarui Lu , Vijil Chenthamarakshan , Aurélie Lozano , Payel Das , Jian Tang

Printed electronics (PE) promises on-demand fabrication, low non-recurring engineering costs, and sub-cent fabrication costs. It also allows for high customization that would be infeasible in silicon, and bespoke architectures prevail to…

机器学习 · 计算机科学 2023-04-04 Giorgos Armeniakos , Georgios Zervakis , Dimitrios Soudris , Mehdi B. Tahoori , Jörg Henkel

Recent advancements in computational chemistry have leveraged the power of trans-former-based language models, such as MoLFormer, pre-trained using a vast amount of simplified molecular-input line-entry system (SMILES) sequences, to…

生物大分子 · 定量生物学 2024-11-05 Tianhao Peng , Yuchen Li , Xuhong Li , Jiang Bian , Zeke Xie , Ning Sui , Shahid Mumtaz , Yanwu Xu , Linghe Kong , Haoyi Xiong

Proteolysis targeting chimeras (PROTACs) are small molecules that trigger the breakdown of traditionally ``undruggable'' proteins by binding simultaneously to their targets and degradation-associated proteins. A key challenge in their…

生物大分子 · 定量生物学 2024-05-14 Bo Qiang , Wenxian Shi , Yuxuan Song , Menghua Wu

Fine-tuning Pre-trained protein language models (PLMs) has emerged as a prominent strategy for enhancing downstream prediction tasks, often outperforming traditional supervised learning approaches. As a widely applied powerful technique in…

计算与语言 · 计算机科学 2024-04-24 Yang Tan , Mingchen Li , Bingxin Zhou , Bozitao Zhong , Lirong Zheng , Pan Tan , Ziyi Zhou , Huiqun Yu , Guisheng Fan , Liang Hong

Recent advances in protein large language models, such as ProtTeX, represent both side-chain amino acids and backbone structure as discrete token sequences of residue length. While this design enables unified modeling of multimodal protein…

机器学习 · 计算机科学 2025-08-19 Chuanliu Fan , Zicheng Ma , Jun Gao , Nan Yu , Jun Zhang , Ziqiang Cao , Yi Qin Gao , Guohong Fu

Less than 1% of protein sequences are structurally and functionally annotated. Natural Language Processing (NLP) community has recently embraced self-supervised learning as a powerful approach to learn representations from unlabeled text,…

生物大分子 · 定量生物学 2020-12-08 Modestas Filipavicius , Matteo Manica , Joris Cadow , Maria Rodriguez Martinez

Vision-language models are integral to computer vision research, yet many high-performing models remain closed-source, obscuring their data, design and training recipe. The research community has responded by using distillation from…

Protein language models (pLMs) have demonstrated success at generating functional proteins across vast sequence spaces but lack the ability to design high-fitness variants on demand. Here, we iteratively guide pLMs toward user-defined…

生物大分子 · 定量生物学 2025-12-01 Filippo Stocco , Maria Artigues-Lleixa , Andrea Hunklinger , Talal Widatalla , Marc Guell , Noelia Ferruz

Free-text crash narratives recorded in real-world crash databases have been shown to play a significant role in improving traffic safety. However, large-scale analyses remain difficult to implement as there are no documented tools that can…

计算与语言 · 计算机科学 2025-10-13 Xixi Wang , Jordanka Kovaceva , Miguel Costa , Shuai Wang , Francisco Camara Pereira , Robert Thomson

As Large Language Models (LLMs) continue to grow in size, storing and transmitting them on edge devices becomes increasingly challenging. Traditional methods like quantization and pruning struggle to achieve extreme compression of LLMs…

机器学习 · 计算机科学 2025-11-25 Ye Tian , Chengcheng Wang , Jing Han , Yehui Tang , Kai Han