English
Related papers

Related papers: T-Cell Receptor Optimization with Reinforcement Le…

200 papers

Trust Region Policy Optimization (TRPO) is a popular and empirically successful policy search algorithm in reinforcement learning (RL). It iteratively solved the surrogate problem which restricts consecutive policies to be close to each…

Machine Learning · Computer Science 2021-10-27 Sahar Roostaie , Mohammad Mehdi Ebadzadeh

Query generation is a critical task for web search engines (e.g. Google, Bing) and recommendation systems. Recently, state-of-the-art query generation methods leverage Large Language Models (LLMs) for their strong capabilities in context…

Diverse T and B cell repertoires play an important role in mounting effective immune responses against a wide range of pathogens and malignant cells. The number of unique T and B cell clones is characterized by T and B cell receptors (TCRs…

Quantitative Methods · Quantitative Biology 2026-05-12 Lucas Böttcher , Sascha Wald , Tom Chou

Legged locomotion in unstructured environments demands not only high-performance control policies but also formal guarantees to ensure robustness under perturbations. Control methods often require carefully designed reference trajectories,…

Robotics · Computer Science 2026-03-23 Vrushabh Zinage , Narek Harutyunyan , Eric Verheyden , Fred Y. Hadaegh , Soon-Jo Chung

Reinforcement learning from verifiable rewards (RLVR), especially with Group Relative Policy Optimization (GRPO), has shown strong potential for improving the reasoning capabilities of large vision-language models (LVLMs). However, in…

Artificial Intelligence · Computer Science 2026-05-11 Bingqing Jiang , Difan Zou

Immune cells learn about their antigenic targets using tactile sense: during recognition, a highly organized yet dynamic motif, named immunological synapse, forms between immune cells and antigen-presenting cells (APCs). Via synapses,…

Biological Physics · Physics 2018-12-12 Miloš Knežević , Shenshen Wang

Reducing operation and maintenance costs is a key objective for advanced reactors in general and microreactors in particular. To achieve this reduction, developing robust autonomous control algorithms is essential to ensure safe and…

Systems and Control · Electrical Eng. & Systems 2024-06-25 Majdi I. Radaideh , Leo Tunkle , Dean Price , Kamal Abdulraheem , Linyu Lin , Moutaz Elias

Modern precision medicine aims to utilize real-world data to provide the best treatment for an individual patient. An individualized treatment rule (ITR) maps each patient's characteristics to a recommended treatment scheme that maximizes…

Applications · Statistics 2025-01-07 Andong Wang , Kelly Wentzlof , Johnny Rajala , Miontranese Green , Yunshu Zhang , Shu Yang

B cells signaling in response to antigen is proportional to antigen affinity, a process known as affinity discrimination. Recent research suggests that B cells can acquire antigen in membrane-bound form on the surface of antigen-presenting…

Cell Behavior · Quantitative Biology 2010-03-09 Philippos K. Tsourkas , Subhadip Raychaudhuri

In safe reinforcement learning (SRL) problems, an agent explores the environment to maximize an expected total reward and meanwhile avoids violation of certain constraints on a number of expected total costs. In general, such SRL problems…

Machine Learning · Computer Science 2021-06-01 Tengyu Xu , Yingbin Liang , Guanghui Lan

In protein biophysics, the separation between the functionally important residues (forming the active site or binding surface) and those that create the overall structure (the fold) is a well-established and fundamental concept. Identifying…

Biomolecules · Quantitative Biology 2023-10-17 Tianxiao Li , Hongyu Guo , Filippo Grazioli , Mark Gerstein , Martin Renqiang Min

This paper provides a self-contained, from-scratch, exposition of key algorithms for instruction tuning of models: SFT, Rejection Sampling, REINFORCE, Trust Region Policy Optimization (TRPO), Proximal Policy Optimization (PPO), Group…

Computation and Language · Computer Science 2025-10-22 Rohit Patel

The problem of \emph{statistical recognition} is considered, as it arises in immunobiology, namely, the discrimination of foreign antigens against a background of the body's own molecules. The precise mechanism of this…

Subcellular Processes · Quantitative Biology 2009-09-30 Florian Lipsmeier , Ellen Baake

Reinforcement Learning (RL) agents can solve diverse tasks but often exhibit unsafe behavior. Constrained Markov Decision Processes (CMDPs) address this by enforcing safety constraints, yet existing methods either sacrifice reward…

Machine Learning · Computer Science 2025-08-18 Nikola Milosevic , Johannes Müller , Nico Scherf

Reinforcement learning (RL) has become a powerful tool for post-training visual generative models, with Group Relative Policy Optimization (GRPO) increasingly used to align generators with human preferences. However, existing GRPO pipelines…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Ziqi Ni , Yuanzhi Liang , Rui Li , Yi Zhou , Haibin Huang , Chi Zhang , Xuelong Li

Biological systems have evolved to amazingly complex states, yet we do not understand in general how evolution operates to generate increasing genetic and functional complexity. Molecular recognition sites are short genome segments or…

Populations and Evolution · Quantitative Biology 2025-01-06 Tom Röschinger , Roberto Morán Tovar , Simone Pompei , Michael Lässig

Retrosynthesis, of which the goal is to find a set of reactants for synthesizing a target product, is an emerging research area of deep learning. While the existing approaches have shown promising results, they currently lack the ability to…

Machine Learning · Computer Science 2021-06-04 Hankook Lee , Sungsoo Ahn , Seung-Woo Seo , You Young Song , Eunho Yang , Sung-Ju Hwang , Jinwoo Shin

A number of image-processing problems can be formulated as optimization problems. The objective function typically contains several terms specifically designed for different purposes. Parameters in front of these terms are used to control…

Medical Physics · Physics 2017-11-02 Chenyang Shen , Yesenia Gonzalez , Liyuan Chen , Steve B. Jiang , Xun Jia

Unsupervised cell type identification is crucial for uncovering and characterizing heterogeneous populations in single cell omics studies. Although a range of clustering methods have been developed, most focus exclusively on intrinsic…

Artificial Intelligence · Computer Science 2025-12-12 Liang Peng , Haopeng Liu , Yixuan Ye , Cheng Liu , Wenjun Shen , Si Wu , Hau-San Wong

Portfolio management is a fundamental problem in finance. It involves periodic reallocations of assets to maximize the expected returns within an appropriate level of risk exposure. Deep reinforcement learning (RL) has been considered a…

Computational Finance · Quantitative Finance 2022-10-05 Hui Niu , Siyuan Li , Jian Li