English
Related papers

Related papers: AnnoDPO: Protein Functional Annotation Learning wi…

200 papers

Diffusion models have achieved state-of-the-art performance across multiple domains, with recent advancements extending their applicability to discrete data. However, aligning discrete diffusion models with task-specific preferences remains…

Machine Learning · Computer Science 2025-04-10 Umberto Borso , Davide Paglieri , Jude Wells , Tim Rocktäschel

The capability of accurate prediction of protein functions and properties is essential in the biotechnology industry, e.g. drug development and artificial protein synthesis, etc. The main challenges of protein function prediction are the…

Quantitative Methods · Quantitative Biology 2021-12-02 Wei-Cheng Tseng , Po-Han Chi , Jia-Hua Wu , Min Sun

Although LLMs have achieved significant success, their reliance on large volumes of human-annotated data has limited their potential for further scaling. In this situation, utilizing self-generated synthetic data has become crucial for…

Computation and Language · Computer Science 2026-03-17 Haoyan Yang , Khiem Le , Ting Hua , Shangqian Gao , Binfeng Xu , Zheng Tang , Jie Xu , Nitesh V. Chawla , Hongxia Jin , Vijay Srinivasan

Diffusion models have achieved remarkable progress in text-to-image generation, yet aligning them with human preference remains challenging due to the presence of multiple, sometimes conflicting, evaluation metrics (e.g., semantic…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Dipesh Tamboli , Souradip Chakraborty , Aditya Malusare , Biplab Banerjee , Amrit Singh Bedi , Vaneet Aggarwal

Preference-based alignment objectives have been widely adopted, from RLHF-style pairwise learning in large language models to emerging applications in recommender systems. Yet, existing work rarely examines how Direct Preference…

Information Retrieval · Computer Science 2026-04-01 Hejin Huang , Jusheng Zhang , Kaitong Cai , Jian Wang , Rong Pan

DPO has become a widely adopted alternative to RLHF for aligning LLMs with human preferences, eliminating the need for a separate reward model or RL loop. Recent theoretical analysis uncovers an asymmetric gradient behavior in DPO: the loss…

Computation and Language · Computer Science 2026-05-28 Shaolong Chen , Madalina Ciobanu , Qingqing Mao , Ritankar Das

The role of reinforcement learning (RL) in enhancing the reasoning of large language models (LLMs) is becoming increasingly significant. Despite the success of RL in many scenarios, there are still many challenges in improving the reasoning…

Artificial Intelligence · Computer Science 2024-12-25 Jiacai Liu , Chaojie Wang , Chris Yuhao Liu , Liang Zeng , Rui Yan , Yiwen Sun , Yang Liu , Yahui Zhou

Background: The increasing volume and variety of genotypic and phenotypic data is a major defining characteristic of modern biomedical sciences. At the same time, the limitations in technology for generating data and the inherently…

Quantitative Methods · Quantitative Biology 2016-12-07 Yuxiang Jiang , Tal Ronnen Oron , Wyatt T Clark , Asma R Bankapur , Daniel D'Andrea , Rosalba Lepore , Christopher S Funk , Indika Kahanda , Karin M Verspoor , Asa Ben-Hur , Emily Koo , Duncan Penfold-Brown , Dennis Shasha , Noah Youngs , Richard Bonneau , Alexandra Lin , Sayed ME Sahraeian , Pier Luigi Martelli , Giuseppe Profiti , Rita Casadio , Renzhi Cao , Zhaolong Zhong , Jianlin Cheng , Adrian Altenhoff , Nives Skunca , Christophe Dessimoz , Tunca Dogan , Kai Hakala , Suwisa Kaewphan , Farrokh Mehryary , Tapio Salakoski , Filip Ginter , Hai Fang , Ben Smithers , Matt Oates , Julian Gough , Petri Törönen , Patrik Koskinen , Liisa Holm , Ching-Tai Chen , Wen-Lian Hsu , Kevin Bryson , Domenico Cozzetto , Federico Minneci , David T Jones , Samuel Chapman , Dukka B K. C. , Ishita K Khan , Daisuke Kihara , Dan Ofer , Nadav Rappoport , Amos Stern , Elena Cibrian-Uhalte , Paul Denny , Rebecca E Foulger , Reija Hieta , Duncan Legge , Ruth C Lovering , Michele Magrane , Anna N Melidoni , Prudence Mutowo-Meullenet , Klemens Pichler , Aleksandra Shypitsyna , Biao Li , Pooya Zakeri , Sarah ElShal , Léon-Charles Tranchevent , Sayoni Das , Natalie L Dawson , David Lee , Jonathan G Lees , Ian Sillitoe , Prajwal Bhat , Tamás Nepusz , Alfonso E Romero , Rajkumar Sasidharan , Haixuan Yang , Alberto Paccanaro , Jesse Gillis , Adriana E Sedeño-Cortés , Paul Pavlidis , Shou Feng , Juan M Cejuela , Tatyana Goldberg , Tobias Hamp , Lothar Richter , Asaf Salamov , Toni Gabaldon , Marina Marcet-Houben , Fran Supek , Qingtian Gong , Wei Ning , Yuanpeng Zhou , Weidong Tian , Marco Falda , Paolo Fontana , Enrico Lavezzo , Stefano Toppo , Carlo Ferrari , Manuel Giollo , Damiano Piovesan , Silvio Tosatto , Angela del Pozo , José M Fernández , Paolo Maietta , Alfonso Valencia , Michael L Tress , Alfredo Benso , Stefano Di Carlo , Gianfranco Politano , Alessandro Savino , Hafeez Ur Rehman , Matteo Re , Marco Mesiti , Giorgio Valentini , Joachim W Bargsten , Aalt DJ van Dijk , Branislava Gemovic , Sanja Glisic , Vladmir Perovic , Veljko Veljkovic , Nevena Veljkovic , Danillo C Almeida-e-Silva , Ricardo ZN Vencio , Malvika Sharan , Jörg Vogel , Lakesh Kansakar , Shanshan Zhang , Slobodan Vucetic , Zheng Wang , Michael JE Sternberg , Mark N Wass , Rachael P Huntley , Maria J Martin , Claire O'Donovan , Peter N Robinson , Yves Moreau , Anna Tramontano , Patricia C Babbitt , Steven E Brenner , Michal Linial , Christine A Orengo , Burkhard Rost , Casey S Greene , Sean D Mooney , Iddo Friedberg , Predrag Radivojac

Aligning large language models (LLMs) with human preferences becomes a key component to obtaining state-of-the-art performance, but it yields a huge cost to construct a large human-annotated preference dataset. To tackle this problem, we…

Machine Learning · Computer Science 2025-03-05 Dongyoung Kim , Kimin Lee , Jinwoo Shin , Jaehyung Kim

Preference optimization is a critical post-training technique used to align large language models (LLMs) with human preferences, typically by fine-tuning on ranked response pairs. While methods like Direct Preference Optimization (DPO) have…

Computation and Language · Computer Science 2025-11-12 Rhitabrat Pokharel , Yufei Tao , Ameeta Agrawal

We introduce ConfPO, a method for preference learning in Large Language Models (LLMs) that identifies and optimizes preference-critical tokens based solely on the training policy's confidence, without requiring any auxiliary models or…

Computation and Language · Computer Science 2025-06-13 Hee Suk Yoon , Eunseop Yoon , Mark Hasegawa-Johnson , Sungwoong Kim , Chang D. Yoo

Large language models have been widely adopted in natural language processing, yet they face the challenge of generating unreliable content. Recent works aim to reduce misinformation and hallucinations by resorting to attribution as a means…

Computation and Language · Computer Science 2024-03-28 Dongfang Li , Zetian Sun , Baotian Hu , Zhenyu Liu , Xinshuo Hu , Xuebo Liu , Min Zhang

Preference-based feedback is important for many applications in machine learning where evaluation of a reward function is not feasible. Notable recent examples arise in preference alignment for large language models, including in…

Preference alignment is pivotal for empowering large language models (LLMs) to generate helpful and harmless responses. However, the performance of preference alignment is highly sensitive to the prevalent noise in the preference data.…

Machine Learning · Computer Science 2024-05-29 Xize Liang , Chao Chen , Shuang Qiu , Jie Wang , Yue Wu , Zhihang Fu , Zhihao Shi , Feng Wu , Jieping Ye

While Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse natural language processing tasks, aligning these models with varying human preferences across multiple objectives remains a significant challenge…

Computation and Language · Computer Science 2025-11-17 Biao Liu , Ning Xu , Junming Yang , Xin Geng

The increasing capabilities of large language models (LLMs) raise opportunities for artificial general intelligence but concurrently amplify safety concerns, such as potential misuse of AI systems, necessitating effective AI alignment.…

Machine Learning · Computer Science 2023-09-29 Chaoqi Wang , Yibo Jiang , Chenghao Yang , Han Liu , Yuxin Chen

Aligning Large Language Models (LLMs) to human preferences in content, style, and presentation is challenging, in part because preferences are varied, context-dependent, and sometimes inherently ambiguous. While successful, Reinforcement…

Machine Learning · Computer Science 2024-10-29 Sam Houliston , Alizée Pace , Alexander Immer , Gunnar Rätsch

Large Language Models (LLMs) rely on Human Preference Alignment (HPA) to ensure the generation of safe content. Due to the heavy cost associated with fine-tuning, fine-tuning-free methods have emerged, typically modifying LLM decoding with…

Computation and Language · Computer Science 2024-02-15 Feifan Song , Yuxuan Fan , Xin Zhang , Peiyi Wang , Houfeng Wang

With the rapid development of Large Language Models (LLMs), numerous Reinforcement Learning from Human Feedback (RLHF) algorithms have been introduced to improve model safety and alignment with human preferences. These algorithms can be…

Machine Learning · Computer Science 2025-02-06 Xuerui Su , Yue Wang , Jinhua Zhu , Mingyang Yi , Feng Xu , Zhiming Ma , Yuting Liu

Direct Preference Optimization is an offline post-SFT method for aligning language models from preference pairs, with strong results in instruction following and summarization. However, DPO's sequence-level implicit reward can be brittle…

Computation and Language · Computer Science 2026-03-03 Samah Fodeh , Linhai Ma , Ganesh Puthiaraju , Srivani Talakokkul , Afshan Khan , Ashley Hagaman , Sarah R. Lowe , Aimee Kendall Roundtree