English
Related papers

Related papers: HYPE-EDIT-1: Benchmark for Measuring Reliability i…

200 papers

Reward models are critical in techniques like Reinforcement Learning from Human Feedback (RLHF) and Inference Scaling Laws, where they guide language model alignment and select optimal responses. Despite their importance, existing reward…

Computation and Language · Computer Science 2024-10-22 Yantao Liu , Zijun Yao , Rui Min , Yixin Cao , Lei Hou , Juanzi Li

Micro-benchmarking offers a solution to the often prohibitive time and cost of language model development: evaluate on a very small subset of existing benchmarks. Can these micro-benchmarks, however, rank models as consistently as the full…

Computation and Language · Computer Science 2026-03-09 Gregory Yauney , Shahzaib Saqib Warraich , Swabha Swayamdipta

High-quality training triplets (instruction, original image, edited image) are essential for instruction-based image editing. Predominant training datasets (e.g., InsPix2Pix) are created using text-to-image generative models (e.g., Stable…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Xin Gu , Ming Li , Libo Zhang , Fan Chen , Longyin Wen , Tiejian Luo , Sijie Zhu

A key challenge in MT evaluation is the inherent noise and inconsistency of human ratings. Regression-based neural metrics struggle with this noise, while prompting LLMs shows promise at system-level evaluation but performs poorly at…

Computation and Language · Computer Science 2025-04-21 Shaomu Tan , Christof Monz

Large language models (LLMs) are widely deployed, but their substantial compute demands make them vulnerable to inference cost attacks that aim to deliberately maximize the output length. In this work, we investigate a distinct attack…

Cryptography and Security · Computer Science 2026-02-24 Xiaobei Yan , Yiming Li , Hao Wang , Han Qiu , Tianwei Zhang

MCP standardizes how LLMs interact with external systems, forming the foundation for general agents. However, existing MCP benchmarks remain narrow in scope: they focus on read-heavy tasks or tasks with limited interaction depth, and fail…

Large language models excel at reasoning but lack key aspects of introspection, including anticipating their own success and the computation required to achieve it. Humans use real-time introspection to decide how much effort to invest,…

Machine Learning · Computer Science 2025-12-24 Rohin Manvi , Joey Hong , Tim Seyde , Maxime Labonne , Mathias Lechner , Sergey Levine

Recent text-guided image editing (TIE) models have achieved remarkable progress, however, many edited results still suffer from artifacts, unintended modifications, and suboptimal aesthetics. Although several benchmarks and evaluation…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Honghua Chen , Zitong Xu , Huiyu Duan , Xinyun Zhang , Xiongkuo Min , Guangtao Zhai

The proliferation of large language models (LLMs) with varying computational costs and performance profiles presents a critical challenge for scalable, cost-effective deployment in real-world applications. We introduce a unified routing…

Automated patent claim validation demands low error tolerance. However, existing approaches face a rigidity-resource dilemma: lightweight encoders cannot track long-range legal dependencies, while exhaustive LLM verification incurs 4-5X…

Computation and Language · Computer Science 2026-05-28 Yongmin Yoo , Qiongkai Xu , Longbing Cao

Large language model (LLM)-based reasoning systems have recently achieved gold medal-level performance in the IMO 2025 competition, writing mathematical proofs where, to receive full credit, each step must be not only correct but also…

Artificial Intelligence · Computer Science 2025-10-16 Shrey Pandit , Austin Xu , Xuan-Phi Nguyen , Yifei Ming , Caiming Xiong , Shafiq Joty

Recently machine unlearning (MU) is proposed to remove the imprints of revoked samples from the already trained model parameters, to solve users' privacy concern. Different from the runtime expensive retraining from scratch, there exist two…

Machine Learning · Computer Science 2024-12-20 Mingxin Li , Yizhen Yu , Ning Wang , Zhigang Wang , Xiaodong Wang , Haipeng Qu , Jia Xu , Shen Su , Zhichao Yin

Automated writing evaluation (AWE) has been shown to be an effective mechanism for quickly providing feedback to students. It has already seen wide adoption in enterprise-scale applications and is starting to be adopted in large-scale…

Computation and Language · Computer Science 2014-12-19 Nicholas Dronen , Peter W. Foltz , Kyle Habermehl

The rapid growth of machine learning has produced an ever-expanding ecosystem of models, making it increasingly challenging to verify the reliability of newly released models on unseen, unlabeled data. Conventional evaluation pipelines…

Machine Learning · Computer Science 2026-05-25 Trinh Pham , Viet Huynh , Hongzhi Yin , Quoc Viet Hung Nguyen , Thanh Tam Nguyen

LLMs have achieved strong results on both function-level code synthesis and repository-level code modification, yet a capability that falls between these two extremes -- compositional code creation, i.e., building a complete, internally…

Software Engineering · Computer Science 2026-04-30 Yeheng Chen , Chaoxiang Xie , Yuling Shi , Wenhao Zeng , Yongpan Wang , Hongyu Zhang , Xiaodong Gu

This work presents an open-source unified benchmarking and evaluation framework for text-to-image generation models, with a particular focus on the impact of metadata augmented prompts. Leveraging the DeepFashion-MultiModal dataset, we…

Graphics · Computer Science 2025-05-09 Kapil Wanaskar , Gaytri Jena , Magdalini Eirinaki

This paper reviews the NTIRE 2022 challenge on efficient single image super-resolution with focus on the proposed solutions and results. The task of the challenge was to super-resolve an input image with a magnification factor of $\times$4…

Computer Vision and Pattern Recognition · Computer Science 2022-05-12 Yawei Li , Kai Zhang , Radu Timofte , Luc Van Gool , Fangyuan Kong , Mingxi Li , Songwei Liu , Zongcai Du , Ding Liu , Chenhui Zhou , Jingyi Chen , Qingrui Han , Zheyuan Li , Yingqi Liu , Xiangyu Chen , Haoming Cai , Yu Qiao , Chao Dong , Long Sun , Jinshan Pan , Yi Zhu , Zhikai Zong , Xiaoxiao Liu , Zheng Hui , Tao Yang , Peiran Ren , Xuansong Xie , Xian-Sheng Hua , Yanbo Wang , Xiaozhong Ji , Chuming Lin , Donghao Luo , Ying Tai , Chengjie Wang , Zhizhong Zhang , Yuan Xie , Shen Cheng , Ziwei Luo , Lei Yu , Zhihong Wen , Qi Wu1 , Youwei Li , Haoqiang Fan , Jian Sun , Shuaicheng Liu , Yuanfei Huang , Meiguang Jin , Hua Huang , Jing Liu , Xinjian Zhang , Yan Wang , Lingshun Long , Gen Li , Yuanfan Zhang , Zuowei Cao , Lei Sun , Panaetov Alexander , Yucong Wang , Minjie Cai , Li Wang , Lu Tian , Zheyuan Wang , Hongbing Ma , Jie Liu , Chao Chen , Yidong Cai , Jie Tang , Gangshan Wu , Weiran Wang , Shirui Huang , Honglei Lu , Huan Liu , Keyan Wang , Jun Chen , Shi Chen , Yuchun Miao , Zimo Huang , Lefei Zhang , Mustafa Ayazoğlu , Wei Xiong , Chengyi Xiong , Fei Wang , Hao Li , Ruimian Wen , Zhijing Yang , Wenbin Zou , Weixin Zheng , Tian Ye , Yuncheng Zhang , Xiangzhen Kong , Aditya Arora , Syed Waqas Zamir , Salman Khan , Munawar Hayat , Fahad Shahbaz Khan , Dandan Gaoand Dengwen Zhouand Qian Ning , Jingzhu Tang , Han Huang , Yufei Wang , Zhangheng Peng , Haobo Li , Wenxue Guan , Shenghua Gong , Xin Li , Jun Liu , Wanjun Wang , Dengwen Zhou , Kun Zeng , Hanjiang Lin , Xinyu Chen , Jinsheng Fang

Rating-based human evaluation has become an essential tool to accurately evaluate the impressive performance of large language models (LLMs). However, current rating systems suffer from several important limitations: first, they fail to…

Computation and Language · Computer Science 2025-02-12 Jasper Dekoninck , Maximilian Baader , Martin Vechev

We present a comprehensive solution to learn and improve text-to-image models from human preference feedback. To begin with, we build ImageReward -- the first general-purpose text-to-image human preference reward model -- to effectively…

Computer Vision and Pattern Recognition · Computer Science 2023-12-29 Jiazheng Xu , Xiao Liu , Yuchen Wu , Yuxuan Tong , Qinkai Li , Ming Ding , Jie Tang , Yuxiao Dong

In the e-commerce realm, compelling advertising images are pivotal for attracting customer attention. While generative models automate image generation, they often produce substandard images that may mislead customers and require…

Computer Vision and Pattern Recognition · Computer Science 2024-08-02 Zhenbang Du , Wei Feng , Haohan Wang , Yaoyu Li , Jingsen Wang , Jian Li , Zheng Zhang , Jingjing Lv , Xin Zhu , Junsheng Jin , Junjie Shen , Zhangang Lin , Jingping Shao
‹ Prev 1 4 5 6 7 8 10 Next ›