English
Related papers

Related papers: Understanding AI Evaluation Patterns: How Differen…

200 papers

Large Language Models (LLMs) are demonstrating remarkable human like capabilities across diverse domains, including psychological assessment. This study evaluates whether LLMs, specifically GPT-4o and GPT-4o mini, can infer Big Five…

Computation and Language · Computer Science 2025-01-14 Jianfeng Zhu , Ruoming Jin , Karin G. Coifman

GPT-4o is an autoregressive omni model that accepts as input any combination of text, audio, image, and video, and generates any combination of text, audio, and image outputs. It's trained end-to-end across text, vision, and audio, meaning…

Computation and Language · Computer Science 2024-10-29 OpenAI , : , Aaron Hurst , Adam Lerer , Adam P. Goucher , Adam Perelman , Aditya Ramesh , Aidan Clark , AJ Ostrow , Akila Welihinda , Alan Hayes , Alec Radford , Aleksander Mądry , Alex Baker-Whitcomb , Alex Beutel , Alex Borzunov , Alex Carney , Alex Chow , Alex Kirillov , Alex Nichol , Alex Paino , Alex Renzin , Alex Tachard Passos , Alexander Kirillov , Alexi Christakis , Alexis Conneau , Ali Kamali , Allan Jabri , Allison Moyer , Allison Tam , Amadou Crookes , Amin Tootoochian , Amin Tootoonchian , Ananya Kumar , Andrea Vallone , Andrej Karpathy , Andrew Braunstein , Andrew Cann , Andrew Codispoti , Andrew Galu , Andrew Kondrich , Andrew Tulloch , Andrey Mishchenko , Angela Baek , Angela Jiang , Antoine Pelisse , Antonia Woodford , Anuj Gosalia , Arka Dhar , Ashley Pantuliano , Avi Nayak , Avital Oliver , Barret Zoph , Behrooz Ghorbani , Ben Leimberger , Ben Rossen , Ben Sokolowsky , Ben Wang , Benjamin Zweig , Beth Hoover , Blake Samic , Bob McGrew , Bobby Spero , Bogo Giertler , Bowen Cheng , Brad Lightcap , Brandon Walkin , Brendan Quinn , Brian Guarraci , Brian Hsu , Bright Kellogg , Brydon Eastman , Camillo Lugaresi , Carroll Wainwright , Cary Bassin , Cary Hudson , Casey Chu , Chad Nelson , Chak Li , Chan Jun Shern , Channing Conger , Charlotte Barette , Chelsea Voss , Chen Ding , Cheng Lu , Chong Zhang , Chris Beaumont , Chris Hallacy , Chris Koch , Christian Gibson , Christina Kim , Christine Choi , Christine McLeavey , Christopher Hesse , Claudia Fischer , Clemens Winter , Coley Czarnecki , Colin Jarvis , Colin Wei , Constantin Koumouzelis , Dane Sherburn , Daniel Kappler , Daniel Levin , Daniel Levy , David Carr , David Farhi , David Mely , David Robinson , David Sasaki , Denny Jin , Dev Valladares , Dimitris Tsipras , Doug Li , Duc Phong Nguyen , Duncan Findlay , Edede Oiwoh , Edmund Wong , Ehsan Asdar , Elizabeth Proehl , Elizabeth Yang , Eric Antonow , Eric Kramer , Eric Peterson , Eric Sigler , Eric Wallace , Eugene Brevdo , Evan Mays , Farzad Khorasani , Felipe Petroski Such , Filippo Raso , Francis Zhang , Fred von Lohmann , Freddie Sulit , Gabriel Goh , Gene Oden , Geoff Salmon , Giulio Starace , Greg Brockman , Hadi Salman , Haiming Bao , Haitang Hu , Hannah Wong , Haoyu Wang , Heather Schmidt , Heather Whitney , Heewoo Jun , Hendrik Kirchner , Henrique Ponde de Oliveira Pinto , Hongyu Ren , Huiwen Chang , Hyung Won Chung , Ian Kivlichan , Ian O'Connell , Ian O'Connell , Ian Osband , Ian Silber , Ian Sohl , Ibrahim Okuyucu , Ikai Lan , Ilya Kostrikov , Ilya Sutskever , Ingmar Kanitscheider , Ishaan Gulrajani , Jacob Coxon , Jacob Menick , Jakub Pachocki , James Aung , James Betker , James Crooks , James Lennon , Jamie Kiros , Jan Leike , Jane Park , Jason Kwon , Jason Phang , Jason Teplitz , Jason Wei , Jason Wolfe , Jay Chen , Jeff Harris , Jenia Varavva , Jessica Gan Lee , Jessica Shieh , Ji Lin , Jiahui Yu , Jiayi Weng , Jie Tang , Jieqi Yu , Joanne Jang , Joaquin Quinonero Candela , Joe Beutler , Joe Landers , Joel Parish , Johannes Heidecke , John Schulman , Jonathan Lachman , Jonathan McKay , Jonathan Uesato , Jonathan Ward , Jong Wook Kim , Joost Huizinga , Jordan Sitkin , Jos Kraaijeveld , Josh Gross , Josh Kaplan , Josh Snyder , Joshua Achiam , Joy Jiao , Joyce Lee , Juntang Zhuang , Justyn Harriman , Kai Fricke , Kai Hayashi , Karan Singhal , Katy Shi , Kavin Karthik , Kayla Wood , Kendra Rimbach , Kenny Hsu , Kenny Nguyen , Keren Gu-Lemberg , Kevin Button , Kevin Liu , Kiel Howe , Krithika Muthukumar , Kyle Luther , Lama Ahmad , Larry Kai , Lauren Itow , Lauren Workman , Leher Pathak , Leo Chen , Li Jing , Lia Guy , Liam Fedus , Liang Zhou , Lien Mamitsuka , Lilian Weng , Lindsay McCallum , Lindsey Held , Long Ouyang , Louis Feuvrier , Lu Zhang , Lukas Kondraciuk , Lukasz Kaiser , Luke Hewitt , Luke Metz , Lyric Doshi , Mada Aflak , Maddie Simens , Madelaine Boyd , Madeleine Thompson , Marat Dukhan , Mark Chen , Mark Gray , Mark Hudnall , Marvin Zhang , Marwan Aljubeh , Mateusz Litwin , Matthew Zeng , Max Johnson , Maya Shetty , Mayank Gupta , Meghan Shah , Mehmet Yatbaz , Meng Jia Yang , Mengchao Zhong , Mia Glaese , Mianna Chen , Michael Janner , Michael Lampe , Michael Petrov , Michael Wu , Michele Wang , Michelle Fradin , Michelle Pokrass , Miguel Castro , Miguel Oom Temudo de Castro , Mikhail Pavlov , Miles Brundage , Miles Wang , Minal Khan , Mira Murati , Mo Bavarian , Molly Lin , Murat Yesildal , Nacho Soto , Natalia Gimelshein , Natalie Cone , Natalie Staudacher , Natalie Summers , Natan LaFontaine , Neil Chowdhury , Nick Ryder , Nick Stathas , Nick Turley , Nik Tezak , Niko Felix , Nithanth Kudige , Nitish Keskar , Noah Deutsch , Noel Bundick , Nora Puckett , Ofir Nachum , Ola Okelola , Oleg Boiko , Oleg Murk , Oliver Jaffe , Olivia Watkins , Olivier Godement , Owen Campbell-Moore , Patrick Chao , Paul McMillan , Pavel Belov , Peng Su , Peter Bak , Peter Bakkum , Peter Deng , Peter Dolan , Peter Hoeschele , Peter Welinder , Phil Tillet , Philip Pronin , Philippe Tillet , Prafulla Dhariwal , Qiming Yuan , Rachel Dias , Rachel Lim , Rahul Arora , Rajan Troll , Randall Lin , Rapha Gontijo Lopes , Raul Puri , Reah Miyara , Reimar Leike , Renaud Gaubert , Reza Zamani , Ricky Wang , Rob Donnelly , Rob Honsby , Rocky Smith , Rohan Sahai , Rohit Ramchandani , Romain Huet , Rory Carmichael , Rowan Zellers , Roy Chen , Ruby Chen , Ruslan Nigmatullin , Ryan Cheu , Saachi Jain , Sam Altman , Sam Schoenholz , Sam Toizer , Samuel Miserendino , Sandhini Agarwal , Sara Culver , Scott Ethersmith , Scott Gray , Sean Grove , Sean Metzger , Shamez Hermani , Shantanu Jain , Shengjia Zhao , Sherwin Wu , Shino Jomoto , Shirong Wu , Shuaiqi , Xia , Sonia Phene , Spencer Papay , Srinivas Narayanan , Steve Coffey , Steve Lee , Stewart Hall , Suchir Balaji , Tal Broda , Tal Stramer , Tao Xu , Tarun Gogineni , Taya Christianson , Ted Sanders , Tejal Patwardhan , Thomas Cunninghman , Thomas Degry , Thomas Dimson , Thomas Raoux , Thomas Shadwell , Tianhao Zheng , Todd Underwood , Todor Markov , Toki Sherbakov , Tom Rubin , Tom Stasi , Tomer Kaftan , Tristan Heywood , Troy Peterson , Tyce Walters , Tyna Eloundou , Valerie Qi , Veit Moeller , Vinnie Monaco , Vishal Kuo , Vlad Fomenko , Wayne Chang , Weiyi Zheng , Wenda Zhou , Wesam Manassra , Will Sheu , Wojciech Zaremba , Yash Patil , Yilei Qian , Yongjik Kim , Youlong Cheng , Yu Zhang , Yuchen He , Yuchen Zhang , Yujia Jin , Yunxing Dai , Yury Malkov

Generative Pre-trained Transformer (GPT) models have exhibited exciting progress in their capabilities, capturing the interest of practitioners and the public alike. Yet, while the literature on the trustworthiness of GPT models remains…

We present a large-scale study of linguistic bias exhibited by ChatGPT covering ten dialects of English (Standard American English, Standard British English, and eight widely spoken non-"standard" varieties from around the world). We…

Computation and Language · Computer Science 2024-09-18 Eve Fleisig , Genevieve Smith , Madeline Bossi , Ishita Rustagi , Xavier Yin , Dan Klein

Generative AI and large language models hold great promise in enhancing programming education by automatically generating individualized feedback for students. We investigate the role of generative AI models in providing human tutor-style…

Artificial intelligence (AI) researchers have been developing and refining large language models (LLMs) that exhibit remarkable capabilities across a variety of domains and tasks, challenging our understanding of learning and cognition. The…

Self-evaluation using large language models (LLMs) has proven valuable not only in benchmarking but also methods like reward modeling, constitutional AI, and self-refinement. But new biases are introduced due to the same LLM acting as both…

Computation and Language · Computer Science 2024-04-23 Arjun Panickssery , Samuel R. Bowman , Shi Feng

The rapid advancement of Large Language Models (LLMs) presents a significant challenge to academic integrity within computing education. As educators seek reliable detection methods, this paper evaluates the capacity of three prominent LLMs…

Computers and Society · Computer Science 2025-12-30 Christopher Burger , Karmece Talley , Christina Trotter

Generative AI (GenAI) holds significant promise for automating everyday image editing tasks, especially following the recent release of GPT-4o on March 25, 2025. However, what subjects do people most often want edited? What kinds of editing…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Mohammad Reza Taesiri , Brandon Collins , Logan Bolton , Viet Dac Lai , Franck Dernoncourt , Trung Bui , Anh Totti Nguyen

The recent swift development of LLMs like GPT-4, Gemini, and GPT-3.5 offers a transformative opportunity in medicine and healthcare, especially in digital diagnostics. This study evaluates each model diagnostic abilities by interpreting a…

Computation and Language · Computer Science 2024-05-14 Gaurav Kumar Gupta , Aditi Singh , Sijo Valayakkad Manikandan , Abul Ehtesham

Recent advances in large language models (LLMs) have enabled general-purpose systems to perform increasingly complex domain-specific reasoning without extensive fine-tuning. In the medical domain, decision-making often requires integrating…

Computation and Language · Computer Science 2025-08-14 Shansong Wang , Mingzhe Hu , Qiang Li , Mojtaba Safari , Xiaofeng Yang

The rapid development of language-based artificial intelligence (AI) offers new possibilities for psychotherapy and assistive systems, particularly benefitting autistic individuals who often respond well to technology. Parents of autistic…

The rapid evolution of artificial intelligence (AI), especially in the domain of Large Language Models (LLMs) and generative AI, has opened new avenues for application across various fields, yet its role in business education remains…

Computation and Language · Computer Science 2024-01-09 Vahid Ashrafimoghari , Necdet Gürkan , Jordan W. Suchow

We investigate the multilingual and multimodal performance of a large language model-based artificial intelligence (AI) system, GPT-4o, using a diverse set of physics concept inventories spanning multiple languages and subject categories.…

Physics Education · Physics 2025-07-14 Gerd Kortemeyer , Marina Babayeva , Giulia Polverini , Ralf Widenhorn , Bor Gregorcic

This paper presents the first systematic measurement of educational alignment in Large Language Models. Using a Delphi-validated instrument comprising 48 items across eight educational-theoretical dimensions, the study reveals that GPT-5.1…

Computers and Society · Computer Science 2026-03-24 Daniel Autenrieth

Artificial intelligence (AI) systems powered by large language models have become increasingly prevalent in modern society, enabling a wide range of applications through natural language interaction. As AI agents proliferate in our daily…

Machine Learning · Computer Science 2025-03-24 J. M. Diederik Kruijssen , Nicholas Emmons

The increasing demand for programming language education and growing class sizes require immediate and personalized feedback. However, traditional code review methods have limitations in providing this level of feedback. As the capabilities…

Software Engineering · Computer Science 2025-06-23 Lee Dong-Kyu

Scientific discovery is an inherently creative and uncertain process, requiring reasoning beyond the recall of known knowledge. While many benchmarks have been proposed to evaluate large language model (LLM) performance on deep research…

Artificial Intelligence · Computer Science 2026-05-29 A. J. Lew , Y. Cao , M. J. Buehler

This study examines the feasibility and potential advantages of using large language models, in particular GPT-4o, to perform partial credit grading of large numbers of student written responses to introductory level physics problems.…

Physics Education · Physics 2025-08-21 Zhongzhou Chen , Tong Wan

Generative artificial intelligences, particularly large language models (LLMs), play an increasingly prominent role in human decision-making contexts, necessitating transparency about their capabilities. While prior studies have shown…

Computation and Language · Computer Science 2026-01-30 Lydia Uhler , Verena Jordan , Jürgen Buder , Markus Huff , Frank Papenmeier