中文
相关论文

相关论文: A Discerning Several Thousand Judgments: GPT-3 Rat…

200 篇论文

Language models learn rare syntactic phenomena, but the extent to which this is attributable to generalization vs. memorization is a major open question. To that end, we iteratively trained transformer language models on systematically…

计算与语言 · 计算机科学 2025-06-26 Kanishka Misra , Kyle Mahowald

It remains debated how well any LM understands natural language or generates reliable metalinguistic judgments. Moreover, relatively little work has demonstrated that LMs can represent and respect subtle relationships between form and…

计算与语言 · 计算机科学 2025-05-15 Nicole Cuneo , Eleanor Graves , Supantho Rakshit , Adele E. Goldberg

Artificial neural networks can generalize productively to novel contexts. Can they also learn exceptions to those productive rules? We explore this question using the case of restrictions on English passivization (e.g., the fact that "The…

计算与语言 · 计算机科学 2023-06-12 Cara Su-Yi Leong , Tal Linzen

Recent advancements in Large Language Models (LLMs) harness linguistic associations in vast natural language data for practical applications. However, their ability to understand the physical world using only language data remains a…

计算与语言 · 计算机科学 2023-05-10 Nigel H. Collier , Fangyu Liu , Ehsan Shareghi

We study semantic construal in grammatical constructions using large language models. First, we project contextual word embeddings into three interpretable semantic spaces, each defined by a different set of psycholinguistic feature norms.…

计算与语言 · 计算机科学 2023-05-31 Gabriella Chronis , Kyle Mahowald , Katrin Erk

Large language models are increasingly capable of generating fluent-appearing text with relatively little task-specific supervision. But can these models accurately explain classification decisions? We consider the task of generating…

计算与语言 · 计算机科学 2022-05-06 Sarah Wiegreffe , Jack Hessel , Swabha Swayamdipta , Mark Riedl , Yejin Choi

This paper investigates the ability of artificial neural networks to judge the grammatical acceptability of a sentence, with the goal of testing their linguistic competence. We introduce the Corpus of Linguistic Acceptability (CoLA), a set…

计算与语言 · 计算机科学 2019-10-03 Alex Warstadt , Amanpreet Singh , Samuel R. Bowman

Levin et al. (2019) show experimentally that the interpretations of novel English noun compounds (e.g., stew skillet), while not fully compositional, are highly predictable based on whether the modifier and head refer to artifacts or…

计算与语言 · 计算机科学 2022-10-19 Siyan Li , Riley Carlson , Christopher Potts

Recent work has demonstrated substantial gains on many NLP tasks and benchmarks by pre-training on a large corpus of text followed by fine-tuning on a specific task. While typically task-agnostic in architecture, this method still requires…

Humans perceive discrete events such as "restaurant visits" and "train rides" in their continuous experience. One important prerequisite for studying human event perception is the ability of researchers to quantify when one event ends and…

计算与语言 · 计算机科学 2023-01-26 Sebastian Michelmann , Manoj Kumar , Kenneth A. Norman , Mariya Toneva

While large language models (LLMs), such as GPT-3, appear to be robust and general, their reasoning ability is not at a level to compete with the best models trained for specific natural language reasoning problems. In this study, we…

计算与语言 · 计算机科学 2023-07-18 Zhun Yang , Adam Ishay , Joohyung Lee

Summary assessment involves evaluating how well a generated summary reflects the key ideas and meaning of the source text, requiring a deep understanding of the content. Large Language Models (LLMs) have been used to automate this process,…

计算与语言 · 计算机科学 2025-12-23 Zahra Sadeghi , Evangelos Milios , Frank Rudzicz

Large language models, particularly GPT-3, are able to produce high quality summaries of general domain news articles in few- and zero-shot settings. However, it is unclear if such models are similarly capable in more specialized,…

计算与语言 · 计算机科学 2023-05-12 Chantal Shaib , Millicent L. Li , Sebastian Joseph , Iain J. Marshall , Junyi Jessy Li , Byron C. Wallace

Large Language Models (LLMs) evaluation is a patchy and inconsistent landscape, and it is becoming clear that the quality of automatic evaluation metrics is not keeping up with the pace of development of generative models. We aim to improve…

计算与语言 · 计算机科学 2023-10-24 Andrea Sottana , Bin Liang , Kai Zou , Zheng Yuan

The recent success of prompting large language models like GPT-3 has led to a paradigm shift in NLP research. In this paper, we study its impact on text summarization, focusing on the classic benchmark domain of news summarization. First,…

计算与语言 · 计算机科学 2023-05-25 Tanya Goyal , Junyi Jessy Li , Greg Durrett

AI large language models have (co-)produced amazing written works from newspaper articles to novels and poetry. These works meet the standards of the standard definition of creativity: being original and useful, and sometimes even the…

人工智能 · 计算机科学 2022-06-22 Claire Stevenson , Iris Smal , Matthijs Baas , Raoul Grasman , Han van der Maas

Large language models (LLMs) have become mainstream technology with their versatile use cases and impressive performance. Despite the countless out-of-the-box applications, LLMs are still not reliable. A lot of work is being done to improve…

计算与语言 · 计算机科学 2023-06-13 Aisha Khatun , Daniel G. Brown

Recent work on evaluating grammatical knowledge in pretrained sentence encoders gives a fine-grained view of a small number of phenomena. We introduce a new analysis dataset that also has broad coverage of linguistic phenomena. We annotate…

计算与语言 · 计算机科学 2020-05-25 Alex Warstadt , Samuel R. Bowman

Open AI's language model, GPT-3, has shown great potential for many NLP tasks, with applications in many different domains. In this work we carry out a first study on GPT-3's capability to communicate musical decisions through textual…

计算与语言 · 计算机科学 2022-06-17 Stephen James Krol , Maria Teresa Llano , Jon McCormack

Human evaluations are typically considered the gold standard in natural language generation, but as models' fluency improves, how well can evaluators detect and judge machine-generated text? We run a study assessing non-experts' ability to…

计算与语言 · 计算机科学 2021-07-08 Elizabeth Clark , Tal August , Sofia Serrano , Nikita Haduong , Suchin Gururangan , Noah A. Smith
‹ 上一页 1 2 3 10 下一页 ›