中文
相关论文

相关论文: The Touch\'e23-ValueEval Dataset for Identifying H…

200 篇论文

Human evaluation is the gold standard for evaluating text generation models. However, it is expensive. In order to fit budgetary constraints, a random subset of the test data is often chosen in practice for human evaluation. However,…

计算与语言 · 计算机科学 2025-06-03 Vilém Zouhar , Peng Cui , Mrinmaya Sachan

Opinion summarization sets itself apart from other types of summarization tasks due to its distinctive focus on aspects and sentiments. Although certain automated evaluation methods like ROUGE have gained popularity, we have found them to…

计算与语言 · 计算机科学 2023-11-14 Yuchen Shen , Xiaojun Wan

Argument mining aims to detect all possible argumentative components and identify their relationships automatically. As a thriving task in natural language processing, there has been a large amount of corpus for academic study and…

信息检索 · 计算机科学 2024-05-31 Huadai Liu , Wenqiang Xu , Xuan Lin , Jingjing Huo , Hong Chen , Zhou Zhao

AI assistants can impart value judgments that shape people's decisions and worldviews, yet little is known empirically about what values these systems rely on in practice. To address this, we develop a bottom-up, privacy-preserving method…

Advanced service robots require superior tactile intelligence to guarantee human-contact safety and to provide essential supplements to visual and auditory information for human-robot interaction, especially when a robot is in physical…

机器人学 · 计算机科学 2021-08-12 Peng Wang , Jixiao Liu , Funing Hou , Dicai Chen , Zihou Xia , Shijie Guo

We present a data-driven, end-to-end approach to transaction-based dialog systems that performs at near-human levels in terms of verbal response quality and factual grounding accuracy. We show that two essential components of the system…

计算与语言 · 计算机科学 2020-12-29 Bill Byrne , Karthik Krishnamoorthi , Saravanan Ganesh , Mihir Sanjay Kale

Scientific document understanding is challenging as the data is highly domain specific and diverse. However, datasets for tasks with scientific text require expensive manual annotation and tend to be small and limited to only one or a few…

计算与语言 · 计算机科学 2021-05-26 Dustin Wright , Isabelle Augenstein

Interpreting individual neurons or directions in activation space is an important topic in mechanistic interpretability. Numerous automated interpretability methods have been proposed to generate such explanations, but it remains unclear…

计算机视觉与模式识别 · 计算机科学 2025-12-04 Tuomas Oikarinen , Ge Yan , Akshay Kulkarni , Tsui-Wei Weng

The importance of benchmarks for assessing the values of language models has been pronounced due to the growing need of more authentic, human-aligned responses. However, existing benchmarks rely on human or machine annotations that are…

计算与语言 · 计算机科学 2025-06-12 Jongwook Han , Dongmin Choi , Woojung Song , Eun-Ju Lee , Yohan Jo

Software development relies heavily on text-based communication, making sentiment analysis a valuable tool for understanding team dynamics and supporting trustworthy AI-driven analytics in requirements engineering. However, existing…

软件工程 · 计算机科学 2025-07-11 Martin Obaidi , Marc Herrmann , Jil Klünder , Kurt Schneider

At the foundation of scientific evaluation is the labor-intensive process of peer review. This critical task requires participants to consume vast amounts of highly technical text. Prior work has annotated different aspects of review…

Evidence data for automated fact-checking (AFC) can be in multiple modalities such as text, tables, images, audio, or video. While there is increasing interest in using images for AFC, previous works mostly focus on detecting manipulated or…

计算与语言 · 计算机科学 2023-01-30 Mubashara Akhtar , Oana Cocarascu , Elena Simperl

In abstract argumentation, multiple argumentation semantics have been proposed that allow to select sets of jointly acceptable arguments from a given argumentation framework, i.e. based only on the attack relation between arguments. The…

人工智能 · 计算机科学 2019-02-28 Marcos Cramer , Mathieu Guillaume

Existing argumentation datasets have succeeded in allowing researchers to develop computational methods for analyzing the content, structure and linguistic features of argumentative text. They have been much less successful in fostering…

计算与语言 · 计算机科学 2019-09-26 Esin Durmus , Claire Cardie

The assessment of argument quality depends on well-established logical, rhetorical, and dialectical properties that are unavoidably subjective: multiple valid assessments may exist, there is no unequivocal ground truth. This aligns with…

计算与语言 · 计算机科学 2025-02-21 Julia Romberg , Maximilian Maurer , Henning Wachsmuth , Gabriella Lapesa

Summarization evaluation remains an open research problem: current metrics such as ROUGE are known to be limited and to correlate poorly with human judgments. To alleviate this issue, recent work has proposed evaluation metrics which rely…

Prior work has commonly defined argument retrieval from heterogeneous document collections as a sentence-level classification task. Consequently, argument retrieval suffers both from low recall and from sentence segmentation errors making…

计算与语言 · 计算机科学 2019-11-22 Dietrich Trautmann , Johannes Daxenberger , Christian Stab , Hinrich Schütze , Iryna Gurevych

Trajectory datasets of road users have become more important in the last years for safety validation of highly automated vehicles. Several naturalistic trajectory datasets with each more than 10.000 tracks were released and others will…

计算机视觉与模式识别 · 计算机科学 2022-04-12 Christoph Glasmacher , Robert Krajewski , Lutz Eckstein

Language comprehension and commonsense knowledge validation by machines are challenging tasks that are still under researched and evaluated for Arabic text. In this paper, we present a benchmark Arabic dataset for commonsense explanation.…

计算与语言 · 计算机科学 2020-12-21 Saja AL-Tawalbeh , Mohammad AL-Smadi

Touch is an important sensing modality for humans, but it has not yet been incorporated into a multimodal generative language model. This is partially due to the difficulty of obtaining natural language labels for tactile data and the…

计算机视觉与模式识别 · 计算机科学 2024-02-21 Letian Fu , Gaurav Datta , Huang Huang , William Chung-Ho Panitch , Jaimyn Drake , Joseph Ortiz , Mustafa Mukadam , Mike Lambeta , Roberto Calandra , Ken Goldberg