中文
相关论文

相关论文: IMLJD: A Computational Dataset for Indian Matrimon…

200 篇论文

The proliferation of Large Language Models (LLMs) has led to an influx of AI-generated content (AIGC) on the internet, transforming the corpus of Information Retrieval (IR) systems from solely human-written to a coexistence with…

信息检索 · 计算机科学 2024-07-03 Sunhao Dai , Weihao Liu , Yuqi Zhou , Liang Pang , Rongju Ruan , Gang Wang , Zhenhua Dong , Jun Xu , Ji-Rong Wen

Question-answering systems have revolutionized information retrieval, but linguistic and cultural boundaries limit their widespread accessibility. This research endeavors to bridge the gap of the absence of efficient QnA datasets in…

计算与语言 · 计算机科学 2024-04-23 Ruturaj Ghatage , Aditya Kulkarni , Rajlaxmi Patil , Sharvi Endait , Raviraj Joshi

We present a collection of open, machine-readable document datasets covering parliamentary proceedings, legal judgments, government publications, news, and tourism statistics from Sri Lanka. The collection currently comprises of 269,194…

计算与语言 · 计算机科学 2026-05-18 Nuwan I. Senaratna

This paper explores the problem of commonsense level vision-knowledge conflict in Multimodal Large Language Models (MLLMs), where visual information contradicts model's internal commonsense knowledge. To study this issue, we introduce an…

计算与语言 · 计算机科学 2025-06-03 Xiaoyuan Liu , Wenxuan Wang , Youliang Yuan , Jen-tse Huang , Qiuzhi Liu , Pinjia He , Zhaopeng Tu

We introduce MMCRICBENCH-3K, a benchmark for Visual Question Answering (VQA) on cricket scorecards, designed to evaluate large vision-language models (LVLMs) on complex numerical and cross-lingual reasoning over semi-structured tabular…

计算机视觉与模式识别 · 计算机科学 2025-08-27 Somraj Gautam , Abhirama Subramanyam Penamakuri , Abhishek Bhandari , Gaurav Harit

The Invertible Bloom Lookup Tables (IBLT) is a data structure which supports insertion, deletion, retrieval and listing operations of the key-value pair. The IBLT can be used to realize efficient set reconciliation for database…

信息论 · 计算机科学 2015-06-12 Daichi Yugawa , Tadashi Wadayama

LLM-based judges have emerged as a scalable alternative to human evaluation and are increasingly used to assess, compare, and improve models. However, the reliability of LLM-based judges themselves is rarely scrutinized. As LLMs become more…

The aim of this work is to develop methods for studying the determinants of marriage incidence using marriage histories collected under two different types of retrospective cross-sectional study designs. These designs are: sampling of ever…

统计方法学 · 统计学 2023-01-25 Sangita Kulathinal , Minna Säävälä , Kari Auranen , Olli Saarela

This paper introduces JurisTCU, a Brazilian Portuguese dataset for legal information retrieval (LIR). The dataset is freely available and consists of 16,045 jurisprudential documents from the Brazilian Federal Court of Accounts, along with…

Boxing and MMA have a longstanding issue with judging, as evidenced by frequent controversial decisions. Like boxing, MMA bouts are scored following the 10-Point Must System, by which judges score each round individually. In the present…

物理与社会 · 物理学 2024-01-09 Vincent Berthet

Offering a promising solution to the scalability challenges associated with human evaluation, the LLM-as-a-judge paradigm is rapidly gaining traction as an approach to evaluating large language models (LLMs). However, there are still many…

Recent NLP advances focus primarily on standardized languages, leaving most low-resource dialects under-served especially in Indian scenarios. In India, the issue is particularly important: despite Hindi being the third most spoken language…

The evaluation of Large Language Models (LLMs) on mathematical reasoning has largely focused on elementary problems, competition-style questions, or formal theorem proving, leaving graduate-level and computational mathematics relatively…

计算与语言 · 计算机科学 2026-03-05 Bianca Raimondi , Francesco Pivi , Davide Evangelista , Maurizio Gabbrielli

In clinical practice, physicians refrain from making decisions when patient information is insufficient. This behavior, known as abstention, is a critical safety mechanism preventing potentially harmful misdiagnoses. Recent investigations…

The opioid crisis represents a significant moment in public health that reveals systemic shortcomings across regulatory systems, healthcare practices, corporate governance, and public policy. Analyzing how these interconnected systems…

Pairwise evaluation of Large Language Models (LLMs) is a common paradigm, but it is prone to preference bias, where judges systematically favor certain outputs, such as their own. This bias leads to inconsistent and skewed rankings across…

人工智能 · 计算机科学 2025-11-18 Yang Zhang , Cunxiang Wang , Lindong Wu , Wenbo Yu , Yidong Wang , Guangsheng Bao , Jie Tang

We introduce the Korean Canonical Legal Benchmark (KCL), a benchmark designed to assess language models' legal reasoning capabilities independently of domain-specific knowledge. KCL provides question-level supporting precedents, enabling a…

计算与语言 · 计算机科学 2026-01-06 Hongseok Oh , Wonseok Hwang , Kyoung-Woon On

LLMs are bound to transform healthcare with advanced decision support and flexible chat assistants. However, LLMs are prone to generate inaccurate medical content. To ground LLMs in high-quality medical knowledge, LLMs have been equipped…

The ICLR conference is unique among the top machine learning conferences in that all submitted papers are openly available. Here we present the ICLR dataset consisting of abstracts of all 24 thousand ICLR submissions from 2017-2024 with…

计算与语言 · 计算机科学 2024-06-06 Rita González-Márquez , Dmitry Kobak

Contemporary vision-language models (VLMs) perform well on existing multimodal reasoning benchmarks (78-85\% accuracy on MMMU, MathVista). Yet, these results fail to sufficiently distinguish true scientific reasoning articulation…

计算与语言 · 计算机科学 2025-11-13 Arka Mukherjee , Shreya Ghosh