English
Related papers

Related papers: IMLJD: A Computational Dataset for Indian Matrimon…

200 papers

Existing research on news summarization primarily focuses on single-language single-document (SLSD), single-language multi-document (SLMD) or cross-language single-document (CLSD). However, in real-world scenarios, news about a…

Computation and Language · Computer Science 2024-10-15 Shengxiang Gao , Fang nan , Yongbing Zhang , Yuxin Huang , Kaiwen Tan , Zhengtao Yu

Advances in machine learning are closely tied to the creation of datasets. While data documentation is widely recognized as essential to the reliability, reproducibility, and transparency of ML, we lack a systematic empirical understanding…

Machine Learning · Computer Science 2024-01-26 Xinyu Yang , Weixin Liang , James Zou

In politically sensitive scenarios like wars, social media serves as a platform for polarized discourse and expressions of strong ideological stances. While prior studies have explored ideological stance detection in general contexts,…

Computation and Language · Computer Science 2026-04-16 Hasin Jawad Ali , Ajwad Abrar , S. M. Hozaifa Hossain , M. Firoz Mridha

Sports have long attracted broad attention as they push the limits of human physical and cognitive capabilities. Amid growing interest in spatial intelligence for vision-language models (VLMs), sports provide a natural testbed for…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Yuchen Yang , Yuqing Shao , Duxiu Huang , Linfeng Dong , Yifei Liu , Suixin Tang , Xiang Zhou , Yuanyuan Gao , Wei Wang , Yue Zhou , Xue Yang , Yanfeng Wang , Xiao Sun , Zhihang Zhong

Missing data imputation in tabular datasets remains a pivotal challenge in data science and machine learning, particularly within socioeconomic research. However, real-world socioeconomic datasets are typically subject to strict data…

Machine Learning · Computer Science 2025-06-11 Siyi Sun , David Antony Selby , Yunchuan Huang , Sebastian Vollmer , Seth Flaxman , Anisoara Calinescu

This paper presents the first dataset for Japanese Legal Judgment Prediction (LJP), the Japanese Tort-case Dataset (JTD), which features two tasks: tort prediction and its rationale extraction. The rationale extraction task identifies the…

Computation and Language · Computer Science 2025-12-02 Hiroaki Yamada , Takenobu Tokunaga , Ryutaro Ohara , Akira Tokutsu , Keisuke Takeshita , Mihoko Sumida

The Scholarly Hybrid Question Answering over Linked Data (QALD) Challenge at the International Semantic Web Conference (ISWC) 2024 focuses on Question Answering (QA) over diverse scholarly sources: DBLP, SemOpenAlex, and Wikipedia-based…

Information Retrieval · Computer Science 2024-12-02 Fomubad Borista Fondi , Azanzi Jiomekong Fidel , Gaoussou Camara

Existing cultural commonsense benchmarks treat nations as monolithic, assuming uniform practices within national boundaries. But does cultural commonsense hold uniformly within a nation, or does it vary at the sub-national level? We…

Computation and Language · Computer Science 2026-04-16 Sangmitra Madhusudan , Trush Shashank More , Steph Buongiorno , Renata Dividino , Jad Kabbara , Ali Emami

Evaluating AI-generated research ideas typically relies on LLM judges or human panels -- both subjective and disconnected from actual research impact. We introduce HindSight, a time-split evaluation framework that measures idea quality by…

Computation and Language · Computer Science 2026-03-18 Bo Jiang

The availability of well-curated datasets has driven the success of Machine Learning (ML) models. Despite the increased access to earth observation data for agriculture, there is a scarcity of curated, labelled datasets, which limits the…

Computer Vision and Pattern Recognition · Computer Science 2022-09-27 Depanshu Sani , Sandeep Mahato , Parichya Sirohi , Saket Anand , Gaurav Arora , Charu Chandra Devshali , Thiagarajan Jayaraman , Harsh Kumar Agarwal

As multi-modal large language models (MLLMs) frequently exhibit errors when solving scientific problems, evaluating the validity of their reasoning processes is critical for ensuring reliability and uncovering fine-grained model weaknesses.…

Artificial Intelligence · Computer Science 2025-03-11 Jiaxin Ai , Pengfei Zhou , Zhaopan Xu , Ming Li , Fanrui Zhang , Zizhen Li , Jianwen Sun , Yukang Feng , Baojin Huang , Zhongyuan Wang , Kaipeng Zhang

"Citizen queries" are questions asked by an individual about government policies, guidance, and services that are relevant to their circumstances, encompassing a range of topics including benefits, taxes, immigration, employment, public…

Computers and Society · Computer Science 2026-02-05 Neil Majithia , Rajat Shinde , Zo Chapman , Prajun Trital , Jordan Decker , Manil Maskey , Elena Simperl , Nigel Shadbolt

Generative AI models, such as the GPT and Llama series, have significant potential to assist laypeople in answering legal questions. However, little prior work focuses on the data sourcing, inference, and evaluation of these models in the…

Computation and Language · Computer Science 2024-09-13 Jonathan Li , Rohan Bhambhoria , Samuel Dahan , Xiaodan Zhu

Existing studies on fairness are largely Western-focused, making them inadequate for culturally diverse countries such as India. To address this gap, we introduce INDIC-BIAS, a comprehensive India-centric benchmark designed to evaluate…

Computation and Language · Computer Science 2025-07-01 Janki Atul Nawale , Mohammed Safi Ur Rahman Khan , Janani D , Mansi Gupta , Danish Pruthi , Mitesh M. Khapra

Citizen reporting platforms help the public and authorities stay informed about sexual harassment incidents. However, the high volume of data shared on these platforms makes reviewing each individual case challenging. Therefore, a…

Computation and Language · Computer Science 2026-04-20 Garima Chhikara , Anurag Sharma , V. Gurucharan , Kripabandhu Ghosh , Abhijnan Chakraborty

Reliable analysis of migration is critically dependent on the quality and consistency of the underlying data. Indian migration data, primarily derived from decennial census records, are affected by systematic gaps arising from uneven…

Applications · Statistics 2026-04-15 Nivedita Batra , Chiranjoy Chattopadhyay , Mayurakshi Chaudhuri

In the absence of neither an effective treatment or vaccine and with an incomplete understanding of the epidemiological cycle, Govt. has implemented a nationwide lockdown to reduce COVID-19 transmission in India. To study the effect of…

Populations and Evolution · Quantitative Biology 2020-08-26 Tridip Sardar , Sk Shahid Nadim , Sourav Rana , Joydev Chattopadhyay

As large language models (LLMs) are increasingly used in legal applications, current evaluation benchmarks tend to focus mainly on factual accuracy while largely neglecting important linguistic quality aspects such as clarity, coherence,…

Computation and Language · Computer Science 2025-11-11 Li yunhan , Wu gengshen

Selecting capable counsel can shape the outcome of litigation, yet evaluating law firm performance remains challenging. Widely used rankings prioritize prestige, size, and revenue rather than empirical litigation outcomes, offering little…

Computers and Society · Computer Science 2025-03-28 Alexandre Mojon , Robert Mahari , Sandro Claudio Lera

Relevance judgment of human assessors is inherently subjective and dynamic when evaluation datasets are created for Information Retrieval (IR) systems. However, a small group of experts' relevance judgment results are usually taken as…

Information Retrieval · Computer Science 2022-08-09 Dengya Zhu , Shastri L Nimmagadda , Kok Wai Wong , Torsten Reiners