中文
相关论文

相关论文: Gender classification by means of online uppercase…

200 篇论文

Modern models for common NLP tasks often employ machine learning techniques and train on journalistic, social media, or other culturally-derived text. These have recently been scrutinized for racial and gender biases, rooting from inherent…

计算与语言 · 计算机科学 2026-01-27 Scott Friedman , Sonja Schmer-Galunder , Anthony Chen , Jeffrey Rye

This paper presents an experiment of automatically scoring handwritten descriptive answers in the trial tests for the new Japanese university entrance examination, which were made for about 120,000 examinees in 2017 and 2018. There are…

机器学习 · 计算机科学 2023-12-04 Hung Tuan Nguyen , Cuong Tuan Nguyen , Haruki Oka , Tsunenori Ishioka , Masaki Nakagawa

Algorithmic classifications of research publications can be used to study many different aspects of the science system, such as the organization of science into fields, the growth of fields, interdisciplinarity, and emerging topics. How to…

数字图书馆 · 计算机科学 2021-01-01 Peter Sjögårde , Per Ahlgren , Ludo Waltman

Air-writing refers to virtually writing linguistic characters through hand gestures in three-dimensional space with six degrees of freedom. This paper proposes a generic video camera-aided convolutional neural network (CNN) based…

计算机视觉与模式识别 · 计算机科学 2025-02-19 Prasun Roy , Subhankar Ghosh , Umapada Pal

This paper investigates the task of writer retrieval, which identifies documents authored by the same individual within a dataset based on handwriting similarities. While existing datasets and methodologies primarily focus on page level…

计算机视觉与模式识别 · 计算机科学 2025-06-10 Marco Peer , Robert Sablatnig , Florian Kleber

The authorship attribution is a problem of considerable practical and technical interest. Several methods have been designed to infer the authorship of disputed documents in multiple contexts. While traditional statistical methods based…

计算与语言 · 计算机科学 2018-03-28 Jeaneth Machicao , Edilson A. Corrêa , Gisele H. B. Miranda , Diego R. Amancio , Odemir M. Bruno

In this paper, we investigate the impact of objects on gender bias in image captioning systems. Our results show that only gender-specific objects have a strong gender bias (e.g., women-lipstick). In addition, we propose a visual…

计算与语言 · 计算机科学 2023-11-21 Ahmed Sabir , Lluís Padró

Online signature verification technologies, such as those available in banks and post offices, rely on dedicated digital devices such as tablets or smart pens to capture, analyze and verify signatures. In this paper, we suggest a novel…

密码学与安全 · 计算机科学 2016-12-20 Ben Nassi , Alona Levy , Yuval Elovici , Erez Shmueli

Social scientists often classify text documents to use the resulting labels as an outcome or a predictor in empirical research. Automated text classification has become a standard tool, since it requires less human coding. However, scholars…

计算与语言 · 计算机科学 2025-05-14 Mitchell Bosley , Saki Kuzushima , Ted Enamorado , Yuki Shiraito

Potential harms of Large Language Models such as mass misinformation and plagiarism can be partially mitigated if there exists a reliable way to detect machine generated text. In this paper, we propose a new watermarking method to detect…

计算与语言 · 计算机科学 2023-12-12 Kaan Efe Keleş , Ömer Kaan Gürbüz , Mucahid Kutlu

Hierarchical text classification aims to categorize each document into a set of classes in a label taxonomy, which is a fundamental web text mining task with broad applications such as web content analysis and semantic indexing. Most…

计算与语言 · 计算机科学 2025-02-06 Yunyi Zhang , Ruozhen Yang , Xueqiang Xu , Rui Li , Jinfeng Xiao , Jiaming Shen , Jiawei Han

Authorship attribution is a natural language processing task that has been widely studied, often by considering small order statistics. In this paper, we explore a complex network approach to assign the authorship of texts based on their…

计算与语言 · 计算机科学 2017-08-08 Vanessa Q. Marinho , Henrique F. de Arruda , Thales S. Lima , Luciano F. Costa , Diego R. Amancio

Statistical watermarking is a common approach for verifying whether text was written by a language model. Most existing schemes assume autoregressive generation, where tokens are produced left to right and contextual hashing is well…

计算与语言 · 计算机科学 2026-05-08 Mohd Ruhul Ameen , Akif Islam , Nadim Mahmud , Md. Ekramul Hamid

Sexism in online media comments is a pervasive challenge that often manifests subtly, complicating moderation efforts as interpretations of what constitutes sexism can vary among individuals. We study monolingual and multilingual…

计算与语言 · 计算机科学 2024-10-03 Florian Bremm , Patrick Gustav Blaneck , Tobias Bornheim , Niklas Grieger , Stephan Bialonski

The number of senses of a given word, or polysemy, is a very subjective notion, which varies widely across annotators and resources. We propose a novel method to estimate polysemy, based on simple geometry in the contextual embedding space.…

计算与语言 · 计算机科学 2023-05-03 Christos Xypolopoulos , Antoine J. -P. Tixier , Michalis Vazirgiannis

Stroke classification remains challenging due to variations in writing style, ambiguous content, and dynamic writing positions. The core challenge in stroke classification is modeling the semantic relationships between strokes. Our…

计算机视觉与模式识别 · 计算机科学 2025-12-09 Yiheng Huang , Shuang She , Zewei Wei , Jianmin Lin , Ming Yang , Wenyin Liu

Offline handwriting recognition (HWR) has improved significantly with the advent of deep learning architectures in recent years. Nevertheless, it remains a challenging problem and practical applications often rely on post-processing…

计算机视觉与模式识别 · 计算机科学 2023-09-20 Andrey Totev , Tomas Ward

Many text corpora exhibit socially problematic biases, which can be propagated or amplified in the models trained on such data. For example, doctor cooccurs more frequently with male pronouns than female pronouns. In this study we (i)…

计算与语言 · 计算机科学 2019-04-08 Shikha Bordia , Samuel R. Bowman

In this paper, we use statistical texture features for handwritten and printed text classification. We primarily aim for word level classification in south Indian scripts. Words are first extracted from the scanned document. For each…

计算机视觉与模式识别 · 计算机科学 2013-04-11 Mallikarjun Hangarge , K. C. Santosh , Srikanth Doddamani , Rajmohan Pardeshi

The rapid advancement of large language models (LLMs) has made it increasingly difficult to distinguish between text written by humans and machines. Addressing this, we propose a novel method for generating watermarks that strategically…

计算与语言 · 计算机科学 2024-05-15 Georg Niess , Roman Kern