English
Related papers

Related papers: Automatically Generating a Large, Culture-Specific…

200 papers

To identify cluster of societies and cultures is not easy in subject to the availability of data. In this study, we propose a novel method to cluster Chinese regional cultures. Using geotagged online-gaming data of Chinese internet users…

Computers and Society · Computer Science 2013-10-03 Xianwen Wang , Wenli Mao , Chen Liu

A broad range of research areas including Internet measurement, privacy, and network security rely on lists of target domains to be analysed; researchers make use of target lists for reasons of necessity or efficiency. The popular Alexa…

Networking and Internet Architecture · Computer Science 2018-09-25 Quirin Scheitle , Oliver Hohlfeld , Julien Gamba , Jonas Jelten , Torsten Zimmermann , Stephen D. Strowes , Narseo Vallina-Rodriguez

We develop a means to detect ongoing per-country anomalies in the daily usage metrics of the Tor anonymous communication network, and demonstrate the applicability of this technique to identifying likely periods of internet censorship and…

Computers and Society · Computer Science 2018-04-13 Joss Wright , Alexander Darer , Oliver Farnan

Over the past decade, Internet centralization and its implications for both people and the resilience of the Internet has become a topic of active debate. While the networking community informally agrees on the definition of centralization,…

Networking and Internet Architecture · Computer Science 2025-04-29 Gautam Akiwate , Kimberly Ruth , Rumaisa Habib , Zakir Durumeric

Existing humor datasets and evaluations predominantly focus on English, leaving limited resources for culturally nuanced humor in non-English languages like Chinese. To address this gap, we construct Chumor, the first Chinese humor…

Computation and Language · Computer Science 2024-12-24 Ruiqi He , Yushu He , Longju Bai , Jiarui Liu , Zhenjie Sun , Zenghao Tang , He Wang , Hanchen Xia , Rada Mihalcea , Naihao Deng

The proliferation of hate speech has inflicted significant societal harm, with its intensity and directionality closely tied to specific targets and arguments. In recent years, numerous machine learning-based methods have been developed to…

Computation and Language · Computer Science 2025-07-16 Zewen Bai , Liang Yang , Shengdi Yin , Yuanyuan Sun , Hongfei Lin

As large language models (LLMs) are increasingly applied to various NLP tasks, their inherent biases are gradually disclosed. Therefore, measuring biases in LLMs is crucial to mitigate its ethical risks. However, most existing bias…

Computation and Language · Computer Science 2025-08-08 Tian Lan , Xiangdong Su , Xu Liu , Ruirui Wang , Ke Chang , Jiang Li , Guanglai Gao

With the increasing pursuit of objective reports, automatically understanding media bias has drawn more attention in recent research. However, most of the previous work examines media bias from Western ideology, such as the left and right…

Computers and Society · Computer Science 2023-11-21 Luyang Lin , Jing Li , Kam-Fai Wong

The advent of natural language understanding (NLU) benchmarks for English, such as GLUE and SuperGLUE allows new NLU models to be evaluated across a diverse set of tasks. These comprehensive benchmarks have facilitated a broad range of…

Current benchmarks for evaluating large language models (LLMs) in social media moderation completely overlook a serious threat: covert advertisements, which disguise themselves as regular posts to deceive and mislead consumers into making…

Machine Learning · Computer Science 2026-04-23 Jingyi Zheng , Tianyi Hu , Yule Liu , Zhen Sun , Zongmin Zhang , Zifan Peng , Wenhan Dong , Xinlei He

The World Wide Web is not only one of the most important platforms of communication and information at present, but also an area of growing interest for scientific research. This motivates a lot of work and projects that require large…

Computer Vision and Pattern Recognition · Computer Science 2021-05-18 Christian Mejia-Escobar , Miguel Cazorla , Ester Martinez-Martin

We examine the behavioral impact of a user location disclosure policy on Sina Weibo, China's largest microblogging platform, using a unique high-frequency dataset of uncensored engagement, including tens of thousands of comments and…

General Economics · Economics 2025-09-03 Leo Yang Yang , Yiqing Xu

Scientific literature serves as a high-quality corpus, supporting a lot of Natural Language Processing (NLP) research. However, existing datasets are centered around the English language, which restricts the development of Chinese…

Computation and Language · Computer Science 2022-09-13 Yudong Li , Yuqing Zhang , Zhe Zhao , Linlin Shen , Weijie Liu , Weiquan Mao , Hui Zhang

Large language models sometimes produce false or misleading responses. Two approaches to this problem are honesty elicitation -- modifying prompts or weights so that the model answers truthfully -- and lie detection -- classifying whether a…

Machine Learning · Computer Science 2026-03-11 Helena Casademunt , Bartosz Cywiński , Khoi Tran , Arya Jakkli , Samuel Marks , Neel Nanda

The proliferation of global censorship has led to the development of a plethora of measurement platforms to monitor and expose it. Censorship of the domain name system (DNS) is a key mechanism used across different countries. It is…

Large Language Models (LLMs) are increasingly relying on web crawling to stay up to date and accurately answer user queries. These crawlers are expected to honor robots.txt files, which govern automated access. In this study, for the first…

Computers and Society · Computer Science 2026-02-24 Nicolas Steinacker-Olsztyn , Devashish Gosain , Ha Dao

With the development of information technology, there is an explosive growth in the number of online comment concerning news, blogs and so on. The massive comments are overloaded, and often contain some misleading and unwelcome information.…

Computation and Language · Computer Science 2018-08-23 Deli Chen , Shuming Ma , Pengcheng Yang , Xu Sun

Large language models (LLMs) excel in high-resource languages but struggle with low-resource languages (LRLs), particularly those spoken by minority communities in China, such as Tibetan, Uyghur, Kazakh, and Mongolian. To systematically…

Computation and Language · Computer Science 2025-06-03 Chen Zhang , Mingxu Tao , Zhiyuan Liao , Yansong Feng

This paper presents a detailed study of the Internet censorship in India. We consolidated a list of potentially blocked websites from various public sources to assess censorship mechanisms used by nine major ISPs. To begin with, we…

Computers and Society · Computer Science 2018-08-07 Tarun Kumar Yadav , Akshat Sinha , Devashish Gosain , Piyush Sharma , Sambuddho Chakravarty

Taxonomies play an important role in machine intelligence. However, most well-known taxonomies are in English, and non-English taxonomies, especially Chinese ones, are still very rare. In this paper, we focus on automatic Chinese taxonomy…

Computation and Language · Computer Science 2019-02-28 Jindong Chen , Ao Wang , Jiangjie Chen , Yanghua Xiao , Zhendong Chu , Jingping Liu , Jiaqing Liang , Wei Wang
‹ Prev 1 3 4 5 6 7 10 Next ›