English
Related papers

Related papers: Automatically Generating a Large, Culture-Specific…

200 papers

For many real-world applications, the user-generated inputs usually contain various noises due to speech recognition errors caused by linguistic variations1 or typographical errors (typos). Thus, it is crucial to test model performance on…

Computation and Language · Computer Science 2023-05-26 Chenglei Si , Zhengyan Zhang , Yingfa Chen , Xiaozhi Wang , Zhiyuan Liu , Maosong Sun

Knowledge graph is a kind of valuable knowledge base which would benefit lots of AI-related applications. Up to now, lots of large-scale knowledge graphs have been built. However, most of them are non-Chinese and designed for general…

Artificial Intelligence · Computer Science 2018-12-18 Feiliang Ren , Yining Hou , Yan Li , Linfeng Pan , Yi Zhang , Xiaobo Liang , Yongkang Liu , Yu Guo , Rongsheng Zhao , Ruicheng Ming , Huiming Wu

Although Google is blocked in China, Chinese provinces export significantly more to foreign countries that recently searched for them (up to 12 months prior). This attention premium is found mainly at the extensive margin of exports, larger…

General Economics · Economics 2025-04-01 Cui Hu , Ben G. Li

Free web proxies promise anonymity and censorship circumvention at no cost. Several websites publish lists of free proxies organized by country, anonymity level, and performance. These lists index hundreds of thousand of hosts discovered…

Networking and Internet Architecture · Computer Science 2017-11-03 Diego Perino , Matteo Varvello , Claudio Soriente

Current alignment evaluation mostly measures whether models encode dangerous concepts and whether they refuse harmful requests. Both miss the layer where alignment often operates: routing from concept detection to behavioral policy. We…

Machine Learning · Computer Science 2026-05-04 Gregory N. Frank

We present GLM-Dialog, a large-scale language model (LLM) with 10B parameters capable of knowledge-grounded conversation in Chinese using a search engine to access the Internet knowledge. GLM-Dialog offers a series of applicable techniques…

Computation and Language · Computer Science 2023-03-01 Jing Zhang , Xiaokang Zhang , Daniel Zhang-Li , Jifan Yu , Zijun Yao , Zeyao Ma , Yiqi Xu , Haohua Wang , Xiaohan Zhang , Nianyi Lin , Sunrui Lu , Juanzi Li , Jie Tang

Lexical simplification has attracted much attention in many languages, which is the process of replacing complex words in a given sentence with simpler alternatives of equivalent meaning. Although the richness of vocabulary in Chinese makes…

Computation and Language · Computer Science 2020-10-15 Jipeng Qiang , Xinyu Lu , Yun Li , Yunhao Yuan , Yang Shi , Xindong Wu

In early January 2020, after China reported the first cases of the new coronavirus (SARS-CoV-2) in the city of Wuhan, unreliable and not fully accurate information has started spreading faster than the virus itself. Alongside this pandemic,…

Machine Learning · Computer Science 2021-06-08 V. Mazzeo , A. Rapisarda , G. Giuffrida

Large language models (LLMs) are increasingly deployed in cost-sensitive and on-device scenarios, and safety guardrails have advanced mainly in English. However, real-world Chinese malicious queries typically conceal intent via homophones,…

Computation and Language · Computer Science 2026-01-06 Zhenhong Zhou , Shilinlu Yan , Chuanpu Liu , Qiankun Li , Kun Wang , Zhigang Zeng

In the current environment, psychological issues are prevalent and widespread, with social media serving as a key outlet for individuals to share their feelings. This results in the generation of vast quantities of data daily, where…

Computation and Language · Computer Science 2024-06-13 Wei Zhai , Hongzhi Qi , Qing Zhao , Jianqiang Li , Ziqi Wang , Han Wang , Bing Xiang Yang , Guanghui Fu

The absence of an appropriate text classification corpus makes the massive amount of online job information unusable for labor market analysis. This paper presents JCTC, a large job posting corpus for text classification. In JCTC…

Information Retrieval · Computer Science 2017-06-13 Haoyu Xu , Chongyang Gu , Han Zhou , Sengpan Kou , Junjie Zhang

Contemporary language models are increasingly multilingual, but Chinese LLM developers must navigate complex political and business considerations of language diversity. Language policy in China aims at influencing the public discourse and…

Computation and Language · Computer Science 2025-08-12 Andrea W Wen-Yi , Unso Eun Seo Jo , Lu Jia Lin , David Mimno

We describe our experience of implementing a news content organization system at Tencent that discovers events from vast streams of breaking news and evolves news story structures in an online fashion. Our real-world system has distinct…

Information Retrieval · Computer Science 2018-03-02 Bang Liu , Di Niu , Kunfeng Lai , Linglong Kong , Yu Xu

Large language models have recently made tremendous progress in a variety of aspects, e.g., cross-task generalization, instruction following. Comprehensively evaluating the capability of large language models in multiple tasks is of great…

Computation and Language · Computer Science 2023-05-23 Chuang Liu , Renren Jin , Yuqi Ren , Linhao Yu , Tianyu Dong , Xiaohan Peng , Shuting Zhang , Jianxiang Peng , Peiyi Zhang , Qingqing Lyu , Xiaowen Su , Qun Liu , Deyi Xiong

The rapid development of Chinese large language models (LLMs) poses big challenges for efficient LLM evaluation. While current initiatives have introduced new benchmarks or evaluation platforms for assessing Chinese LLMs, many of these…

Computation and Language · Computer Science 2024-03-20 Chuang Liu , Linhao Yu , Jiaxuan Li , Renren Jin , Yufei Huang , Ling Shi , Junhui Zhang , Xinmeng Ji , Tingting Cui , Tao Liu , Jinwang Song , Hongying Zan , Sun Li , Deyi Xiong

As artificial intelligence systems increasingly mediate consumer information discovery, brands face algorithmic invisibility. This study investigates Cultural Encoding in Large Language Models (LLMs) -- systematic differences in brand…

Artificial Intelligence · Computer Science 2026-01-06 Huang Junyao , Situ Ruimin , Ye Renqin

Great research interests have been attracted to devise AI services that are able to provide mental health support. However, the lack of corpora is a main obstacle to this research, particularly in Chinese language. In this paper, we propose…

Computation and Language · Computer Science 2021-06-04 Hao Sun , Zhenru Lin , Chujie Zheng , Siyang Liu , Minlie Huang

Wikipedia is the world's largest online encyclopedia, but maintaining article quality through collaboration is challenging. Wikipedia designed a quality scale, but with such a manual assessment process, many articles remain unassessed. We…

Computation and Language · Computer Science 2023-10-04 Pedro Miguel Moás , Carla Teixeira Lopes

While large language models (LLMs) have showcased impressive capabilities, they struggle with addressing legal queries due to the intricate complexities and specialized expertise required in the legal field. In this paper, we introduce…

Computation and Language · Computer Science 2024-06-24 Zhiwei Fei , Songyang Zhang , Xiaoyu Shen , Dawei Zhu , Xiao Wang , Maosong Cao , Fengzhe Zhou , Yining Li , Wenwei Zhang , Dahua Lin , Kai Chen , Jidong Ge

The rapid development of advanced large language models (LLMs) has made AI-generated text indistinguishable from human-written text. Previous work on detecting AI-generated text has made effective progress, but has not involved modern…

Computation and Language · Computer Science 2025-09-03 Shanshan Wang , Junchao Wu , Fengying Ye , Jingming Yao , Lidia S. Chao , Derek F. Wong