中文
相关论文

相关论文: Automatically Generating a Large, Culture-Specific…

200 篇论文

For many real-world applications, the user-generated inputs usually contain various noises due to speech recognition errors caused by linguistic variations1 or typographical errors (typos). Thus, it is crucial to test model performance on…

计算与语言 · 计算机科学 2023-05-26 Chenglei Si , Zhengyan Zhang , Yingfa Chen , Xiaozhi Wang , Zhiyuan Liu , Maosong Sun

Knowledge graph is a kind of valuable knowledge base which would benefit lots of AI-related applications. Up to now, lots of large-scale knowledge graphs have been built. However, most of them are non-Chinese and designed for general…

Although Google is blocked in China, Chinese provinces export significantly more to foreign countries that recently searched for them (up to 12 months prior). This attention premium is found mainly at the extensive margin of exports, larger…

综合经济学 · 经济学 2025-04-01 Cui Hu , Ben G. Li

Free web proxies promise anonymity and censorship circumvention at no cost. Several websites publish lists of free proxies organized by country, anonymity level, and performance. These lists index hundreds of thousand of hosts discovered…

网络与互联网体系结构 · 计算机科学 2017-11-03 Diego Perino , Matteo Varvello , Claudio Soriente

Current alignment evaluation mostly measures whether models encode dangerous concepts and whether they refuse harmful requests. Both miss the layer where alignment often operates: routing from concept detection to behavioral policy. We…

机器学习 · 计算机科学 2026-05-04 Gregory N. Frank

We present GLM-Dialog, a large-scale language model (LLM) with 10B parameters capable of knowledge-grounded conversation in Chinese using a search engine to access the Internet knowledge. GLM-Dialog offers a series of applicable techniques…

Lexical simplification has attracted much attention in many languages, which is the process of replacing complex words in a given sentence with simpler alternatives of equivalent meaning. Although the richness of vocabulary in Chinese makes…

计算与语言 · 计算机科学 2020-10-15 Jipeng Qiang , Xinyu Lu , Yun Li , Yunhao Yuan , Yang Shi , Xindong Wu

In early January 2020, after China reported the first cases of the new coronavirus (SARS-CoV-2) in the city of Wuhan, unreliable and not fully accurate information has started spreading faster than the virus itself. Alongside this pandemic,…

机器学习 · 计算机科学 2021-06-08 V. Mazzeo , A. Rapisarda , G. Giuffrida

Large language models (LLMs) are increasingly deployed in cost-sensitive and on-device scenarios, and safety guardrails have advanced mainly in English. However, real-world Chinese malicious queries typically conceal intent via homophones,…

计算与语言 · 计算机科学 2026-01-06 Zhenhong Zhou , Shilinlu Yan , Chuanpu Liu , Qiankun Li , Kun Wang , Zhigang Zeng

In the current environment, psychological issues are prevalent and widespread, with social media serving as a key outlet for individuals to share their feelings. This results in the generation of vast quantities of data daily, where…

计算与语言 · 计算机科学 2024-06-13 Wei Zhai , Hongzhi Qi , Qing Zhao , Jianqiang Li , Ziqi Wang , Han Wang , Bing Xiang Yang , Guanghui Fu

The absence of an appropriate text classification corpus makes the massive amount of online job information unusable for labor market analysis. This paper presents JCTC, a large job posting corpus for text classification. In JCTC…

信息检索 · 计算机科学 2017-06-13 Haoyu Xu , Chongyang Gu , Han Zhou , Sengpan Kou , Junjie Zhang

Contemporary language models are increasingly multilingual, but Chinese LLM developers must navigate complex political and business considerations of language diversity. Language policy in China aims at influencing the public discourse and…

计算与语言 · 计算机科学 2025-08-12 Andrea W Wen-Yi , Unso Eun Seo Jo , Lu Jia Lin , David Mimno

We describe our experience of implementing a news content organization system at Tencent that discovers events from vast streams of breaking news and evolves news story structures in an online fashion. Our real-world system has distinct…

信息检索 · 计算机科学 2018-03-02 Bang Liu , Di Niu , Kunfeng Lai , Linglong Kong , Yu Xu

Large language models have recently made tremendous progress in a variety of aspects, e.g., cross-task generalization, instruction following. Comprehensively evaluating the capability of large language models in multiple tasks is of great…

The rapid development of Chinese large language models (LLMs) poses big challenges for efficient LLM evaluation. While current initiatives have introduced new benchmarks or evaluation platforms for assessing Chinese LLMs, many of these…

As artificial intelligence systems increasingly mediate consumer information discovery, brands face algorithmic invisibility. This study investigates Cultural Encoding in Large Language Models (LLMs) -- systematic differences in brand…

人工智能 · 计算机科学 2026-01-06 Huang Junyao , Situ Ruimin , Ye Renqin

Great research interests have been attracted to devise AI services that are able to provide mental health support. However, the lack of corpora is a main obstacle to this research, particularly in Chinese language. In this paper, we propose…

计算与语言 · 计算机科学 2021-06-04 Hao Sun , Zhenru Lin , Chujie Zheng , Siyang Liu , Minlie Huang

Wikipedia is the world's largest online encyclopedia, but maintaining article quality through collaboration is challenging. Wikipedia designed a quality scale, but with such a manual assessment process, many articles remain unassessed. We…

计算与语言 · 计算机科学 2023-10-04 Pedro Miguel Moás , Carla Teixeira Lopes

While large language models (LLMs) have showcased impressive capabilities, they struggle with addressing legal queries due to the intricate complexities and specialized expertise required in the legal field. In this paper, we introduce…

The rapid development of advanced large language models (LLMs) has made AI-generated text indistinguishable from human-written text. Previous work on detecting AI-generated text has made effective progress, but has not involved modern…

计算与语言 · 计算机科学 2025-09-03 Shanshan Wang , Junchao Wu , Fengying Ye , Jingming Yao , Lidia S. Chao , Derek F. Wong