中文
相关论文

相关论文: Automatic register identification for the open web…

200 篇论文

We introduce a Content-based Document Alignment approach (CDA), an efficient method to align multilingual web documents based on content in creating parallel training data for machine translation (MT) systems operating at the industrial…

计算与语言 · 计算机科学 2021-02-23 Thuy Vu , Alessandro Moschitti

Logging statements are central to debugging, failure diagnosis, and production observability, yet writing them requires developers to decide where to place a logging statement, which API and severity level to use, and what runtime…

软件工程 · 计算机科学 2026-04-21 Renyi Zhong , Yichen Li , Yulun Wu , Jinxi Kuang , Yintong Huo , Michael R. Lyu

This paper introduces the Multi-Genre Natural Language Inference (MultiNLI) corpus, a dataset designed for use in the development and evaluation of machine learning models for sentence understanding. In addition to being one of the largest…

计算与语言 · 计算机科学 2018-02-21 Adina Williams , Nikita Nangia , Samuel R. Bowman

Claims are the central component of an argument. Detecting claims across different domains or data sets can often be challenging due to their varying conceptualization. We propose to alleviate this problem by fine tuning a language model…

计算与语言 · 计算机科学 2019-05-20 Tuhin Chakrabarty , Christopher Hidey , Kathleen McKeown

Currently, publicly available models for website classification do not offer an embedding method and have limited support for languages beyond English. We release a dataset of more than two million category-labeled websites in 92 languages…

计算与语言 · 计算机科学 2022-04-11 Sylvain Lugeon , Tiziano Piccardi , Robert West

Security vulnerabilities present in a code that has been written in diverse programming languages are among the most critical yet complicated aspects of source code to detect. Static analysis tools based on rule-based patterns usually do…

密码学与安全 · 计算机科学 2025-08-19 Hael Abdulhakim Ali Humran , Ferdi Sonmez

In today's global digital landscape, misinformation transcends linguistic boundaries, posing a significant challenge for moderation systems. Most approaches to misinformation detection are monolingual, focused on high-resource languages,…

计算与语言 · 计算机科学 2025-04-01 Xinyu Wang , Wenbo Zhang , Sarah Rajtmajer

As open-ended human-chatbot interaction becomes commonplace, sensitive content detection gains importance. In this work, we propose a two stage semi-supervised approach to bootstrap large-scale data for automatic sensitive language…

计算与语言 · 计算机科学 2018-12-03 Chandra Khatri , Behnam Hedayatnia , Rahul Goel , Anushree Venkatesh , Raefer Gabriel , Arindam Mandal

The study of register in computational language research has historically been divided into register analysis, seeking to determine the registerial character of a text or corpus, and register synthesis, seeking to generate a text in a…

计算与语言 · 计算机科学 2019-01-10 Shlomo Engelson Argamon

Continued pretraining and instruction tuning on large-scale multilingual data have proven to be effective in scaling large language models (LLMs) to low-resource languages. However, the unaligned nature of such data limits its ability to…

计算与语言 · 计算机科学 2025-10-22 Yingli Shen , Wen Lai , Shuo Wang , Ge Gao , Kangyang Luo , Alexander Fraser , Maosong Sun

Multilingual dense retrieval aims to retrieve relevant documents across different languages based on a unified retriever model. The challenge lies in aligning representations of different languages in a shared vector space. The common…

信息检索 · 计算机科学 2025-09-12 Chao Huang , Fengran Mo , Yufeng Chen , Changhao Guan , Zhenrui Yue , Xinyu Wang , Jinan Xu , Kaiyu Huang

Pre-training state-of-the-art large language models (LLMs) requires vast amounts of clean and diverse text data. While the open development of large high-quality English pre-training datasets has seen substantial recent progress, training…

Language development experts need tools that can automatically identify languages from fluent, conversational speech, and provide reliable estimates of usage rates at the level of an individual recording. However, language identification…

音频与语音处理 · 电气工程与系统科学 2023-05-31 Suzy J. Styles , Victoria Y. H. Chua , Fei Ting Woon , Hexin Liu , Leibny Paola Garcia Perera , Sanjeev Khudanpur , Andy W. H. Khong , Justin Dauwels

India's linguistic landscape is one of the most diverse in the world, comprising over 120 major languages and approximately 1,600 additional languages, with 22 officially recognized as scheduled languages in the Indian Constitution. Despite…

This paper measures the stability of cross-linguistic register variation. A register is a variety of a language that is associated with extra-linguistic context. The relationship between a register and its context is functional: the…

计算与语言 · 计算机科学 2022-09-21 Haipeng Li , Jonathan Dunn , Andrea Nini

Sentence Boundary Detection (SBD) is one of the foundational building blocks of Natural Language Processing (NLP), with incorrectly split sentences heavily influencing the output quality of downstream tasks. It is a challenging task for…

计算与语言 · 计算机科学 2023-05-03 Tobias Brugger , Matthias Stürmer , Joel Niklaus

Google's multilingual speech recognition system combines low-level acoustic signals with language-specific recognizer signals to better predict the language of an utterance. This paper presents our experience with different signal…

机器学习 · 计算机科学 2019-11-05 Shengye Wang , Li Wan , Yang Yu , Ignacio Lopez Moreno

This paper proposes a novel framework for digital curation of Web corpora in order to provide robust estimation of their parameters, such as their composition and the lexicon. In recent years language models pre-trained on large corpora…

计算与语言 · 计算机科学 2020-03-16 Serge Sharoff

While multilingual language models successfully transfer factual and syntactic knowledge across languages, it remains unclear whether they process culture-specific pragmatic registers, such as slang, as isolated language-specific…

计算与语言 · 计算机科学 2026-03-30 Uri Z. Kialy , Avi Shtarkberg , Ayal Klein
‹ 上一页 1 8 9 10 下一页 ›