中文
相关论文

相关论文: Telenor Nordics Customer Service self-help corpus

200 篇论文

We present MLSUM, the first large-scale MultiLingual SUMmarization dataset. Obtained from online newspapers, it contains 1.5M+ article/summary pairs in five different languages -- namely, French, German, Spanish, Russian, Turkish. Together…

计算与语言 · 计算机科学 2020-05-01 Thomas Scialom , Paul-Alexis Dray , Sylvain Lamprier , Benjamin Piwowarski , Jacopo Staiano

High-quality data resources play a crucial role in learning large language models (LLMs), particularly for low-resource languages like Cantonese. Despite having more than 85 million native speakers, Cantonese is still considered a…

Personal assistants, automatic speech recognizers and dialogue understanding systems are becoming more critical in our interconnected digital world. A clear example is air traffic control (ATC) communications. ATC aims at guiding aircraft…

Natural Language Processing (NLP) is revolutionising the way both professionals and laypersons operate in the legal field. The considerable potential for NLP in the legal sector, especially in developing computational assistance tools for…

计算与语言 · 计算机科学 2025-12-12 Farid Ariai , Joel Mackenzie , Gianluca Demartini

Manually curated biomedical repositories -- spanning bioactivity, genomics, and chemistry -- are expensive to maintain, lag behind primary literature, and discard experimental context, obscuring nuances needed to assess data correctness and…

Fact-checkers are often hampered by the sheer amount of online content that needs to be fact-checked. NLP can help them by retrieving already existing fact-checks relevant to the content being investigated. This paper introduces a new…

In general, Terms of Service (ToS) and other policy documents are verbose and full of legal jargon, which poses challenges for users to understand. To improve user accessibility and transparency, the "Terms of Service; Didn't Read" (ToS;DR)…

人机交互 · 计算机科学 2025-02-14 Shikha Soneji , Sourav Panda , Sameer Neve , Jonathan Dodge

Traditionally, Text Simplification is treated as a monolingual translation task where sentences between source texts and their simplified counterparts are aligned for training. However, especially for longer input documents, summarizing the…

计算与语言 · 计算机科学 2022-07-29 Dennis Aumiller , Michael Gertz

We present a collection of open, machine-readable document datasets covering parliamentary proceedings, legal judgments, government publications, news, and tourism statistics from Sri Lanka. The collection currently comprises of 269,194…

计算与语言 · 计算机科学 2026-05-18 Nuwan I. Senaratna

Text data is inherently temporal. The meaning of words and phrases changes over time, and the context in which they are used is constantly evolving. This is not just true for social media data, where the language used is rapidly influenced…

计算与语言 · 计算机科学 2025-03-05 Kai-Robin Lange , Niklas Benner , Lars Grönberg , Aymane Hachcham , Imene Kolli , Jonas Rieger , Carsten Jentsch

LLMs are ubiquitous in modern NLP, and while their applicability extends to texts produced for democratic activities such as online deliberations or large-scale citizen consultations, ethical questions have been raised for their usage as…

计算与语言 · 计算机科学 2026-04-21 Pierre-Antoine Lequeu , Léo Labat , Laurène Cave , Gaël Lejeune , François Yvon , Benjamin Piwowarski

We introduce the Speak & Improve Corpus 2025, a dataset of L2 learner English data with holistic scores and language error annotation, collected from open (spontaneous) speaking tests on the Speak & Improve learning platform. The aim of the…

计算与语言 · 计算机科学 2024-12-18 Kate Knill , Diane Nicholls , Mark J. F. Gales , Mengjie Qian , Pawel Stroinski

Countless terms of service (ToS) are being signed everyday by users all over the world while interacting with all kinds of apps and websites. More often than not, these online contracts spanning double-digit pages are signed blindly by…

计算与语言 · 计算机科学 2024-09-09 Mirgita Frasheri , Arian Bakhtiarnia , Lukas Esterle , Alexandros Iosifidis

Recent advances in natural language processing (NLP) can be largely attributed to the advent of pre-trained language models such as BERT and RoBERTa. While these models demonstrate remarkable performance on general datasets, they can…

This paper presents a hierarchical multi-agent LLM architecture to bridge communication gaps between non-technical end users and telecommunications domain experts in private network environments. We propose a cross-domain query translation…

网络与互联网体系结构 · 计算机科学 2026-05-19 Nguyen Phuc Tran , Brigitte Jaumard , Karthikeyan Premkumar , Salman Memon

We present a scalable, modular pipeline for automatic neologism detection that combines rule-based filtering with LLM classification. The pipeline is grounded in two complementary word-formation frameworks, grammatical and extra-grammatical…

计算与语言 · 计算机科学 2026-05-08 Diego Rossini , Lonneke van der Plas

Social Media platforms have offered invaluable opportunities for linguistic research. The availability of up-to-date data, coming from any part in the world, and coming from natural contexts, has allowed researchers to study language in…

计算与语言 · 计算机科学 2024-07-23 Simon Gonzalez

We present Multilingual Open Text (MOT), a new multilingual corpus containing text in 44 languages, many of which have limited existing text resources for natural language processing. The first release of the corpus contains over 2.8…

计算与语言 · 计算机科学 2022-06-10 Chester Palen-Michel , June Kim , Constantine Lignos

How can a text corpus stored in a customer relationship management (CRM) database be used for data mining and segmentation? In order to answer this question we inherited the state of the art methods commonly used in natural language…

计算与语言 · 计算机科学 2021-06-10 Şükrü Ozan