中文
相关论文

相关论文: hinglishNorm -- A Corpus of Hindi-English Code Mix…

200 篇论文

Cross-lingual information retrieval is a challenging task in the absence of aligned parallel corpora. In this paper, we address this problem by considering topically aligned corpora designed for evaluating an IR setup. To emphasize, we…

信息检索 · 计算机科学 2018-04-13 Mitodru Niyogi , Kripabandhu Ghosh , Arnab Bhattacharya

Recent research has focused on literary machine translation (MT) as a new challenge in MT. However, the evaluation of literary MT remains an open problem. We contribute to this ongoing discussion by introducing LITEVAL-CORPUS, a…

计算与语言 · 计算机科学 2025-02-26 Ran Zhang , Wei Zhao , Steffen Eger

In this paper, we describe our submission to the WMT19 low-resource parallel corpus filtering shared task. Our main approach is based on the LASER toolkit (Language-Agnostic SEntence Representations), which uses an encoder-decoder…

计算与语言 · 计算机科学 2019-06-24 Vishrav Chaudhary , Yuqing Tang , Francisco Guzmán , Holger Schwenk , Philipp Koehn

Text similarity detection aims at measuring the degree of similarity between a pair of texts. Corpora available for text similarity detection are designed to evaluate the algorithms to assess the paraphrase level among documents. In this…

信息检索 · 计算机科学 2017-03-14 Juan-Manuel Torres-Moreno , Gerardo Sierra , Peter Peinl

We present a parallel machine translation training corpus for English and Akuapem Twi of 25,421 sentence pairs. We used a transformer-based translator to generate initial translations in Akuapem Twi, which were later verified and corrected…

We introduce COMI-LINGUA, the largest manually annotated Hindi-English code-mixed dataset, comprising 125K+ high-quality instances across five core NLP tasks: Matrix Language Identification, Token-level Language Identification,…

计算与语言 · 计算机科学 2025-09-18 Rajvee Sheth , Himanshu Beniwal , Mayank Singh

Multilingual encoder-based language models are widely adopted for code-mixed analysis tasks, yet we know surprisingly little about how they represent code-mixed inputs internally - or whether those representations meaningfully connect to…

计算与语言 · 计算机科学 2026-03-23 Debajyoti Mazumder , Divyansh Pathak , Prashant Kodali , Jasabanta Patro

With a large amount of parallel data, neural machine translation systems are able to deliver human-level performance for sentence-level translation. However, it is costly to label a large amount of parallel data by humans. In contrast,…

计算与语言 · 计算机科学 2020-09-21 Guokun Lai , Zihang Dai , Yiming Yang

Visual Genome is a dataset connecting structured image information with English language. We present ``Hindi Visual Genome'', a multimodal dataset consisting of text and images suitable for English-Hindi multimodal machine translation task…

计算与语言 · 计算机科学 2019-07-23 Shantipriya Parida , Ondřej Bojar , Satya Ranjan Dash

We propose an unsupervised method to obtain cross-lingual embeddings without any parallel data or pre-trained word embeddings. The proposed model, which we call multilingual neural language models, takes sentences of multiple languages as…

计算与语言 · 计算机科学 2018-09-10 Takashi Wada , Tomoharu Iwata

We show that margin-based bitext mining in a multilingual sentence space can be applied to monolingual corpora of billions of sentences. We are using ten snapshots of a curated common crawl corpus (Wenzek et al., 2019) totalling 32.7…

计算与语言 · 计算机科学 2020-05-04 Holger Schwenk , Guillaume Wenzek , Sergey Edunov , Edouard Grave , Armand Joulin

Translationese refers to linguistic properties that usually occur in translated texts. Previous works study translationese by framing it as a binary classification between original texts and translated texts. In this paper, we argue that…

计算与语言 · 计算机科学 2025-09-22 Yikang Liu , Wanyang Zhang , Yiming Wang , Jialong Tang , Pei Zhang , Baosong Yang , Fei Huang , Rui Wang , Hai Hu

The development of automated approaches to linguistic acceptability has been greatly fostered by the availability of the English CoLA corpus, which has also been included in the widely used GLUE benchmark. However, this kind of research for…

计算与语言 · 计算机科学 2022-10-14 Daniela Trotta , Raffaele Guarasci , Elisa Leonardelli , Sara Tonelli

The Parallel Meaning Bank is a corpus of translations annotated with shared, formal meaning representations comprising over 11 million words divided over four languages (English, German, Italian, and Dutch). Our approach is based on…

The entropy rate of printed English is famously estimated to be about one bit per character, a benchmark that modern large language models (LLMs) have only recently approached. This entropy rate implies that English contains nearly 80…

计算与语言 · 计算机科学 2026-02-19 Weishun Zhong , Doron Sivan , Tankut Can , Mikhail Katkov , Misha Tsodyks

This paper introduces a new corpus of Mandarin-English code-switching speech recognition--TALCS corpus, suitable for training and evaluating code-switching speech recognition systems. TALCS corpus is derived from real online one-to-one…

计算与语言 · 计算机科学 2022-06-28 Chengfei Li , Shuhao Deng , Yaoping Wang , Guangjing Wang , Yaguang Gong , Changbin Chen , Jinfeng Bai

Multilingual Large Language Models (LLMs) have demonstrated exceptional performance in Machine Translation (MT) tasks. However, their MT abilities in the context of code-switching (the practice of mixing two or more languages in an…

计算与语言 · 计算机科学 2024-10-16 Ayushman Gupta , Akhil Bhogal , Kripabandhu Ghosh

Existing graph- and hypergraph-based algorithms for document summarization represent the sentences of a corpus as the nodes of a graph or a hypergraph in which the edges represent relationships of lexical similarities between sentences.…

计算与语言 · 计算机科学 2019-04-17 Hadrien Van Lierde , Tommy W. S. Chow

Cross-Language Information Retrieval (CLIR) has become an important problem to solve in the recent years due to the growth of content in multiple languages in the Web. One of the standard methods is to use query translation from source to…

计算与语言 · 计算机科学 2016-08-05 Paheli Bhattacharya , Pawan Goyal , Sudeshna Sarkar

Automated image captioning using the content from the image is very appealing when done by harnessing the capability of computer vision and natural language processing. Extensive research has been done in the field with a major focus on the…

计算机视觉与模式识别 · 计算机科学 2026-02-17 Wasim Akram Khan , Anil Kumar Vuppala
‹ 上一页 1 8 9 10 下一页 ›