English
Related papers

Related papers: TuGeBiC: A Turkish German Bilingual Code-Switching…

200 papers

In 1993, the homophonic quotient groups for French and English (the quotient of the free group generated by the French (respectively English) alphabet determined by relations representing standard pronunciation rules) were explicitly…

Group Theory · Mathematics 2018-12-31 Herbert Gangl , Gizem Karaali , Woohyung Lee

Recent works in spoken language translation (SLT) have attempted to build end-to-end speech-to-text translation without using source language transcription during learning or decoding. However, while large quantities of parallel texts (such…

Computation and Language · Computer Science 2018-02-12 Ali Can Kocabiyikoglu , Laurent Besacier , Olivier Kraif

Code-Switching (CS) multilingual Automatic Speech Recognition (ASR) models can transcribe speech containing two or more alternating languages during a conversation. This paper proposes (1) a new method for creating code-switching ASR…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-19 Kunal Dhawan , Dima Rekesh , Boris Ginsburg

We present SpeechMatrix, a large-scale multilingual corpus of speech-to-speech translations mined from real speech of European Parliament recordings. It contains speech alignments in 136 language pairs with a total of 418 thousand hours of…

Computation and Language · Computer Science 2022-11-10 Paul-Ambroise Duquenne , Hongyu Gong , Ning Dong , Jingfei Du , Ann Lee , Vedanuj Goswani , Changhan Wang , Juan Pino , Benoît Sagot , Holger Schwenk

Synthesizing voice with the help of machine learning techniques has made rapid progress over the last years [1] and first high profile fraud cases have been recently reported [2]. Given the current increase in using conferencing tools for…

Audio and Speech Processing · Electrical Eng. & Systems 2021-08-04 Vanessa Barnekow , Dominik Binder , Niclas Kromrey , Pascal Munaretto , Andreas Schaad , Felix Schmieder

This study presents an LLM-assisted annotation pipeline for the sociolinguistic and topical analysis of bilingual discourse in two typologically distinct contexts: Spanish-English and Spanish-Guaran\'i. Using large language models, we…

Computation and Language · Computer Science 2025-12-04 Nemika Tyagi , Nelvin Licona Guevara , Olga Kellert

We present SwissGPC v1.0, the first mid-to-large-scale corpus of spontaneous Swiss German speech, developed to support research in ASR, TTS, dialect identification, and related fields. The dataset consists of links to talk shows and…

Computation and Language · Computer Science 2025-09-25 Samuel Stucki , Mark Cieliebak , Jan Deriu

Crafting an effective Automatic Speech Recognition (ASR) solution for dialects demands innovative approaches that not only address the data scarcity issue but also navigate the intricacies of linguistic diversity. In this paper, we address…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-26 Ahmed Amine Ben Abdallah , Ata Kabboudi , Amir Kanoun , Salah Zaiem

Code-Switching, a common phenomenon in written text and conversation, has been studied over decades by the natural language processing (NLP) research community. Initially, code-switching is intensively explored by leveraging linguistic…

Computation and Language · Computer Science 2023-05-26 Genta Indra Winata , Alham Fikri Aji , Zheng-Xin Yong , Thamar Solorio

Tokenization is a fundamental preprocessing step in Natural Language Processing (NLP), significantly impacting the capability of large language models (LLMs) to capture linguistic and semantic nuances. This study introduces a novel…

Computation and Language · Computer Science 2025-08-19 M. Ali Bayram , Ali Arda Fincan , Ahmet Semih Gümüş , Sercan Karakaş , Banu Diri , Savaş Yıldırım

The "MEG-MASC" dataset provides a curated set of raw magnetoencephalography (MEG) recordings of 27 English speakers who listened to two hours of naturalistic stories. Each participant performed two identical sessions, involving listening to…

Quantitative Methods · Quantitative Biology 2022-08-25 Laura Gwilliams , Graham Flick , Alec Marantz , Liina Pylkkanen , David Poeppel , Jean-Remi King

We present a cross-lingual summarisation corpus with long documents in a source language associated with multi-sentence summaries in a target language. The corpus covers twelve language pairs and directions for four European languages,…

Computation and Language · Computer Science 2022-02-22 Laura Perez-Beltrachini , Mirella Lapata

Code-switching (CS), the alternation between two or more languages within a single speaker's utterances, is common in real-world conversations and poses significant challenges for multilingual speech technology. However, systems capable of…

Computation and Language · Computer Science 2025-08-22 Sangmin Lee , Woojin Chung , Seyun Um , Hong-Goo Kang

The development of robust, multilingual speaker recognition systems is hindered by a lack of large-scale, publicly available and multilingual datasets, particularly for the read-speech style crucial for applications like anti-spoofing. To…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-26 Aref Farhadipour , Jan Marquenie , Srikanth Madikeri , Eleanor Chodroff

We investigate how transformer models represent complex verb paradigms in Turkish and Modern Hebrew, concentrating on how tokenization strategies shape this ability. Using the Blackbird Language Matrices task on natural data, we show that…

Computation and Language · Computer Science 2026-02-06 Giuseppe Samo , Paola Merlo

Recent studies have demonstrated how to assess the stereotypical bias in pre-trained English language models. In this work, we extend this branch of research in multiple different dimensions by systematically investigating (a) mono- and…

Spoken language translation has recently witnessed a resurgence in popularity, thanks to the development of end-to-end models and the creation of new corpora, such as Augmented LibriSpeech and MuST-C. Existing datasets involve language…

Computation and Language · Computer Science 2020-06-11 Changhan Wang , Juan Pino , Anne Wu , Jiatao Gu

Topic models are widely used in natural language processing, allowing researchers to estimate the underlying themes in a collection of documents. Most topic models use unsupervised methods and hence require the additional step of attaching…

Computation and Language · Computer Science 2018-08-28 Alexander Herzog , Peter John , Slava Jankin Mikhaylov

Tokenization is a pivotal design choice for neural language modeling in morphologically rich languages (MRLs) such as Turkish, where productive agglutination challenges both vocabulary efficiency and morphological fidelity. Prior studies…

Computation and Language · Computer Science 2026-02-09 Duygu Altinok

Tokenization is an important text preprocessing step to prepare input tokens for deep language models. WordPiece and BPE are de facto methods employed by important models, such as BERT and GPT. However, the impact of tokenization can be…

Computation and Language · Computer Science 2023-03-28 Cagri Toraman , Eyup Halit Yilmaz , Furkan Şahinuç , Oguzhan Ozcelik
‹ Prev 1 3 4 5 6 7 10 Next ›