English
Related papers

Related papers: TuGeBiC: A Turkish German Bilingual Code-Switching…

200 papers

Code-mixing involves the seamless integration of linguistic elements from multiple languages within a single discourse, reflecting natural multilingual communication patterns. Despite its prominence in informal interactions such as social…

Computation and Language · Computer Science 2025-06-17 Svetlana Churina , Akshat Gupta , Insyirah Mujtahid , Kokil Jaidka

This paper introduces GigaST, a large-scale pseudo speech translation (ST) corpus. We create the corpus by translating the text in GigaSpeech, an English ASR corpus, into German and Chinese. The training set is translated by a strong…

Computation and Language · Computer Science 2023-06-07 Rong Ye , Chengqi Zhao , Tom Ko , Chutong Meng , Tao Wang , Mingxuan Wang , Jun Cao

This paper reports on the preliminary phase of our ongoing research towards developing an intelligent tutoring environment for Turkish grammar. One of the components of this environment is a corpus search tool which, among other aspects of…

cmp-lg · Computer Science 2016-08-31 H. Altay Guvenir , Kemal Oflazer

Tokenisation - "the process of splitting text into atomic parts" (Brezina & Timperley, 2017: 1) - is a crucial step for corpus linguistics, as it provides the basis for any applicable quantitative method (e.g. collocations) while ensuring…

Computation and Language · Computer Science 2025-07-03 Matteo Di Cristofaro

To support machine learning of cross-language prosodic mappings and other ways to improve speech-to-speech translation, we present a protocol for collecting closely matched pairs of utterances across languages, a description of the…

Computation and Language · Computer Science 2023-07-17 Nigel G. Ward , Jonathan E. Avila , Emilia Rivas , Divette Marco

In this paper we present the final result of a project on Tunisian Arabic encoded in Arabizi, the Latin-based writing system for digital conversations. The project led to the creation of two integrated and independent resources: a corpus…

Computation and Language · Computer Science 2022-07-12 Elisa Gugliotta , Marco Dinarelli

In this study, we develop and assess new corpus selection and training methodologies to improve the effectiveness of Turkish language models. Specifically, we adapted Large Language Model generated datasets and translated English datasets…

Computation and Language · Computer Science 2024-12-05 H. Toprak Kesgin , M. Kaan Yuce , Eren Dogan , M. Egemen Uzun , Atahan Uz , Elif Ince , Yusuf Erdem , Osama Shbib , Ahmed Zeer , M. Fatih Amasyali

Open-access corpora are essential for advancing natural language processing (NLP) and computational social science (CSS). However, large-scale resources for German remain limited, restricting research on linguistic trends and societal…

Computation and Language · Computer Science 2025-06-09 Stefanie Urchs , Veronika Thurner , Matthias Aßenmacher , Christian Heumann , Stephanie Thiemichen

Large language model development relies on large-scale training corpora, yet most contain data of unclear licensing status, limiting the development of truly open models. This problem is exacerbated for non-English languages, where openly…

Recent state-of-the-art neural text-to-speech (TTS) synthesis models have dramatically improved intelligibility and naturalness of generated speech from text. However, building a good bilingual or code-switched TTS for a particular voice is…

Sound · Computer Science 2020-10-19 Shengkui Zhao , Trung Hieu Nguyen , Hao Wang , Bin Ma

In this paper, we present a corpus for use in automatic readability assessment and automatic text simplification of German. The corpus is compiled from web sources and consists of approximately 211,000 sentences. As a novel contribution, it…

Computation and Language · Computer Science 2019-09-20 Alessia Battisti , Sarah Ebling

The use of Project Gutenberg (PG) as a text corpus has been extremely popular in statistical analysis of language for more than 25 years. However, in contrast to other major linguistic datasets of similar importance, no consensual full…

Computation and Language · Computer Science 2018-12-20 Martin Gerlach , Francesc Font-Clos

Code-Switching (CS) is referred to the phenomenon of alternately using words and phrases from different languages. While today's neural end-to-end (E2E) models deliver state-of-the-art performances on the task of automatic speech…

Computation and Language · Computer Science 2023-07-04 Enes Yavuz Ugan , Christian Huber , Juan Hussain , Alexander Waibel

Natural language processing is a branch of computer science that combines artificial intelligence with linguistics. It aims to analyze a language element such as writing or speaking with software and convert it into information. Considering…

Computation and Language · Computer Science 2021-01-28 Kadir Tohma , Yakup Kutlu

The goal of voice anonymization is to modify an audio such that the true identity of its speaker is hidden. Research on this task is typically limited to the same English read speech datasets, thus the efficacy of current methods for other…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-03 Sarina Meyer , Ekaterina Kolos , Ngoc Thang Vu

Multilingual speakers tend to alternate between languages within a conversation, a phenomenon referred to as "code-switching" (CS). CS is a complex phenomenon that not only encompasses linguistic challenges, but also contains a great deal…

Computation and Language · Computer Science 2021-12-14 Injy Hamed , Alia El Bolock , Nader Rizk , Cornelia Herbert , Slim Abdennadher , Ngoc Thang Vu

This work aims to build a multilingual text-to-speech (TTS) synthesis system for ten lower-resourced Turkic languages: Azerbaijani, Bashkir, Kazakh, Kyrgyz, Sakha, Tatar, Turkish, Turkmen, Uyghur, and Uzbek. We specifically target the…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-26 Rustem Yeshpanov , Saida Mussakhojayeva , Yerbolat Khassanov

The popularity of automatic speech-to-speech translation for human conversations is growing, but the quality varies significantly depending on the language pair. In a context of community interpreting for low-resource languages, namely…

Computation and Language · Computer Science 2025-06-03 Andrei Popescu-Belis , Alexis Allemann , Teo Ferrari , Gopal Krishnamani

The application of natural language processing on political texts as well as speeches has become increasingly relevant in political sciences due to the ability to analyze large text corpora which cannot be read by a single person. But such…

Computation and Language · Computer Science 2024-10-24 Kai-Robin Lange , Carsten Jentsch

Spoken dialog systems need to be able to handle both multiple languages and multilinguality inside a conversation (\textit{e.g} in case of code-switching). In this work, we introduce new pretraining losses tailored to learn multilingual…

Computation and Language · Computer Science 2021-09-10 Emile Chapuis , Pierre Colombo , Matthieu Labeau , Chloe Clavel