English
Related papers

Related papers: Curation of a Palaeohispanic Dataset for Machine L…

200 papers

Even in highly-developed countries, as many as 15-30\% of the population can only understand texts written using a basic vocabulary. Their understanding of everyday texts is limited, which prevents them from taking an active role in society…

Computation and Language · Computer Science 2022-09-13 Sanja Stajner , Daniel Ferres , Matthew Shardlow , Kai North , Marcos Zampieri , Horacio Saggion

Endangered languages, such as Navajo - the most widely spoken Native American language - are significantly underrepresented in contemporary language technologies, exacerbating the challenges of their preservation and revitalization. This…

Computation and Language · Computer Science 2025-02-12 Ivory Yang , Weicheng Ma , Chunhui Zhang , Soroush Vosoughi

Machine translation from polysynthetic to fusional languages is a challenging task, which gets further complicated by the limited amount of parallel text available. Thus, translation performance is far from the state of the art for…

Computation and Language · Computer Science 2018-07-03 Manuel Mager , Elisabeth Mager , Alfonso Medina-Urrea , Ivan Meza , Katharina Kann

This paper presents methods to discriminate between languages and dialects written in Cuneiform script, one of the first writing systems in the world. We report the results obtained by the PZ team in the Cuneiform Language Identification…

Computation and Language · Computer Science 2019-04-30 Gustavo Henrique Paetzold , Marcos Zampieri

This paper proposes a novel framework for understanding large language models (LLMs) by reconceptualizing them as semiotic machines rather than as imitations of human cognition. Drawing from structuralist and post-structuralist theories of…

Artificial Intelligence · Computer Science 2024-10-18 Elad Vromen

Zero-shot Voice Cloning (VC) and Text-to-Speech (TTS) methods have advanced rapidly, enabling the generation of highly realistic synthetic speech and raising serious concerns about their misuse. While numerous detectors have been developed…

Machine Learning · Computer Science 2025-09-12 Maria Risques , Kratika Bhagtani , Amit Kumar Singh Yadav , Edward J. Delp

Multilingual semantic parsing is a cost-effective method that allows a single model to understand different languages. However, researchers face a great imbalance of availability of training data, with English being resource rich, and other…

Computation and Language · Computer Science 2021-06-15 Menglin Xia , Emilio Monti

Large Pre-trained Language Models (PLMs) have become ubiquitous in the development of language understanding technology and lie at the heart of many artificial intelligence advances. While advances reported for English using PLMs are…

Computation and Language · Computer Science 2021-04-12 Amit Seker , Elron Bandel , Dan Bareket , Idan Brusilovsky , Refael Shaked Greenfeld , Reut Tsarfaty

Generative Pre-trained Transformers (GPTs) have recently been scaled to unprecedented sizes in the history of machine learning. These models, solely trained on the language modeling objective, have been shown to exhibit outstanding few-shot…

Computation and Language · Computer Science 2021-08-31 Jordi Armengol-Estapé , Ona de Gibert Bonet , Maite Melero

This paper presents Translatotron 3, a novel approach to unsupervised direct speech-to-speech translation from monolingual speech-text datasets by combining masked autoencoder, unsupervised embedding mapping, and back-translation.…

Computation and Language · Computer Science 2024-01-17 Eliya Nachmani , Alon Levkovitch , Yifan Ding , Chulayuth Asawaroengchai , Heiga Zen , Michelle Tadmor Ramanovich

This paper presents a new database collected from a bilingual speakers set (49), in two different languages: Spanish and Catalan. Phonetically there are significative differences between both languages. These differences have let us to…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-07 Antonio Satue-Villar , Marcos Faundez-Zanuy

Mexico is a country with a large number of indigenous languages, among which the most widely spoken is Nawatl, with more than two million people currently speaking it (mainly in North and Central America). Despite its rich cultural…

The work in this paper describes the training and evaluation of machine learning (ML) techniques for the classification of cuneiform signs. There is a lot of variability in cuneiform signs, depending on where they come from, for what and by…

Machine Learning · Computer Science 2025-07-21 Eli Verwimp , Gustav Ryberg Smidt , Hendrik Hameeuw , Katrien De Graef

We have so many languages to communicate with others as humans. There are approximately 7000 languages in the world, and many are becoming extinct for a variety of reasons. In order to preserve and prevent the extinction of these languages,…

Digital Libraries · Computer Science 2023-10-09 Udaya Varadarajan , Sneha Bharti

Large pretrained multilingual models, trained on dozens of languages, have delivered promising results due to cross-lingual learning capabilities on variety of language tasks. Further adapting these models to specific languages, especially…

Computation and Language · Computer Science 2022-11-24 Fahim Faisal , Antonios Anastasopoulos

We present models which complete missing text given transliterations of ancient Mesopotamian documents, originally written on cuneiform clay tablets (2500 BCE - 100 CE). Due to the tablets' deterioration, scholars often rely on contextual…

Computation and Language · Computer Science 2021-10-26 Koren Lazar , Benny Saret , Asaf Yehudai , Wayne Horowitz , Nathan Wasserman , Gabriel Stanovsky

Identification of the languages written using cuneiform symbols is a difficult task due to the lack of resources and the problem of tokenization. The Cuneiform Language Identification task in VarDial 2019 addresses the problem of…

Computation and Language · Computer Science 2020-09-24 Ehsan Doostmohammadi , Minoo Nassajian

Dictionaries and phrase tables are the basis of modern statistical machine translation systems. This paper develops a method that can automate the process of generating and extending dictionaries and phrase tables. Our method can translate…

Computation and Language · Computer Science 2013-09-18 Tomas Mikolov , Quoc V. Le , Ilya Sutskever

The growing interest in Large Language Models (LLMs) and in particular in conversational models with which users can interact has led to the development of a large number of open-source chat LLMs. These models are evaluated on a wide range…

Large language models, such as the well-known ChatGPT, have brought about an unexpected revolution in the field of artificial intelligence. On the one hand, they have numerous practical applications and enormous potential still to be…

Computation and Language · Computer Science 2025-02-26 Carlos Gómez-Rodríguez
‹ Prev 1 3 4 5 6 7 10 Next ›