English
Related papers

Related papers: A Hackathon for Classical Tibetan

200 papers

Automatic text analysis methods, such as Topic Modelling, are gaining much attention in Humanities. However, scholars need to have extensive coding skills to use such methods appropriately. The need of having this technical expertise…

Digital Libraries · Computer Science 2020-11-30 Ivan Heibi , Silvio Peroni , Luca Pareschi , Paolo Ferri

Across almost all scientific disciplines, the instruments that record our experimental data and the methods required for storage and data analysis are rapidly increasing in complexity. This gives rise to the need for scientific communities…

Physics Education · Physics 2018-09-26 Daniela Huppenkothen , Anthony Arendt , David W. Hogg , Karthik Ram , Jake VanderPlas , Ariel Rokem

This paper focuses on AI tutors in foreign language learning, a field of application of AI tutors with great development, especially during the last years, when great advances in natural language understanding and processing in real time,…

Human-Computer Interaction · Computer Science 2025-08-08 Nikolaos Avouris

We consider three major text sources about the Tang Dynasty of China in our experiments that aim to segment text written in classical Chinese. These corpora include a collection of Tang Tomb Biographies, the New Tang Book, and the Old Tang…

Computation and Language · Computer Science 2020-07-23 Chao-Lin Liu , Chang-Ting Chu , Wei-Ting Chang , Ti-Yong Zheng

This paper reports the findings of a study conducted on the application of an on-line human-computer dialog system with natural language (chatbot) on the teaching of foreign languages. A keywords-based human-computer dialog system makes it…

Computers and Society · Computer Science 2007-05-23 Jiyou Jia

There is no consensus on the state-of-the-art approach to historical text normalization. Many techniques have been proposed, including rule-based methods, distance metrics, character-based statistical machine translation, and neural…

Computation and Language · Computer Science 2019-10-15 Marcel Bollmann

Equipment shortages in Africa undermine Science, Technology, Engineering and Mathematics (STEM) Education. We have pioneered the LabHackathon (LabHack): a novel initiative that adapts the conventional hackathon and draws on insights from…

Computers and Society · Computer Science 2019-04-04 Helena Webb , Jason R. C. Nurse , Louise Bezuidenhout , Marina Jirotka

What are the best methods of capturing thematic similarity between literary texts? Knowing the answer to this question would be useful for automatic clustering of book genres, or any other thematic grouping. This paper compares a variety of…

Computation and Language · Computer Science 2023-05-22 Oleg Sobchuk , Artjoms Šeļa

Collaborative competitions have gained popularity in the scientific and technological fields. These competitions involve defining tasks, selecting evaluation scores, and devising result verification methods. In the standard scenario,…

Machine Learning · Computer Science 2024-08-22 Sergio Nava-Muñoz , Mario Graff , Hugo Jair Escalante

A virtual hands-on computer laboratory has been designed within the Gather online meeting platform. Gather's features such as spatial audio, private spaces and interactable objects offer scope for great improvements over currently used…

Human-Computer Interaction · Computer Science 2022-09-07 Rika Kobayashi , Sarah Jaffa , Jiachen Dong , Roger D. Amos , Jeremy Cohen , Emily F. Kerrison

This paper presents our process for developing a sample-efficient language model for a conversational Hinglish chatbot. Hinglish, a code-mixed language that combines Hindi and English, presents a unique computational challenge due to…

Computation and Language · Computer Science 2025-04-29 Sakshi Singh , Abhinav Prakash , Aakriti Shah , Chaitanya Sachdeva , Sanjana Dumpala

The problem of short text matching is formulated as follows: given a pair of sentences or questions, a matching model determines whether the input pair mean the same or not. Models that can automatically identify questions with the same…

Computation and Language · Computer Science 2018-11-15 Asmelash Teka Hadgu

The Digital Corpus of Sanskrit records around 650,000 sentences along with their morphological and lexical tagging. But inconsistencies in morphological analysis, and in providing crucial information like the segmented word, urges the need…

Computation and Language · Computer Science 2020-05-15 Sriram Krishnan , Amba Kulkarni , Gérard Huet

This technical report describes the methods and results of a three-week sprint to produce deployable speech recognition models for 31 under-served languages of the Common Voice project. We outline the preprocessing steps, hyperparameter…

Computation and Language · Computer Science 2021-05-12 Francis M. Tyers , Josh Meyer

We present iNLTK, an open-source NLP library consisting of pre-trained language models and out-of-the-box support for Data Augmentation, Textual Similarity, Sentence Embeddings, Word Embeddings, Tokenization and Text Generation in 13 Indic…

Computation and Language · Computer Science 2021-02-15 Gaurav Arora

This paper discusses Centre for Development of Advanced Computing Mumbai's (CDACM) submission to the NLP Tools Contest on Statistical Machine Translation in Indian Languages (ILSMT) 2014 (collocated with ICON 2014). The objective of the…

Computation and Language · Computer Science 2016-10-25 Raj Nath Patel , Prakash B. Pimpale , Sasikumar M

Adapting large language models (LLMs) to low-resource languages remains a major challenge due to data scarcity and cross-lingual drift. This work presents a two-stage adaptation of Qwen2.5-3B to Tibetan, a morphologically rich and…

Computation and Language · Computer Science 2025-12-04 Lifeng Chen , Ryan Lai , Tianming Liu

Past research has identified a rich set of handcrafted linguistic features that can potentially assist various tasks. However, their extensive number makes it difficult to effectively select and utilize existing handcrafted features.…

Computation and Language · Computer Science 2023-06-02 Bruce W. Lee , Jason Hyung-Jong Lee

The present study provides the first-ever report on the language shift from Tibetan to Arabic among descendants of Tibetan families who migrated from the Tibet region to Saudi Arabia around 70 years ago. The aim of this study was to…

Computation and Language · Computer Science 2025-02-17 Sumaiyah Turkistani Mohammad Almoaily