English
Related papers

Related papers: Indus script corpora, archaeo-metallurgy and Meluh…

200 papers

The paper presents a new script classification method for the discrimination of the South Slavic medieval labels. It consists in the textural analysis of the script types. In the first step, each letter is coded by the equivalent script…

Computer Vision and Pattern Recognition · Computer Science 2015-09-08 Darko Brodic , Alessia Amelio , Zoran N. Milivojevic

The earliest extant Chinese characters originate from oracle bone inscriptions, which are closely related to other East Asian languages. These inscriptions hold immense value for anthropology and archaeology. However, deciphering oracle…

Artificial Intelligence · Computer Science 2024-02-14 Haisu Guan , Jinpeng Wan , Yuliang Liu , Pengjie Wang , Kaile Zhang , Zhebin Kuang , Xinyu Wang , Xiang Bai , Lianwen Jin

Communication plays a vital role in human interaction. Studying language is a worthwhile task and more recently has become quantitative in nature with developments of fields like quantitative comparative linguistics and lexicostatistics.…

Applications · Statistics 2024-05-13 Garett Ordway , Vic Patrangenaru

For any deep computational processing of language we need evidences, and one such set of evidences is corpus. This paper describes the development of a text-based corpus for the Bishnupriya Manipuri language. A Corpus is considered as a…

Computation and Language · Computer Science 2013-12-12 Nayan Jyoti Kalita , Navanath Saharia , Smriti Kumar Sinha

This study presents a multi-modal multi-granularity tokenizer specifically designed for analyzing ancient Chinese scripts, focusing on the Chu bamboo slip (CBS) script used during the Spring and Autumn and Warring States period (771-256…

Computation and Language · Computer Science 2024-09-04 Yingfa Chen , Chenlong Hu , Cong Feng , Chenyang Song , Shi Yu , Xu Han , Zhiyuan Liu , Maosong Sun

This paper presents an end-to-end methodology for collecting datasets to recognize handwritten English alphabets by utilizing Inertial Measurement Units (IMUs) and leveraging the diversity present in the Indian writing style. The IMUs are…

Computer Vision and Pattern Recognition · Computer Science 2023-07-06 Hari Prabhat Gupta , Rahul Mishra

Chinese text processing systems are using Double Byte Coding , while almost all existing Sanskrit Based Indian Languages have been using Single Byte coding for text processing. Through observation, Chinese Information Processing Technique…

cmp-lg · Computer Science 2008-02-03 Md Maruf Hasan

Writing systems of Indic languages have orthographic syllables, also known as complex graphemes, as unique horizontal units. A prominent feature of these languages is these complex grapheme units that comprise consonants/consonant…

I describe my experience writing the first original, modern Computer Science research paper expressed entirely in an Indian language. The paper is in Telugu, a language with approximately 100 million speakers. The paper is in the field of…

General Literature · Computer Science 2026-04-07 Siddhartha Visveswara Jayanti

The Digital Corpus of Sanskrit records around 650,000 sentences along with their morphological and lexical tagging. But inconsistencies in morphological analysis, and in providing crucial information like the segmented word, urges the need…

Computation and Language · Computer Science 2020-05-15 Sriram Krishnan , Amba Kulkarni , Gérard Huet

Known by more than 1.5 billion people in the Indian subcontinent, Indic languages present unique challenges and opportunities for natural language processing (NLP) research due to their rich cultural heritage, linguistic diversity, and…

Computation and Language · Computer Science 2025-01-29 Sankalp KJ , Ashutosh Kumar , Laxmaan Balaji , Nikunj Kotecha , Vinija Jain , Aman Chadha , Sreyoshi Bhaduri

Full-duplex spoken dialogue systems can model natural conversational behaviours such as interruptions, overlaps, and backchannels, yet such systems remain largely unexplored for Indian languages. We present the first open, reproducible…

Computation and Language · Computer Science 2026-05-26 Bhaskar Singh , Shobhit Banga , Mahima Manik , Pranav Sharma

As one of the earliest ancient languages, Oracle Bone Script (OBS) encapsulates the cultural records and intellectual expressions of ancient civilizations. Despite the discovery of approximately 4,500 OBS characters, only about 1,600 have…

Computer Vision and Pattern Recognition · Computer Science 2025-06-27 Caoshuo Li , Zengmao Ding , Xiaobin Hu , Bang Li , Donghao Luo , AndyPian Wu , Chaoyang Wang , Chengjie Wang , Taisong Jin , SevenShu , Yunsheng Wu , Yongge Liu , Rongrong Ji

This study demonstrates how hybrid neural-symbolic methods can yield significant new insights into the evolution of a morphologically rich, low-resource language. We challenge the naive assumption that linguistic change is simplification by…

Computation and Language · Computer Science 2025-12-08 Ananth Hariharan , David Mortensen

Evaluating Large Language Models (LLMs) in low-resource and linguistically diverse languages remains a significant challenge in NLP, particularly for languages using non-Latin scripts like those spoken in India. Existing benchmarks…

Computation and Language · Computer Science 2025-02-05 Sshubam Verma , Mohammed Safi Ur Rahman Khan , Vishwajeet Kumar , Rudra Murthy , Jaydeep Sen

In a multilingual country like India where 12 different official scripts are in use, automatic identification of handwritten script facilitates many important applications such as automatic transcription of multilingual documents, searching…

Computer Vision and Pattern Recognition · Computer Science 2020-09-17 Pawan Kumar Singh , Iman Chatterjee , Ram Sarkar , Mita Nasipuri

Urdu is a challenging language because of, first, its Perso-Arabic script and second, its morphological system having inherent grammatical forms and vocabulary of Arabic, Persian and the native languages of South Asia. This paper describes…

Computation and Language · Computer Science 2022-04-08 Muhammad Humayoun , Harald Hammarström , Aarne Ranta

A cornerstone in AI research has been the creation and adoption of standardized training and test datasets to earmark the progress of state-of-the-art models. A particularly successful example is the GLUE dataset for training and evaluating…

Computation and Language · Computer Science 2022-12-16 Tahir Javed , Kaushal Santosh Bhogale , Abhigyan Raman , Anoop Kunchukuttan , Pratyush Kumar , Mitesh M. Khapra

Midrash collections are complex rabbinic works that consist of text in multiple languages, which evolved through long processes of unstable oral and written transmission. Determining the origin of a given passage in such a compilation is…

Computation and Language · Computer Science 2024-02-14 Shlomo Tannor , Nachum Dershowitz , Moshe Lavee

Recognition of text on word or line images, without the need for sub-word segmentation has become the mainstream of research and development of text recognition for Indian languages. Modelling unsegmented sequences using Connectionist…

Computer Vision and Pattern Recognition · Computer Science 2024-12-19 Minesh Mathew , Ajoy Mondal , CV Jawahar