English
Related papers

Related papers: Indus script corpora, archaeo-metallurgy and Meluh…

200 papers

Oracle bone script, one of the earliest known forms of ancient Chinese writing, presents invaluable research materials for scholars studying the humanities and geography of the Shang Dynasty, dating back 3,000 years. The immense historical…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Pengjie Wang , Kaile Zhang , Xinyu Wang , Shengwei Han , Yongge Liu , Jinpeng Wan , Haisu Guan , Zhebin Kuang , Lianwen Jin , Xiang Bai , Yuliang Liu

Digital humanities are significantly transforming how Egyptologists study ancient Egyptian texts. The OCR-PT-CT project proposes a recognition method for hieroglyphs based on images of Coffin Texts (CT) from Adriaan de Buck (1935-1961) and…

In this paper, we introduce Neural Information Retrieval resources for 11 widely spoken Indian Languages (Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Oriya, Punjabi, Tamil, and Telugu) from two major Indian language…

Information Retrieval · Computer Science 2023-12-18 Saiful Haq , Ashutosh Sharma , Pushpak Bhattacharyya

Purpose: This study investigates the transcription principles underlying Hu\`it\'onggu\v{a}nx\`i Hu\'ay\'iy\`iy\v{u} (HHY), a series of multilingual glossaries compiled by the Ming government between the fifteenth and sixteenth centuries…

Computation and Language · Computer Science 2026-05-27 Ji-eun Kim

We have summarized here the astronomical knowledge of the ancient Hindu astronomers. This knowledge was accumulated from before 1500 B.C. up to around 1200 A.D. In Section \ref{equiv} we have correlated terms used by the Hindu astronomers…

History and Philosophy of Physics · Physics 2009-03-11 P. Rudra

We present models which complete missing text given transliterations of ancient Mesopotamian documents, originally written on cuneiform clay tablets (2500 BCE - 100 CE). Due to the tablets' deterioration, scholars often rely on contextual…

Computation and Language · Computer Science 2021-10-26 Koren Lazar , Benny Saret , Asaf Yehudai , Wayne Horowitz , Nathan Wasserman , Gabriel Stanovsky

Part-of-speech (POS) tagging remains a foundational component in natural language processing pipelines, particularly critical for historical text analysis at the intersection of computational linguistics and digital humanities. Despite…

The evolution of languages closely resembles the evolution of haploid organisms. This similarity has been recently exploited \cite{GA,GJ} to construct language trees. The key point is the definition of a distance among all pairs of…

Physics and Society · Physics 2009-11-13 Maurizio Serva , Filippo Petroni

Language Identification (LI) is crucial for various natural language processing tasks, serving as a foundational step in applications such as sentiment analysis, machine translation, and information retrieval. In multilingual societies like…

Computation and Language · Computer Science 2025-03-13 Aniket Deroy , Subhankar Maity

UniGlyph is a constructed language (conlang) designed to create a universal transliteration system using a script derived from seven-segment characters. The goal of UniGlyph is to facilitate cross-language communication by offering a…

Computation and Language · Computer Science 2024-10-14 G. V. Bency Sherin , A. Abijesh Euphrine , A. Lenora Moreen , L. Arun Jose

This study aims to develop a semi-automatically labelled prosody database for Hindi, for enhancing the intonation component in ASR and TTS systems, which is also helpful for building Speech to Speech Machine Translation systems. Although no…

Computation and Language · Computer Science 2021-12-14 Esha Banerjee , Atul Kr. Ojha , Girish Nath Jha

Historical palm-leaf manuscript and early paper documents from Indian subcontinent form an important part of the world's literary and cultural heritage. Despite their importance, large-scale annotated Indic manuscript image datasets do not…

Computer Vision and Pattern Recognition · Computer Science 2019-12-17 Abhishek Prusty , Sowmya Aitha , Abhishek Trivedi , Ravi Kiran Sarvadevabhatla

India has 1369 languages of which 22 are official. About 13 different scripts are used to represent these languages. A Common Label Set (CLS) was developed based on phonetics to address the issue of large vocabulary of units required in the…

Computation and Language · Computer Science 2025-02-24 Utkarsh P

Communication is defined as the act of sharing or exchanging information, ideas or feelings. To establish communication between two people, both of them are required to have knowledge and understanding of a common language. But in the case…

Computer Vision and Pattern Recognition · Computer Science 2022-03-04 Sharvani Srivastava , Amisha Gangwar , Richa Mishra , Sudhakar Singh

Tamil, a Dravidian language of South Asia, is a highly diglossic language with two very different registers in everyday use: Literary Tamil (preferred in writing and formal communication) and Spoken Tamil (confined to speech and informal…

Computation and Language · Computer Science 2023-11-15 Kabilan Prasanna , Aryaman Arora

Reduplication and repetition, though similar in form, serve distinct linguistic purposes. Reduplication is a deliberate morphological process used to express grammatical, semantic, or pragmatic nuances, while repetition is often…

Computation and Language · Computer Science 2024-07-12 Arif Ahmad , Mothika Gayathri Khyathi , Pushpak Bhattacharyya

Culture and language evolve together. The old literary form of Tamil is used commonly for writing and the contemporary colloquial Tamil is used for speaking. Human-computer interaction applications require Colloquial Tamil (CT) to make it…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-27 M. Nanmalar , P. Vijayalakshmi , T. Nagarajan

A Kannada OCR, named Lipi Gnani, has been designed and developed from scratch, with the motivation of it being able to convert printed text or poetry in Kannada script, without any restriction on vocabulary. The training and test sets have…

Computer Vision and Pattern Recognition · Computer Science 2019-01-03 Shiva Kumar H R , Ramakrishnan A G

This paper reports a preliminary study on quantitative frequency domain rhythm cues for classifying five Indian languages: Bengali, Kannada, Malayalam, Marathi, and Tamil. We employ rhythm formant (R-formants) analysis, a technique…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-10 Parismita Gogoi , Sishir Kalita , Priyankoo Sarmah , S. R Mahadeva Prasanna

The work presented here involves the design of a Multi Layer Perceptron (MLP) based pattern classifier for recognition of handwritten Bangla digits using a 76 element feature vector. Bangla is the second most popular script and language in…

Computer Vision and Pattern Recognition · Computer Science 2012-03-06 Subhadip Basu , Nibaran Das , Ram Sarkar , Mahantapas Kundu , Mita Nasipuri , Dipak Kumar Basu
‹ Prev 1 8 9 10 Next ›