English
Related papers

Related papers: Indus script corpora, archaeo-metallurgy and Meluh…

200 papers

The Houma Alliance Book, one of history's earliest calligraphic examples, was unearthed in the 1970s. These artifacts were meticulously organized, reproduced, and copied by the Shanxi Provincial Institute of Cultural Relics. However,…

Computer Vision and Pattern Recognition · Computer Science 2024-08-27 Xiaoyu Yuan , Xiaohua Huang , Zibo Zhang , Yabo Sun

Despite 230 million speakers, Urdu remains critically under-resourced in speech technology. We introduce UrduSpeech: a large high-fidelity Urdu corpus comprising 156 hours of audio with 12-dimension paralinguistic metadata, encompassing…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-19 Attia Nafees ul Haq , Zeyu Zhu , Jingbin Hu , ChunJiang He , Lei Xie

Large Language Models (LLMs) are now capable of generating text that closely resembles human writing, making them powerful tools for content creation, but this growing ability has also made it harder to tell whether a piece of text was…

Computation and Language · Computer Science 2025-10-21 Muhammad Ammar , Hadiya Murad Hadi , Usman Majeed Butt

Originating from China's Shang Dynasty approximately 3,000 years ago, the Oracle Bone Script (OBS) is a cornerstone in the annals of linguistic history, predating many established writing systems. Despite the discovery of thousands of…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Haisu Guan , Huanxin Yang , Xinyu Wang , Shengwei Han , Yongge Liu , Lianwen Jin , Xiang Bai , Yuliang Liu

Dictionaries are essence of any language providing vital linguistic recourse for the language learners, researchers and scholars. This paper focuses on the methodology and techniques used in developing software architecture for a UBSESD…

Computation and Language · Computer Science 2014-01-14 Imdad Ali Ismaili , Zeeshan Bhatti , Azhar Ali Shah

This review paper provides a comprehensive overview of large language model (LLM) research directions within Indic languages. Indic languages are those spoken in the Indian subcontinent, including India, Pakistan, Bangladesh, Sri Lanka,…

Computation and Language · Computer Science 2024-06-17 Sankalp KJ , Vinija Jain , Sreyoshi Bhaduri , Tamoghna Roy , Aman Chadha

Gestural language is used by deaf & mute communities to communicate through hand gestures & body movements that rely on visual-spatial patterns known as sign languages. Sign languages, which rely on visual-spatial patterns of hand gestures…

Computer Vision and Pattern Recognition · Computer Science 2025-08-01 Rajat Singhal , Jatin Gupta , Akhil Sharma , Anushka Gupta , Navya Sharma

This paper presents first benchmark corpus of Sanskrit Pratyaya (suffix) and inflectional words (padas) formed due to suffixes along with neural network based approaches to process the formation and splitting of inflectional words.…

Computation and Language · Computer Science 2024-09-05 Arun Kumar Singh , Sushant Dave , Prathosh A. P. , Brejesh Lall , Shresth Mehta

A line of a bilingual document page may contain text words in regional language and numerals in English. For Optical Character Recognition (OCR) of such a document page, it is necessary to identify different script forms before running an…

Computer Vision and Pattern Recognition · Computer Science 2011-07-05 B. V. Dhandra , Mallikarjun Hangarge

In this age of information technology, information access in a convenient manner has gained importance. Since speech is a primary mode of communication among human beings, it is natural for people to expect to be able to carry out spoken…

Computation and Language · Computer Science 2013-05-14 Neema Mishra , Urmila Shrawankar , V M Thakare

Script identification plays a vital role in applications that involve handwriting and document analysis within a multi-script and multi-lingual environment. Moreover, it exhibits a profound connection with human cognition. This paper…

Computer Vision and Pattern Recognition · Computer Science 2024-05-30 Miguel A. Ferrer , Abhijit Das , Moises Diaz , Aythami Morales , Cristina Carmona-Duarte , Umapada Pal

There are a lot of intensive researches on handwritten character recognition (HCR) for almost past four decades. The research has been done on some of popular scripts such as Roman, Arabic, Chinese and Indian. In this paper we present a…

Computer Vision and Pattern Recognition · Computer Science 2013-08-28 Aini Najwa Azmi , Dewi Nasien , Siti Mariyam Shamsuddin

Understanding humanity's earliest writing systems is crucial for reconstructing civilization's origins, yet many ancient scripts remain undeciphered. Oracle Bone Script (OBS) from China's Shang dynasty exemplifies this challenge: only…

Information Retrieval · Computer Science 2026-04-14 Yin Wu , Gangjian Zhang , Jiayu Chen , Chang Xu , Yuyu Luo , Nan Tang , Hui Xiong

Sign Language Recognition is one of the most growing fields of research today. Many new techniques have been developed recently in these fields. Here in this paper, we have proposed a system using Eigen value weighted Euclidean distance as…

Computer Vision and Pattern Recognition · Computer Science 2013-03-05 Joyeeta Singha , Karen Das

It is well known that translations of songs and poems not only break rhythm and rhyming patterns, but can also result in loss of semantic information. The Bhagavad Gita is an ancient Hindu philosophical text originally written in Sanskrit…

Computation and Language · Computer Science 2022-02-16 Rohitash Chandra , Venkatesh Kulkarni

We argue that there was a link between Indus Valley India and the Mayans of Central America which is brought out by astronomical references. The former used a Jovian calendar while the latter had perfected a calendar based on Venus. This…

General Physics · Physics 2007-05-23 B. G. Sidharth

Large Language Models (LLMs) have made significant progress in incorporating Indic languages within multilingual models. However, it is crucial to quantitatively assess whether these languages perform comparably to globally dominant ones,…

Computation and Language · Computer Science 2024-10-31 Pritika Rohera , Chaitrali Ginimav , Akanksha Salunke , Gayatri Sawant , Raviraj Joshi

We present the IndicNLP corpus, a large-scale, general-domain corpus containing 2.7 billion words for 10 Indian languages from two language families. We share pre-trained word embeddings trained on these corpora. We create news article…

Computation and Language · Computer Science 2020-05-04 Anoop Kunchukuttan , Divyanshu Kakwani , Satish Golla , Gokul N. C. , Avik Bhattacharyya , Mitesh M. Khapra , Pratyush Kumar

Monument classification can be performed on the basis of their appearance and shape from coarse to fine categories. Although there is much semantic information present in the monuments which is reflected in the eras they were built, its…

Multimedia · Computer Science 2020-09-01 Ronak Gupta , Prerana Mukherjee , Brejesh Lall , Varshul Gupta

Language identification is used as the first step in many data collection and crawling efforts because it allows us to sort online text into language-specific buckets. However, many modern languages, such as Konkani, Kashmiri, Punjabi etc.,…

Computation and Language · Computer Science 2024-06-27 Milind Agarwal , Joshua Otten , Antonios Anastasopoulos