English
Related papers

Related papers: Using n-aksaras to model Sanskrit and Sanskrit-adj…

200 papers

The rendering of Sanskrit poetry from text to speech is a problem that has not been solved before. One reason may be the complications in the language itself. We present unique algorithms based on extensive empirical analysis, to synthesize…

Computation and Language · Computer Science 2014-09-16 Rama N. , Meenakshi Lakshmanan

In recent advances in automatic text recognition (ATR), deep neural networks have demonstrated the ability to implicitly capture language statistics, potentially reducing the need for traditional language models. This study directly…

Computer Vision and Pattern Recognition · Computer Science 2024-05-01 Solène Tarride , Christopher Kermorvant

Maximizing the likelihood of the next token is an established, statistically sound objective for pre-training language models. In this paper we show that we can train better models faster by pre-aggregating the corpus with a collapsed…

Computation and Language · Computer Science 2024-07-04 Ashutosh Sathe , Sunita Sarawagi

Text-based foundation models have become an important part of scientific discovery, with molecular foundation models accelerating advancements in material science and molecular design.However, existing models are constrained by…

Machine Learning · Computer Science 2026-01-29 Alexius Wadell , Anoushka Bhutani , Venkatasubramanian Viswanathan

Data-driven approaches for dependency parsing have been of great interest in Natural Language Processing for the past couple of decades. However, Sanskrit still lacks a robust purely data-driven dependency parser, probably with an exception…

Computation and Language · Computer Science 2020-04-20 Amrith Krishna , Ashim Gupta , Deepak Garasangi , Jivnesh Sandhan , Pavankumar Satuluri , Pawan Goyal

Searching for words in Sanskrit E-text is a problem that is accompanied by complexities introduced by features of Sanskrit such as euphonic conjunctions or sandhis. A word could occur in an E-text in a transformed form owing to the…

Computation and Language · Computer Science 2014-09-16 S. V. Kasmir Raja , V. Rajitha , Meenakshi Lakshmanan

In Sanskrit, small words (morphemes) are combined to form compound words through a process known as Sandhi. Sandhi splitting is the process of splitting a given compound word into its constituent morphemes. Although rules governing word…

Computation and Language · Computer Science 2019-07-16 Rahul Aralikatte , Neelamadhav Gantayat , Naveen Panwar , Anush Sankaran , Senthil Mani

This study addresses the problem of authorship attribution for Romanian texts using the ROST corpus, a standard benchmark in the field. We systematically evaluate six machine learning techniques: Support Vector Machine (SVM), Logistic…

Computation and Language · Computer Science 2025-06-30 Dana Lupsa , Sanda-Maria Avram , Radu Lupsa

Recent advancements in recurrent neural networks (RNNs) have reinvigorated interest in their application to natural language processing tasks, particularly with the development of more efficient and parallelizable variants known as state…

Computation and Language · Computer Science 2025-03-11 Vinoth Nandakumar , Qiang Qu , Peng Mi , Tongliang Liu

We present Chandoj\~n\=anam, a web-based Sanskrit meter (Chanda) identification and utilization system. In addition to the core functionality of identifying meters, it sports a friendly user interface to display the scansion, which is a…

Software Engineering · Computer Science 2023-10-13 Hrishikesh Terdalkar , Arnab Bhattacharya

This study demonstrates how hybrid neural-symbolic methods can yield significant new insights into the evolution of a morphologically rich, low-resource language. We challenge the naive assumption that linguistic change is simplification by…

Computation and Language · Computer Science 2025-12-08 Ananth Hariharan , David Mortensen

In multimedia, text or bioinformatics databases, applications query sequences of n consecutive symbols called n-grams. Estimating the number of distinct n-grams is a view-size estimation problem. While view sizes can be estimated by…

Databases · Computer Science 2014-02-05 Daniel Lemire , Owen Kaser

Beyond traditional binary relational facts, n-ary relational knowledge graphs (NKGs) are comprised of n-ary relational facts containing more than two entities, which are closer to real-world facts with broader applications. However, the…

Artificial Intelligence · Computer Science 2024-10-31 Haoran Luo , Haihong E , Yuhao Yang , Tianyu Yao , Yikai Guo , Zichen Tang , Wentai Zhang , Kaiyang Wan , Shiyao Peng , Meina Song , Wei Lin , Yifan Zhu , Luu Anh Tuan

Statistical language models are powerful tools which have been used for many tasks within natural language processing. Recently, they have been used for other sequential data such as source code.(Ray et al., 2015) showed that it is possible…

Software Engineering · Computer Science 2018-03-26 Jack Lanchantin , Ji Gao

The anusaaraka system (a kind of machine translation system) makes text in one Indian language accessible through another Indian language. The machine presents an image of the source text in a language close to the target language. In the…

Computation and Language · Computer Science 2007-05-23 Akshar Bharati , Vineet Chaitanya , Amba P. Kulkarni , Rajeev Sangal

Constituency parsing is a fundamental and important task for natural language understanding, where a good representation of contextual information can help this task. N-grams, which is a conventional type of feature for contextual…

Computation and Language · Computer Science 2020-10-16 Yuanhe Tian , Yan Song , Fei Xia , Tong Zhang

We propose a new approach to text semantic analysis and general corpus analysis using, as termed in this article, a "bi-gram graph" representation of a corpus. The different attributes derived from graph theory are measured and analyzed as…

Machine Learning · Computer Science 2021-07-30 Thomas Konstantinovsky , Matan Mizrachi

We present a neural Sanskrit Natural Language Processing (NLP) toolkit named SanskritShala (a school of Sanskrit) to facilitate computational linguistic analyses for several tasks such as word segmentation, morphological tagging, dependency…

Computation and Language · Computer Science 2023-05-30 Jivnesh Sandhan , Anshul Agarwal , Laxmidhar Behera , Tushar Sandhan , Pawan Goyal

Next word prediction is an input technology that simplifies the process of typing by suggesting the next word to a user to select, as typing in a conversation consumes time. A few previous studies have focused on the Kurdish language,…

Computation and Language · Computer Science 2020-08-05 Hozan K. Hamarashid , Soran A. Saeed , Tarik A. Rashid

The presence of sarcasm in conversational systems and social media like chatbots, Facebook, Twitter, etc. poses several challenges for downstream NLP tasks. This is attributed to the fact that the intended meaning of a sarcastic text is…

Computation and Language · Computer Science 2022-02-08 Aditya Shah , Chandresh Kumar Maurya