English
Related papers

Related papers: Kannudi -- A Reference Editor for Kannada

200 papers

Extracting Handwritten text is one of the most important components of digitizing information and making it available for large scale setting. Handwriting Optical Character Reader (OCR) is a research problem in computer vision and natural…

Computer Vision and Pattern Recognition · Computer Science 2022-06-10 Mohammad Daniyal Shaiq , Musa Dildar Ahmed Cheema , Ali Kamal

We propose Object-oriented Neural Programming (OONP), a framework for semantically parsing documents in specific domains. Basically, OONP reads a document and parses it into a predesigned object-oriented data structure (referred to as…

Machine Learning · Computer Science 2018-07-26 Zhengdong Lu , Xianggen Liu , Haotian Cui , Yukun Yan , Daqi Zheng

Hybrid Retrieval systems, combining Sparse and Dense Retrieval methods, struggle with Traditional Chinese non-narrative documents due to their complex formatting, rich vocabulary, and the insufficient understanding of Chinese synonyms by…

Information Retrieval · Computer Science 2025-05-02 Hsin-Ling Hsu , Ping-Sheng Lin , Jing-Di Lin , Jengnan Tzeng

Knowledge distillation (KD) is a widely adopted technique for transferring knowledge from large language models to smaller student models; however, conventional supervised KD often suffers from a distribution mismatch between training and…

Machine Learning · Computer Science 2026-04-21 Ijun Jang , Jewon Yeom , Juan Yeo , Hyunggu Lim , Taesup Kim

Speech Translation has always been about giving source text or audio input and waiting for system to give translated output in desired form. In this paper, we present the Acoustic Dialect Decoder (ADD) - a voice to voice ear-piece…

Computation and Language · Computer Science 2016-12-01 Hans Krupakar , Keerthika Rajvel , Bharathi B , Angel Deborah S , Vallidevi Krishnamurthy

Ontology interoperability is one of the complicated issues that restricts the use of ontologies in knowledge graphs (KGs). Different ontologies with conflicting and overlapping concepts make it difficult to design, develop, and deploy an…

Information Retrieval · Computer Science 2026-03-23 Zhangcheng Qiang

Drawing inspiration from prompt tuning techniques applied to Large Language Models, recent methods based on pre-trained ViT networks have achieved remarkable results in the field of Continual Learning. Specifically, these approaches propose…

Machine Learning · Computer Science 2024-11-22 Quyen Tran , Hoang Phan , Lam Tran , Khoat Than , Toan Tran , Dinh Phung , Trung Le

With the surging inclination towards carrying out tasks on computational devices and digital mediums, any method that converts a task that was previously carried out manually, to a digitized version, is always welcome. Irrespective of the…

Computer Vision and Pattern Recognition · Computer Science 2023-04-19 Pranav Guruprasad , Sujith Kumar S , Vigneswaran C , V. Srinivasa Chakravarthy

Automated language processing is central to the drive to enable facilitated referencing of increasingly available Sanskrit E texts. The first step towards processing Sanskrit text involves the handling of Sanskrit compound words that are an…

Computation and Language · Computer Science 2009-11-05 N. Rama , Meenakshi Lakshmanan

These are the guidelines for the application of SNACS (Semantic Network of Adposition and Case Supersenses; Schneider et al. 2018) to Modern Standard Hindi of Delhi. SNACS is an inventory of 50 supersenses (semantic labels) for labelling…

Computation and Language · Computer Science 2021-03-03 Aryaman Arora , Nitin Venkateswaran , Nathan Schneider

Out-of-vocabulary (OOV) words can pose serious challenges for machine translation (MT) tasks, and in particular, for low-resource language (LRL) pairs, i.e., language pairs for which few or no parallel corpora exist. Our work adapts…

Computation and Language · Computer Science 2021-04-21 Saurav Jha , Akhilesh Sudhakar , Anil Kumar Singh

The annotation of music content is a complex process to represent due to its inherent multifaceted, subjectivity, and interdisciplinary nature. Numerous systems and conventions for annotating music have been developed as independent…

Artificial Intelligence · Computer Science 2023-04-04 Jacopo de Berardinis , Albert Meroño-Peñuela , Andrea Poltronieri , Valentina Presutti

We propose Okapi, a simple, efficient, and general method for robust semi-supervised learning based on online statistical matching. Our method uses a nearest-neighbours-based matching procedure to generate cross-domain views for a…

Computer Vision and Pattern Recognition · Computer Science 2022-11-11 Myles Bartlett , Sara Romiti , Viktoriia Sharmanska , Novi Quadrianto

Knowledge distillation (KD) is a key technique for compressing large-scale language models (LLMs), yet prevailing logit-based methods typically employ static strategies that are misaligned with the dynamic learning process of student…

Computation and Language · Computer Science 2025-10-14 Xurong Xie , Zhucun Xue , Jiafu Wu , Jian Li , Yabiao Wang , Xiaobin Hu , Yong Liu , Jiangning Zhang

ICD Coding aims to assign a wide range of medical codes to a medical text document, which is a popular and challenging task in the healthcare domain. To alleviate the problems of long-tail distribution and the lack of annotations of…

Computation and Language · Computer Science 2025-05-27 Xu Zhang , Kun Zhang , Wenxin Ma , Rongsheng Wang , Chenxu Wu , Yingtai Li , S. Kevin Zhou

Ontohub is a repository engine for managing distributed heterogeneous ontologies. The distributed nature enables communities to share and exchange their contributions easily. The heterogeneous nature makes it possible to integrate…

Artificial Intelligence · Computer Science 2016-12-16 Mihai Codescu , Eugen Kuksa , Oliver Kutz , Till Mossakowski , Fabian Neuhaus

The most commonly used Japanese alphabets are Kanji, Hiragana and Katakana. The Kanji alphabet includes pictographs or ideographic characters that were adopted from the Chinese alphabet. Hiragana is used to spell words of Japanese origin,…

Human-Computer Interaction · Computer Science 2013-10-14 Umakant Mishra

Long-context multiple-choice question answering tasks require robust reasoning over extensive text sources. Since most of the pre-trained transformer models are restricted to processing only a few hundred words at a time, successful…

Information Retrieval · Computer Science 2025-01-28 Manish Singh , Manish Shrivastava

The intelligent robotics community usually organizes knowledge into symbolic and sub-symbolic levels. These two levels establish the set of symbols and rules for manipulating knowledge based on their (symbol system - dictionary). Thus, the…

Knowledge Distillation (KD) is extensively used to compress and deploy large pre-trained language models on edge devices for real-world applications. However, one neglected area of research is the impact of noisy (corrupted) labels on KD.…

Computation and Language · Computer Science 2021-09-22 Shivendra Bhardwaj , Abbas Ghaddar , Ahmad Rashid , Khalil Bibi , Chengyang Li , Ali Ghodsi , Philippe Langlais , Mehdi Rezagholizadeh