English
Related papers

Related papers: IndicDLP: A Foundational Dataset for Multi-Lingual…

200 papers

The rapid advancement of large language models(LLMs) has intensified the need for domain and culture specific evaluation. Existing benchmarks are largely Anglocentric and domain-agnostic, limiting their applicability to India-centric…

Multi-domain learning (MDL) refers to simultaneously constructing a model or a set of models on datasets collected from different domains. Conventional approaches emphasize domain-shared information extraction and domain-private information…

Machine Learning · Computer Science 2023-07-31 Rui He , Shengcai Liu , Jiahao Wu , Shan He , Ke Tang

Recent advancements in large multimodal models (LMMs) have leveraged extensive multimodal datasets to enhance capabilities in complex knowledge-driven tasks. However, persistent challenges in perceptual and reasoning errors limit their…

Transliteration is very important in the Indian language context due to the usage of multiple scripts and the widespread use of romanized inputs. However, few training and evaluation sets are publicly available. We introduce Aksharantar,…

Computation and Language · Computer Science 2023-10-27 Yash Madhani , Sushane Parthan , Priyanka Bedekar , Gokul NC , Ruchi Khapra , Anoop Kunchukuttan , Pratyush Kumar , Mitesh M. Khapra

We introduce Multi-SimLex, a large-scale lexical resource and evaluation benchmark covering datasets for 12 typologically diverse languages, including major languages (e.g., Mandarin Chinese, Spanish, Russian) as well as less-resourced ones…

In the rapidly evolving digital era, there is an increasing demand for concise information as individuals seek to distil key insights from various sources. Recent attention from researchers on Multi-document Summarisation (MDS) has resulted…

Computation and Language · Computer Science 2024-09-19 Kushan Hewapathirana , Nisansa de Silva , C. D. Athuraliya

Advancements in Large Language Models (LLMs) have significantly enhanced instruction-following capabilities. However, most Instruction Fine-Tuning (IFT) datasets are predominantly in English, limiting model performance in other languages.…

Computation and Language · Computer Science 2024-07-03 Sathish Reddy Indurthi , Wenxuan Zhou , Shamil Chollampatt , Ravi Agrawal , Kaiqiang Song , Lingxiao Zhao , Chenguang Zhu

Acronym extraction is the task of identifying acronyms and their expanded forms in texts that is necessary for various NLP applications. Despite major progress for this task in recent years, one limitation of existing AE research is that…

Computation and Language · Computer Science 2022-02-22 Amir Pouran Ben Veyseh , Nicole Meister , Seunghyun Yoon , Rajiv Jain , Franck Dernoncourt , Thien Huu Nguyen

Scientific literature contain important information related to cutting-edge innovations in diverse domains. Advances in natural language processing have been driving the fast development in automated information extraction from scientific…

Information Retrieval · Computer Science 2021-06-29 Antonio Jimeno Yepes , Xu Zhong , Douglas Burdick

Transforming documents into machine-processable representations is a challenging task due to their complex structures and variability in formats. Recovering the layout structure and content from PDF files or scanned material has remained a…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Christoph Auer , Ahmed Nassar , Maksym Lysak , Michele Dolfi , Nikolaos Livathinos , Peter Staar

Legal multi-label classification is a critical task for organizing and accessing the vast amount of legal documentation. Despite its importance, it faces challenges such as the complexity of legal language, intricate label dependencies, and…

Computation and Language · Computer Science 2025-04-15 Emily Johnson , Xavier Holt , Noah Wilson

India's linguistic landscape, spanning 22 scheduled languages and hundreds of marginalized dialects, has driven rapid growth in NLP datasets, benchmarks, and pretrained models. However, no dedicated survey consolidates resources developed…

Computation and Language · Computer Science 2026-04-21 Raghvendra Kumar , Devankar Raj , Sriparna Saha

The pervasive influence of social biases in language data has sparked the need for benchmark datasets that capture and evaluate these biases in Large Language Models (LLMs). Existing efforts predominantly focus on English language and the…

Computation and Language · Computer Science 2024-04-04 Nihar Ranjan Sahoo , Pranamya Prashant Kulkarni , Narjis Asad , Arif Ahmad , Tanu Goyal , Aparna Garimella , Pushpak Bhattacharyya

We introduce Docling, an easy-to-use, self-contained, MIT-licensed, open-source toolkit for document conversion, that can parse several types of popular document formats into a unified, richly structured representation. It is powered by…

Datasets are foundational to many breakthroughs in modern artificial intelligence. Many recent achievements in the space of natural language processing (NLP) can be attributed to the finetuning of pre-trained models on a diverse set of…

We present an ongoing initiative to provide open, very large, high-quality, and richly annotated textual datasets for almost 200 languages. At 30 trillion tokens, this is likely the largest generally available multilingual collection of LLM…

Digitization of newspapers is of interest for many reasons including preservation of history, accessibility and search ability, etc. While digitization of documents such as scientific articles and magazines is prevalent in literature, one…

Computer Vision and Pattern Recognition · Computer Science 2022-02-04 Wenzhen Zhu , Negin Sokhandan , Guang Yang , Sujitha Martin , Suchitra Sathyanarayana

We introduce INDOTABVQA, a benchmark for evaluating cross-lingual Table Visual Question Answering (VQA) on real-world document images in Bahasa Indonesia. The dataset comprises 1,593 document images across three visual styles (bordered,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Somraj Gautam , Anathapindika Dravichi , Gaurav Harit

The exponential growth of scientific literature in PDF format necessitates advanced tools for efficient and accurate document understanding, summarization, and content optimization. Traditional methods fall short in handling complex layouts…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Kun Qian , Wenjie Li , Tianyu Sun , Wenhong Wang , Wenhan Luo