English
Related papers

Related papers: BHAAV- A Text Corpus for Emotion Analysis from Hin…

200 papers

Human language is colored by a broad range of topics, but existing text analysis tools only focus on a small number of them. We present Empath, a tool that can generate and validate new lexical categories on demand from a small set of seed…

Computation and Language · Computer Science 2016-02-24 Ethan Fast , Binbin Chen , Michael Bernstein

Work in Computational Affective Science and Computational Social Science explores a wide variety of research questions about people, emotions, behavior, and health. Such work often relies on language data that is first labeled with relevant…

Computation and Language · Computer Science 2026-04-03 Jan Philip Wahle , Krishnapriya Vishnubhotla , Bela Gipp , Saif M. Mohammad

This paper presents StoryDB - a broad multi-language dataset of narratives. StoryDB is a corpus of texts that includes stories in 42 different languages. Every language includes 500+ stories. Some of the languages include more than 20 000…

Computation and Language · Computer Science 2022-11-15 Alexey Tikhonov , Igor Samenko , Ivan P. Yamshchikov

In this paper, we discuss the development of a multilingual annotated corpus of misogyny and aggression in Indian English, Hindi, and Indian Bangla as part of a project on studying and automatically identifying misogyny and communalism on…

Computation and Language · Computer Science 2020-03-18 Shiladitya Bhattacharya , Siddharth Singh , Ritesh Kumar , Akanksha Bansal , Akash Bhagat , Yogesh Dawer , Bornini Lahiri , Atul Kr. Ojha

Exponential growths of social media and micro-blogging sites not only provide platforms for empowering freedom of expressions and individual voices but also enables people to express anti-social behaviour like online harassment,…

Computation and Language · Computer Science 2020-04-21 Md. Rezaul Karim , Bharathi Raja Chakravarthi , John P. McCrae , Michael Cochez

The development of robust language models for low-resource languages is frequently bottlenecked by the scarcity of high-quality, coherent, and domain-appropriate training corpora. In this paper, we introduce the Multilingual TinyStories…

Computation and Language · Computer Science 2026-03-17 Deepon Halder , Angira Mukherjee

It is important for machines to interpret human emotions properly for better human-machine communications, as emotion is an essential part of human-to-human communications. One aspect of emotion is reflected in the language we use. How to…

Computation and Language · Computer Science 2018-08-23 Ji Ho Park

Reading scene text, that is, text appearing in images, has numerous application areas, including assistive technology, search, and e-commerce. Although scene text recognition in English has advanced significantly and is often considered…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Anik De , Abhirama Subramanyam Penamakuri , Rajeev Yadav , Aditya Rathore , Harshiv Shah , Devesh Sharma , Sagar Agarwal , Pravin Kumar , Anand Mishra

We present our shared task on text-based emotion detection, covering more than 30 languages from seven distinct language families. These languages are predominantly low-resource and are spoken across various continents. The data instances…

We describe EmoBank, a corpus of 10k English sentences balancing multiple genres, which we annotated with dimensional emotion metadata in the Valence-Arousal-Dominance (VAD) representation format. EmoBank excels with a bi-perspectival and…

Computation and Language · Computer Science 2022-05-05 Sven Buechel , Udo Hahn

The availability of large, high-quality emotional speech databases is essential for advancing speech emotion recognition (SER) in real-world scenarios. However, many existing databases face limitations in size, emotional balance, and…

It is well known that translations of songs and poems not only break rhythm and rhyming patterns, but can also result in loss of semantic information. The Bhagavad Gita is an ancient Hindu philosophical text originally written in Sanskrit…

Computation and Language · Computer Science 2022-02-16 Rohitash Chandra , Venkatesh Kulkarni

News media often shape the public mood not only by what they report but by how they frame it. The same event can appear calm in one outlet and alarming in another, reflecting subtle emotional bias in reporting. Negative or emotionally…

Computation and Language · Computer Science 2025-10-21 Mohd Ruhul Ameen , Akif Islam , Abu Saleh Musa Miah , Ayesha Siddiqua , Jungpil Shin

Availability of naturalistic affective stimuli is needed for creating the affective technological solution as well as making progress in affective science. Although a lot of progress in the collection of affective multimedia stimuli has…

Image and Video Processing · Electrical Eng. & Systems 2022-10-19 Sudhakar Mishra , Narayanan Srinivasan , Uma Shanker Tiwary

This research introduces a bilingual dataset comprising 23,456 entries for Arabic and 10,036 entries for English, annotated for emotions and hope speech, addressing the scarcity of multi-emotion (Emotion and hope) datasets. The dataset…

Computation and Language · Computer Science 2025-05-22 Wajdi Zaghouani , Md. Rafiul Biswas

We present the IIT Bombay English-Hindi Parallel Corpus. The corpus is a compilation of parallel corpora previously available in the public domain as well as new parallel corpora we collected. The corpus contains 1.49 million parallel…

Computation and Language · Computer Science 2018-05-22 Anoop Kunchukuttan , Pratik Mehta , Pushpak Bhattacharyya

This paper describes the development of a multilingual, manually annotated dataset for three under-resourced Dravidian languages generated from social media comments. The dataset was annotated for sentiment analysis and offensive language…

This paper presents the InScript corpus (Narrative Texts Instantiating Script structure). InScript is a corpus of 1,000 stories centered around 10 different scenarios. Verbs and noun phrases are annotated with event and participant types,…

Computation and Language · Computer Science 2017-03-16 Ashutosh Modi , Tatjana Anikina , Simon Ostermann , Manfred Pinkal

We present sentence aligned parallel corpora across 10 Indian Languages - Hindi, Telugu, Tamil, Malayalam, Gujarati, Urdu, Bengali, Oriya, Marathi, Punjabi, and English - many of which are categorized as low resource. The corpora are…

Computation and Language · Computer Science 2020-07-16 Shashank Siripragada , Jerin Philip , Vinay P. Namboodiri , C V Jawahar

The detection of negative emotions through daily activities such as handwriting is useful for promoting well-being. The spread of human-machine interfaces such as tablets makes the collection of handwriting samples easier. In this context,…

Computer Vision and Pattern Recognition · Computer Science 2022-02-25 Laurence Likforman-Sulem , Anna Esposito , Marcos Faundez-Zanuy , Stephan Clemençon , Gennaro Cordasco