English
Related papers

Related papers: Development of POS tagger for English-Bengali Code…

200 papers

The sentiment analysis task in Tamil-English code-mixed texts has been explored using advanced transformer-based models. Challenges from grammatical inconsistencies, orthographic variations, and phonetic ambiguities have been addressed. The…

Computation and Language · Computer Science 2025-04-01 Mikhail Krasitskii , Olga Kolesnikova , Liliana Chanona Hernandez , Grigori Sidorov , Alexander Gelbukh

Code-mixing is a well-studied linguistic phenomenon when two or more languages are mixed in text or speech. Several datasets have been build with the goal of training computational models for code-mixing. Although it is very common to…

Computation and Language · Computer Science 2023-11-30 Md Nishat Raihan , Dhiman Goswami , Antara Mahmud , Antonios Anastasopoulos , Marcos Zampieri

Social media often serves as a breeding ground for various hateful and offensive content. Identifying such content on social media is crucial due to its impact on the race, gender, or religion in an unprejudiced society. However, while…

Computation and Language · Computer Science 2022-10-10 Mithun Das , Somnath Banerjee , Punyajoy Saha , Animesh Mukherjee

Part-of-Speech (POS) tagging is an important component of the NLP pipeline, but many low-resource languages lack labeled data for training. An established method for training a POS tagger in such a scenario is to create a labeled training…

Computation and Language · Computer Science 2022-11-01 Ayyoob Imani , Silvia Severini , Masoud Jalili Sabet , François Yvon , Hinrich Schütze

Sentiment Analysis is the process of deciphering what a sentence emotes and classifying them as either positive, negative, or neutral. In recent times, India has seen a huge influx in the number of active social media users and this has led…

Computation and Language · Computer Science 2020-09-07 Subhra Jyoti Baroi , Nivedita Singh , Ringki Das , Thoudam Doren Singh

Mixed language data is one of the difficult yet less explored domains of natural language processing. Most research in fields like machine translation or sentiment analysis assume monolingual input. However, people who are capable of using…

Neural and Evolutionary Computing · Computer Science 2014-12-23 Joseph Chee Chang , Chu-Cheng Lin

This study proposes a language-agnostic transformer-based POS tagging framework designed for low-resource languages, using Bangla and Hindi as case studies. With only three lines of framework-specific code, the model was adapted from Bangla…

Computation and Language · Computer Science 2025-12-02 Md Abdullah Al Kafi , Sumit Kumar Banshal

We present an algorithm that automatically learns context constraints using statistical decision trees. We then use the acquired constraints in a flexible POS tagger. The tagger is able to use information of any degree: n-grams,…

cmp-lg · Computer Science 2008-02-03 Lluis Marquez , Lluis Padro

In this paper, we give an overview for the shared task at the 4th CCF Conference on Natural Language Processing \& Chinese Computing (NLPCC 2015): Chinese word segmentation and part-of-speech (POS) tagging for micro-blog texts. Different…

Computation and Language · Computer Science 2015-07-01 Xipeng Qiu , Peng Qian , Liusong Yin , Shiyu Wu , Xuanjing Huang

Building a multilingual Automated Speech Recognition (ASR) system in a linguistically diverse country like India can be a challenging task due to the differences in scripts and the limited availability of speech data. This problem can be…

Computation and Language · Computer Science 2023-06-01 Kaousheik Jayakumar , Vrunda N. Sukhadia , A Arunkumar , S. Umesh

Part-of-speech (POS) taggers for low-resource languages which are exclusively based on various forms of weak supervision - e.g., cross-lingual transfer, type-level supervision, or a combination thereof - have been reported to perform almost…

Computation and Language · Computer Science 2020-04-29 Katharina Kann , Ophélie Lacroix , Anders Søgaard

Using code-mixed data in natural language processing (NLP) research currently gets a lot of attention. Language identification of social media code-mixed text has been an interesting problem of study in recent years due to the advancement…

Computation and Language · Computer Science 2022-11-29 Atnafu Lambebo Tonja , Mesay Gemeda Yigezu , Olga Kolesnikova , Moein Shahiki Tash , Grigori Sidorov , Alexander Gelbuk

Bangla-English code-mixing is widespread across South Asian social media, yet resources for implicit meaning identification in this setting remain scarce. Existing sentiment and sarcasm models largely focus on monolingual English or…

Computation and Language · Computer Science 2026-02-26 Kazi Samin Yasar Alam , Md Tanbir Chowdhury , Tamim Ahmed , Ajwad Abrar , Md Rafid Haque

Offensive content moderation is vital in social media platforms to support healthy online discussions. However, their prevalence in codemixed Dravidian languages is limited to classifying whole comments without identifying part of it…

Social networking platforms provide a conduit to disseminate our ideas, views and thoughts and proliferate information. This has led to the amalgamation of English with natively spoken languages. Prevalence of Hindi-English code-mixed data…

Computation and Language · Computer Science 2021-05-12 Ananya Srivastava , Mohammed Hasan , Bhargav Yagnik , Rahee Walambe , Ketan Kotecha

Automatic text tagging is an important component in higher level analysis of text corpora, and its output can be used in many natural language processing applications. In languages like Turkish or Finnish, with agglutinative morphology,…

cmp-lg · Computer Science 2008-02-03 Kemal Oflazer , Ilker Kuruoz

In this paper, we describe a research method that generates Bangla word clusters on the basis of relating to meaning in language and contextual similarity. The importance of word clustering is in parts of speech (POS) tagging, word sense…

Computation and Language · Computer Science 2017-01-31 Dipaloke Saha , Md Saddam Hossain , MD. Saiful Islam , Sabir Ismail

Cross-lingual embeddings represent the meaning of words from different languages in the same vector space. Recent work has shown that it is possible to construct such representations by aligning independently learned monolingual embedding…

Identifying the language of social media messages is an important first step in linguistic processing. Existing models for Twitter focus on content analysis, which is successful for dissimilar language pairs. We propose a label propagation…

Computation and Language · Computer Science 2016-07-20 Will Radford , Matthias Galle

Sense tagging, the automatic assignment of the appropriate sense from some lexicon to each of the words in a text, is a specialised instance of the general problem of semantic tagging by category or type. We discuss which recent word sense…

cmp-lg · Computer Science 2008-02-03 Yorick Wilks , Mark Stevenson
‹ Prev 1 3 4 5 6 7 10 Next ›