English
Related papers

Related papers: Combining Deep Learning and String Kernels for the…

200 papers

In this work, we describe our approach addressing the Social Media Variety Geolocation task featured in the 2021 VarDial Evaluation Campaign. We focus on the second subtask, which is based on a data set formed of approximately 30 thousand…

Computation and Language · Computer Science 2021-03-02 Mihaela Gaman , Sebastian Cojocariu , Radu Tudor Ionescu

We propose a simple yet effective text- based user geolocation model based on a neural network with one hidden layer, which achieves state of the art performance over three Twitter benchmark geolocation datasets, in addition to producing…

Computation and Language · Computer Science 2017-04-28 Afshin Rahimi , Trevor Cohn , Timothy Baldwin

We present a machine learning approach that ranked on the first place in the Arabic Dialect Identification (ADI) Closed Shared Tasks of the 2018 VarDial Evaluation Campaign. The proposed approach combines several kernels using multiple…

Computation and Language · Computer Science 2018-07-31 Andrei M. Butnaru , Radu Tudor Ionescu

In this paper, we present a kernel-based learning approach for the 2018 Complex Word Identification (CWI) Shared Task. Our approach is based on combining multiple low-level features, such as character n-grams, with high-level semantic…

Computation and Language · Computer Science 2018-05-23 Andrei M. Butnaru , Radu Tudor Ionescu

Predicting the geographical location of users of social media like Twitter has found several applications in health surveillance, emergency monitoring, content personalization, and social studies in general. In this work we contribute to…

Social and Information Networks · Computer Science 2021-12-15 Federico M. Funes , José Ignacio Alvarez-Hamelin , Mariano G. Beiró

Deep learning mechanisms are prevailing approaches in recent days for the various tasks in natural language processing, speech recognition, image processing and many others. To leverage this we use deep learning based mechanism specifically…

Computation and Language · Computer Science 2019-01-03 Vidya Prasad K , Akarsh S , Vinayakumar R , Soman KP

This paper presents our approach for SwissText & KONVENS 2020 shared task 2, which is a multi-stage neural model for Swiss German (GSW) identification on Twitter. Our model outputs either GSW or non-GSW and is not meant to be used as a…

Computation and Language · Computer Science 2020-06-08 Mohammadreza Banaei , Rémi Lebret , Karl Aberer

Positive, supportive online communication in social media (candy speech) has the potential to foster civility, yet automated detection of such language remains underexplored, limiting systematic analysis of its impact. We investigate how…

Computation and Language · Computer Science 2025-09-17 Christian Rene Thelen , Patrick Gustav Blaneck , Tobias Bornheim , Niklas Grieger , Stephan Bialonski

We propose a deep learning approach for discovering kernels tailored to identifying clusters over sample data. Our neural network produces sample embeddings that are motivated by--and are at least as expressive as--spectral clustering. Our…

Machine Learning · Computer Science 2020-01-03 Chieh Wu , Zulqarnain Khan , Yale Chang , Stratis Ioannidis , Jennifer Dy

Recently, string kernels have obtained state-of-the-art results in various text classification tasks such as Arabic dialect identification or native language identification. In this paper, we apply two simple yet effective transductive…

Computation and Language · Computer Science 2018-09-03 Radu Tudor Ionescu , Andrei M. Butnaru

While Self-supervised Learning (SSL) has significantly improved Spoken Language Identification (LID), existing models often struggle to consistently classify dialects and accents of the same language as a unified class. To address this…

Computation and Language · Computer Science 2025-08-26 Qingzheng Wang , Hye-jin Shim , Jiancheng Sun , Shinji Watanabe

The problem of predicting the location of users on large social networks like Twitter has emerged from real-life applications such as social unrest detection and online marketing. Twitter user geolocation is a difficult and active research…

Machine Learning · Computer Science 2017-12-22 Tien Huu Do , Duc Minh Nguyen , Evaggelia Tsiligianni , Bruno Cornelis , Nikos Deligiannis

In this paper we present the GDI_classification entry to the second German Dialect Identification (GDI) shared task organized within the scope of the VarDial Evaluation Campaign 2018. We present a system based on SVM classifier ensembles…

Computation and Language · Computer Science 2018-07-24 Alina Maria Ciobanu , Shervin Malmasi , Liviu P. Dinu

For many text classification tasks, there is a major problem posed by the lack of labeled data in a target domain. Although classifiers for a target domain can be trained on labeled text data from a related source domain, the accuracy of…

Computation and Language · Computer Science 2018-11-06 Radu Tudor Ionescu , Andrei M. Butnaru

We propose a method for embedding two-dimensional locations in a continuous vector space using a neural network-based model incorporating mixtures of Gaussian distributions, presenting two model variants for text-based geolocation and…

Computation and Language · Computer Science 2017-08-16 Afshin Rahimi , Timothy Baldwin , Trevor Cohn

We describe a machine learning approach for the 2017 shared task on Native Language Identification (NLI). The proposed approach combines several kernels using multiple kernel learning. While most of our kernels are based on character…

Computation and Language · Computer Science 2017-08-07 Radu Tudor Ionescu , Marius Popescu

We introduce a neural network-based system of Word Sense Disambiguation (WSD) for German that is based on SenseFitting, a novel method for optimizing WSD. We outperform knowledge-based WSD methods by up to 25% F1-score and produce a new…

Computation and Language · Computer Science 2019-08-01 Manuel Stoeckel , Sajawel Ahmed , Alexander Mehler

Sequential hypothesis testing is a desirable decision making strategy in any time sensitive scenario. Compared with fixed sample-size testing, sequential testing is capable of achieving identical probability of error requirements using less…

Machine Learning · Statistics 2017-11-17 Diyan Teng , Emre Ertin

This research is aimed to solve the tweet/user geolocation prediction task and provide a flexible methodology for the geotagging of textual big data. The suggested approach implements neural networks for natural language processing (NLP) to…

Computation and Language · Computer Science 2025-01-13 Kateryna Lutsai , Christoph H. Lampert

In the last few years, microblogging platforms such as Twitter have given rise to a deluge of textual data that can be used for the analysis of informal communication between millions of individuals. In this work, we propose an…

Computation and Language · Computer Science 2021-11-17 Gonzalo Donoso , David Sanchez
‹ Prev 1 2 3 10 Next ›