中文
相关论文

相关论文: Combining Deep Learning and String Kernels for the…

200 篇论文

In this work, we describe our approach addressing the Social Media Variety Geolocation task featured in the 2021 VarDial Evaluation Campaign. We focus on the second subtask, which is based on a data set formed of approximately 30 thousand…

计算与语言 · 计算机科学 2021-03-02 Mihaela Gaman , Sebastian Cojocariu , Radu Tudor Ionescu

We propose a simple yet effective text- based user geolocation model based on a neural network with one hidden layer, which achieves state of the art performance over three Twitter benchmark geolocation datasets, in addition to producing…

计算与语言 · 计算机科学 2017-04-28 Afshin Rahimi , Trevor Cohn , Timothy Baldwin

We present a machine learning approach that ranked on the first place in the Arabic Dialect Identification (ADI) Closed Shared Tasks of the 2018 VarDial Evaluation Campaign. The proposed approach combines several kernels using multiple…

计算与语言 · 计算机科学 2018-07-31 Andrei M. Butnaru , Radu Tudor Ionescu

In this paper, we present a kernel-based learning approach for the 2018 Complex Word Identification (CWI) Shared Task. Our approach is based on combining multiple low-level features, such as character n-grams, with high-level semantic…

计算与语言 · 计算机科学 2018-05-23 Andrei M. Butnaru , Radu Tudor Ionescu

Predicting the geographical location of users of social media like Twitter has found several applications in health surveillance, emergency monitoring, content personalization, and social studies in general. In this work we contribute to…

社会与信息网络 · 计算机科学 2021-12-15 Federico M. Funes , José Ignacio Alvarez-Hamelin , Mariano G. Beiró

Deep learning mechanisms are prevailing approaches in recent days for the various tasks in natural language processing, speech recognition, image processing and many others. To leverage this we use deep learning based mechanism specifically…

计算与语言 · 计算机科学 2019-01-03 Vidya Prasad K , Akarsh S , Vinayakumar R , Soman KP

This paper presents our approach for SwissText & KONVENS 2020 shared task 2, which is a multi-stage neural model for Swiss German (GSW) identification on Twitter. Our model outputs either GSW or non-GSW and is not meant to be used as a…

计算与语言 · 计算机科学 2020-06-08 Mohammadreza Banaei , Rémi Lebret , Karl Aberer

Positive, supportive online communication in social media (candy speech) has the potential to foster civility, yet automated detection of such language remains underexplored, limiting systematic analysis of its impact. We investigate how…

计算与语言 · 计算机科学 2025-09-17 Christian Rene Thelen , Patrick Gustav Blaneck , Tobias Bornheim , Niklas Grieger , Stephan Bialonski

We propose a deep learning approach for discovering kernels tailored to identifying clusters over sample data. Our neural network produces sample embeddings that are motivated by--and are at least as expressive as--spectral clustering. Our…

机器学习 · 计算机科学 2020-01-03 Chieh Wu , Zulqarnain Khan , Yale Chang , Stratis Ioannidis , Jennifer Dy

Recently, string kernels have obtained state-of-the-art results in various text classification tasks such as Arabic dialect identification or native language identification. In this paper, we apply two simple yet effective transductive…

计算与语言 · 计算机科学 2018-09-03 Radu Tudor Ionescu , Andrei M. Butnaru

While Self-supervised Learning (SSL) has significantly improved Spoken Language Identification (LID), existing models often struggle to consistently classify dialects and accents of the same language as a unified class. To address this…

计算与语言 · 计算机科学 2025-08-26 Qingzheng Wang , Hye-jin Shim , Jiancheng Sun , Shinji Watanabe

The problem of predicting the location of users on large social networks like Twitter has emerged from real-life applications such as social unrest detection and online marketing. Twitter user geolocation is a difficult and active research…

机器学习 · 计算机科学 2017-12-22 Tien Huu Do , Duc Minh Nguyen , Evaggelia Tsiligianni , Bruno Cornelis , Nikos Deligiannis

In this paper we present the GDI_classification entry to the second German Dialect Identification (GDI) shared task organized within the scope of the VarDial Evaluation Campaign 2018. We present a system based on SVM classifier ensembles…

计算与语言 · 计算机科学 2018-07-24 Alina Maria Ciobanu , Shervin Malmasi , Liviu P. Dinu

For many text classification tasks, there is a major problem posed by the lack of labeled data in a target domain. Although classifiers for a target domain can be trained on labeled text data from a related source domain, the accuracy of…

计算与语言 · 计算机科学 2018-11-06 Radu Tudor Ionescu , Andrei M. Butnaru

We propose a method for embedding two-dimensional locations in a continuous vector space using a neural network-based model incorporating mixtures of Gaussian distributions, presenting two model variants for text-based geolocation and…

计算与语言 · 计算机科学 2017-08-16 Afshin Rahimi , Timothy Baldwin , Trevor Cohn

We describe a machine learning approach for the 2017 shared task on Native Language Identification (NLI). The proposed approach combines several kernels using multiple kernel learning. While most of our kernels are based on character…

计算与语言 · 计算机科学 2017-08-07 Radu Tudor Ionescu , Marius Popescu

We introduce a neural network-based system of Word Sense Disambiguation (WSD) for German that is based on SenseFitting, a novel method for optimizing WSD. We outperform knowledge-based WSD methods by up to 25% F1-score and produce a new…

计算与语言 · 计算机科学 2019-08-01 Manuel Stoeckel , Sajawel Ahmed , Alexander Mehler

Sequential hypothesis testing is a desirable decision making strategy in any time sensitive scenario. Compared with fixed sample-size testing, sequential testing is capable of achieving identical probability of error requirements using less…

机器学习 · 统计学 2017-11-17 Diyan Teng , Emre Ertin

This research is aimed to solve the tweet/user geolocation prediction task and provide a flexible methodology for the geotagging of textual big data. The suggested approach implements neural networks for natural language processing (NLP) to…

计算与语言 · 计算机科学 2025-01-13 Kateryna Lutsai , Christoph H. Lampert

In the last few years, microblogging platforms such as Twitter have given rise to a deluge of textual data that can be used for the analysis of informal communication between millions of individuals. In this work, we propose an…

计算与语言 · 计算机科学 2021-11-17 Gonzalo Donoso , David Sanchez
‹ 上一页 1 2 3 10 下一页 ›