English
Related papers

Related papers: EEVEE: An Easy Annotation Tool for Natural Languag…

200 papers

Natural language (NL) toolkits enable visualization developers, who may not have a background in natural language processing (NLP), to create natural language interfaces (NLIs) for end-users to flexibly specify and interact with…

Human-Computer Interaction · Computer Science 2022-08-16 Rishab Mitra , Arpit Narechania , Alex Endert , John Stasko

We introduce the MuSe-Toolbox - a Python-based open-source toolkit for creating a variety of continuous and discrete emotion gold standards. In a single framework, we unify a wide range of fusion methods and propose the novel Rater Aligned…

Computation and Language · Computer Science 2021-10-22 Lukas Stappen , Lea Schumann , Benjamin Sertolli , Alice Baird , Benjamin Weigel , Erik Cambria , Björn W. Schuller

Learning robust representations of authorial style is crucial for authorship attribution and AI-generated text detection. However, existing methods often struggle with content-style entanglement, where models learn spurious correlations…

Computation and Language · Computer Science 2026-04-24 Hieu Man , Van-Cuong Pham , Nghia Trung Ngo , Franck Dernoncourt , Thien Huu Nguyen

Extracting structured and grounded fact triples from raw text is a fundamental task in Information Extraction (IE). Existing IE datasets are typically collected from Wikipedia articles, using hyperlinks to link entities to the Wikidata…

Computation and Language · Computer Science 2023-06-16 Chenxi Whitehouse , Clara Vania , Alham Fikri Aji , Christos Christodoulopoulos , Andrea Pierleoni

Automatic medical image segmentation plays a critical role in scientific research and medical care. Existing high-performance deep learning methods typically rely on large training datasets with high-quality manual annotations, which are…

Image and Video Processing · Electrical Eng. & Systems 2021-11-17 Shanshan Wang , Cheng Li , Rongpin Wang , Zaiyi Liu , Meiyun Wang , Hongna Tan , Yaping Wu , Xinfeng Liu , Hui Sun , Rui Yang , Xin Liu , Jie Chen , Huihui Zhou , Ismail Ben Ayed , Hairong Zheng

While the general analysis of named entities has received substantial research attention on unstructured as well as structured data, the analysis of relations among named entities has received limited focus. In fact, a review of the…

Computation and Language · Computer Science 2018-10-05 Saeedeh Shekarpour , Faisal Alshargi , Valerie Shalin , Krishnaprasad Thirunarayan , Amit P. Sheth

Business documents come in a variety of structures, formats and information needs which makes information extraction a challenging task. Due to these variations, having a document generic model which can work well across all types of…

Computation and Language · Computer Science 2022-11-10 Neelesh K Shukla , Msp Raja , Raghu Katikeri , Amit Vaid

The usefulness of annotated corpora is greatly increased if there is an associated tool that can allow various kinds of operations to be performed in a simple way. Different kinds of annotation frameworks and many query languages for them…

Computation and Language · Computer Science 2011-08-10 Anil Kumar Singh

Annotated data have traditionally been used to provide the input for training a supervised machine learning (ML) model. However, current pre-trained ML models for natural language processing (NLP) contain embedded linguistic information…

Computation and Language · Computer Science 2022-04-05 Abe Kazemzadeh

This demonstration paper presents StreamSide, an open-source toolkit for annotating multiple kinds of meaning representations. StreamSide supports frame-based annotation schemes e.g., Abstract Meaning Representation (AMR) and frameless…

Computation and Language · Computer Science 2021-09-22 Jinho D. Choi , Gregor Williamson

We present skweak, a versatile, Python-based software toolkit enabling NLP developers to apply weak supervision to a wide range of NLP tasks. Weak supervision is an emerging machine learning paradigm based on a simple idea: instead of…

Computation and Language · Computer Science 2021-08-18 Pierre Lison , Jeremy Barnes , Aliaksandr Hubin

The processing of entities in natural language is essential to many medical NLP systems. Unfortunately, existing datasets vastly under-represent the entities required to model public health relevant texts such as health advice often found…

Computation and Language · Computer Science 2022-10-10 Joseph Gatto , Parker Seegmiller , Garrett Johnston , Sarah M. Preum

DataFlow has been emerging as a new paradigm for building task-oriented chatbots due to its expressive semantic representations of the dialogue tasks. Despite the availability of a large dataset SMCalFlow and a simplified syntax, the…

Computation and Language · Computer Science 2022-12-19 Han He , Song Feng , Daniele Bonadiman , Yi Zhang , Saab Mansour

The paper describes the ALVIS annotation format designed for the indexing of large collections of documents in topic-specific search engines. This paper is exemplified on the biological domain and on MedLine abstracts, as developing a…

Artificial Intelligence · Computer Science 2016-08-16 Adeline Nazarenko , Erick Alphonse , Julien Derivière , Thierry Hamon , Guillaume Vauvert , Davy Weissenbacher

With the surging inclination towards carrying out tasks on computational devices and digital mediums, any method that converts a task that was previously carried out manually, to a digitized version, is always welcome. Irrespective of the…

Computer Vision and Pattern Recognition · Computer Science 2023-04-19 Pranav Guruprasad , Sujith Kumar S , Vigneswaran C , V. Srinivasa Chakravarthy

Item categorization is a machine learning task which aims at classifying e-commerce items, typically represented by textual attributes, to their most suitable category from a predefined set of categories. An accurate item categorization…

Machine Learning · Computer Science 2021-10-25 Yonatan Hadar , Erez Shmueli

Current pre-training works in natural language generation pay little attention to the problem of exposure bias on downstream tasks. To address this issue, we propose an enhanced multi-flow sequence to sequence pre-training and fine-tuning…

Computation and Language · Computer Science 2020-06-09 Dongling Xiao , Han Zhang , Yukun Li , Yu Sun , Hao Tian , Hua Wu , Haifeng Wang

Recent advancements in machine learning and adaptive cognitive systems are driving a growing demand for large and richly annotated multimodal data. A prominent example of this trend are fusion models, which increasingly incorporate multiple…

Software Engineering · Computer Science 2025-10-20 Rathi Adarshi Rammohan , Moritz Meier , Dennis Küster , Tanja Schultz

With the recent progress in machine learning, boosted by techniques such as deep learning, many tasks can be successfully solved once a large enough dataset is available for training. Nonetheless, human-annotated datasets are often…

Computation and Language · Computer Science 2019-08-19 Daniel Specht Menezes , Pedro Savarese , Ruy Luiz Milidiú

Having a comprehensive, high-quality dataset of road sign annotation is critical to the success of AI-based Road Sign Recognition (RSR) systems. In practice, annotators often face difficulties in learning road sign systems of different…

Artificial Intelligence · Computer Science 2020-12-07 Ji Eun Kim , Cory Henson , Kevin Huang , Tuan A. Tran , Wan-Yi Lin