English
Related papers

Related papers: WIKIR: A Python toolkit for building a large-scale…

200 papers

In order to disseminate the exponential extent of knowledge being produced in the form of scientific publications, it would be best to design mechanisms that connect it with already existing rich repository of concepts -- the Wikipedia. Not…

Information Retrieval · Computer Science 2017-05-10 Abhik Jana , Sruthi Mooriyath , Animesh Mukherjee , Pawan Goyal

In the realm of web agent research, achieving both generalization and accuracy remains a challenging problem. Due to high variance in website structure, existing approaches often fail. Moreover, existing fine-tuning and in-context learning…

Computation and Language · Computer Science 2024-04-10 Michael Lutz , Arth Bohra , Manvel Saroyan , Artem Harutyunyan , Giovanni Campagna

We introduce WikiLingua, a large-scale, multilingual dataset for the evaluation of crosslingual abstractive summarization systems. We extract article and summary pairs in 18 languages from WikiHow, a high quality, collaborative resource of…

Computation and Language · Computer Science 2020-10-08 Faisal Ladhak , Esin Durmus , Claire Cardie , Kathleen McKeown

Predicting which words are considered hard to understand for a given target population is a vital step in many NLP applications such as text simplification. This task is commonly referred to as Complex Word Identification (CWI). With a few…

Computation and Language · Computer Science 2020-06-12 Matthew Shardlow , Michael Cooper , Marcos Zampieri

The exponential increase in the usage of Wikipedia as a key source of scientific knowledge among the researchers is making it absolutely necessary to metamorphose this knowledge repository into an integral and self-contained source of…

Computation and Language · Computer Science 2018-06-18 Abhik Jana , Pranjal Kanojiya , Pawan Goyal , Animesh Mukherjee

The amount of scientific papers published every day is daunting and constantly increasing. Keeping up with literature represents a challenge. If one wants to start exploring new topics it is hard to have a big picture without reading lots…

Information Retrieval · Computer Science 2020-11-10 Alberto Calderone

We introduce Repro, an open-source library which aims at improving the reproducibility and usability of research code. The library provides a lightweight Python API for running software released by researchers within Docker containers which…

Computation and Language · Computer Science 2022-05-02 Daniel Deutsch , Dan Roth

OpenMatch is a Python-based library that serves for Neural Information Retrieval (Neu-IR) research. It provides self-contained neural and traditional IR modules, making it easy to build customized and higher-capacity IR systems. In order to…

Information Retrieval · Computer Science 2021-05-07 Zhenghao Liu , Kaitao Zhang , Chenyan Xiong , Zhiyuan Liu , Maosong Sun

In this paper, we present WildlifeDatasets (https://github.com/WildlifeDatasets/wildlife-datasets) - an open-source toolkit intended primarily for ecologists and computer-vision / machine-learning researchers. The WildlifeDatasets is…

Computer Vision and Pattern Recognition · Computer Science 2023-12-15 Vojtěch Čermák , Lukas Picek , Lukáš Adam , Kostas Papafitsoros

Data-driven design and innovation is a process to reuse and provide valuable and useful information. However, existing semantic networks for design innovation is built on data source restricted to technological and scientific information.…

Computation and Language · Computer Science 2022-11-22 Haoyu Zuo , Qianzhi Jing , Tianqi Song , Huiting Liu , Lingyun Sun , Peter Childs , Liuqing Chen

Wikidata is an open knowledge graph built by a global community of volunteers. As it advances in scale, it faces substantial challenges around editor engagement. These challenges are in terms of both attracting new editors to keep up with…

Information Retrieval · Computer Science 2021-08-02 Kholoud AlGhamdi , Miaojing Shi , Elena Simperl

Large knowledge graphs like DBpedia and YAGO are always based on the same source, i.e., Wikipedia. But there are more wikis that contain information about long-tail entities such as wiki hosting platforms like Fandom. In this paper, we…

Information Retrieval · Computer Science 2022-10-07 Sven Hertling , Heiko Paulheim

We propose a tool for experts finding applied to academic data generated by the start-up DSRT in the context of its application Peerus. A user may submit the title, the abstract and optionnally the authors and the journal of publication of…

Information Retrieval · Computer Science 2018-07-11 Robin Brochier , Adrien Guille , Julien Velcin , Benjamin Rothan , Di Cioccio

The proliferation of open-source scientific software for science and research presents opportunities and challenges. In this paper, we introduce the SciCat dataset -- a comprehensive collection of Free-Libre Open Source Software (FLOSS)…

Software Engineering · Computer Science 2023-12-12 Addi Malviya-Thakur , Reed Milewicz , Lavinia Paganini , Ahmed Samir Imam Mahmoud , Audris Mockus

Wikipedia is a free Internet encyclopedia with an enormous amount of content. This encyclopedia is written by volunteers with various backgrounds in a collective fashion; anyone can access and edit most of the articles. This open-editing…

Physics and Society · Physics 2016-01-26 Jinhyuk Yun , Sang Hoon Lee , Hawoong Jeong

Retrieving paragraphs to populate a Wikipedia article is a challenging task. The new TREC Complex Answer Retrieval (TREC CAR) track introduces a comprehensive dataset that targets this retrieval scenario. We present early results from a…

Information Retrieval · Computer Science 2017-05-16 Federico Nanni , Bhaskar Mitra , Matt Magnusson , Laura Dietz

Phenotyping consists in applying algorithms to identify individuals associated with a specific, potentially complex, trait or condition, typically out of a collection of Electronic Health Records (EHRs). Because a lot of the clinical…

Keyphrase extraction is the task of extracting a small set of phrases that best describe a document. Most existing benchmark datasets for the task typically have limited numbers of annotated documents, making it challenging to train…

Computation and Language · Computer Science 2020-10-26 Tuan Manh Lai , Trung Bui , Doo Soon Kim , Quan Hung Tran

Split and rephrase is the task of breaking down a sentence into shorter ones that together convey the same meaning. We extract a rich new dataset for this task by mining Wikipedia's edit history: WikiSplit contains one million naturally…

Computation and Language · Computer Science 2018-08-30 Jan A. Botha , Manaal Faruqui , John Alex , Jason Baldridge , Dipanjan Das

In this paper, we present MADE-WIC, a large dataset of functions and their comments with multiple annotations for technical debt and code weaknesses leveraging different state-of-the-art approaches. It contains about 860K code functions and…

Software Engineering · Computer Science 2025-01-28 Moritz Mock , Jorge Melegati , Max Kretschmann , Nicolás E. Díaz Ferreyra , Barbara Russo