English
Related papers

Related papers: Python vs. R: A Text Mining Approach for analyzing…

200 papers

In order to face the complexity of business environments and detect priorities while triggering contingency strategies, we propose a new methodological approach that combines text mining, social network and big data analytics, with the…

Social and Information Networks · Computer Science 2021-05-26 M. A. Barchiesi , A. Fronzetti Colladon

A major factor in the recent success of large language models is the use of enormous and ever-growing text datasets for unsupervised pre-training. However, naively training a model on all available data may not be optimal (or feasible), as…

Reading text from images (either natural scenes or documents) has been a long-standing research topic for decades, due to the high technical challenge and wide application range. Previously, individual specialist models are developed to…

Computer Vision and Pattern Recognition · Computer Science 2024-08-28 Peng Wang , Zhaohai Li , Jun Tang , Humen Zhong , Fei Huang , Zhibo Yang , Cong Yao

Recent waves of technological transformation are reshaping work in uncertain and hard-to-predict ways. However, jobs at the forefront of the digitizing economy offer an early glimpse of these changes and leave rich activity traces. We…

General Economics · Economics 2026-04-10 Xiangnan Feng , Johannes Wachs , Simone Daniotti , Frank Neffke

Texts reveal the subjects of interest in research fields, and the values, beliefs, and practices of researchers. In this study, texts are examined through bibliometric mapping and topic modeling to provide a birds eye view of the social…

Social and Information Networks · Computer Science 2014-01-29 Laura Sheble , Annie T. Chen

Many of quality approaches are described in hundreds of textual pages. Manual processing of information consumes plenty of resources. In this report we present a text mining approach applied on CMMI, one well known and widely known quality…

Software Engineering · Computer Science 2013-11-12 Zádor Dániel Kelemen , Rob Kusters , Jos Trienekens , Katalin Balla

This study examines how large language models categorize sentences from scientific papers using prompt engineering. We use two advanced web-based models, GPT-4o (by OpenAI) and DeepSeek R1, to classify sentences into predefined relationship…

Computation and Language · Computer Science 2025-03-05 Aniruddha Maiti , Samuel Adewumi , Temesgen Alemayehu Tikure , Zichun Wang , Niladri Sengupta , Anastasiia Sukhanova , Ananya Jana

Scientists are increasingly overwhelmed by the volume of articles being published. Total articles indexed in Scopus and Web of Science have grown exponentially in recent years; in 2022 the article total was approximately ~47% higher than in…

Digital Libraries · Computer Science 2024-12-12 Mark A. Hanson , Pablo Gómez Barreiro , Paolo Crosetto , Dan Brockington

With the rapid evolution of cross-strait situation, "Mainland China" as a subject of social science study has evoked the voice of "Rethinking China Study" among intelligentsia recently. This essay tried to apply an automatic content…

Digital Libraries · Computer Science 2023-06-22 Hsuan-Lei Shao , Sieh-Chuen Huang , Yun-Cheng Tsai

The exponential growth of digital content has generated massive textual datasets, necessitating the use of advanced analytical approaches. Large Language Models (LLMs) have emerged as tools that are capable of processing and extracting…

Computation and Language · Computer Science 2024-05-24 Benjamin M. Ampel , Chi-Heng Yang , James Hu , Hsinchun Chen

The trend toward open science increases the pressure on authors to provide access to the source code and data they used to compute the results reported in their scientific papers. Since sharing materials reproducibly is challenging, several…

Digital Libraries · Computer Science 2020-07-15 Markus Konkol , Daniel Nüst , Laura Goulier

In the vision domain, dataset distillation arises as a technique to condense a large dataset into a smaller synthetic one that exhibits a similar result in the training process. While image data presents an extensive literature of…

This paper highlights the challenges, current trends, and open issues related to the representation, querying and analytics of content extracted from texts. The internet contains vast text-based information on various subjects, including…

Databases · Computer Science 2023-10-11 Genoveva Vargas-Solar , Mirian Halfeld Ferrari Alves , Anne-Lyse Minard Forst

Supervised text classification is a classical and active area of ML research. In large enterprise, solutions to this problem has significant importance. This is specifically true in ticketing systems where prediction of the type and subtype…

Information Retrieval · Computer Science 2020-12-02 Nabarun Mondal , Mrunal Lohia

The rapid expansion of research across machine learning, vision, and language has produced a volume of publications that is increasingly difficult to synthesize. Traditional bibliometric tools rely mainly on metadata and offer limited…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Zhucun Xue , Jiangning Zhang , Juntao Jiang , Jinzhuo Liu , Haoyang He , Teng Hu , Xiaobin Hu , Yong Liu , Shuicheng Yan

The sheer volume of scientific experimental results and complex technical statements, often presented in tabular formats, presents a formidable barrier to individuals acquiring preferred information. The realms of scientific reasoning and…

Computation and Language · Computer Science 2024-03-28 Zhixin Guo , Jianping Zhou , Jiexing Qi , Mingxuan Yan , Ziwei He , Guanjie Zheng , Zhouhan Lin , Xinbing Wang , Chenghu Zhou

The study of trajectories is often a core task in several research fields. In environmental modelling, trajectories are crucial to study fluid pollution, animal migrations, oil slick patterns or land movements. In this contribution, we…

Computation · Statistics 2022-09-23 A. Reyes , G. Viera-López , J. J. Morgado-Vega , E. Altshuler

Social and technical trends have significantly changed methods for evaluating and disseminating computing research. Traditional venues for reviewing and publishing, such as conferences and journals, worked effectively in the past. Recently,…

Computers and Society · Computer Science 2020-07-03 Benjamin Zorn , Tom Conte , Keith Marzullo , Suresh Venkatasubramanian

Against the background of what has been termed a reproducibility crisis in science, the NLP field is becoming increasingly interested in, and conscientious about, the reproducibility of its results. The past few years have seen an…

Computation and Language · Computer Science 2021-03-23 Anya Belz , Shubham Agarwal , Anastasia Shimorina , Ehud Reiter

The workshop "Mining Scientific Papers: Computational Linguistics and Bibliometrics" (CLBib 2015), co-located with the 15th International Society of Scientometrics and Informetrics Conference (ISSI 2015), brought together researchers in…

Computation and Language · Computer Science 2015-06-18 Iana Atanassova , Marc Bertin , Philipp Mayr