English
Related papers

Related papers: From data to corpus: semiotic and documentary issu…

200 papers

In this manifesto, we put forward the idea of data alchemy as a narrative device to discuss storytelling and transdisciplinarity in visualization. If data is the prima materia of modern science, how does one perform the Great Work? We use…

Human-Computer Interaction · Computer Science 2022-10-06 Victor Schetinger , Velitchko Filipov , Ignacio Pérez-Messina , Ethan Smith , Rodrigo Oliveira de Oliveira

The paper illustrates the research result of the application of semantic technology to ease the use and reuse of digital contents exposed as Linked Data on the web. It focuses on the specific issue of explorative research for the resource…

Digital Libraries · Computer Science 2011-10-12 Riccardo Albertoni , Monica De Martino

We explore methods for content selection and address the issue of coherence in the context of the generation of multimedia artifacts. We use audio and video to present two case studies: generation of film tributes, and lecture-driven…

Artificial Intelligence · Computer Science 2015-08-14 Paulo Figueiredo , Marta Aparício , David Martins de Matos , Ricardo Ribeiro

Many disciplines pose natural-language research questions over large document collections whose answers typically require structured evidence, traditionally obtained by manually designing an annotation schema and exhaustively labeling the…

Computation and Language · Computer Science 2026-04-13 Shahar Levy , Eliya Habba , Reshef Mintz , Barak Raveh , Renana Keydar , Gabriel Stanovsky

Purpose: We seek to explore the realm of literature about digital libraries. We specifically seek to ascertain how interest in this subject has evolved, its impact, the most productive journals and countries, the number of occurrences of…

Digital Libraries · Computer Science 2021-06-28 Mathieu Andro , Marc Maisonneuve

Unstructured text from legal, medical, and administrative sources offers a rich but underutilized resource for research in public health and the social sciences. However, large-scale analysis is hampered by two key challenges: the presence…

Computation and Language · Computer Science 2025-07-16 Anders Ledberg , Anna Thalén

My research focuses on the analysis, recovery, and generation of 4D content, where 4D includes three spatial dimensions (x, y, z) and a temporal dimension t, such as shape and motion. This focus goes beyond static objects to include dynamic…

Computer Vision and Pattern Recognition · Computer Science 2024-09-24 Zhiyang Dou

With the rapid development of Internet and multimedia services in the past decade, a huge amount of user-generated and service provider-generated multimedia data become available. These data are heterogeneous and multi-modal in nature,…

Multimedia · Computer Science 2020-01-07 Wenwu Zhu , Xin Wang , Hongzhi Li

Electronic Healthcare Records contain large volumes of unstructured data, including extensive free text. Yet this source of detailed information often remains under-used because of a lack of methodologies to extract interpretable content in…

Computation and Language · Computer Science 2018-07-10 M. Tarik Altuncu , Erik Mayer , Sophia N. Yaliraki , Mauricio Barahona

In traditional production plants, current technologies do not provide sufficient context to support information integration and interpretation. Digital transformation technologies have the potential to support contextualization, but it is…

Human-Computer Interaction · Computer Science 2024-08-20 Romy Müller , Franziska Kessler , David W. Humphrey , Julian Rahm

Parallel texts are a relatively rare language resource, however, they constitute a very useful research material with a wide range of applications. This study presents and analyses new methodologies we developed for obtaining such data from…

Computation and Language · Computer Science 2016-03-23 Krzysztof Wołk , Emilia Rejmund , Krzysztof Marasek

The number of documents available into Internet moves each day up. For this reason, processing this amount of information effectively and expressibly becomes a major concern for companies and scientists. Methods that represent a textual…

Information Retrieval · Computer Science 2017-03-21 Mohamed Morchid , Juan-Manuel Torres-Moreno , Richard Dufour , Javier Ramírez-Rodríguez , Georges Linarès

Text data is often seen as "take-away" materials with little noise and easy to process information. Main questions are how to get data and transform them into a good document format. But data can be sensitive to noise oftenly called…

Computation and Language · Computer Science 2023-06-22 Nicolas Turenne

The need for discovering knowledge from XML documents according to both structure and content features has become challenging, due to the increase in application contexts for which handling both structure and content information in XML data…

Databases · Computer Science 2015-04-17 Olfa Arfaoui , Minyar Sassi Hidri

We present an analysis pipeline and best practice guidelines for building and curating corpora of everyday conversation in diverse languages. Surveying language documentation corpora and other resources that cover 67 languages and varieties…

Computation and Language · Computer Science 2022-05-11 Andreas Liesenfeld , Mark Dingemanse

The opacity of machine learning data is a significant threat to ethical data work and intelligible systems. Previous research has addressed this issue by proposing standardized checklists to document datasets. This paper expands that field…

Human-Computer Interaction · Computer Science 2022-08-11 Milagros Miceli , Tianling Yang , Adriana Alvarado Garcia , Julian Posada , Sonja Mei Wang , Marc Pohl , Alex Hanna

Legacy procedures for topic modelling have generally suffered problems of overfitting and a weakness towards reconstructing sparse topic structures. With motivation from a consumer-generated corpora, this paper proposes semiparametric topic…

Computation and Language · Computer Science 2025-03-05 Dominic B. Dayta , Erniel B. Barrios

The process of documenting and describing the world's languages is undergoing radical transformation with the rapid uptake of new digital technologies for capture, storage, annotation and dissemination. However, uncritical adoption of new…

Computation and Language · Computer Science 2007-05-23 Steven Bird , Gary Simons

The methodology of context-sensitive access to e-documents considers context as a problem model based on the knowledge extracted from the application domain, and presented in the form of application ontology. Efficient access to an…

Information Retrieval · Computer Science 2007-05-23 A. V. Smirnov , T. V. Levashova , M. P. Pashkin , N. G. Shilov , A. A. Krizhanovsky , A. M. Kashevnik , A. S. Komarova

A growing body of work shows that many problems in fairness, accountability, transparency, and ethics in machine learning systems are rooted in decisions surrounding the data collection and annotation process. In spite of its fundamental…

Machine Learning · Computer Science 2019-12-24 Eun Seo Jo , Timnit Gebru
‹ Prev 1 4 5 6 7 8 10 Next ›