中文
相关论文

相关论文: Open Data Portal Germany (OPAL) Projektergebnisse

200 篇论文

This paper presents a framework for assessing data and metadata quality within Open Data portals. Although a few benchmark frameworks already exist for this purpose, they are not yet detailed enough in both breadth and depth to make valid…

信息检索 · 计算机科学 2021-06-18 Lisa Wenige , Claus Stadler , Michael Martin , Richard Figura , Robert Sauter , Christopher W. Frank

This paper presents a comprehensive overview of the data preparation pipeline developed for the OpenGPT-X project, a large-scale initiative aimed at creating open and high-performance multilingual large language models (LLMs). The project…

The availability of both structured and unstructured databases, such as electronic health data, social media data, patent data, and surveys that are often updated in real time, among others, has grown rapidly over the past decade. With this…

数据库 · 计算机科学 2023-07-26 Rebecca C. Steorts

This paper presents an approach for metadata reconciliation, curation and linking for Open Governamental Data Portals (ODPs). ODPs have been lately the standard solution for governments willing to put their public data available for the…

信息检索 · 计算机科学 2015-10-16 Alan Tygel , Sören Auer , Jeremy Debattista , Fabrizio Orlandi , Maria Luiza Machado Campos

Scaling data quantity is essential for large language models (LLMs), yet recent findings show that data quality can significantly boost performance and training efficiency. We introduce a German-language dataset curation pipeline that…

Software systems increasingly include AI components based on deep learning (DL). Reliable testing of such systems requires near-perfect test-input validity and label accuracy, with minimal human effort. Yet, the DL community has largely…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Mohammad Hossein Amini , Mehrdad Sabetzadeh , Shiva Nejati

Currently, a variety of pipeline tools are available for use in data engineering. Data scientists can use these tools to resolve data wrangling issues associated with data and accomplish some data engineering tasks from data ingestion…

机器学习 · 计算机科学 2024-06-21 Anthony Mbata , Yaji Sripada , Mingjun Zhong

Analysing educational data sets is fundamental to many fields of research focusing on improving student learning. However, large educational data sets are complex and can involve intensive preprocessing. These obstacles can be overcome…

统计计算 · 统计学 2025-01-17 Emma Howard

Forms are our gates to the web. They enable us to access the deep content of web sites. Automatic form understanding provides applications, ranging from crawlers over meta-search engines to service integrators, with a key to this content.…

数据库 · 计算机科学 2012-10-23 Tim Furche , Georg Gottlob , Giovanni Grasso , Xiaonan Guo , Giorgio Orsi , Christian Schallhart

The Open Dataset of Audio Quality (ODAQ) was recently introduced to address the scarcity of openly available audio datasets with corresponding subjective quality scores. The dataset, released under permissive licenses, comprises audio…

音频与语音处理 · 电气工程与系统科学 2025-04-02 Sascha Dick , Christoph Thompson , Chih-Wei Wu , Matteo Torcoli , Pablo Delgado , Phillip A. Williams , Emanuel Habets

We present OCR-Quality, a comprehensive human-annotated dataset designed for evaluating and developing OCR quality assessment methods. The dataset consists of 1,000 PDF pages converted to PNG images at 300 DPI, sampled from diverse…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Yulong Zhang

Open data has been around for many years but with the advancement of technology and its steady adoption by businesses and governments it promises to create new opportunities for the advancement of society as a whole. Many popular open data…

计算机与社会 · 计算机科学 2016-06-21 Cherlton Millette , Patrick Hosein

Multidimensional databases are a great asset for decision making. Their users express complex OLAP (On-Line Analytical Processing) queries, often returning huge volumes of facts, sometimes providing little or no information. Furthermore,…

数据库 · 计算机科学 2012-08-02 Saida Aissi , Mohamed Salah Gouider

We demonstrate a novel table discovery pipeline called DIALITE that allows users to discover, integrate and analyze open data tables. DIALITE has three main stages. First, it allows users to discover tables from open data platforms using…

数据库 · 计算机科学 2023-04-18 Aamod Khatiwada , Roee Shraga , Renée J. Miller

Open Government Data (OGD) initiatives aim to enhance public participation and collaboration by making government data accessible to diverse stakeholders, fostering social, environmental, and economic benefits through public value…

人机交互 · 计算机科学 2024-06-14 Fillip Molodtsov , Anastasija Nikiforova

This paper presents maplet, an open-source R package for the creation of highly customizable, fully reproducible statistical pipelines for omics data analysis, with a special focus on metabolomics-based methods. It builds on the…

Healthcare data are generated in many different formats, which makes it difficult to integrate and reuse across institutions and studies. Standardisation is required to enable consistent large-scale analysis. The OMOP-CDM, developed by the…

定量方法 · 定量生物学 2025-11-13 Jacob Desmond , Ryan Wartmann , Chng Wei Lau , Steven Thomas , Paul M. Middleton , Jeewani Anupama Ginige

The Extract, Transform, Load (ETL) workflow is fundamental for populating and maintaining data warehouses and other data stores accessed by analysts for downstream tasks. A major shortcoming of modern ETL solutions is the extensive need for…

软件工程 · 计算机科学 2025-08-01 Mattia Di Profio , Mingjun Zhong , Yaji Sripada , Marcel Jaspars

The term Data Space, understood as the secure exchange of data in distributed systems, ensuring openness, transparency, decentralization, sovereignty, and interoperability of information, has gained importance during the last years.…

数据库 · 计算机科学 2024-02-13 Javier Conde , Alejandro Pozo , Andrés Munoz-Arcentales , Johnny Choque , Álvaro Alonso

This paper introduces a pipeline to parametrically sample and render multi-task vision datasets from comprehensive 3D scans from the real world. Changing the sampling parameters allows one to "steer" the generated datasets to emphasize…

计算机视觉与模式识别 · 计算机科学 2021-10-12 Ainaz Eftekhar , Alexander Sax , Roman Bachmann , Jitendra Malik , Amir Zamir
‹ 上一页 1 2 3 10 下一页 ›