中文
相关论文

相关论文: The mapKurator System: A Complete Pipeline for Ext…

200 篇论文

Mordecai3 is a new end-to-end text geoparser and event geolocation system. The system performs toponym resolution using a new neural ranking model to resolve a place name extracted from a document to its entry in the Geonames gazetteer. It…

计算与语言 · 计算机科学 2023-03-27 Andrew Halterman

While many NLP pipelines assume raw, clean texts, many texts we encounter in the wild, including a vast majority of legal documents, are not so clean, with many of them being visually structured documents (VSDs) such as PDFs. Conventional…

计算与语言 · 计算机科学 2021-11-09 Yuta Koreeda , Christopher D. Manning

Tables have been an ever-existing structure to store data. There exist now different approaches to store tabular data physically. PDFs, images, spreadsheets, and CSVs are leading examples. Being able to parse table structures and extract…

计算机视觉与模式识别 · 计算机科学 2022-01-06 Susie Xi Rao , Johannes Rausch , Peter Egger , Ce Zhang

Image compositions are helpful in the study of image structures and assist in discovering the semantics of the underlying scene portrayed across art forms and styles. With the digitization of artworks in recent years, thousands of images of…

计算机视觉与模式识别 · 计算机科学 2022-06-23 Prathmesh Madhu , Tilman Marquart , Ronak Kosti , Dirk Suckow , Peter Bell , Andreas Maier , Vincent Christlein

We introduce DocSCAN, a completely unsupervised text classification approach using Semantic Clustering by Adopting Nearest-Neighbors (SCAN). For each document, we obtain semantically informative vectors from a large pre-trained language…

计算与语言 · 计算机科学 2022-10-05 Dominik Stammbach , Elliott Ash

Effective data-driven biomedical discovery requires data curation: a time-consuming process of finding, organizing, distilling, integrating, interpreting, annotating, and validating diverse information into a structured form suitable for…

Information Extraction from visually rich documents is a challenging task that has gained a lot of attention in recent years due to its importance in several document-control based applications and its widespread commercial value. The…

计算机视觉与模式识别 · 计算机科学 2023-05-03 Mohamed Dhouib , Ghassen Bettaieb , Aymen Shabou

Information visualization is essential in making sense out of large data sets. Often, high-dimensional data are visualized as a collection of points in 2-dimensional space through dimensionality reduction techniques. However, these…

计算几何 · 计算机科学 2009-07-16 Emden R. Gansner , Yifan Hu , Stephen G. Kobourov

In a highly multilingual and multicultural environment such as in the European Commission with soon over twenty official languages, there is an urgent need for text analysis tools that use minimal linguistic knowledge so that they can be…

计算与语言 · 计算机科学 2007-05-23 Camelia Ignat , Bruno Pouliquen , Antonio Ribeiro , Ralf Steinberger

With the need of fast retrieval speed and small memory footprint, document hashing has been playing a crucial role in large-scale information retrieval. To generate high-quality hashing code, both semantics and neighborhood information are…

信息检索 · 计算机科学 2021-05-28 Zijing Ou , Qinliang Su , Jianxing Yu , Bang Liu , Jingwen Wang , Ruihui Zhao , Changyou Chen , Yefeng Zheng

Maps are fundamental medium to visualize and represent the real word in a simple and 16 philosophical way. The emergence of the 3rd wave information has made a proportion of maps are available to be generated ubiquitously, which would…

计算机视觉与模式识别 · 计算机科学 2023-12-15 Xiran Zhou , Yi Wen , Honghao Li , Kaiyuan Li , Zhenfeng Shao , Zhigang Yan , Xiao Xie

Text-to-image generative models can be tremendously valuable in supporting creative tasks by providing inspirations and enabling quick exploration of different design ideas. However, one common challenge is that users may still not be able…

人机交互 · 计算机科学 2025-10-06 Yuhan Guo , Xingyou Liu , Xiaoru Yuan , Kai Xu

The amount of information available on the Web grows at an incredible high rate. Systems and procedures devised to extract these data from Web sources already exist, and different approaches and techniques have been investigated during the…

人工智能 · 计算机科学 2012-02-13 Emilio Ferrara , Robert Baumgartner

In this paper, we develop a high-dimensional map building technique that incorporates raw pixelated semantic measurements into the map representation. The proposed technique uses Gaussian Processes (GPs) multi-class classification for map…

机器人学 · 计算机科学 2017-07-07 Maani Ghaffari Jadidi , Lu Gan , Steven A. Parkison , Jie Li , Ryan M. Eustice

Deep learning continues to push state-of-the-art performance for the semantic segmentation of color (i.e., RGB) imagery; however, the lack of annotated data for many remote sensing sensors (i.e. hyperspectral imagery (HSI)) prevents…

机器学习 · 统计学 2018-04-03 Ronald Kemker , Utsav B. Gewali , Christopher Kanan

Multimodal approaches have shown great promise for searching and navigating digital collections held by libraries, archives, and museums. In this paper, we introduce map-RAS: a retrieval-augmented search system for historic maps. In…

信息检索 · 计算机科学 2025-10-30 Jamie Mahowald , Benjamin Charles Germain Lee

Recent research has shown that transformer networks can be used as differentiable search indexes by representing each document as a sequences of document ID tokens. These generative retrieval models cast the retrieval problem to a document…

信息检索 · 计算机科学 2023-11-16 Hansi Zeng , Chen Luo , Bowen Jin , Sheikh Muhammad Sarwar , Tianxin Wei , Hamed Zamani

Rule-based information extraction has lately received a fair amount of attention from the database community, with several languages appearing in the last few years. Although information extraction systems are intended to deal with…

数据库 · 计算机科学 2018-01-01 Francisco Maturana , Cristian Riveros , Domagoj Vrgoč

Historical maps are invaluable sources of information about the past, and scanned historical maps are increasingly accessible in online libraries. To retrieve maps from these large libraries that contain specific places of interest,…

信息检索 · 计算机科学 2024-10-23 Rhett Olson , Jina Kim , Yao-Yi Chiang

This paper describes a machine learning approach for annotating and analyzing data curation work logs at ICPSR, a large social sciences data archive. The systems we studied track curation work and coordinate team decision-making at ICPSR.…

计算与语言 · 计算机科学 2024-10-28 Sara Lafia , Andrea Thomer , David Bleckley , Dharma Akmon , Libby Hemphill