中文
相关论文

相关论文: Fundus: A Simple-to-Use News Scraper Optimized for…

200 篇论文

This paper presents the design and implementation of a user-friendly, automated web application that simplifies and optimizes the web scraping process for non-technical users. The application breaks down the complex task of web scraping…

信息检索 · 计算机科学 2025-10-28 Alok Dutta , Nilanjana Roy , Rhythm Sen , Sougata Dutta , Prabhat Das

We introduce an advanced information extraction pipeline to automatically process very large collections of unstructured textual data for the purpose of investigative journalism. The pipeline serves as a new input processor for the upcoming…

计算与语言 · 计算机科学 2018-09-17 Gregor Wiedemann , Seid Muhie Yimam , Chris Biemann

Large-scale news corpora support a wide range of research in Computational Social Science and NLP, yet access remains constrained: commercial archives impose prohibitive costs and licensing restrictions, while open alternatives like Common…

计算与语言 · 计算机科学 2026-05-19 Ruggero Marino Lazzaroni , Jana Lasser , Kirill Solovev

Recently, the term "fake news" has been broadly and extensively utilized for disinformation, misinformation, hoaxes, propaganda, satire, rumors, click-bait, and junk news. It has become a serious problem around the world. We present a new…

社会与信息网络 · 计算机科学 2020-10-06 Jiawei Xu , Vladimir Zadorozhny , Danchen Zhang , John Grant

This paper proposes OCR++, an open-source framework designed for a variety of information extraction tasks from scholarly articles including metadata (title, author names, affiliation and e-mail), structure (section headings and body text,…

The web contains large-scale, diverse, and abundant information to satisfy the information-seeking needs of humans. Through meticulous data collection, preprocessing, and curation, webpages can be used as a fundamental data resource for…

计算与语言 · 计算机科学 2024-06-18 Zhipeng Xu , Zhenghao Liu , Yukun Yan , Zhiyuan Liu , Ge Yu , Chenyan Xiong

News editors need to find the photos that best illustrate a news piece and fulfill news-media quality standards, while being pressed to also find the most recent photos of live events. Recently, it became common to use social-media content…

信息检索 · 计算机科学 2018-10-10 Gonçalo Marcelino , Ricardo Pinto , João Magalhães

Modern web scraping struggles with dynamic, interactive websites that require more than static HTML parsing. Current methods are often brittle and require manual customization for each site. To address this, we introduce Webscraper, a…

人工智能 · 计算机科学 2026-04-01 Guan-Lun Huang , Yuh-Jzer Joung

This study develops and evaluates a systematic methodology for constructing news datasets from Google News, combining automated web scraping, large language model (LLM)-based metadata extraction, and SCImago Media Rankings enrichment. Using…

数字图书馆 · 计算机科学 2026-04-30 Victor Herrero-Solana

This paper presents a pipeline with minimal human influence for scraping and detecting bias on college newspaper archives. This paper introduces a framework for scraping complex archive sites that automated tools fail to grab data from, and…

计算与语言 · 计算机科学 2023-09-14 Adam M. Lehavi , William McCormack , Noah Kornfeld , Solomon Glazer

Access to diverse perspectives is essential for understanding real-world events, yet most news retrieval systems prioritize textual relevance, leading to redundant results and limited viewpoint exposure. We propose NEWSCOPE, a two-stage…

计算与语言 · 计算机科学 2025-09-01 Yixuan Tang , Yuanyuan Shi , Yiqun Sun , Anthony Kum Hoe Tung

Document content analysis has been a crucial research area in computer vision. Despite significant advancements in methods such as OCR, layout detection, and formula recognition, existing open-source solutions struggle to consistently…

计算机视觉与模式识别 · 计算机科学 2024-09-30 Bin Wang , Chao Xu , Xiaomeng Zhao , Linke Ouyang , Fan Wu , Zhiyuan Zhao , Rui Xu , Kaiwen Liu , Yuan Qu , Fukai Shang , Bo Zhang , Liqun Wei , Zhihao Sui , Wei Li , Botian Shi , Yu Qiao , Dahua Lin , Conghui He

In today's day and age where information is rapidly spread through online platforms, the rise of fake news poses an alarming threat to the integrity of public discourse, societal trust, and reputed news sources. Classical machine learning…

计算与语言 · 计算机科学 2024-10-15 Arjun Shah , Hetansh Shah , Vedica Bafna , Charmi Khandor , Sindhu Nair

In this era of the Internet, the amount of news articles added every minute of everyday is humongous. As a result of this explosive amount of news articles, news retrieval systems are required to process the news articles frequently and…

分布式、并行与集群计算 · 计算机科学 2012-04-17 Arockia Anand Raj , T. Mala

We present a software tool that employs state-of-the-art natural language processing (NLP) and machine learning techniques to help newspaper editors compose effective headlines for online publication. The system identifies the most salient…

计算与语言 · 计算机科学 2019-05-21 Terrence Szymanski , Claudia Orellana-Rodriguez , Mark T. Keane

For the many journalists who use data and computation to report the news, data wrangling is an integral part of their work.Despite an abundance of literature on data wrangling in the context of enterprise data analysis, little is known…

人机交互 · 计算机科学 2020-09-24 Stephen Kasica , Charles Berret , Tamara Munzner

The profusion of online news articles makes it difficult to find interesting articles, a problem that can be assuaged by using a recommender system to bring the most relevant news stories to readers. However, news recommendation is…

信息检索 · 计算机科学 2014-11-04 Florent Garcin , Christos Dimitrakakis , Boi Faltings

Automated news verification requires structured claim extraction, but existing approaches either lack schema compliance or generalize poorly across domains. This paper presents NewsScope, a cross-domain dataset, benchmark, and fine-tuned…

计算与语言 · 计算机科学 2026-01-15 Nidhi Pandya

Fake news detection is a challenging task aiming to reduce human time and effort to check the truthfulness of news. Automated approaches to combat fake news, however, are limited by the lack of labeled benchmark datasets, especially in…

计算与语言 · 计算机科学 2021-03-02 Inna Vogel , Jeong-Eun Choi , Meghana Meghana

Existing full text datasets of U.S. public domain newspapers do not recognize the often complex layouts of newspaper scans, and as a result the digitized content scrambles texts from articles, headlines, captions, advertisements, and other…

‹ 上一页 1 2 3 10 下一页 ›