中文
相关论文

相关论文: Papeos: Augmenting Research Papers with Talk Video…

200 篇论文

Interacting and understanding with text heavy visual content with multiple images is a major challenge for traditional vision models. This paper is on enhancing vision models' capability to comprehend or understand and learn from images…

计算机视觉与模式识别 · 计算机科学 2024-08-31 Adithya TG , Adithya SK , Abhinav R Bharadwaj , Abhiram HA , Surabhi Narayan

This paper investigates the effects of seeing ideas presented in-person when they are easily accessible online. Presentations may increase the diffusion of ideas intentionally (when one attends the presentation of an idea of interest) and…

数字图书馆 · 计算机科学 2024-01-22 Misha Teplitskiy , Soya Park , Neil Thompson , David Karger

When video is shot in noisy environment, the voice of a speaker seen in the video can be enhanced using the visible mouth movements, reducing background noise. While most existing methods use audio-only inputs, improved performance is…

计算机视觉与模式识别 · 计算机科学 2018-06-14 Aviv Gabbay , Asaph Shamir , Shmuel Peleg

Mediated by today's visual displays, information space allows users to discover, access and interact with a wide range of digital and physical information. The information presented in this space may be digital, physical or a blend of both,…

人机交互 · 计算机科学 2025-06-06 Chen Chen

Publishing open-source academic video recordings is an emergent and prevalent approach to sharing knowledge online. Such videos carry rich multimodal information including speech, the facial and body movements of the speakers, as well as…

计算与语言 · 计算机科学 2024-06-05 Zhe Chen , Heyang Liu , Wenyi Yu , Guangzhi Sun , Hongcheng Liu , Ji Wu , Chao Zhang , Yu Wang , Yanfeng Wang

Text-level discourse parsing aims to unmask how two sentences in the text are related to each other. We propose the task of Visual Discourse Parsing, which requires understanding discourse relations among scenes in a video. Here we use the…

计算机视觉与模式识别 · 计算机科学 2022-01-25 Arjun R. Akula , Song-Chun Zhu

Text-to-Speech (TTS) synthesis faces the inherent challenge of producing multiple speech outputs with varying prosody given a single text input. While previous research has addressed this by predicting prosodic information from both text…

计算与语言 · 计算机科学 2025-08-19 Shumin Que , Anton Ragni

The paper is concerned with the design and development of a P2P Presentation System (P2P-PS) for live streaming of lectures coupled with a shareable whiteboard. Video streaming uses a data-driven modified mesh architecture which supports…

分布式、并行与集群计算 · 计算机科学 2019-03-14 Nikita Bhagatkar , Kapil Dolas , R. K. Ghosh

Speech as a natural signal is composed of three parts - visemes (visual part of speech), phonemes (spoken part of speech), and language (the imposed structure). However, video as a medium for the delivery of speech and a multimedia…

计算与语言 · 计算机科学 2020-06-17 Dhruva Sahrawat , Yaman Kumar , Shashwat Aggarwal , Yifang Yin , Rajiv Ratn Shah , Roger Zimmermann

Speechreading or lipreading is the technique of understanding and getting phonetic features from a speaker's visual features such as movement of lips, face, teeth and tongue. It has a wide range of multimedia applications such as in…

Our conferences face a growing crisis: an overwhelming flood of submissions, increased reviewing burdens, and diminished opportunities for meaningful engagement. With AI making paper generation easier than ever, we must ask whether the…

计算机与社会 · 计算机科学 2025-09-10 Daniel Russo , Margaret-Anne Storey

Slide-based teaching is widely used in higher education, yet in online, hybrid, and asynchronous contexts, slides often lose instructor presence, narrative continuity, and expressive framing that help learners connect with course content.…

人机交互 · 计算机科学 2026-05-26 Xinxing Wu

We investigate methods of segmenting, visualizing, and indexing presentation videos by separately considering audio and visual data. The audio track is segmented by speaker, and augmented with key phrases which are extracted using an…

多媒体 · 计算机科学 2007-05-23 Alexander Haubold , John R. Kender

We present a review and analysis of scientific paper embellishments -- simple visual elements that are deeply integrated into the text of scientific publications. These embellishments are increasingly used in research papers, which have the…

数字图书馆 · 计算机科学 2026-03-24 Jiayi Hong , Yixuan Wang , Petra Isenberg , Ross Maciejewski

Text-to-Video (T2V) retrieval aims to identify the most relevant item from a gallery of videos based on a user's text query. Traditional methods rely solely on aligning video and text modalities to compute the similarity and retrieve…

计算机视觉与模式识别 · 计算机科学 2025-06-02 Adriano Fragomeni , Dima Damen , Michael Wray

Physical computing infrastructure, data gathering, and algorithms have recently had significant advances to extract information from images and videos. The growth has been especially outstanding in image captioning and video captioning.…

计算机视觉与模式识别 · 计算机科学 2022-01-25 Soheyla Amirian , Thiab R. Taha , Khaled Rasheed , Hamid R. Arabnia

Presenters commonly use slides as visual aids for informative talks. When presenters fail to verbally describe the content on their slides, blind and visually impaired audience members lose access to necessary content, making the…

人机交互 · 计算机科学 2021-03-29 Yi-Hao Peng , JiWoong Jang , Jeffrey P. Bigham , Amy Pavel

Literature reviews can be time-consuming and tedious to complete. By cataloging and refactoring three state-of-the-art active learning techniques from evidence-based medicine and legal electronic discovery, this paper finds and implements…

软件工程 · 计算机科学 2018-03-09 Zhe Yu , Nicholas A. Kraft , Tim Menzies

Documents containing mathematical content remain largely inaccessible to blind and visually impaired readers because they are predominantly published as untagged PDF which does not include the semantic data necessary for effective…

人机交互 · 计算机科学 2022-02-04 Rynhardt Kruger , Febe de Wet , Thomas Niesler

Human visual attention is susceptible to social influences. In education, peer effects impact student learning, but their precise role in modulating attention remains unclear. Our experiment (N=311) demonstrates that displaying peer visual…

人机交互 · 计算机科学 2023-12-06 Songlin Xu , Dongyin Hu , Ru Wang , Xinyu Zhang