English
Related papers

Related papers: A Novel Frame Structure for Cloud-Based Audio-Visu…

200 papers

Recent advancements in polymer microwave fiber (PMF) technology have created significant opportunities for robust, low-cost, and high-speed sub-terahertz (THz) radio-over-fiber communications. Recognizing these potential benefits, this…

Signal Processing · Electrical Eng. & Systems 2025-12-08 Dexin Kong , Diana Pamela Moya Osorio , Erik G. Larsson

Traditional physical (PHY) layer protocols contain chains of signal processing blocks that have been mathematically optimized to transmit information bits efficiently over noisy channels. Unfortunately, this same optimality encourages…

Signal Processing · Electrical Eng. & Systems 2019-08-30 Adam Anderson , Steven R. Young , F. Kyle Reed , Jason M. Vann

Photoacoustic imaging (PAI) is an emerging non-invasive imaging modality combining the advantages of deep ultrasound penetration and high optical contrast. Image reconstruction is an essential topic in PAI, which is unfortunately an…

Image and Video Processing · Electrical Eng. & Systems 2019-08-06 Hengrong Lan , Daohuai Jiang , Changchun Yang , Fei Gao

Audio-Visual Segmentation (AVS) aims to segment sound-producing objects in video frames based on the associated audio signal. Prevailing AVS methods typically adopt an audio-centric Transformer architecture, where object queries are derived…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Shaofei Huang , Rui Ling , Tianrui Hui , Hongyu Li , Xu Zhou , Shifeng Zhang , Si Liu , Richang Hong , Meng Wang

We propose Mobile Audio Streaming Networks (MASnet) for efficient low-latency speech enhancement, which is particularly suitable for mobile devices and other applications where computational capacity is a limitation. MASnet processes…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-18 Michał Romaniuk , Piotr Masztalski , Karol Piaskowski , Mateusz Matuszewski

Ensuring intelligible speech communication for hearing assistive devices in low-latency scenarios presents significant challenges in terms of speech enhancement, coding and transmission. In this paper, we propose novel solutions for…

Audio and Speech Processing · Electrical Eng. & Systems 2024-05-01 Mohammad Bokaei , Jesper Jensen , Simon Doclo , Jan Østergaard

Visual information can serve as an effective cue for target speaker extraction (TSE) and is vital to improving extraction performance. In this paper, we propose AV-SepFormer, a SepFormer-based attention dual-scale model that utilizes cross-…

Satellite communications face severe bottlenecks in supporting high-fidelity synchronized audiovisual services, as conventional schemes struggle with cross-modal coherence under fluctuating channel conditions, limited bandwidth, and long…

Image and Video Processing · Electrical Eng. & Systems 2026-03-12 Fangyu Liu , Peiwen Jiang , Wenjin Wang , Chao-Kai Wen , Xiao Li , Shi Jin

This paper presents a novel neural vocoder named APNet which reconstructs speech waveforms from acoustic features by predicting amplitude and phase spectra directly. The APNet vocoder is composed of an amplitude spectrum predictor (ASP) and…

Sound · Computer Science 2023-05-16 Yang Ai , Zhen-Hua Ling

Given the rapid development of 3D scanners, point clouds are becoming popular in AI-driven machines. However, point cloud data is inherently sparse and irregular, causing significant difficulties for machine perception. In this work, we…

Computer Vision and Pattern Recognition · Computer Science 2022-10-04 Shi Qiu , Saeed Anwar , Nick Barnes

Analog joint source-channel coding (JSCC) has demonstrated superior performance for semantic communications through graceful degradation across channel conditions. However, a fundamental hardware-software mismatch prevents deployment on…

Information Theory · Computer Science 2026-03-11 Shumin Yao , Hao Chen , Yaping Sun , Nan Ma , Xiaodong Xu , Qinglin Zhao , Shuguang Cui

The goal of this work is to develop a meeting transcription system that can recognize speech even when utterances of different speakers are overlapped. While speech overlaps have been regarded as a major obstacle in accurately transcribing…

Audio and Speech Processing · Electrical Eng. & Systems 2018-10-10 Takuya Yoshioka , Hakan Erdogan , Zhuo Chen , Xiong Xiao , Fil Alleva

We introduce a framework for designing multi-scale, adaptive, shift-invariant frames and bi-frames for representing signals. The new framework, called AdaFrame, improves over dictionary learning-based techniques in terms of computational…

Computer Vision and Pattern Recognition · Computer Science 2015-07-20 Cheng Tai , Weinan E

This paper proposes a novel Semantic Communication (SemCom) framework for real-time adaptive-bitrate video streaming by integrating Latent Diffusion Models (LDMs) within the FFmpeg techniques. This solution addresses the challenges of high…

Multimedia · Computer Science 2025-08-04 Zijiang Yan , Jianhua Pei , Hongda Wu , Hina Tabassum , Ping Wang

We introduce a state-of-the-art real-time, high-fidelity, audio codec leveraging neural networks. It consists in a streaming encoder-decoder architecture with quantized latent space trained in an end-to-end fashion. We simplify and speed-up…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-25 Alexandre Défossez , Jade Copet , Gabriel Synnaeve , Yossi Adi

Hearing-impaired individuals often face significant barriers in daily communication due to the inherent challenges of producing clear speech. To address this, we introduce the Omni-Model paradigm into assistive technology and present…

Computation and Language · Computer Science 2025-11-17 Zhiming Ma , Shiyu Gan , Junhao Zhao , Xianming Li , Qingyun Pan , Peidong Wang , Mingjun Pan , Yuhao Mo , Jiajie Cheng , Chengxin Chen , Zhonglun Cao , Chonghan Liu , Shi Cheng

Full-duplex speech interaction, as the most natural and intuitive mode of human communication, is driving artificial intelligence toward more human-like conversational systems. Traditional cascaded speech processing pipelines suffer from…

Artificial Intelligence · Computer Science 2026-05-01 Yadong Li , Guoxin Wu , Haiping Hou , Biye Li

Recent advances in video diffusion models have significantly improved visual quality, yet ultra-high-resolution (UHR) video generation remains a formidable challenge due to the compounded difficulties of motion modeling, semantic planning,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-28 Chen Zhao , Jiawei Chen , Hongyu Li , Zhuoliang Kang , Shilin Lu , Xiaoming Wei , Kai Zhang , Jian Yang , Ying Tai

In this paper, a cloud radio access network (Cloud-RAN) based collaborative edge AI inference architecture is proposed. Specifically, geographically distributed devices capture real-time noise-corrupted sensory data samples and extract the…

Information Theory · Computer Science 2024-04-10 Pengfei Zhang , Dingzhu Wen , Guangxu Zhu , Qimei Chen , Kaifeng Han , Yuanming Shi

Speech Foundation Models have gained significant attention recently. Prior works have shown that the fusion of representations from multiple layers of the same model or the fusion of multiple models can improve performance on downstream…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-12 Yi-Jen Shih , David Harwath