中文
相关论文

相关论文: CaST: A Toolchain for Creating and Characterizing …

200 篇论文

Advanced auditory models are useful in designing signal-processing algorithms for hearing-loss compensation or speech enhancement. Such auditory models provide rich and detailed descriptions of the auditory pathway, and might allow for…

音频与语音处理 · 电气工程与系统科学 2024-03-18 Peter Leer , Jesper Jensen , Zheng-Hua Tan , Jan Østergaard , Lars Bramsløw

We present a novel machine-learning (ML) approach (EM-GANSim) for real-time electromagnetic (EM) propagation that is used for wireless communication simulation in 3D indoor environments. Our approach uses a modified conditional Generative…

机器学习 · 计算机科学 2025-04-29 Ruichen Wang , Dinesh Manocha

In this work, we focus on front-end design for speech deepfake detectors, the component that determines the discriminative acoustic cues provided to the classifier. Existing approaches are primarily categorized into two types. Hand-crafted…

音频与语音处理 · 电气工程与系统科学 2026-05-01 Xi Xuan , Davide Carbone , Wenxin Zhang , Ruchi Pandey , Tomi H. Kinnunen

The evolution of the Radio Access Network (RAN) in 5G and 6G technologies marks a shift toward open, programmable, and softwarized architectures, driven by the Open RAN paradigm. This approach emphasizes open interfaces for telemetry…

网络与互联网体系结构 · 计算机科学 2026-01-28 Davide Villa

With the development of large-scale diffusion-based and language-modeling-based generative models, impressive progress has been achieved in text-to-audio generation. Despite producing high-quality outputs, existing text-to-audio models…

声音 · 计算机科学 2026-04-28 Yi Yuan , Xubo Liu , Haohe Liu , Xiyuan Kang , Zhuo Chen , Yuxuan Wang , Mark D. Plumbley , Wenwu Wang

This paper presents a simulation platform, namely CIMulator, for quantifying the efficacy of various synaptic devices in neuromorphic accelerators for different neural network architectures. Nonvolatile memory devices, such as resistive…

In this paper, we design a new class of high-efficiency deep joint source-channel coding methods to achieve end-to-end video transmission over wireless channels. The proposed methods exploit nonlinear transform and conditional coding…

计算机视觉与模式识别 · 计算机科学 2022-11-03 Sixian Wang , Jincheng Dai , Zijian Liang , Kai Niu , Zhongwei Si , Chao Dong , Xiaoqi Qin , Ping Zhang

Strong authentication in an interconnected wireless environment continues to be an important, but sometimes elusive goal. Research in physical-layer authentication using channel features holds promise as a technique to improve network…

信号处理 · 电气工程与系统科学 2020-09-01 Ken St. Germain , Frank Kragh

We propose Cross-Attention in Audio, Space, and Time (CA^2ST), a transformer-based method for holistic video recognition. Recognizing actions in videos requires both spatial and temporal understanding, yet most existing models lack a…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Jongseo Lee , Joohyun Chang , Dongho Lee , Jinwoo Choi

Streaming speech enhancement is a crucial task for real-time applications such as online meetings, smart home appliances, and hearing aids. Deep neural network-based approaches achieve exceptional performance while demanding substantial…

音频与语音处理 · 电气工程与系统科学 2025-09-29 Sunghwan Ahn , Jinmo Han , Beom Jun Woo , Nam Soo Kim

In this paper, we present an advanced channel sounding system designed for sensing and propagation experiments in all types of cellular deployment scenarios. The system's exceptional adaptability, high resolution, and sensitivity makes it…

信息论 · 计算机科学 2025-10-02 K. F. Nieman , O. Kanhere , R. Shiu , W. Xu , C. Duan , S. S. Ghassemzadeh

A significant amount of research literature is dedicated to interference mitigation in Wireless Mesh Networks (WMNs), with a special emphasis on designing channel allocation (CA) schemes which alleviate the impact of interference on WMN…

网络与互联网体系结构 · 计算机科学 2018-08-20 Srikant Manas Kala , Vanlin Sathya , M Pavan Kumar Reddy , Betty Lala , Bheemarjuna Reddy Tamma

Fluid antenna systems (FAS) have emerged as a promising technology for next-generation wireless systems. However, practical multiuser multiple-input multiple-output FAS (MIMO-FAS) faces two inherently coupled challenges: acquiring accurate…

信息论 · 计算机科学 2026-05-29 Erqiang Tang , Wei Guo , Hengtao He , Shenghui Song , Jun Zhang , Khaled B. Letaief

Simultaneous speech translation (SimulST) systems must balance translation quality with response time, making latency measurement crucial for evaluating their real-world performance. However, there has been a longstanding belief that…

计算与语言 · 计算机科学 2024-10-22 Xi Xu , Wenda Xu , Siqi Ouyang , Lei Li

Developing algorithms for sound classification, detection, and localization requires large amounts of flexible and realistic audio data, especially when leveraging modern machine learning and beamforming techniques. However, most existing…

音频与语音处理 · 电气工程与系统科学 2026-01-23 Luca Barbisan , Marco Levorato , Fabrizio Riente

This document provides a brief description of the National Institute of Standards and Technology (NIST) speaker recognition evaluation (SRE) conversational telephone speech (CTS) Superset. The CTS Superset has been created in an attempt to…

声音 · 计算机科学 2021-08-17 Seyed Omid Sadjadi

This paper introduces a novel task in generative speech processing, Acoustic Scene Transfer (AST), which aims to transfer acoustic scenes of speech signals to diverse environments. AST promises an immersive experience in speech perception…

音频与语音处理 · 电气工程与系统科学 2024-06-19 Miseul Kim , Soo-Whan Chung , Youna Ji , Hong-Goo Kang , Min-Seok Choi

Brain-computer interface (BCI) research, while promising, has largely been confined to static and fixed environments, limiting real-world applicability. To move towards practical BCI, we introduce a real-time wireless imagined speech…

人工智能 · 计算机科学 2025-11-12 Ji-Ha Park , Heon-Gyu Kwak , Gi-Hwan Shin , Yoo-In Jeon , Sun-Min Park , Ji-Yeon Hwang , Seong-Whan Lee

Audio foundation models learn general-purpose audio representations that facilitate a wide range of downstream tasks. While the performance of these models has greatly increased for conventional single-channel, dry audio clips, their…

声音 · 计算机科学 2026-02-05 Goksenin Yuksel , Marcel van Gerven , Kiki van der Heijden

Recovering high-quality 3D scenes from a single RGB image is a challenging task in computer graphics. Current methods often struggle with domain-specific limitations or low-quality object generation. To address these, we propose CAST…

计算机视觉与模式识别 · 计算机科学 2025-05-14 Kaixin Yao , Longwen Zhang , Xinhao Yan , Yan Zeng , Qixuan Zhang , Wei Yang , Lan Xu , Jiayuan Gu , Jingyi Yu