English
Related papers

Related papers: STASE: A spatialized text-to-audio synthesis engin…

200 papers

Simultaneous localization and mapping (SLAM) is a critical technology that enables autonomous robots to be aware of their surrounding environment. With the development of deep learning, SLAM systems can achieve a higher level of perception…

Recent advances in large language models have demonstrated impressive capabilities in task-oriented applications, yet building emotionally intelligent chatbots that can engage in natural, strategic conversations remains a challenge. We…

Computation and Language · Computer Science 2025-07-02 Yizhe Zhang , Navdeep Jaitly

High-quality singing annotations are fundamental to modern Singing Voice Synthesis (SVS) systems. However, obtaining these annotations at scale through manual labeling is unrealistic due to the substantial labor and musical expertise…

Sound · Computer Science 2026-05-07 Yukun Chen , Tianrui Wang , Zhaoxi Mu , Xinyu Yang , EngSiong Chng

Stance detection plays a pivotal role in enabling an extensive range of downstream applications, from discourse parsing to tracing the spread of fake news and the denial of scientific facts. While most stance classification models rely on…

Computation and Language · Computer Science 2024-12-13 Guy Barel , Oren Tsur , Dan Vilenchik

In this paper, we introduce ConversaSynth, a framework designed to generate synthetic conversation audio using large language models (LLMs) with multiple persona settings. The framework first creates diverse and coherent text-based…

Sound · Computer Science 2025-07-08 Kaung Myat Kyaw , Jonathan Hoyin Chan

Multimodal large language models (MLLMs) play a pivotal role in advancing the quest for general artificial intelligence. However, achieving unified target for multimodal understanding and generation remains challenging due to optimization…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Jie Qin , Jiancheng Huang , Limeng Qiao , Lin Ma

Large Language Models (LLMs) have shown strong general capabilities in many applications. However, how to make them reliable tools for some specific tasks such as automated short answer grading (ASAG) remains a challenge. We present SteLLA…

Computation and Language · Computer Science 2025-05-26 Hefei Qiu , Brian White , Ashley Ding , Reinaldo Costa , Ali Hachem , Wei Ding , Ping Chen

Electroencephalogram (EEG) technology, particularly high-density EEG (HD EEG) devices, is widely used in fields such as neuroscience. HD EEG devices improve the spatial resolution of EEG by placing more electrodes on the scalp, which meet…

Signal Processing · Electrical Eng. & Systems 2025-02-25 Shuqiang Wang , Tong Zhou , Yanyan Shen , Ye Li , Guoheng Huang , Yong Hu

While the spatial directivity of multichannel speech enhancement algorithms improves with the number of microphones, fitting large capture arrays into real-world edge devices is typically limited by physical constraints. To overcome this…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-08 Dongheon Lee , Ashutosh Pandey , Sanjeel Parekh , Daniel Wong , Jacob Donley , Buye Xu , Juan Azcarreta

Sound modelling is the process of developing algorithms that generate sound under parametric control. There are a few distinct approaches that have been developed historically including modelling the physics of sound production and…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-26 M. Huzaifah , L. Wyse

Large audio-language models (LALMs) can generate reasoning chains for their predictions, but it remains unclear whether these reasoning chains remain grounded in the input audio. In this paper, we propose an RL-based strategy that grounds…

Sound · Computer Science 2026-03-23 Jihoon Jeong , Pooneh Mousavi , Mirco Ravanelli , Cem Subakan

LLM-based automatic speech recognition models demonstrate strong performance by connecting audio encoders and LLMs. However, data scarcity of paired speech and transcription often hinders their adaptation to new domains, making text-only…

Sound · Computer Science 2026-05-15 Ryo Magoshi , Takashi Maekaku , Yusuke Shinohara

Knowledge editing enables targeted updates without retraining, but prior work focuses on textual or visual facts, leaving abstract auditory perceptual knowledge underexplored. We introduce SAKE, the first benchmark for editing perceptual…

A fundamental challenge in embodied AI is verifying if agents build internal models of spatial structure or merely learn to mimic task-specific expert trajectories. This is critical as foundational approaches rooted in action-centric tasks…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Jinzhou Tang , Sidi Liu , Waikit Xiu , Weixing Chen , Keze Wang

Recent advancements in Latent Diffusion Models (LDMs) have propelled them to the forefront of various generative tasks. However, their iterative sampling process poses a significant computational burden, resulting in slow generation speeds…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-10 Huadai Liu , Rongjie Huang , Yang Liu , Hengyuan Cao , Jialei Wang , Xize Cheng , Siqi Zheng , Zhou Zhao

The recent success in StyleGAN demonstrates that pre-trained StyleGAN latent space is useful for realistic video generation. However, the generated motion in the video is usually not semantically meaningful due to the difficulty of…

Computer Vision and Pattern Recognition · Computer Science 2022-10-24 Seung Hyun Lee , Gyeongrok Oh , Wonmin Byeon , Chanyoung Kim , Won Jeong Ryoo , Sang Ho Yoon , Hyunjun Cho , Jihyun Bae , Jinkyu Kim , Sangpil Kim

Speech understanding is essential for interpreting the diverse forms of information embedded in spoken language, including linguistic, paralinguistic, and non-linguistic cues that are vital for effective human-computer interaction. The…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-08 Jing Peng , Yucheng Wang , Bohan Li , Yiwei Guo , Hankun Wang , Yangui Fang , Yu Xi , Haoyu Li , Xu Li , Ke Zhang , Shuai Wang , Kai Yu

Large language models deliver strong generative performance but at the cost of massive parameter counts, memory use, and decoding latency. Prior work has shown that pruning and structured sparsity can preserve accuracy under substantial…

Computation and Language · Computer Science 2026-04-17 Andrew Kiruluta

With recent advances of AIGC, video generation have gained a surge of research interest in both academia and industry (e.g., Sora). However, it remains a challenge to produce temporally aligned audio to synchronize the generated video,…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-24 Yuchen Hu , Yu Gu , Chenxing Li , Rilin Chen , Dong Yu

Humans can easily isolate a single speaker from a complex acoustic environment, a capability referred to as the "Cocktail Party Effect." However, replicating this ability has been a significant challenge in the field of target speaker…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-08 Xiang Hao , Jibin Wu , Jianwei Yu , Chenglin Xu , Kay Chen Tan
‹ Prev 1 8 9 10 Next ›