English
Related papers

Related papers: PodAgent: A Comprehensive Framework for Podcast Ge…

200 papers

Automatically generating presentations from documents is a challenging task that requires accommodating content quality, visual appeal, and structural coherence. Existing methods primarily focus on improving and evaluating the content…

Artificial Intelligence · Computer Science 2025-02-24 Hao Zheng , Xinyan Guan , Hao Kong , Jia Zheng , Weixiang Zhou , Hongyu Lin , Yaojie Lu , Ben He , Xianpei Han , Le Sun

This paper analyses AI-generated podcasts produced by Google's NotebookLM, which generates audio podcasts with two chatty AI hosts discussing whichever documents a user uploads. While AI-generated podcasts have been discussed as tools, for…

Computers and Society · Computer Science 2026-05-26 Jill Walker Rettberg

Audio is an essential part of our life, but creating it often requires expertise and is time-consuming. Research communities have made great progress over the past year advancing the performance of large scale audio generative models for a…

Despite recent breakthroughs, audio foundation models struggle in processing complex multi-source acoustic scenes. We refer to this challenging domain as audio stories, which can have multiple speakers and background/foreground sound…

High-quality speech dialogue datasets are crucial for Speech-LLM development, yet existing acquisition methods face significant limitations. Human recordings incur high costs and privacy concerns, while synthetic approaches often lack…

Computation and Language · Computer Science 2025-04-01 Minghan Wang , Ye Bai , Yuxia Wang , Thuy-Trang Vu , Ehsan Shareghi , Gholamreza Haffari

Omnimodal large language models have made significant strides in unifying audio and visual modalities; however, they often face challenges in fine-grained cross-modal understanding and have difficulty with multimodal alignment. To address…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Keda Tao , Wenjie Du , Bohan Yu , Weiqiang Wang , Jian Liu , Huan Wang

AI-empowered music processing is a diverse field that encompasses dozens of tasks, ranging from generation tasks (e.g., timbre synthesis) to comprehension tasks (e.g., music classification). For developers and amateurs, it is very difficult…

Computation and Language · Computer Science 2023-10-26 Dingyao Yu , Kaitao Song , Peiling Lu , Tianyu He , Xu Tan , Wei Ye , Shikun Zhang , Jiang Bian

Text-to-audio (TTA) system has recently gained attention for its ability to synthesize general audio based on text descriptions. However, previous studies in TTA have limited generation quality with high computational costs. In this study,…

Sound · Computer Science 2023-09-12 Haohe Liu , Zehua Chen , Yi Yuan , Xinhao Mei , Xubo Liu , Danilo Mandic , Wenwu Wang , Mark D. Plumbley

Although audio generation shares commonalities across different types of audio, such as speech, music, and sound effects, designing models for each type requires careful consideration of specific objectives and biases that can significantly…

Human conversation involves language, speech, and visual cues, with each medium providing complementary information. For instance, speech conveys a vibe or tone not fully captured by text alone. While multimodal LLMs focus on generating…

Human-Computer Interaction · Computer Science 2025-09-19 Taesoo Kim , Yongsik Jo , Hyunmin Song , Taehwan Kim

Currently, high-quality, synchronized audio is synthesized using various multi-modal joint learning frameworks, leveraging video and optional text inputs. In the video-to-audio benchmarks, video-to-audio quality, semantic alignment, and…

Computer Vision and Pattern Recognition · Computer Science 2025-03-31 Haomin Zhang , Chang Liu , Junjie Zheng , Zihao Chen , Chaofan Ding , Xinhan Di

News podcasts are a popular medium to stay informed and dive deep into news topics. Today, most podcasts are handcrafted by professionals. In this work, we advance the state-of-the-art in automatically generated podcasts, making use of…

Human-Computer Interaction · Computer Science 2022-02-16 Philippe Laban , Elicia Ye , Srujay Korlakunta , John Canny , Marti A. Hearst

While previous approaches to 3D human motion generation have achieved notable success, they often rely on extensive training and are limited to specific tasks. To address these challenges, we introduce Motion-Agent, an efficient…

Computer Vision and Pattern Recognition · Computer Science 2024-10-08 Qi Wu , Yubo Zhao , Yifan Wang , Xinhang Liu , Yu-Wing Tai , Chi-Keung Tang

In this paper we introduce ResearchCodeAgent, a novel multi-agent system leveraging large language models (LLMs) agents to automate the codification of research methodologies described in machine learning literature. The system bridges the…

Software Engineering · Computer Science 2025-05-06 Shubham Gandhi , Dhruv Shah , Manasi Patwardhan , Lovekesh Vig , Gautam Shroff

Effective prompt design is essential for improving the planning capabilities of large language model (LLM)-driven agents. However, existing structured prompting strategies are typically limited to single-agent, plan-only settings, and often…

Artificial Intelligence · Computer Science 2025-07-08 Bruce Yang , Xinfeng He , Huan Gao , Yifan Cao , Xiaofan Li , David Hsu

Text-to-image generation has advanced rapidly, but existing models still struggle with faithfully composing multiple objects and preserving their attributes in complex scenes. We propose coDrawAgents, an interactive multi-agent dialogue…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Chunhan Li , Qifeng Wu , Jia-Hui Pan , Ka-Hei Hui , Jingyu Hu , Yuming Jiang , Bin Sheng , Xihui Liu , Wenjuan Gong , Zhengzhe Liu

Listening head generation aims to synthesize a non-verbal responsive listener head by modeling the correlation between the speaker and the listener in dynamic conversion.The applications of listener agent generation in virtual interaction…

Computer Vision and Pattern Recognition · Computer Science 2024-04-01 Xi Liu , Ying Guo , Cheng Zhen , Tong Li , Yingying Ao , Pengfei Yan

Recent advancements in AI-driven conversational agents have exhibited immense potential of AI applications. Effective response generation is crucial to the success of these agents. While extensive research has focused on leveraging multiple…

Computation and Language · Computer Science 2025-03-26 Junfeng Liu , Christopher T. Symons , Ranga Raju Vatsavai

High-quality personalized question banks are crucial for supporting adaptive learning and individualized assessment. Manually designing questions is time-consuming and often fails to meet diverse learning needs, making automated question…

Computers and Society · Computer Science 2025-11-18 Rui Jia , Min Zhang , Fengrui Liu , Bo Jiang , Kun Kuang , Zhongxiang Dai

Recently, the AI community has made significant strides in developing powerful foundation models, driven by large-scale multimodal datasets. However, for audio representation learning, existing datasets suffer from limitations in the…

Sound · Computer Science 2024-09-10 Luoyi Sun , Xuenan Xu , Mengyue Wu , Weidi Xie