中文
相关论文

相关论文: SalsaAgent: A multimodal embodied language model f…

200 篇论文

Despite advancements in Large Language Models (LLMs) and Large Multimodal Models (LMMs), their integration into language-grounded, human-like embodied agents remains incomplete, hindering complex real-life task performance in physical…

计算与语言 · 计算机科学 2024-08-20 Zhili Cheng , Zhitong Wang , Jinyi Hu , Shengding Hu , An Liu , Yuge Tu , Pengkai Li , Lei Shi , Zhiyuan Liu , Maosong Sun

Immersive rooms are increasingly popular augmented reality systems that support multi-agent interactions within a virtual world. However, despite extensive content creation and technological developments, insights about perceptually-driven…

人机交互 · 计算机科学 2025-12-22 Jerry M. Huang , Stefan T. Radev

We investigate the use of Large Language Models (LLMs) to equip neural robotic agents with human-like social and cognitive competencies, for the purpose of open-ended human-robot conversation and collaboration. We introduce a modular and…

机器人学 · 计算机科学 2024-09-30 Philipp Allgeuer , Hassan Ali , Stefan Wermter

Humans possess the innate ability to extract latent visuo-lingual cues to infer context through human interaction. During collaboration, this enables proactive prediction of the underlying intention of a series of tasks. In contrast,…

机器人学 · 计算机科学 2023-10-05 Pranay Mathur

Traditional visual storytelling is complex, requiring specialized knowledge and substantial resources, yet often constrained by human creativity and creation precision. While Large Language Models (LLMs) enhance visual storytelling, current…

计算机视觉与模式识别 · 计算机科学 2024-08-22 Yuzhou Huang , Yiran Qin , Shunlin Lu , Xintao Wang , Rui Huang , Ying Shan , Ruimao Zhang

A key component of dyadic spoken interactions is the contextually relevant non-verbal gestures, such as head movements that reflect a listener's response to the interlocutor's speech. Although significant progress has been made in the…

机器人学 · 计算机科学 2024-10-01 Bishal Ghosh , Emma Li , Tanaya Guha

Recent progress in large language model (LLM) technology has significantly enhanced the interaction experience between humans and voice assistants (VAs). This project aims to explore a user's continuous interaction with LLM-based VA…

人机交互 · 计算机科学 2024-09-04 Szeyi Chan , Shihan Fu , Jiachen Li , Bingsheng Yao , Smit Desai , Mirjana Prpa , Dakuo Wang

Speech is essential for human communication, yet millions of people face impairments such as dysarthria, stuttering, and aphasia conditions that often lead to social isolation and reduced participation. Despite recent progress in automatic…

系统与控制 · 电气工程与系统科学 2025-10-24 Haowei Lou , Chengkai Huang , Hye-young Paik , Yongquan Hu , Aaron Quigley , Wen Hu , Lina Yao

Recent text-to-image (T2I) models have made remarkable progress in generating visually realistic and semantically coherent images. However, they still suffer from randomness and inconsistency with the given prompts, particularly when…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Kaishen Wang , Ruibo Chen , Tong Zheng , Heng Huang

Video generation has advanced rapidly, producing photorealistic videos from text or image prompts. Meanwhile, film production and social robotics increasingly demand multi-person videos with rich social interactions, including…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Liangyang Ouyang , Ruicong Liu , Caixin Kang , Yifei Huang , Yoichi Sato

This paper proposes a multi-agent artificial intelligence system that generates response-oriented media content in real time based on audio-derived emotional signals. Unlike conventional speech emotion recognition studies that focus…

人工智能 · 计算机科学 2026-01-21 HyeYoung Lee

Bilingual text-to-motion generation, which synthesizes 3D human motions from bilingual text inputs, holds immense potential for cross-linguistic applications in gaming, film, and robotics. However, this task faces critical challenges: the…

计算机视觉与模式识别 · 计算机科学 2025-08-04 Wanjiang Weng , Xiaofeng Tan , Hongsong Wang , Pan Zhou

The Agent and AIGC (Artificial Intelligence Generated Content) technologies have recently made significant progress. We propose AesopAgent, an Agent-driven Evolutionary System on Story-to-Video Production. AesopAgent is a practical…

计算机视觉与模式识别 · 计算机科学 2024-03-14 Jiuniu Wang , Zehua Du , Yuyuan Zhao , Bo Yuan , Kexiang Wang , Jian Liang , Yaxi Zhao , Yihen Lu , Gengliang Li , Junlong Gao , Xin Tu , Zhenyu Guo

We present a novel Speech Augmented Language Model (SALM) with {\em multitask} and {\em in-context} learning capabilities. SALM comprises a frozen text LLM, a audio encoder, a modality adapter module, and LoRA layers to accommodate speech…

Despite recent advances in multimodal large language models (MLLMs), their ability to understand and interact with music remains limited. Music understanding requires grounded reasoning over symbolic scores and expressive performance audio,…

多媒体 · 计算机科学 2026-01-21 Qihao Zhao , Yunqi Cao , Yangyu Huang , Hui Yi Leong , Fan Zhang , Kim-Hui Yap , Wei Hu

Although there has been rapid progress in endowing robots with the ability to solve complex manipulation tasks, generating control policies for bimanual robots to solve tasks involving two hands is still challenging because of the…

机器人学 · 计算机科学 2024-10-11 Kun Chu , Xufeng Zhao , Cornelius Weber , Mengdi Li , Wenhao Lu , Stefan Wermter

Deploying humanoid robots in real-world settings is fundamentally challenging, as it demands tight integration of perception, locomotion, and manipulation under partial-information observations and dynamically changing environments. As well…

机器人学 · 计算机科学 2026-02-05 Yu Bai , MingMing Yu , Chaojie Li , Ziyi Bai , Xinlong Wang , Börje F. Karlsson

Significant progress has been made in vision-language models. However, language-conditioned robotic manipulation for contact-rich tasks remains underexplored, particularly in terms of tactile sensing. To address this gap, we introduce the…

机器人学 · 计算机科学 2025-03-12 Peng Hao , Chaofan Zhang , Dingzhe Li , Xiaoge Cao , Xiaoshuai Hao , Shaowei Cui , Shuo Wang

Language-guided scene-aware human motion generation has great significance for entertainment and robotics. In response to the limitations of existing datasets, we introduce LaserHuman, a pioneering dataset engineered to revolutionize…

计算机视觉与模式识别 · 计算机科学 2024-03-22 Peishan Cong , Ziyi Wang , Zhiyang Dou , Yiming Ren , Wei Yin , Kai Cheng , Yujing Sun , Xiaoxiao Long , Xinge Zhu , Yuexin Ma

Designing realistic multi-object scenes requires not only generating images, but also planning spatial layouts that respect semantic relations and physical plausibility. On one hand, while recent advances in diffusion models have enabled…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Zezhong Fan , Xiaohan Li , Luyi Ma , Kai Zhao , Liang Peng , Topojoy Biswas , Evren Korpeoglu , Kaushiki Nag , Kannan Achan