中文
相关论文

相关论文: DRIVE: Disfluency-Rich Synthetic Dialog Data Gener…

200 篇论文

Compared with CrossWOZ (Chinese) and MultiWOZ (English) dataset which have coarse-grained information, there is no dataset which handle fine-grained and hierarchical level information properly. In this paper, we publish a first Cantonese…

计算与语言 · 计算机科学 2021-12-15 Hongru Wang , Min Li , Zimo Zhou , Gabriel Pui Cheong Fung , Kam-Fai Wong

Recent advancements in autonomous vehicles (AVs) use Large Language Models (LLMs) to perform well in normal driving scenarios. However, ensuring safety in dynamic, high-risk environments and managing safety-critical long-tail events remain…

人工智能 · 计算机科学 2024-12-20 Zhiyuan Zhou , Heye Huang , Boqi Li , Shiyue Zhao , Yao Mu , Jianqiang Wang

Large-scale labelled driving video data is essential for training autonomous driving systems. Although simulation offers scalable and fully annotated data, the domain gap between synthetic and real-world driving videos significantly limits…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Haonan Zhao , Yiting Wang , Jingkun Chen , Valentina Donzella , Thomas Bashford-Rogers , Kurt Debattista

Pragmatic reasoning plays a pivotal role in deciphering implicit meanings that frequently arise in real-life conversations and is essential for the development of communicative social agents. In this paper, we introduce a novel challenge,…

计算与语言 · 计算机科学 2023-06-21 Hengli Li , Song-Chun Zhu , Zilong Zheng

We release SVIRO, a synthetic dataset for sceneries in the passenger compartment of ten different vehicles, in order to analyze machine learning-based approaches for their generalization capacities and reliability when trained on a limited…

计算机视觉与模式识别 · 计算机科学 2020-01-13 Steve Dias Da Cruz , Oliver Wasenmüller , Hans-Peter Beise , Thomas Stifter , Didier Stricker

Using generative models to synthesize new data has become a de-facto standard in autonomous driving to address the data scarcity issue. Though existing approaches are able to boost perception models, we discover that these approaches fail…

计算机视觉与模式识别 · 计算机科学 2024-06-07 Enhui Ma , Lijun Zhou , Tao Tang , Zhan Zhang , Dong Han , Junpeng Jiang , Kun Zhan , Peng Jia , Xianpeng Lang , Haiyang Sun , Di Lin , Kaicheng Yu

In the design of traffic monitoring solutions for optimizing the urban mobility infrastructure, acoustic vehicle counting models have received attention due to their cost effectiveness and energy efficiency. Although deep learning has…

音频与语音处理 · 电气工程与系统科学 2024-01-18 Stefano Damiano , Luca Bondi , Shabnam Ghaffarzadegan , Andre Guntoro , Toon van Waterschoot

Learning-based planners generate natural human-like driving behaviors by learning to reason about nuanced interactions from data, overcoming the rigid behaviors that arise from rule-based planners. Nonetheless, data-driven approaches often…

机器人学 · 计算机科学 2025-06-02 Wenhao Ding , Sushant Veer , Yuxiao Chen , Yulong Cao , Chaowei Xiao , Marco Pavone

Autonomous driving necessitates the ability to reason about future interactions between traffic agents and to make informed evaluations for planning. This paper introduces the \textit{Gen-Drive} framework, which shifts from the traditional…

机器人学 · 计算机科学 2024-10-10 Zhiyu Huang , Xinshuo Weng , Maximilian Igl , Yuxiao Chen , Yulong Cao , Boris Ivanovic , Marco Pavone , Chen Lv

As sharing images in an instant message is a crucial factor, there has been active research on learning an image-text multi-modal dialogue models. However, training a well-generalized multi-modal dialogue model remains challenging due to…

计算机视觉与模式识别 · 计算机科学 2024-04-01 Young-Jun Lee , Byungsoo Ko , Han-Gyu Kim , Jonghwan Hyeon , Ho-Jin Choi

Text-to-image synthesis models require the ability to generate diverse images while maintaining stability. To overcome this challenge, a number of methods have been proposed, including the collection of prompt-image datasets and the…

计算机视觉与模式识别 · 计算机科学 2025-11-17 Keunwoo Park , Jihye Chae , Joong Ho Ahn , Jihoon Kweon

Recent advancements in video diffusion models have significantly enhanced audio-driven portrait animation. However, current methods still suffer from flickering, identity drift, and poor audio-visual synchronization. These issues primarily…

计算机视觉与模式识别 · 计算机科学 2025-12-19 Zhenjie Liu , Jianzhang Lu , Renjie Lu , Cong Liang , Shangfei Wang

Disfluencies is an under-studied topic in NLP, even though it is ubiquitous in human conversation. This is largely due to the lack of datasets containing disfluencies. In this paper, we present a new challenge question answering dataset,…

计算与语言 · 计算机科学 2021-06-09 Aditya Gupta , Jiacheng Xu , Shyam Upadhyay , Diyi Yang , Manaal Faruqui

With increasing focus on privacy protection, alternative methods to identify vehicle operator without the use of biometric identifiers have gained traction for automotive data analysis. The wide variety of sensors installed on modern…

机器学习 · 计算机科学 2021-02-11 Jingbo Yang , Ruge Zhao , Meixian Zhu , David Hallac , Jaka Sodnik , Jure Leskovec

Evaluating and training autonomous driving systems require diverse and scalable corner cases. However, most existing scene generation methods lack controllability, accuracy, and versatility, resulting in unsatisfactory generation results.…

机器人学 · 计算机科学 2024-10-11 Sheng Wang , Ge Sun , Fulong Ma , Tianshuai Hu , Qiang Qin , Yongkang Song , Lei Zhu , Junwei Liang

End-to-End (E2E) solutions have emerged as a mainstream approach for autonomous driving systems, with Vision-Language-Action (VLA) models representing a new paradigm that leverages pre-trained multimodal knowledge from Vision-Language…

机器人学 · 计算机科学 2025-09-25 Pengxiang Li , Yinan Zheng , Yue Wang , Huimin Wang , Hang Zhao , Jingjing Liu , Xianyuan Zhan , Kun Zhan , Xianpeng Lang

Traffic collision reconstruction traditionally relies on human expertise and can be accurate, but pre-crash reconstruction is more challenging. This study develops a multi-agent AI framework that reconstructs pre-crash scenarios and infers…

人工智能 · 计算机科学 2026-04-03 Gerui Xu , Boyou Chen , Huizhong Guo , Dave LeBlanc , Arpan Kusari , Efe Yarbasi , Ananna Ahmed , Zhaonan Sun , Shan Bao

Full-duplex, spontaneous conversational data are essential for enhancing the naturalness and interactivity of synthesized speech in conversational TTS systems. We present two open-source dual-track conversational speech datasets, one in…

声音 · 计算机科学 2025-09-05 Zhitong Zhou , Qingqing Zhang , Lei Luo , Jiechen Liu , Ruohua Zhou

Autonomous systems conducting schema-grounded information-gathering dialogues face an instrumentation gap, lacking turn-level observables for monitoring acquisition efficiency and detecting when questioning becomes unproductive. We…

计算与语言 · 计算机科学 2026-01-15 Dimitris Panagopoulos , Adolfo Perrusquia , Weisi Guo

Physicians spend significant time documenting clinical encounters, a burden that contributes to professional burnout. To address this, robust automation tools for medical documentation are crucial. We introduce MedSynth -- a novel dataset…

计算与语言 · 计算机科学 2025-08-05 Ahmad Rezaie Mianroodi , Amirali Rezaie , Niko Grisel Todorov , Cyril Rakovski , Frank Rudzicz