中文
相关论文

相关论文: D-ORCA: Dialogue-Centric Optimization for Robust A…

200 篇论文

Audio captioning aims to generate text descriptions of audio clips. In the real world, many objects produce similar sounds. How to accurately recognize ambiguous sounds is a major challenge for audio captioning. In this work, inspired by…

音频与语音处理 · 电气工程与系统科学 2023-05-30 Xubo Liu , Qiushi Huang , Xinhao Mei , Haohe Liu , Qiuqiang Kong , Jianyuan Sun , Shengchen Li , Tom Ko , Yu Zhang , Lilian H. Tang , Mark D. Plumbley , Volkan Kılıç , Wenwu Wang

Dense video captioning (DVC) aims to generate multi-sentence descriptions to elucidate the multiple events in the video, which is challenging and demands visual consistency, discoursal coherence, and linguistic diversity. Existing methods…

计算机视觉与模式识别 · 计算机科学 2021-11-22 Xu Yan , Zhengcong Fei , Shuhui Wang , Qingming Huang , Qi Tian

Grounding language in the physical world requires AI systems to interpret references that emerge dynamically during conversation. While current vision-language models (VLMs) excel at static image tasks, they struggle to resolve ambiguous…

计算机视觉与模式识别 · 计算机科学 2026-05-22 Anna Deichler , Jim O'Regan , Fethiye Irmak Dogan , Lubos Marcinek , Anna Klezovich , Iolanda Leite , Jonas Beskow

Automated audio captioning (AAC), a task that mimics human perception as well as innovatively links audio processing and natural language processing, has overseen much progress over the last few years. AAC requires recognizing contents such…

声音 · 计算机科学 2023-11-17 Xuenan Xu , Zeyu Xie , Mengyue Wu , Kai Yu

Generative AI has significantly changed industries by enabling text-driven image generation, yet challenges remain in achieving high-resolution outputs that align with fine-grained user preferences. Consequently, multi-round interactions…

计算机视觉与模式识别 · 计算机科学 2025-04-28 Kun Li , Jianhui Wang , Yangfan He , Xinyuan Song , Ruoyu Wang , Hongyang He , Wenxin Zhang , Jiaqi Chen , Keqin Li , Sida Li , Miao Zhang , Tianyu Shi , Xueqian Wang

In text-video retrieval, auxiliary captions are often used to enhance video understanding, bridging the gap between the modalities. While recent advances in multi-modal large language models (MLLMs) have enabled strong zero-shot caption…

计算机视觉与模式识别 · 计算机科学 2025-09-23 Ji Soo Lee , Byungoh Ko , Jaewon Cho , Howoong Lee , Jaewoon Byun , Hyunwoo J. Kim

Recent focus in video captioning has been on designing architectures that can consume both video and text modalities, and using large-scale video datasets with text transcripts for pre-training, such as HowTo100M. Though these approaches…

计算机视觉与模式识别 · 计算机科学 2023-06-23 Yuhan Shen , Linjie Yang , Longyin Wen , Haichao Yu , Ehsan Elhamifar , Heng Wang

Goal-oriented proactive dialogue systems are designed to guide user conversations seamlessly towards specific objectives by planning a goal-oriented path. However, previous research has focused predominantly on optimizing these paths while…

计算与语言 · 计算机科学 2025-06-19 Didi Zhang , Yaxin Fan , Peifeng Li , Qiaoming Zhu

Measurement of interaction quality is a critical task for the improvement of spoken dialog systems. Existing approaches to dialog quality estimation either focus on evaluating the quality of individual turns, or collect dialog-level quality…

The proliferation of hour-long videos (e.g., lectures, podcasts, documentaries) has intensified demand for efficient content structuring. However, existing approaches are constrained by small-scale training with annotations that are typical…

计算机视觉与模式识别 · 计算机科学 2025-11-19 Junfu Pu , Teng Wang , Yixiao Ge , Yuying Ge , Chen Li , Ying Shan

Classifying the general intent of the user utterance in a conversation, also known as Dialogue Act (DA), e.g., open-ended question, statement of opinion, or request for an opinion, is a key step in Natural Language Understanding (NLU) for…

计算与语言 · 计算机科学 2020-05-29 Ali Ahmadvand , Jason Ingyu Choi , Eugene Agichtein

Open-vocabulary scene understanding aims to localize and recognize unseen categories beyond the annotated label space. The recent breakthrough of 2D open-vocabulary perception is largely driven by Internet-scale paired image-text data with…

计算机视觉与模式识别 · 计算机科学 2023-03-23 Runyu Ding , Jihan Yang , Chuhui Xue , Wenqing Zhang , Song Bai , Xiaojuan Qi

Classroom discourse is an essential vehicle through which teaching and learning take place. Assessing different characteristics of discursive practices and linking them to student learning achievement enhances the understanding of teaching…

计算机与社会 · 计算机科学 2025-05-14 Ruikun Hou , Babette Bühler , Tim Fütterer , Efe Bozkir , Peter Gerjets , Ulrich Trautwein , Enkelejda Kasneci

When beginners learn to speak a non-native language, it is difficult for them to judge for themselves whether they are speaking well. Therefore, computer-assisted pronunciation training systems are used to detect learner mispronunciations.…

音频与语音处理 · 电气工程与系统科学 2022-12-12 Kazuki Kawamura , Jun Rekimoto

Multi-source localization is an important and challenging technique for multi-talker conversation analysis. This paper proposes a novel supervised learning method using deep neural networks to estimate the direction of arrival (DOA) of all…

音频与语音处理 · 电气工程与系统科学 2021-11-30 Aswin Shanmugam Subramanian , Chao Weng , Shinji Watanabe , Meng Yu , Dong Yu

Video captioning is a challenging task as it needs to accurately transform visual understanding into natural language description. To date, state-of-the-art methods inadequately model global-local representation across video frames for…

计算机视觉与模式识别 · 计算机科学 2023-03-01 Liqi Yan , Qifan Wang , Yiming Cui , Fuli Feng , Xiaojun Quan , Xiangyu Zhang , Dongfang Liu

Recent Video Large Language Models (Video-LLMs) have shown strong multimodal reasoning capabilities, yet remain challenged by video understanding tasks that require consistent temporal ordering and causal coherence. Many parameter-efficient…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Zhengjian Kang , Qi Chen , Rui Liu , Kangtong Mo , Xingyu Zhang , Xiaoyu Deng , Ye Zhang

Video captioning combines video understanding and language generation. Different from image captioning that describes a static image with details of almost every object, video captioning usually considers a sequence of frames and biases…

计算与语言 · 计算机科学 2022-01-11 Fenglin Liu , Xuancheng Ren , Xian Wu , Bang Yang , Shen Ge , Yuexian Zou , Xu Sun

We present CAT-V (Caption AnyThing in Video), a training-free framework for fine-grained object-centric video captioning that enables detailed descriptions of user-selected objects through time. CAT-V integrates three key components: a…

Tuning language models for dialogue generation has been a prevalent paradigm for building capable dialogue agents. Yet, traditional tuning narrowly views dialogue generation as resembling other language generation tasks, ignoring the role…

计算与语言 · 计算机科学 2024-05-31 Jian Wang , Chak Tou Leong , Jiashuo Wang , Dongding Lin , Wenjie Li , Xiao-Yong Wei