中文
相关论文

相关论文: Data mining Mandarin tone contour shapes

200 篇论文

Linguistic typology aims to capture structural and semantic variation across the world's languages. A large-scale typology could provide excellent guidance for multilingual Natural Language Processing (NLP), particularly for languages that…

Much work in the space of NLP has used computational methods to explore sociolinguistic variation in text. In this paper, we argue that memes, as multimodal forms of language comprised of visual templates and text, also exhibit meaningful…

计算与语言 · 计算机科学 2023-11-16 Naitian Zhou , David Jurgens , David Bamman

Lip reading aims at decoding texts from the movement of a speaker's mouth. In recent years, lip reading methods have made great progress for English, at both word-level and sentence-level. Unlike English, however, Chinese Mandarin is a…

计算机视觉与模式识别 · 计算机科学 2019-12-02 Ya Zhao , Rui Xu , Mingli Song

The notion of "in-domain data" in NLP is often over-simplistic and vague, as textual data varies in many nuanced linguistic aspects such as topic, style or level of formality. In addition, domain labels are many times unavailable, making it…

计算与语言 · 计算机科学 2020-05-04 Roee Aharoni , Yoav Goldberg

Generative Large Language Models have shown impressive in-context learning abilities, performing well across various tasks with just a prompt. Previous melody-to-lyric research has been limited by scarce high-quality aligned data and…

计算与语言 · 计算机科学 2024-10-03 Hong-Hsiang Liu , Yi-Wen Liu

The goal of dialogue topic shift detection is to identify whether the current topic in a conversation has changed or needs to change. Previous work focused on detecting topic shifts using pre-trained models to encode the utterance, failing…

计算与语言 · 计算机科学 2023-05-24 Jiangyi Lin , Yaxin Fan , Xiaomin Chu , Peifeng Li , Qiaoming Zhu

To make sense of massive data, we often fit simplified models and then interpret the parameters; for example, we cluster the text embeddings and then interpret the mean parameters of each cluster. However, these parameters are often…

人工智能 · 计算机科学 2025-01-14 Ruiqi Zhong , Heng Wang , Dan Klein , Jacob Steinhardt

Text summarization is an interesting area for researchers to develop new techniques to provide human like summaries for vast amounts of information. Summarization techniques tend to focus on providing accurate representation of content, and…

信息检索 · 计算机科学 2018-02-28 Mayank Chaudhari , Aakash Nelson Mattukoyya

The Lombard effect refers to individuals' unconscious modulation of vocal effort in response to variations in the ambient noise levels, intending to enhance speech intelligibility. The impact of different decibel levels and types of…

声音 · 计算机科学 2023-09-15 Qingmu Liu , Yuhong Yang , Baifeng Li , Hongyang Chen , Weiping Tu , Song Lin

While idiosyncrasies of the Chinese classifier system have been a richly studied topic among linguists (Adams and Conklin, 1973; Erbaugh, 1986; Lakoff, 1986), not much work has been done to quantify them with statistical methods. In this…

计算与语言 · 计算机科学 2020-05-26 Shijia Liu , Hongyuan Mei , Adina Williams , Ryan Cotterell

For text-to-speech (TTS) synthesis, prosodic structure prediction (PSP) plays an important role in producing natural and intelligible speech. Although inter-utterance linguistic information can influence the speech interpretation of the…

声音 · 计算机科学 2023-09-01 Jie Chen , Changhe Song , Deyi Tuo , Xixin Wu , Shiyin Kang , Zhiyong Wu , Helen Meng

This study investigates whether the phonological features derived from the Featurally Underspecified Lexicon model can be applied in text-to-speech systems to generate native and non-native speech in English and Mandarin. We present a…

计算与语言 · 计算机科学 2022-04-18 Cong Zhang , Huinan Zeng , Huang Liu , Jiewen Zheng

Quantification is a fundamental component of everyday language use, yet little is known about how speakers decide whether and how to quantify in naturalistic production. We investigate quantification in Mandarin Chinese using a…

计算与语言 · 计算机科学 2026-02-11 Yayun Zhang , Guanyi Chen , Fahime Same , Saad Mahamood , Tingting He

The impressive achievements of transformers force NLP researchers to delve into how these models represent the underlying structure of natural language. In this paper, we propose a novel standpoint to investigate the above issue: using…

Speech production and perception are the main ways humans communicate daily. Prior brain-to-text decoding studies have largely focused on a single modality and alphabetic languages. Here, we present a unified brain-to-sentence decoding…

神经元与认知 · 定量生物学 2026-03-16 Zhizhang Yuan , Yang Yang , Gaorui Zhang , Baowen Cheng , Zehan Wu , Yuhao Xu , Xiaoying Liu , Liang Chen , Ying Mao , Meng Li

This paper models the fundamental frequency contours on both Mandarin and Cantonese speech with decision trees and DNNs (deep neural networks). Different kinds of f0 representations and model architectures are tested for decision trees and…

计算与语言 · 计算机科学 2018-07-05 Weidong Yuan , Alan W Black

As NLP tools become ubiquitous in today's technological landscape, they are increasingly applied to languages with a variety of typological structures. However, NLP research does not focus primarily on typological differences in its…

计算与语言 · 计算机科学 2020-05-04 Sophie Groenwold , Samhita Honnavalli , Lily Ou , Aesha Parekh , Sharon Levy , Diba Mirza , William Yang Wang

Machine learning models are trained to find patterns in data. NLP models can inadvertently learn socially undesirable patterns when training on gender biased text. In this work, we propose a general framework that decomposes gender bias in…

计算与语言 · 计算机科学 2020-05-05 Emily Dinan , Angela Fan , Ledell Wu , Jason Weston , Douwe Kiela , Adina Williams

In recent years linguistic typology, which classifies the world's languages according to their functional and structural properties, has been widely used to support multilingual NLP. While the growing importance of typological information…

计算与语言 · 计算机科学 2016-10-12 Helen O'Horan , Yevgeni Berzak , Ivan Vulić , Roi Reichart , Anna Korhonen

Developing explainability methods for Natural Language Processing (NLP) models is a challenging task, for two main reasons. First, the high dimensionality of the data (large number of tokens) results in low coverage and in turn small…

计算与语言 · 计算机科学 2023-03-08 Peyman Jalali , Nengfeng Zhou , Yufei Yu