English
Related papers

Related papers: Asymmetric Proxy Loss for Multi-View Acoustic Word…

200 papers

Previous researches on acoustic word embeddings used in query-by-example spoken term detection have shown remarkable performance improvements when using a triplet network. However, the triplet network is trained using only a limited…

Audio and Speech Processing · Electrical Eng. & Systems 2018-11-29 Hyungjun Lim , Younggwan Kim , Youngmoon Jung , Myunghun Jung , Hoirin Kim

Anomalous sound detection (ASD) typically involves self-supervised proxy tasks to learn feature representations from normal sound data, owing to the scarcity of anomalous samples. In ASD research, proxy tasks such as AutoEncoders operate…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-14 Seunghyeon Shin , Seokjin Lee

Ensembling word embeddings to improve distributed word representations has shown good success for natural language processing tasks in recent years. These approaches either carry out straightforward mathematical operations over a set of…

Computation and Language · Computer Science 2018-08-14 James O' Neill , Danushka Bollegala

Word embedding models learn semantically rich vector representations of words and are widely used to initialize natural processing language (NLP) models. The popular continuous bag-of-words (CBOW) model of word2vec learns a vector embedding…

Computation and Language · Computer Science 2020-06-02 Shashank Sonkar , Andrew E. Waters , Richard G. Baraniuk

We address the problem of distance metric learning (DML), defined as learning a distance consistent with a notion of semantic similarity. Traditionally, for this problem supervision is expressed in the form of sets of points that follow an…

Computer Vision and Pattern Recognition · Computer Science 2017-08-03 Yair Movshovitz-Attias , Alexander Toshev , Thomas K. Leung , Sergey Ioffe , Saurabh Singh

We propose a new unsupervised model for mapping a variable-duration speech segment to a fixed-dimensional representation. The resulting acoustic word embeddings can form the basis of search, discovery, and indexing systems for low- and…

Audio and Speech Processing · Electrical Eng. & Systems 2020-12-07 Puyuan Peng , Herman Kamper , Karen Livescu

In the past few years, triplet loss-based metric embeddings have become a de-facto standard for several important computer vision problems, most no-tably, person reidentification. On the other hand, in the area of speech recognition the…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-08 Roman Vygon , Nikolay Mikhaylovskiy

How do neural networks "perceive" speech sounds from unknown languages? Does the typological similarity between the model's training language (L1) and an unknown language (L2) have an impact on the model representations of L2 speech…

Computation and Language · Computer Science 2021-09-22 Badr M. Abdullah , Iuliia Zaitova , Tania Avgustinova , Bernd Möbius , Dietrich Klakow

Speech embeddings are fixed-size acoustic representations of variable-length speech sequences. They are increasingly used for a variety of tasks ranging from information retrieval to unsupervised term discovery and speech segmentation.…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-09 Robin Algayres , Mohamed Salah Zaiem , Benoit Sagot , Emmanuel Dupoux

The technique of Cross-Lingual Word Embedding (CLWE) plays a fundamental role in tackling Natural Language Processing challenges for low-resource languages. Its dominant approaches assumed that the relationship between embeddings could be…

Computation and Language · Computer Science 2022-06-14 Xutan Peng , Mark Stevenson , Chenghua Lin , Chen Li

Speech encodes multiple simultaneous attributes -- linguistic content, speaker identity, dialect, gender --that conventional single-vector embeddings conflate. We present a factor-partitioned embedding framework that maps each utterance…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-11 Jim O'Regan , Jens Edlund

Cross-modal representation learning learns a shared embedding between two or more modalities to improve performance in a given task compared to using only one of the modalities. Cross-modal representation learning from different data types…

Machine Learning · Computer Science 2023-09-12 Felix Ott , David Rügamer , Lucas Heublein , Bernd Bischl , Christopher Mutschler

For text enrollment-based open-vocabulary keyword spotting (KWS), acoustic and text embeddings are typically compared at either the phoneme or utterance level. To facilitate this, we optimize acoustic and text encoders using deep metric…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-26 Youngmoon Jung , Yong-Hyeok Lee , Myunghun Jung , Jaeyoung Roh , Chang Woo Han , Hoon-Young Cho

The mainstream researche in deep metric learning can be divided into two genres: proxy-based and pair-based methods. Proxy-based methods have attracted extensive attention due to the lower training complexity and fast network convergence.…

Information Retrieval · Computer Science 2023-04-19 Xinyue Li , Jian Wang , Wei Song , Yanling Du , Zhixiang Liu

Multimodal sentence embedding models typically leverage image-caption pairs in addition to textual data during training. However, such pairs often contain noise, including redundant or irrelevant information on either the image or caption…

Computation and Language · Computer Science 2025-08-04 Kaiyan Zhao , Zhongtao Miao , Yoshimasa Tsuruoka

To simplify the generation process, several text-to-speech (TTS) systems implicitly learn intermediate latent representations instead of relying on predefined features (e.g., mel-spectrogram). However, their generation quality is…

Sound · Computer Science 2023-08-29 Hyungchan Yoon , Seyun Um , Changwhan Kim , Hong-Goo Kang

Learning sentence embeddings from dialogues has drawn increasing attention due to its low annotation cost and high domain adaptability. Conventional approaches employ the siamese-network for this task, which obtains the sentence embeddings…

Computation and Language · Computer Science 2021-09-28 Che Liu , Rui Wang , Jinghua Liu , Jian Sun , Fei Huang , Luo Si

Language-audio joint representation learning frameworks typically depend on deterministic embeddings, assuming a one-to-one correspondence between audio and text. In real-world settings, however, the language-audio relationship is…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-22 Toranosuke Manabe , Yuchi Ishikawa , Hokuto Munakata , Tatsuya Komatsu

We propose to learn acoustic word embeddings with temporal context for query-by-example (QbE) speech search. The temporal context includes the leading and trailing word sequences of a word. We assume that there exist spoken word pairs in…

Computation and Language · Computer Science 2018-06-19 Yougen Yuan , Cheung-Chi Leung , Lei Xie , Hongjie Chen , Bin Ma , Haizhou Li

Recent studies have been revisiting whole words as the basic modelling unit in speech recognition and query applications, instead of phonetic units. Such whole-word segmental systems rely on a function that maps a variable-length speech…

Computation and Language · Computer Science 2016-01-11 Herman Kamper , Weiran Wang , Karen Livescu