Do audio and visual tokenizers capture backchannels?

Loading...
Thumbnail Image

Date

Authors

Favre, B. Benoit
Boudin, A.R.J.B. Auriane

Journal Title

Journal ISSN

Volume Title

Publisher

Stroudsburg, PA : Association for Computational Linguistics

Research Projects

Organizational Units

Journal Issue

Abstract

Audio and video tokenizers are autoencoders trained to represent the content of recordings as a sequence of vectors. They are prevalently used to interface large language models with non-textual modalities. While they allow advanced applications such as video generation, the envelope of their limitations is not known in the context of multimodal conversation. This work focuses on backchannels, which listeners use to signal to the speaker that they are listening. This feedback is essential to maintain the conversation flow. We evaluate whether a representative set of audio and video tokenizers encode backchannels using linear probing. Results show that although audio tokenizers capture the phenomenon relatively well, backchannels are not linearly separated by video tokenizers. However, joint representations resulting from concatenating representations in both modalities improve accuracy significantly over audio-only representations, suggesting to train multimodal tokenizers

Description

Keywords

Citation

Endorsement

Review

Supplemented By

Referenced By