SpeechCLIP: Integrating Speech with Pre-Trained Vision and Language Model, Accepted to IEEE SLT 2022
-
Updated
Nov 25, 2022 - Python
SpeechCLIP: Integrating Speech with Pre-Trained Vision and Language Model, Accepted to IEEE SLT 2022
SpeechCLIP+: Self-supervised multi-task representation learning for speech via CLIP and speech-image data. Accepted to ICASSP 2024, Self-supervision in Audio, Speech, and Beyond (SASB) workshop.
Baselines for the Zero-Resources Speech Challenge using VisuallyGrounded Models of Spoken Language, 2021 edition
Library for training visually-grounded models of spoken language understanding.
Code for the paper "Textual supervision for visually grounded spoken language understanding".
Code used in my Master's thesis
To associate your repository with the visually-grounded-speech topic, visit your repo's landing page and select "manage topics."