Hello, everyone.

When people hear an animal call or keyboard typing, they can infer something about the source and setting. If we turn those sounds into numerical vectors, does their distance preserve the same kind of meaning?

Today, I will export the CLAP audio encoder to ONNX and compare its environmental sound embeddings across the 50 classes in ESC-50.

What I Tested

CLAP (Contrastive Language-Audio Pretraining) maps audio and text into a shared embedding space. This experiment runs only the audio encoder from laion/clap-htsat-unfused with ONNX Runtime and asks: