Hello, everyone.

Today is Tanabata in Japan. Events like Tanabata, summer dance festivals, and Christmas all seem to have their own sound images. They also feel like composites of multiple sound categories, such as music, effects, voices, and ambient noise, forming one overall impression.

Today, I am testing how well the YAMNet TFLite model can tag the ESC-50 environmental sound dataset.

What I Tested

YAMNet is an audio tagging model that predicts 521 AudioSet sound event classes. In this experiment, I run the TFLite version published on TensorFlow Hub from Python and map its labels to the 50 categories in ESC-50.