Today we’ve added something new to our leaderboard - scores on new audio benchmarks for models doing ASR and Synthetic Voice Detection. After our initial leaderboard launch, we received many requests to cover other models and chose audio next because it corresponds to today’s launch of a new audio ensemble model by Modulate.ai that wins on most audio tasks in the performance vs parameters category.
The model is called Modulate Velma 2, and is an ensemble of many different small models. The ensemble approach matches the data we have an Neurometric showing that for most tasks, ensembles of models win.
We’ve included Modulate’s submission data here and are in the process of performing our own verification. Stay tuned for more upgrades and improvements to the audio leaderboard in coming weeks.



