โ† All examples

ModelEthio-ASR Amharic (600M)1Ethio-ASR Multi (300M)6Ethio-ASR Multi (600M)5Ethio-ASR Multi (94M)7Meta MMS-1B13OmniASR-CTC (1B)4OmniASR-CTC (300M)3OmniASR-CTC (3B)3OmniASR-LLM (1B)3OmniASR-LLM (300M)1OmniASR-LLM (3B)2OpenAI Whisper-small18
Audio qualityClean Studio13Ambient Room Noise18Harsh Environment13Distant / Low-Volume Mic13Wideband Phone (VoLTE)23Randomized Phone Channel18Narrowband 2G + Packet Loss14

What did the model get wrong?

OpenAI Whisper-small Randomized Phone Channel   Example 13 of 20  ยท  speaker: Female

๐ŸŽง Listen โ€” Randomized Phone Channel
Phone-call audio with random extra distortions layered on top.
โœ… What was actually said
แ‹ซแŠ›แ‹ แˆตแˆแŠญ แˆŒแˆ‹ แˆ˜แˆณแˆชแ‹ซ แˆ‹แ‹ญ แАแ‹ แ‹ซแˆˆแ‹ แŠจแ‹šแˆ…แŠ›แ‹ แ‰แŒฅแˆญ แˆ‹แ‹ญ แˆ†แŠ˜ แˆ˜แŒจแˆจแˆต แŠ แˆแ‰ฝแˆแˆ
๐Ÿค– What this model heard
here is only a place where you can find a place to live and a place to live
Correct
Wrong (different word)
Missed (skipped)
Extra (added)
0
Correct
13
Wrong
0
Missed
5
Extra
0%
Words right
Out of 13 spoken words, this model got 0 right (0%). Errors: 13 words wrong; 5 extra.

Try switching Audio quality โ€” see the same model degrade as the audio gets noisier.