โ† All examples

ModelEthio-ASR Amharic (600M)0Ethio-ASR Multi (300M)6Ethio-ASR Multi (600M)5Ethio-ASR Multi (94M)6Meta MMS-1B13OmniASR-CTC (1B)2OmniASR-CTC (300M)5OmniASR-CTC (3B)3OmniASR-LLM (1B)4OmniASR-LLM (300M)2OmniASR-LLM (3B)2OpenAI Whisper-small13
Audio qualityClean Studio13Ambient Room Noise18Harsh Environment13Distant / Low-Volume Mic13Wideband Phone (VoLTE)23Randomized Phone Channel18Narrowband 2G + Packet Loss14

What did the model get wrong?

OpenAI Whisper-small Harsh Environment   Example 13 of 20  ยท  speaker: Female

๐ŸŽง Listen โ€” Harsh Environment
Very heavy background noise โ€” much harder than typical audio.
โœ… What was actually said
แ‹ซแŠ›แ‹ แˆตแˆแŠญ แˆŒแˆ‹ แˆ˜แˆณแˆชแ‹ซ แˆ‹แ‹ญ แАแ‹ แ‹ซแˆˆแ‹ แŠจแ‹šแˆ…แŠ›แ‹ แ‰แŒฅแˆญ แˆ‹แ‹ญ แˆ†แŠ˜ แˆ˜แŒจแˆจแˆต แŠ แˆแ‰ฝแˆแˆ
๐Ÿค– What this model heard
this is a very important part of the video
Correct
Wrong (different word)
Missed (skipped)
Extra (added)
0
Correct
9
Wrong
4
Missed
0
Extra
0%
Words right
Out of 13 spoken words, this model got 0 right (0%). Errors: 9 words wrong; 4 missed.

Try switching Audio quality โ€” see the same model degrade as the audio gets noisier.