โ† All examples

ModelEthio-ASR Amharic (600M)0Ethio-ASR Multi (300M)6Ethio-ASR Multi (600M)1Ethio-ASR Multi (94M)4Meta MMS-1B7OmniASR-CTC (1B)1OmniASR-CTC (300M)4OmniASR-CTC (3B)1OmniASR-LLM (1B)3OmniASR-LLM (300M)3OmniASR-LLM (3B)3OpenAI Whisper-small7
Audio qualityClean Studio3Ambient Room Noise1Harsh Environment4Distant / Low-Volume Mic4Wideband Phone (VoLTE)9Randomized Phone Channel10Narrowband 2G + Packet Loss9

What did the model get wrong?

OmniASR-CTC (300M) Harsh Environment   Example 10 of 20  ยท  speaker: Male

๐ŸŽง Listen โ€” Harsh Environment
Very heavy background noise โ€” much harder than typical audio.
โœ… What was actually said
แˆตแˆˆ แŒฅแ‰…แˆŽแ‰น แ‹แˆญแ‹แˆญ แˆ˜แˆจแŒƒ แ‹จแ‰ต แАแ‹ แ‹จแˆ›แŒˆแŠ˜แ‹
๐Ÿค– What this model heard
แˆตแˆˆ แŒฅแ‰…แˆŽแ‰ฑ แ‹แˆญแ‹แˆญ แˆ˜แ‹ฐแˆญแŒƒแ‹ญแ‹ แАแ‹ แˆ›แŒˆแŠ˜แ‹
Correct
Wrong (different word)
Missed (skipped)
Extra (added)
3
Correct
3
Wrong
1
Missed
0
Extra
43%
Words right
Out of 7 spoken words, this model got 3 right (43%). Errors: 3 words wrong; 1 missed.

Try switching Audio quality โ€” see the same model degrade as the audio gets noisier.