← Back to examples

Does the model treat men and women equally?

Word Error Rate broken down by speaker gender. If the two bars are similar height, the model is fair; if one is much taller, that group is being under-served.

Across all models & audio conditions, average WER is 72.7% for male speakers and 79.0% for female speakers — a gap of 6.2pp (higher (worse) for female).

📊 Male vs Female WER — averaged across all models
124%93%62%31%0%56%60%Clean Studio73%77%Ambient Room Noise75%79%Harsh Environment56%60%Distant / Low-Volume Mic74%81%Wideband Phone (VoLTE)75%82%Randomized Phone Channel99%113%Narrowband 2G + Packet LossMale speakersFemale speakers
🔢 Per-model breakdown (WER, Δ = Female − Male)
ModelClean StudioAmbient Room NoiseHarsh EnvironmentDistant / Low-Volume MicWideband Phone (VoLTE)Randomized Phone ChannelNarrowband 2G + Packet Loss
MFΔMFΔMFΔMFΔMFΔMFΔMFΔ
Ethio-ASR Amharic (600M)35.2%37.3%+2.137.0%39.8%+2.939.5%43.7%+4.235.4%37.4%+2.040.0%48.4%+8.441.2%50.1%+8.967.1%79.6%+12.5
Ethio-ASR Multi (300M)58.8%65.3%+6.562.8%72.2%+9.469.8%81.6%+11.858.9%65.4%+6.474.0%92.9%+18.976.9%95.4%+18.5123.1%131.7%+8.6
Ethio-ASR Multi (600M)46.6%49.2%+2.648.9%53.4%+4.452.3%58.4%+6.046.8%49.5%+2.753.3%65.2%+11.954.7%66.5%+11.783.4%93.5%+10.1
Ethio-ASR Multi (94M)55.8%61.2%+5.458.8%65.5%+6.763.5%71.9%+8.355.9%61.2%+5.466.1%78.7%+12.668.2%80.6%+12.493.3%99.3%+6.0
Meta MMS-1B100.5%100.7%+0.2100.1%100.6%+0.4100.1%100.6%+0.5100.5%100.7%+0.2100.4%101.0%+0.6100.4%101.3%+0.9103.2%105.9%+2.8
OmniASR-CTC (1B)41.7%44.9%+3.261.3%62.2%+0.962.3%65.1%+2.841.5%44.3%+2.862.2%67.2%+5.063.1%68.9%+5.888.6%94.0%+5.4
OmniASR-CTC (300M)52.0%55.5%+3.463.8%66.7%+2.965.6%70.0%+4.452.2%55.6%+3.367.0%75.3%+8.368.7%77.3%+8.693.5%97.1%+3.6
OmniASR-CTC (3B)31.9%34.5%+2.656.8%56.9%+0.156.2%57.4%+1.331.7%34.3%+2.657.6%58.2%+0.658.6%60.7%+2.269.4%84.2%+14.8
OmniASR-LLM (1B)32.4%34.2%+1.878.0%75.4%-2.678.0%71.9%-6.232.5%33.9%+1.479.5%78.4%-1.179.4%77.7%-1.782.9%100.0%+17.1
OmniASR-LLM (300M)42.4%42.3%-0.181.8%83.3%+1.583.6%86.1%+2.546.5%47.3%+0.881.3%81.7%+0.481.4%81.8%+0.488.2%99.2%+11.0
OmniASR-LLM (3B)26.6%27.8%+1.278.2%72.5%-5.776.4%65.1%-11.326.3%27.9%+1.667.7%57.4%-10.365.2%59.7%-5.582.4%96.7%+14.3
OpenAI Whisper-small145.2%172.4%+27.2148.8%173.4%+24.6157.2%177.9%+20.8149.0%168.3%+19.4138.3%165.8%+27.5148.0%165.0%+17.1215.1%272.5%+57.4
Δ (delta) shows how many percentage points worse (positive, red) or better (negative, green) the model is on female speech compared to male. Values near 0 mean the model is roughly gender-fair.

← Back to example gallery