Real Ethio Telecom recordings can't be shared for privacy reasons β so we rebuilt a representative dataset from the ground up, in seven stages, with two rounds of expert-in-the-loop review.
We first listened to real customerβagent conversations from Ethio Telecom to catalogue the acoustic fingerprints of the channel: narrow bandwidth, 50 Hz mains hum, packet loss, AGC compression, and codec artifacts.
Gathered internal Ethio Telecom documents β service manuals, customer-support FAQs, and agent-training material β to seed realistic domain vocabulary and typical customer scenarios.
Used the source documents to synthesize realistic multi-turn customer β agent dialogues that follow the flow of real telecom support calls β greetings, problem statements, verification, resolution, closing.
Ethio Telecom staff read every generated dialogue and gave feedback on realism, terminology, and correctness. We iterated on the text until domain experts confirmed the scripts sounded like calls they actually take.
Trained a 50/50 male/female voice-actor team to perform the dialogues in a conversational, un-scripted register β including hesitations, back-channels, and overlapping turns. Total 30 hours of speech collected.
Ethio Telecom workers listened to the recordings and verified that the audio conversations matched the tone, pacing, and behavior of real interactions. Any takes that felt stilted or off-domain were re-recorded.
The clean recordings are passed through seven controlled conditions that simulate the artifacts observed in step 1 β from Clean Studio all the way to Narrowband 2G + Packet Loss β producing the final benchmark splits used for evaluation.