How does the same voice look under different conditions?
Each row below is the same spoken sentence passed through every noise / channel condition.
The images are mel spectrograms — bright colors mean strong sound at that frequency (vertical)
at that moment (horizontal).
Watch how high frequencies disappear in the phone conditions, how noise fills the background,
and how quiet segments get swallowed in the low-volume version.
What you're seeing. Speech energy concentrates in a few frequency bands (formants).
Phone / narrowband codecs chop off everything above ~3–4 kHz, so the top of the spectrogram goes dark.
Heavy background noise adds a haze everywhere. Low-audio simply dims the whole picture.