/

/

Background noise ≠ background noise: A joint piece from Coval and ai-coustics

/

/

Background noise ≠ background noise: A joint piece from Coval and ai-coustics

Background noise ≠ background noise: A joint piece from Coval and ai-coustics

Background noise ≠ background noise: A joint piece from Coval and ai-coustics

Photo of Fabian Seipel, CEO of ai-coustics

Written by

Fabian Seipel

,

CEO and Co-founder

Research

/

Photo of Fabian Seipel, CEO of ai-coustics

Fabian Seipel

Written by

,

CEO and Co-founder

Research

/

A joint piece from Coval and ai-coustics. Coval is a testing and simulation platform that helps teams validate how their voice AI agents actually perform under real-world conditions. ai-coustics builds real-time audio models for voice AI, including Tyto, an audio insight model, and Quail Voice Focus, which isolates the primary speaker.

The myth

Most voice AI teams treat "audio quality" as one problem: suppress anything that is unpleasant or disturbing to a human listener, and you're safe. ai-coustics' benchmarking data says that's the wrong mental model. Worse, teams often spend engineering effort suppressing the one thing that was never going to break their agent, while missing the thing that actually does.

Benchmarking real voice AI audio shows something different. Based on those findings, ai-coustics built Tyto: a lightweight model that analyzes the real-world audio flowing into a voice AI stack (a voice agent, a speech-to-text pipeline, whatever's listening) and predicts whether that audio is likely to cause a downstream failure. Its main output is a risk score between 0 and 1: low means the audio is fine; high means a failure is likely.

Tyto backs its headline risk score with six qualitative dimensions that explain what kind of problem is present: noise, interfering speech, speaker reverb, packet loss, codec degradation, speaker loudness.

What the data shows

To find out which of these six actually break voice AI systems, ai-coustics tested each degradation type by artificially adding it to clean speech recordings and measuring how downstream speech-to-text and voice activity detection models respond.

Commercial STT models aren't cast from the same mold. The overall trend below holds across providers, but how much a given model actually degrades under these conditions can vary a lot from one to the next.

Noise

Total WER — measured
Clean speech (each engine)
typical call0%20%40%60%80%100%02040Total WER (%)SNR (dB)
WER components — measured · average
InsertionsDeletionsSubstitutions
0%20%40%60%02040WER components (%)SNR (dB)
VAD accuracy — measured
VAD accuracyClean speech
65%70%75%80%85%90%95%100%02040VAD accuracy (%)SNR (dB)
Tyto — predicted · average
overall risk scorenamed dimension
0.00.20.40.60.81.0risk_score vs WER+0.50risk_score vs VAD error rate+0.68noise vs SNR+0.93Spearman rho per clip
Method

Five degradations swept over one corpus (59 clips, human transcripts and human frame-level voice-activity labels). Rows: noise · room reverb · Opus lowdelay codec · packet dropouts · a second talker at a controlled SIR.

SNR, SIR and DRR follow aicpy conventions (ITU-R BS.1770 integrated loudness; DRR set on the impulse response). Every provider streams at 1.0× real time. The shaded band is the range a real call occupies; dotted lines are each engine's score on the clean recording.

Voice-activity detection is measured on the audio, not the transcript, so that column does not depend on the provider. The average is the mean over the three providers and is not a product. Correlations are Spearman rho computed per clip, never on condition means.

Noise

Total WER — measured
Clean speech (each engine)
typical call0%20%40%60%80%100%02040Total WER (%)SNR (dB)
WER components — measured · average
InsertionsDeletionsSubstitutions
0%20%40%60%02040WER components (%)SNR (dB)
VAD accuracy — measured
VAD accuracyClean speech
65%70%75%80%85%90%95%100%02040VAD accuracy (%)SNR (dB)
Tyto — predicted · average
overall risk scorenamed dimension
0.00.20.40.60.81.0risk_score vs WER+0.50risk_score vs VAD error rate+0.68noise vs SNR+0.93Spearman rho per clip
Method

Five degradations swept over one corpus (59 clips, human transcripts and human frame-level voice-activity labels). Rows: noise · room reverb · Opus lowdelay codec · packet dropouts · a second talker at a controlled SIR.

SNR, SIR and DRR follow aicpy conventions (ITU-R BS.1770 integrated loudness; DRR set on the impulse response). Every provider streams at 1.0× real time. The shaded band is the range a real call occupies; dotted lines are each engine's score on the clean recording.

Voice-activity detection is measured on the audio, not the transcript, so that column does not depend on the provider. The average is the mean over the three providers and is not a product. Correlations are Spearman rho computed per clip, never on condition means.

Noise

Total WER — measured
Clean speech (each engine)
typical call0%20%40%60%80%100%02040Total WER (%)SNR (dB)
WER components — measured · average
InsertionsDeletionsSubstitutions
0%20%40%60%02040WER components (%)SNR (dB)
VAD accuracy — measured
VAD accuracyClean speech
65%70%75%80%85%90%95%100%02040VAD accuracy (%)SNR (dB)
Tyto — predicted · average
overall risk scorenamed dimension
0.00.20.40.60.81.0risk_score vs WER+0.50risk_score vs VAD error rate+0.68noise vs SNR+0.93Spearman rho per clip
Method

Five degradations swept over one corpus (59 clips, human transcripts and human frame-level voice-activity labels). Rows: noise · room reverb · Opus lowdelay codec · packet dropouts · a second talker at a controlled SIR.

SNR, SIR and DRR follow aicpy conventions (ITU-R BS.1770 integrated loudness; DRR set on the impulse response). Every provider streams at 1.0× real time. The shaded band is the range a real call occupies; dotted lines are each engine's score on the clean recording.

Voice-activity detection is measured on the audio, not the transcript, so that column does not depend on the provider. The average is the mean over the three providers and is not a product. Correlations are Spearman rho computed per clip, never on condition means.

Four of the five hold up better than expected. Ambient noise barely affects transcription until it gets extreme. A typical phone call sits around 20–40 dB SNR. Good STT models handle 10 dB with only a small accuracy penalty, and only really fall apart below roughly 6 dB, a level where a human listener would already be straining to follow along. STT tolerates this kind of noise with essentially no quality loss, arguably better than a person would, who'd need more concentration and would tire over time. Reverb matters less for word accuracy and more for knowing when someone's actually finished speaking, since voice activity detection is more sensitive to it than transcription is. Compression artifacts and packet loss both need to get genuinely severe, well past what a typical call ever encounters, before either causes real damage.

Interfering speech breaks the pattern. STT models try to transcribe every voice they can hear, so a second person talking in the background, or a TV, gets inserted into the transcript as if the caller had said it. It also trips voice activity detection into thinking the caller has stopped speaking, or that someone is interrupting. Either way, it breaks the flow of the conversation.

Testing for it: the Coval approach

A transcription error can become a wrong order. Testing the full conversation shows whether the agent catches it, recovers, and completes the task correctly.

Picture a drive-thru: an engine idles as a customer orders one burger. In the next lane, someone orders three. A few seconds later, the agent confidently confirms three burgers with the first customer. The customer replies, “No, I said one.” Does the agent correct the order, or carry the mistake into checkout?

Coval simulates conversations with voice agents so teams can test these situations and compare the results. A test case gives the simulated caller a task, such as placing an order. A persona defines how that caller speaks and the sounds around them. Together, they let you ask a practical question: how does the same agent handle this order with an engine running, with another person talking, and with no background sound at all?

Coval's Background Noise panel listing sounds such as Office, Crowd Talking, Airport Boarding and Kids Playing labelled "Ambient", and Doorbell labelled "Point source", with a volume slider at 80%.

Choose from Coval’s background sounds or upload your own recording, then adjust the volume to compare how different audio conditions affect the same task.

Start with those three conditions, keeping the caller’s voice, volume, and behavior fixed. Repeat each across the same set of orders. Now you can see whether the agent struggles broadly with noise or specifically with competing speech or accents. Custom background recordings let you bring in the sounds of a particular restaurant, call center, or other setting.

The first result to check is the final order, including the quantity sent to the ordering tool when that trace is available. Does the agent meet the expected behavior? Coval’s metrics then help explain how the conversation unfolded: did the agent hesitate, repeat a question, talk over the caller, or need another prompt? Transcription checks highlight discrepancies worth reviewing. Check for dropouts or repeated audio when the recording suggests those problems.

Table of six caller personas, such as Louisiana Creole accent, Rush and Confused Customer, with 51 calls each. Artifact scores and completion rates stay high for every persona, drive-up abandonment is 0%, and successful upsells range from 0% to 45%.

Example drive-up evaluation grouped by persona. Coval brings audio-artifact measurements, recovery, and task outcomes into one report. The same view can compare personas configured with different background sounds.

If you find a failed order, follow it from the report into its recording, transcript, and available traces. Listen for the competing order and check where the agent’s response diverges, using its own STT trace if available. Then follow the customer’s correction through to the final action. ai-coustics’ Tyto’s analysis of the same caller audio can add context about competing speech around that moment, giving you a hypothesis to investigate.

Call transcript where the caller says "I'm Dana Patel" and the agent replies "Sure thing, Jada," next to a trace of the book_reservation call made under "Jada Patel" that returned success.

Restaurant demo: In a simulated restaurant call with interfering speech, a misheard name carries through to the reservation.

That makes the next decision easier to evaluate. Try another STT configuration against the same orders and audio conditions. Use a statistical significant number of simulations to check whether your agent captures the right quantities, keeps the conversation moving, and still performs well on clean calls. Save those cases and their pass criteria for the next release, then expand coverage to different callers and combinations of conditions.

Fixing it: how Quail Voice Focus helps

Coval's tests show where and why an agent struggles. Quail Voice Focus is ai-coustics' answer to what to do about it.

Voice Focus is a real-time speaker isolation model built for single-speaker human-to-agent scenarios. It identifies the primary speaker and suppresses everything else in the signal, including interfering speakers, background media, residual echo and ambient noise, before that audio reaches a VAD or STT model downstream. It runs zero-shot, so it needs no enrollment or reference audio. It also works across languages.

Voice Focus sits upstream of whatever STT a team chooses. That makes it a complement to Coval's stack comparisons, useful no matter which provider gets tested.

That scope lines up with what the benchmarking data found. Voice Focus targets interfering speech and background media speech directly, both shown earlier to break transcripts and trip false VAD triggers.

Suppression isn't always the right call, though. In fraud detection, for instance, a second voice coaching a caller through a loan application is a signal you want to detect. Suppressing it defeats the purpose. That's an argument for applying Voice Focus deliberately, per use case.

Closing

Most voice AI teams have a hard time tracing a pipeline failure back to audio. Coval and ai-coustics close that gap from three different angles. Coval's simulations show where an agent struggles under real conditions, and Tyto explains why, breaking a risk score into the six dimensions covered above. Quail Voice Focus fixes what's fixable, applied deliberately depending on the use case.

Final logo

Bring real-time audio intelligence into your voice AI stack

Bring real-time audio intelligence into your voice AI stack

Bring real-time audio intelligence into your voice AI stack