A joint piece from Coval and ai-coustics. Coval is a testing and simulation platform that helps teams validate how their voice AI agents actually perform under real-world conditions. ai-coustics builds real-time audio models for voice AI, including Tyto, an audio insight model, and Quail Voice Focus, which isolates the primary speaker.
The myth
Most voice AI teams treat "audio quality" as one problem: suppress anything that is unpleasant or disturbing to a human listener, and you're safe. ai-coustics' benchmarking data says that's the wrong mental model. Worse, teams often spend engineering effort suppressing the one thing that was never going to break their agent, while missing the thing that actually does.
Benchmarking real voice AI audio shows something different. Based on those findings, ai-coustics built Tyto: a lightweight model that analyzes the real-world audio flowing into a voice AI stack (a voice agent, a speech-to-text pipeline, whatever's listening) and predicts whether that audio is likely to cause a downstream failure. Its main output is a risk score between 0 and 1: low means the audio is fine; high means a failure is likely.
Tyto backs its headline risk score with six qualitative dimensions that explain what kind of problem is present: noise, interfering speech, speaker reverb, packet loss, codec degradation, speaker loudness.
What the data shows
To find out which of these six actually break voice AI systems, ai-coustics tested each degradation type by artificially adding it to clean speech recordings and measuring how downstream speech-to-text and voice activity detection models respond.
Commercial STT models aren't cast from the same mold. The overall trend below holds across providers, but how much a given model actually degrades under these conditions can vary a lot from one to the next.
Four of the five hold up better than expected. Ambient noise barely affects transcription until it gets extreme. A typical phone call sits around 20–40 dB SNR. Good STT models handle 10 dB with only a small accuracy penalty, and only really fall apart below roughly 6 dB, a level where a human listener would already be straining to follow along. STT tolerates this kind of noise with essentially no quality loss, arguably better than a person would, who'd need more concentration and would tire over time. Reverb matters less for word accuracy and more for knowing when someone's actually finished speaking, since voice activity detection is more sensitive to it than transcription is. Compression artifacts and packet loss both need to get genuinely severe, well past what a typical call ever encounters, before either causes real damage.
Interfering speech breaks the pattern. STT models try to transcribe every voice they can hear, so a second person talking in the background, or a TV, gets inserted into the transcript as if the caller had said it. It also trips voice activity detection into thinking the caller has stopped speaking, or that someone is interrupting. Either way, it breaks the flow of the conversation.
Testing for it: the Coval approach
A transcription error can become a wrong order. Testing the full conversation shows whether the agent catches it, recovers, and completes the task correctly.
Picture a drive-thru: an engine idles as a customer orders one burger. In the next lane, someone orders three. A few seconds later, the agent confidently confirms three burgers with the first customer. The customer replies, “No, I said one.” Does the agent correct the order, or carry the mistake into checkout?
Coval simulates conversations with voice agents so teams can test these situations and compare the results. A test case gives the simulated caller a task, such as placing an order. A persona defines how that caller speaks and the sounds around them. Together, they let you ask a practical question: how does the same agent handle this order with an engine running, with another person talking, and with no background sound at all?

Choose from Coval’s background sounds or upload your own recording, then adjust the volume to compare how different audio conditions affect the same task.
Start with those three conditions, keeping the caller’s voice, volume, and behavior fixed. Repeat each across the same set of orders. Now you can see whether the agent struggles broadly with noise or specifically with competing speech or accents. Custom background recordings let you bring in the sounds of a particular restaurant, call center, or other setting.
The first result to check is the final order, including the quantity sent to the ordering tool when that trace is available. Does the agent meet the expected behavior? Coval’s metrics then help explain how the conversation unfolded: did the agent hesitate, repeat a question, talk over the caller, or need another prompt? Transcription checks highlight discrepancies worth reviewing. Check for dropouts or repeated audio when the recording suggests those problems.

Example drive-up evaluation grouped by persona. Coval brings audio-artifact measurements, recovery, and task outcomes into one report. The same view can compare personas configured with different background sounds.
If you find a failed order, follow it from the report into its recording, transcript, and available traces. Listen for the competing order and check where the agent’s response diverges, using its own STT trace if available. Then follow the customer’s correction through to the final action. ai-coustics’ Tyto’s analysis of the same caller audio can add context about competing speech around that moment, giving you a hypothesis to investigate.

Restaurant demo: In a simulated restaurant call with interfering speech, a misheard name carries through to the reservation.
That makes the next decision easier to evaluate. Try another STT configuration against the same orders and audio conditions. Use a statistical significant number of simulations to check whether your agent captures the right quantities, keeps the conversation moving, and still performs well on clean calls. Save those cases and their pass criteria for the next release, then expand coverage to different callers and combinations of conditions.
Fixing it: how Quail Voice Focus helps
Coval's tests show where and why an agent struggles. Quail Voice Focus is ai-coustics' answer to what to do about it.
Voice Focus is a real-time speaker isolation model built for single-speaker human-to-agent scenarios. It identifies the primary speaker and suppresses everything else in the signal, including interfering speakers, background media, residual echo and ambient noise, before that audio reaches a VAD or STT model downstream. It runs zero-shot, so it needs no enrollment or reference audio. It also works across languages.
Voice Focus sits upstream of whatever STT a team chooses. That makes it a complement to Coval's stack comparisons, useful no matter which provider gets tested.
That scope lines up with what the benchmarking data found. Voice Focus targets interfering speech and background media speech directly, both shown earlier to break transcripts and trip false VAD triggers.
Suppression isn't always the right call, though. In fraud detection, for instance, a second voice coaching a caller through a loan application is a signal you want to detect. Suppressing it defeats the purpose. That's an argument for applying Voice Focus deliberately, per use case.
Closing
Most voice AI teams have a hard time tracing a pipeline failure back to audio. Coval and ai-coustics close that gap from three different angles. Coval's simulations show where an agent struggles under real conditions, and Tyto explains why, breaking a risk score into the six dimensions covered above. Quail Voice Focus fixes what's fixable, applied deliberately depending on the use case.

