Independent benchmarks
Independent benchmarks
ai-coustics speech enhancement models benchmarked by SLNG against Maxine and Krisp
ai-coustics speech enhancement models benchmarked by SLNG against Maxine and Krisp
Word error rates on commercial STT models
Raw
Quail Voice Focus 2.2 S
Quail Voice Focus 2.2 L
Left part: Deletions / Middle: Substitutions / Right: Insertions
lower % is better
Independent benchmark - Quail Multi Speaker vs. Krisp NC & NVIDIA Maxine BNR
Δ WER vs. no noise cancellation baseline, background noise, full enhancement intensity - English & Hindi
Δ WER vs. no noise cancellation baseline, background noise, full enhancement intensity - English & Hindi
Quail Multi Speaker
Krisp NC
NVIDIA Maxine BNR
lower is better
Quail Multi Speaker is the only model that holds up in both English and Hindi
Against NVIDIA Maxine BNR and Krisp NC, Quail Multi Speaker was the only model that reduced WER under background noise - by 0.94pp (English) and 1.08pp (Hindi). NVIDIA Maxine and Krisp NC increased WER instead, by 2.40–4.29pp across both languages.
Benchmark conducted by SLNG, the execution layer for voice AI.
Independent benchmark - Quail Voice Focus vs. Krisp NC & NVIDIA Maxine BNR
Δ WER vs. no noise cancellation baseline, English competing speech (single talker), full enhancement intensity
Quail Voice Focus
Krisp NC
NVIDIA Maxine BNR
lower is better
Quail Voice Focus is the clear winner on English speaker isolation
On English competing speech, Voice Focus reduced WER by 4.23pp at full intensity - up to 6.96pp at its optimal 0.75 setting. NVIDIA Maxine BNR showed no reliable effect; Krisp NC raised WER by over 10pp.
Benchmark conducted by SLNG, the execution layer for voice AI.
ai-coustics benchmarks
ai-coustics benchmarks
Measured on live audio streams to quantify WER, background speaker suppression, and VAD stability under real-world conditions.
Measured on live audio streams to quantify WER, background speaker suppression, and VAD stability under real-world conditions.
Word error rates on commercial STT models
Raw
Quail Voice Focus 2.2 S
Quail Voice Focus 2.2 L
lower % is better
Left part: Deletions / Middle: Substitutions / Right: Insertions
| AssemblyAI | Deepgram | Soniox | Mistral | Cartesia | Gladia | Speechmatics | Gradium | |
|---|---|---|---|---|---|---|---|---|
| Raw | 50.1% | 74.8% | 80.0% | 40.5% | 51.4% | 50.8% | 68.5% | 53.0% |
| Voice Focus 2.2 S | 17.0% | 17.3% | 13.1% | 13.3% | 16.3% | 13.2% | 12.6% | 18.1% |
| Voice Focus 2.2 L | 14.8% | 16.4% | 12.5% | 12.5% | 15.3% | 13.0% | 11.7% | 16.1% |
Raw: unprocessed microphone input with no enhancement. Voice Focus 2.2 S: our speech enhancement model optimized for machine understanding - 10x smaller than 2.0, built for high call volumes and edge deployments. Voice Focus 2.2 L: our speech enhancement model optimized for machine understanding - best-in-class quality at 25% lower compute than 2.0.
Word error rates on commercial STT models
English subset
Raw
Krisp
Quail Multi Speaker
Top part: Insertions / Middle: Substitutions / Bottom: Deletions
lower % is better
Word error rates on commercial STT models
English subset
Raw
Krisp
Quail Multi Speaker
lower % is better
Top part: Insertions / Middle: Substitutions / Bottom: Deletions
| Deepgram Nova | Cartesia Ink-Whisper | Gladia Solaria | AssemblyAI Universal-2 | ElevenLabs Scribe v1 | |
|---|---|---|---|---|---|
| Raw | 8.8% | 12.7% | 9.8% | 7.5% | 9.6% |
| Krisp | 11.2% | 10.9% | 9.7% | 9.5% | 9.7% |
| Quail | 8.7% | 8.8% | 7.7% | 7.8% | 9.1% |
Raw: unprocessed microphone input with no enhancement. Krisp: a perceptual denoiser optimized for human listening. Quail Voice Focus 2.0: our speech enhancement model optimized for machine understanding.
Word error rates on commercial STT models
German subset
Raw
Krisp
Quail Multi Speaker
Top part: Insertions / Middle: Substitutions / Bottom: Deletions
lower % is better
Word error rates on commercial STT models
German subset
Raw
Krisp
Quail Multi Speaker
lower % is better
Top part: Insertions / Middle: Substitutions / Bottom: Deletions
| Deepgram Nova | Cartesia Ink-Whisper | Gladia Solaria | AssemblyAI Universal-2 | ElevenLabs Scribe v1 | |
|---|---|---|---|---|---|
| Raw | 21.3% | 16.9% | 17.5% | 12.4% | 18.5% |
| Krisp | 22.0% | 19.9% | 20.6% | 15.4% | 22.9% |
| Quail | 19.4% | 15.1% | 12.1% | 12.3% | 19.2% |
Raw: unprocessed microphone input with no enhancement. Krisp: a perceptual denoiser optimized for human listening. Quail Voice Focus 2.0: our speech enhancement model optimized for machine understanding.
VAD Performance Comparison
Silero VAD
VAD 2.1
VAD Performance Comparison
Silero VAD
VAD 2.1
| Accuracy | Recall | |
|---|---|---|
| Silero VAD | 0.67 | 0.48 |
| VAD 2.1 | 0.82 | 0.83 |
Silero VAD: an open-source voice activity detector. VAD 2.1: our voice activity detection model optimized for real-time voice AI pipelines.
Independent benchmark -
Quail Multi Speaker vs. Krisp NC & NVIDIA Maxine BNR
Δ WER vs. no noise cancellation baseline, background noise, full enhancement intensity - English & Hindi
Quail Multi Speaker
Krisp NC
Maxine BNR
lower is better
Against NVIDIA Maxine BNR and Krisp NC, Quail Multi Speaker was the only model that reduced WER under background noise - by 0.94pp (English) and 1.08pp (Hindi). NVIDIA Maxine BNR and Krisp NC increased WER instead, by 2.40–4.29pp across both languages.
Benchmark conducted by SLNG, the execution layer for voice AI.
Independent benchmark -
Quail Voice Focus vs. Krisp NC & NVIDIA Maxine BNR
Δ WER vs. no-NoiseCancellation baseline, English competing speech (single talker), full enhancement intensity
Quail Voice Focus
Krisp NC
Maxine BNR
lower is better
On English competing speech, Voice Focus reduced WER by 4.23pp at full intensity - up to 6.96pp at its optimal 0.75 setting. NVIDIA Maxine BNR showed no reliable effect; Krisp NC raised WER by over 10pp.
Benchmark conducted by SLNG, the execution layer for voice AI.
