Independent benchmarks

Independent benchmarks

ai-coustics speech enhancement models benchmarked by SLNG against Maxine and Krisp

ai-coustics speech enhancement models benchmarked by SLNG against Maxine and Krisp

Word error rates on commercial STT models

Raw

Quail Voice Focus 2.2 S

Quail Voice Focus 2.2 L

Left part: Deletions / Middle: Substitutions / Right: Insertions

lower % is better

Independent benchmark - Quail Multi Speaker vs. Krisp NC & NVIDIA Maxine BNR

Δ WER vs. no noise cancellation baseline, background noise, full enhancement intensity - English & Hindi

Δ WER vs. no noise cancellation baseline, background noise, full enhancement intensity - English & Hindi

Quail Multi Speaker

Krisp NC

NVIDIA Maxine BNR

lower is better

-1no-NC+1+2+3+4+5-0.94Quail MS+3.39Krisp NC+2.40Maxine BNREnglish-1.08Quail MS+4.29Krisp NC+2.76Maxine BNRHindi
-1no-NC+1+2+3+4+5-0.94Quail+3.39Krisp NC+2.40Maxine BNREnglish-1.08Quail+4.29Krisp NC+2.76Maxine BNRHindi

Quail Multi Speaker is the only model that holds up in both English and Hindi

Against NVIDIA Maxine BNR and Krisp NC, Quail Multi Speaker was the only model that reduced WER under background noise - by 0.94pp (English) and 1.08pp (Hindi). NVIDIA Maxine and Krisp NC increased WER instead, by 2.40–4.29pp across both languages.

Benchmark conducted by SLNG, the execution layer for voice AI.

Independent benchmark - Quail Voice Focus vs. Krisp NC & NVIDIA Maxine BNR

Δ WER vs. no noise cancellation baseline, English competing speech (single talker), full enhancement intensity

Quail Voice Focus

Krisp NC

NVIDIA Maxine BNR

lower is better

-5-2.5no-NC+2.5+5+7.5+10-4.23Quail Voice Focus+10.03Krisp NC-0.20*Maxine BNR
* not reliably different from zero (95% CI crosses zero)
-5-2.5no-NC+2.5+5+7.5+10-4.23Quail Voice Focus+10.03Krisp NC-0.20*Maxine BNR
* not reliably different from zero (95% CI crosses zero)

Quail Voice Focus is the clear winner on English speaker isolation

On English competing speech, Voice Focus reduced WER by 4.23pp at full intensity - up to 6.96pp at its optimal 0.75 setting. NVIDIA Maxine BNR showed no reliable effect; Krisp NC raised WER by over 10pp.

Benchmark conducted by SLNG, the execution layer for voice AI.

ai-coustics benchmarks

ai-coustics benchmarks

Measured on live audio streams to quantify WER, background speaker suppression, and VAD stability under real-world conditions.

Measured on live audio streams to quantify WER, background speaker suppression, and VAD stability under real-world conditions.

Word error rates on commercial STT models

Raw

Quail Voice Focus 2.2 S

Quail Voice Focus 2.2 L

lower % is better

Left part: Deletions / Middle: Substitutions / Right: Insertions

AssemblyAIDeepgramSonioxMistralCartesiaGladiaSpeechmaticsGradium
Raw50.1%74.8%80.0%40.5%51.4%50.8%68.5%53.0%
Voice Focus 2.2 S17.0%17.3%13.1%13.3%16.3%13.2%12.6%18.1%
Voice Focus 2.2 L14.8%16.4%12.5%12.5%15.3%13.0%11.7%16.1%

Raw: unprocessed microphone input with no enhancement. Voice Focus 2.2 S: our speech enhancement model optimized for machine understanding - 10x smaller than 2.0, built for high call volumes and edge deployments. Voice Focus 2.2 L: our speech enhancement model optimized for machine understanding - best-in-class quality at 25% lower compute than 2.0.

Word error rates on commercial STT models

English subset

Raw

Krisp

Quail Multi Speaker

Top part: Insertions / Middle: Substitutions / Bottom: Deletions

lower % is better

Word error rates on commercial STT models

English subset

Raw

Krisp

Quail Multi Speaker

lower % is better

Top part: Insertions / Middle: Substitutions / Bottom: Deletions

Deepgram NovaCartesia Ink-WhisperGladia SolariaAssemblyAI Universal-2ElevenLabs Scribe v1
Raw8.8%12.7%9.8%7.5%9.6%
Krisp11.2%10.9%9.7%9.5%9.7%
Quail8.7%8.8%7.7%7.8%9.1%

Raw: unprocessed microphone input with no enhancement. Krisp: a perceptual denoiser optimized for human listening. Quail Voice Focus 2.0: our speech enhancement model optimized for machine understanding.

Word error rates on commercial STT models

German subset

Raw

Krisp

Quail Multi Speaker

Top part: Insertions / Middle: Substitutions / Bottom: Deletions

lower % is better

Word error rates on commercial STT models

German subset

Raw

Krisp

Quail Multi Speaker

lower % is better

Top part: Insertions / Middle: Substitutions / Bottom: Deletions

Deepgram NovaCartesia Ink-WhisperGladia SolariaAssemblyAI Universal-2ElevenLabs Scribe v1
Raw21.3%16.9%17.5%12.4%18.5%
Krisp22.0%19.9%20.6%15.4%22.9%
Quail19.4%15.1%12.1%12.3%19.2%

Raw: unprocessed microphone input with no enhancement. Krisp: a perceptual denoiser optimized for human listening. Quail Voice Focus 2.0: our speech enhancement model optimized for machine understanding.

VAD Performance Comparison

Silero VAD

VAD 2.1

VAD Performance Comparison

Silero VAD

VAD 2.1

AccuracyRecall
Silero VAD0.670.48
VAD 2.10.820.83

Silero VAD: an open-source voice activity detector. VAD 2.1: our voice activity detection model optimized for real-time voice AI pipelines.

Independent benchmark -
Quail Multi Speaker vs. Krisp NC & NVIDIA Maxine BNR

Δ WER vs. no noise cancellation baseline, background noise, full enhancement intensity - English & Hindi

Quail Multi Speaker

Krisp NC

Maxine BNR

lower is better

no-NCEnglishQuail-0.94Krisp NC+3.39Maxine BNR+2.40HindiQuail-1.08Krisp NC+4.29Maxine BNR+2.76

Against NVIDIA Maxine BNR and Krisp NC, Quail Multi Speaker was the only model that reduced WER under background noise - by 0.94pp (English) and 1.08pp (Hindi). NVIDIA Maxine BNR and Krisp NC increased WER instead, by 2.40–4.29pp across both languages.

Benchmark conducted by SLNG, the execution layer for voice AI.

Independent benchmark -
Quail Voice Focus vs. Krisp NC & NVIDIA Maxine BNR

Δ WER vs. no-NoiseCancellation baseline, English competing speech (single talker), full enhancement intensity

Quail Voice Focus

Krisp NC

Maxine BNR

lower is better

no-NCQuail Voice Focus-4.23Krisp NC+10.03Maxine BNR-0.20*
* not reliably different from zero (95% CI crosses zero)

On English competing speech, Voice Focus reduced WER by 4.23pp at full intensity - up to 6.96pp at its optimal 0.75 setting. NVIDIA Maxine BNR showed no reliable effect; Krisp NC raised WER by over 10pp.

Benchmark conducted by SLNG, the execution layer for voice AI.

Final logo

Bring real-time audio intelligence into your voice AI stack

Bring real-time audio intelligence into your voice AI stack

Bring real-time audio intelligence into your voice AI stack