Research

A cross-model benchmark of streaming noise cancellation on STT accuracy

We tested four commercial streaming NC models before Deepgram Nova-3 on background noise and competing speech, across 3,600 noisy cases in English and Hindi. On background noise, only one model reliably reduced WER in both languages. On competing speech, the largest reduction cut English WER by 30%. Several configurations made transcription worse.

Introduction

Noise cancellation removes noise from a speech-to-text model's input, to improve transcription accuracy. But perceptual quality and recognition accuracy are different objectives, and a model tuned to sound clean to a human can raise word error rate (WER). We tested ai-coustics Quail, Quail Voice Focus, NVIDIA Maxine BNR, and Krisp NC against a no-NC baseline on identical audio.


Methodology

4 models · 2 languages · 6 noise types · 6 SNRs · 3,600 cases · Deepgram Nova-3

Each model ran in streaming mode at multiple intensity settings. Clips came from the FLEURS corpus, 50 per language, mixed with four background-noise and two competing-speech conditions at six SNRs.

Exact model versions, configurations, confidence intervals and per-SNR results are in the PDF.

Findings

Key findings from 3,600 noise conditions.
Carousel image
Carousel image
Carousel image

What we deployed


We ran this benchmark to evaluate NC models for the SLNG Execution layer, and selected the ai-coustics Quail models. Quail was the only model that reduced WER on background noise in both languages. Quail Voice Focus produced the largest reduction in the benchmark.

Quail models now runs in the SLNG Execution Layer as a single request header. Off by default, on where it measurably helps. Noise-type detection is in testing.

Noise cancellation should not be a global toggle

No model we tested reduced WER across all conditions. An NC stage must be validated against:

Noise profile

Background noise and competing speech drive different error modes and require models tuned for each.

Suppression strength

Sweep intensity. Stronger settings can suppress target speech. The best setting is rarely the maximum.

Language

Results in one language cannot be assumed to transfer to another. NC must be validated per deployment language.

Downstream STT

Test against the recognizer that will run in production.

Get the full paper

Full methodology, model configurations, confidence intervals, error breakdown and per-SNR results.

Stay in the loop

LinkedInXGitHub

Define. Execute. Govern. Verify.

Logo
Noise Cancellation Benchmark // SLNG