/

/

How Beam uses ai-coustics to make real-time interpretation work for the world's most vulnerable users

/

/

How Beam uses ai-coustics to make real-time interpretation work for the world's most vulnerable users

How Beam uses ai-coustics to make real-time interpretation work for the world's most vulnerable users

How Beam uses ai-coustics to make real-time interpretation work for the world's most vulnerable users

Written by

Elgun Guliyev

,

Digital Marketing Manager

Case studies

/

Elgun Guliyev

Written by

,

Digital Marketing Manager

Case studies

/

Beam Interpret lets social workers and citizens who share no common language have a real conversation. Getting the audio right turned out to be the hardest part of building it.

Picture a caseworker in a UK local authority office. Across the table sits a recently arrived refugee who doesn't speak English. Neither can understand the other, until Beam Interpret turns the conversation into a live, two-way exchange. It's made out of a stack of speech-to-text, translation and text-to-speech, working back and forth in real time.

Beam powers thousands of conversations like that one, across 20+ languages. Founded by Alex Stephany, Beam started by delivering frontline services to people experiencing homelessness, helping over 6,000 people into jobs and homes. Today, it's grown into a full AI platform for frontline social workers, caseworkers and clinicians, deployed across local authorities and healthcare organisations in the UK, US and Australia, used by over 100,000 frontline workers. Interpret is one of its newest and fastest-growing products.

But a translation product is only as good as what it hears. When testing Interpret, Beam's engineering team found that the audio layer, more than the language models, was the hardest problem to solve. Solving it meant bringing in ai-coustics' real-time audio insight model, Tyto, to catch problems before they broke a conversation.

Real environments are harder than the lab

"When we first started working on Interpret, we were imagining conversations taking place in a meeting room," says Joel Holmes, Lead Product Engineer at Beam. "But some of our customers wanted to use Interpret in busy offices and call centres, with a constant hum of background noise, people talking right next to them. It’s crucial that we capture what people are saying accurately, so they can receive the right support, but it was really tough to do this with so much background noise."

Noise and interfering speech affected voice activity detection, transcription accuracy and latency all at once. That made it impossible to cleanly detect when someone starts and stops speaking.

"A lot of our users are refugees, coming to a new country, in unfamiliar environments," Joel says. "They're often shy, low-confidence, and speak really quietly."

Given who's on the other end of these conversations, audio quality carries far more weight in Interpret than it would in most products.

Finding ai-coustics

Joel discovered ai-coustics after seeing co-founder Fabian Seipel speak at a Voice AI meetup. A quick Slack conversation later, he used ai-coustics' developer platform himself to build something quick and lightweight into Beam's existing architecture. The results were immediate.

"We got that up and running and definitely saw improvements," he says. "On telephony, it was night and day."

The result changed more than just audio quality - it improved Beam's architecture. Bringing ai-coustics in prompted the team to rethink how they handled audio generally, moving processing from the client to the server for a more stable, production-ready pipeline.

Engineering the fix

Today, every piece of incoming audio in Interpret passes through ai-coustics' Quail Multi Speaker model before it ever reaches Beam's speech-to-text tool. Alongside it, Tyto, ai-coustics' audio insight model, runs on a rolling 6-second buffer of stored audio, scoring overall audio quality, inaudible speech and overspeaking.

That scoring matters because bad audio has a real cost downstream. When quality drops, so does speech-to-text confidence, and Beam ends up either skipping a turn and asking the user to try again, or catching and discarding a transcription error. Either way, the conversation loses its flow. Tyto lets Beam get ahead of that, flagging degrading audio before it derails a turn and pushing guidance back to the client in real time over WebSocket. In in-person mode, that guidance shows up as an on-screen alert: move closer to the mic, change rooms, check your internet connection. Telephony has no screen to alert, so Beam is exploring an audio chime to do the same job. And for the noisiest environments, the fix is physical: Beam sends clients dual headsets with a splitter, so two people can use headsets from a single laptop.

"Our feedback and sentiment went from around 70% positive to now consistently around 90% across all conversations," Joel says. "Some of that's user training and awareness, but ai-coustics has definitely been part of helping us get there."

What's next: three stages

Beam's roadmap for Tyto is built around meeting users at the moment they have the most attention to actually fix a problem:

  • Real-time (live today): guidance during the call itself, though attention is naturally limited when someone's mid-conversation.

  • Staging area (in progress): on the setup page, before a conversation starts, Tyto begins analyzing the user's mic audio right away, flagging background noise or a weak signal so they can fix it before the call begins, while they still have more room to make a change than they would mid-conversation.

  • Post-conversation (planned): after the call ends, when attention is highest, Interpret will surface tips for next time, including simply making users aware that audio quality is something they can control at all.

Beam is also looking to expand its use of ai-coustics beyond Interpret to its other products. "It's easily one of the top two biggest factors in how good a conversation is," Joel says.

"The better transcript we get out, the more we can do with AI, everything flows down from there."

Final logo

Bring real-time audio intelligence into your voice AI stack

Bring real-time audio intelligence into your voice AI stack

Bring real-time audio intelligence into your voice AI stack