I

AI in practice

Making Hebrew speech recognition production-safe

A 52-minute mixed-language meeting crashed our local model. The fix was a routing policy, not a bigger server.

The short answer

Local, Hebrew-tuned speech recognition is accurate and private, but a CPU-only container cannot safely transcribe long recordings. The production-safe design routes each recording by language and length: short Hebrew meetings go to the local model, and long or mixed-language ones go to a cloud model. An automatic fallback surfaces a quality warning instead of failing silently.

Why run it locally at all

Most AI note-takers are English-first. Hebrew, and especially Hebrew mixed with English in the same sentence, degrades their transcripts. Names come out wrong, and right-to-left text is mangled.

MeetSum, our meeting-intelligence platform, uses a Hebrew-tuned Whisper model (ivrit-ai) that runs on our own server. The recordings of client calls never leave infrastructure we control.

The incident

On 17 May 2026, a 52-minute Hebrew and English recording went through the local model. The model loaded in 13.7 seconds, then ran about ten minutes of CPU inference. The weights had been converted from float16 to float32 on a machine without a GPU, memory grew, and the process crashed inside the container.

The meeting still completed, because the pipeline fell back to the cloud model automatically. More importantly, users saw that it had fallen back.

Why the fallback mattered more than the crash

Every stage of the pipeline records the provider, model, latency and confidence behind it. So the fallback was not a silent swap. It showed up as a quality warning on that meeting, and an operator could see exactly which engine produced the transcript.

In an AI system, the dangerous failure is not the crash. It is the result that looks fine and was produced by something other than you think.

The fix: route by risk

Recordings under 15 minutes that are mostly Hebrew go to the local model. Anything longer, or clearly mixed-language, goes to the cloud model. Any failure falls back automatically, with a warning.

The next step is chunking long recordings before local transcription, capping memory and shortening timeouts, so the fallback is fast rather than slow. A built-in word-error-rate harness compares engines on private Hebrew samples, so routing decisions rest on measurement rather than preference.

The general lesson

Multi-model systems need a written policy: which engine handles which job, what happens when it fails, and how that failure becomes visible. The MeetSum case study describes the full pipeline and its integrations.

FROM NOTE TO VENTURE

What physical system are you seeing?

Share the opportunity