Companion code for Building Production Voice AI Agents (Book 5, Production AI Agent Engineering series) - IAgentic-LLC/voice-agents
1 comment
Call a voice agent and press digits on your keypad. In the pipeline I tested, if the agent listens through a speech-to-text model, it hears nothing: across 6 real calls of nothing but touch-tone keypresses, the transcriber produced zero transcript and never even triggered voice-activity detection. Not "got it wrong," got literally nothing.
Then I ran the same tone-decoder (a Goertzel filter over 20ms blocks, the actual DTMF detection algorithm) against a normal spoken sentence instead of tones. It hallucinated 109 phantom keypresses that nobody ever pressed.
The robust fix for our SIP/RTP path was RFC 4733: carry DTMF as named RTP telephone events instead of asking a speech model or an audio Goertzel detector to infer keypad input. Once I wired that up on a real LiveKit + Twilio SIP call, the same kind of calls that produced zero transcript for tones instead surfaced correct keypresses through room.on("sip_dtmf_received"), no STT model involved, 12/12 correct across real calls.
Full writeup with the actual traces is chapter 13/14 of "Building Production Voice AI Agents" (https://www.amazon.com/dp/B0HL53Y5HK), but this repo and its tests stand on their own.
Read the full thread on Hacker News →
Related stories
- Hacker News · 7 points · 8 days ago
- So Zero It's ... Negative? (Zero-Copy #3)manishearth.github.ioLobsters · 14 points · about 4 years ago
- The Verge · 0 points · 8 days ago
- Hacker News · 1 points · 7 days ago
- Show HN: OpenDecision – a 400M zero-shot model makes local decisions, plays Doomdeepanwadhwa.github.ioHacker News · 6 points · 10 days ago
- Hacker News · 2 points · 9 days ago