<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Voice Agents on Mitja Martini</title><link>https://mitjamartini.com/en/tags/voice-agents/</link><description>Recent content in Voice Agents on Mitja Martini</description><generator>Hugo -- gohugo.io</generator><language>en</language><copyright>© 2026 Mitja Martini</copyright><lastBuildDate>Fri, 09 May 2025 00:00:00 +0000</lastBuildDate><atom:link href="https://mitjamartini.com/en/tags/voice-agents/index.xml" rel="self" type="application/rss+xml"/><item><title>Deploying Voice Agents to Production</title><link>https://mitjamartini.com/en/posts/deploying-voice-agents-to-production/</link><pubDate>Fri, 09 May 2025 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/deploying-voice-agents-to-production/</guid><description>&lt;p&gt;Here are my notes on the session about deploying voice agents to production which is part of the &lt;a
href="https://maven.com/pipecat/voice-ai-and-voice-agents-a-technical-deep-dive"
target="_blank"
&gt;Voice Agents Course&lt;/a&gt;.&lt;/p&gt;
&lt;!-- more --&gt;
&lt;div
class="flex px-4 py-3 rounded-md bg-primary-100 dark:bg-primary-900"
&gt;
&lt;span
class="text-primary-400 ltr:pr-3 rtl:pl-3 flex items-center"
&gt;
&lt;span class="relative block icon"&gt;&lt;svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 384 512"&gt;&lt;path fill="currentColor" d="M112.1 454.3c0 6.297 1.816 12.44 5.284 17.69l17.14 25.69c5.25 7.875 17.17 14.28 26.64 14.28h61.67c9.438 0 21.36-6.401 26.61-14.28l17.08-25.68c2.938-4.438 5.348-12.37 5.348-17.7L272 415.1h-160L112.1 454.3zM191.4 .0132C89.44 .3257 16 82.97 16 175.1c0 44.38 16.44 84.84 43.56 115.8c16.53 18.84 42.34 58.23 52.22 91.45c.0313 .25 .0938 .5166 .125 .7823h160.2c.0313-.2656 .0938-.5166 .125-.7823c9.875-33.22 35.69-72.61 52.22-91.45C351.6 260.8 368 220.4 368 175.1C368 78.61 288.9-.2837 191.4 .0132zM192 96.01c-44.13 0-80 35.89-80 79.1C112 184.8 104.8 192 96 192S80 184.8 80 176c0-61.76 50.25-111.1 112-111.1c8.844 0 16 7.159 16 16S200.8 96.01 192 96.01z"/&gt;&lt;/svg&gt;
&lt;/span&gt;
&lt;/span&gt;
&lt;span
class="dark:text-neutral-300"
&gt;The course if held by kwindla und swyx, the CTO und an investor of Daily.co, a WebRTC und Voice AI infrastructure provider. Some of their recommendations might be predisposed. I still state them as is as I trust them and because I don&amp;rsquo;t have enough experience with voice agents in production.&lt;/span&gt;
&lt;/div&gt;
&lt;h2 class="relative group"&gt;TL;DR
&lt;div id="tldr" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#tldr" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Use a voice AI provider for simple, scalable deployment for production.&lt;/li&gt;
&lt;li&gt;Use a single VM or your homelab for demos.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 class="relative group"&gt;Differences between voice agents and traditional web apps
&lt;div id="differences-between-voice-agents-and-traditional-web-apps" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#differences-between-voice-agents-and-traditional-web-apps" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;are mostly in the transport&lt;/li&gt;
&lt;li&gt;persistent connnection (minutes)&lt;/li&gt;
&lt;li&gt;bidirectional streaming&lt;/li&gt;
&lt;li&gt;stateful sessions&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 class="relative group"&gt;Voice agents in production need
&lt;div id="voice-agents-in-production-need" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#voice-agents-in-production-need" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;A http service for
&lt;ul&gt;
&lt;li&gt;API endpoints,&lt;/li&gt;
&lt;li&gt;a website, and&lt;/li&gt;
&lt;li&gt;webhooks,&lt;/li&gt;
&lt;li&gt;spawning bots.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;A media transport layer or service:
&lt;ul&gt;
&lt;li&gt;WebRTC based for client-to-server (udp), or&lt;/li&gt;
&lt;li&gt;websocket based for server-to-server (tcp).&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Bots (udp or tcp, connect to media transport layer)&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 class="relative group"&gt;Bots
&lt;div id="bots" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#bots" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Are instances of the agents.&lt;/li&gt;
&lt;li&gt;Can be written in Python with PipeCat.&lt;/li&gt;
&lt;li&gt;Use STT, LLM, and TTS providers, which also are the main cost and latency drivers.&lt;/li&gt;
&lt;li&gt;Usually come packaged with small models, eg. for voice activity detection (VAD)&lt;/li&gt;
&lt;li&gt;Each spawned bot serves one session and needs allocated resources during the whole session:
&lt;ul&gt;
&lt;li&gt;0,5 vCPU&lt;/li&gt;
&lt;li&gt;1 GB RAM&lt;/li&gt;
&lt;li&gt;40kbps for WebRTC audio (in 30-60 kbps range)&lt;/li&gt;
&lt;li&gt;video requires more CPU (eg. 1 vCPU), and bandwidth&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Need to be quickly available. Target time-to-first-word:
&lt;ul&gt;
&lt;li&gt;2-3 secs (web),&lt;/li&gt;
&lt;li&gt;3-5 secs (phone)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 class="relative group"&gt;Ways to solve the &amp;ldquo;fast start challenge&amp;rdquo;
&lt;div id="ways-to-solve-the-fast-start-challenge" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#ways-to-solve-the-fast-start-challenge" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;percentage-based warm pool&lt;/li&gt;
&lt;li&gt;fast startup times (caching, pre-loading)&lt;/li&gt;
&lt;li&gt;proactive/predictive scheduling&lt;/li&gt;
&lt;li&gt;fallbacks from reactive world (eg. UX based solutions, not just silent fails)&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 class="relative group"&gt;Infra providers
&lt;div id="infra-providers" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#infra-providers" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;need to support tcp and udp&lt;/li&gt;
&lt;li&gt;voice ai providers are easiest (Pipecat Cloud, Daily, Vapi, Layercode)&lt;/li&gt;
&lt;li&gt;Fly.io (and potentially other container platforms) are good if they support udp (Fly does)&lt;/li&gt;
&lt;li&gt;ML focused provides are good for converged bots with larger models included (gpu clouds)&lt;/li&gt;
&lt;li&gt;hyperscalers are flexible but complex&lt;/li&gt;
&lt;li&gt;BTW: CloudRun does not support udp&lt;/li&gt;
&lt;li&gt;demos can run on single VMs or even be served from a home lab&lt;/li&gt;
&lt;li&gt;by serving everything converged, time-to-first word can get down to 500ms&lt;/li&gt;
&lt;li&gt;otherwise 800-1000 ms is good enough and achievable&lt;/li&gt;
&lt;li&gt;proximity to users matters (Daily plans global regions for PipeCat cloud, currently only us-west)&lt;/li&gt;
&lt;li&gt;conn between servers can be implemented with WebSockets&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 class="relative group"&gt;What&amp;rsquo;s next?
&lt;div id="whats-next" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#whats-next" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;p&gt;After the session I have looked at the PipeCat examples and realized it should be easy enough to run a basic voice agent with PipeCat&amp;rsquo;s &lt;a
href="https://docs.pipecat.ai/server/services/transport/small-webrt"
target="_blank"
&gt;SmallWebRTCTransport&lt;/a&gt; on a virtual server hosted in Europe and then switch the transport and deploy it to production on PipeCat Cloud.&lt;/p&gt;
&lt;p&gt;I will probably try that to see if the latency between US based PipeCat cloud and users in Europe is low enough for a good user experience.&lt;/p&gt;</description></item><item><title>An Overview of the Voice AI Landscape (Session Notes)</title><link>https://mitjamartini.com/en/posts/overview-of-voice-ai-landscape/</link><pubDate>Thu, 08 May 2025 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/overview-of-voice-ai-landscape/</guid><description>&lt;p&gt;I&amp;rsquo;m so happy to be part of the &lt;a
href="https://maven.com/pipecat/voice-ai-and-voice-agents-a-technical-deep-dive"
target="_blank"
&gt;Voice Agents Course&lt;/a&gt; by Kwindla and swyx. Yesterday, Kwindla kicked it off with an overview of the voice AI landscape. The pace, insights, and questions from the audience were just great.&lt;/p&gt;
&lt;p&gt;Here are my personal notes, probably incomplete and maybe not always correct. For a more authoritative overview of the Voice AI landscape, check out their free online book &lt;a
href="https://voiceaiandvoiceagents.com"
target="_blank"
&gt;Voice AI &amp;amp; Voice Agents - An Illustrated Primer&lt;/a&gt;.&lt;/p&gt;
&lt;!-- more --&gt;
&lt;ul&gt;
&lt;li&gt;Voice AI has highly valuable use cases with actual real business value.&lt;/li&gt;
&lt;li&gt;Benefits:
&lt;ul&gt;
&lt;li&gt;Today: lower cost.&lt;/li&gt;
&lt;li&gt;Soon: Better. (peak load response, better answers than most humans can give)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;RAG is still important&lt;/li&gt;
&lt;li&gt;Some challenges:
&lt;ul&gt;
&lt;li&gt;latency
&lt;ul&gt;
&lt;li&gt;measure regularly end-to-end from/to clients,&lt;/li&gt;
&lt;li&gt;record conversation with mic,&lt;/li&gt;
&lt;li&gt;visually look at gaps in waveforms,&lt;/li&gt;
&lt;li&gt;test calls from different regions and cell phone providers&lt;/li&gt;
&lt;li&gt;aim for 800ms, tough but possible to hit with hosted inference,&lt;/li&gt;
&lt;li&gt;very optimized/limited deployments can hit 500ms with quality compromises&lt;/li&gt;
&lt;li&gt;1000ms is not uncommon (still makes users happy)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;turn detection
&lt;ul&gt;
&lt;li&gt;from VAD to semantic&lt;/li&gt;
&lt;li&gt;OpenAI shipped a good text mode TDM&lt;/li&gt;
&lt;li&gt;Gemini flash in audio is ok, but needs to run as parallel flow (&amp;ldquo;greedily&amp;rdquo;)&lt;/li&gt;
&lt;li&gt;LifeKit vs. PipeCat is text vs. audio / end of speech, no audio cues vs. with audio cues&lt;/li&gt;
&lt;li&gt;both a hard ML challenge and important for experience&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;interruption handling&lt;/li&gt;
&lt;li&gt;context management&lt;/li&gt;
&lt;li&gt;function calling, tool use&lt;/li&gt;
&lt;li&gt;sripting, instruction following&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;standard architecture
&lt;ul&gt;
&lt;li&gt;today: 3 models (STT - LLM - STT), easier to achieve robust results with LLMs in text mode&lt;/li&gt;
&lt;li&gt;future probably converged&lt;/li&gt;
&lt;li&gt;3 models, because
&lt;ul&gt;
&lt;li&gt;LLMs&amp;rsquo; text mode is their mode&lt;/li&gt;
&lt;li&gt;we ride on the edge what the best models can do&lt;/li&gt;
&lt;li&gt;today, we need to use their best mode&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;other use cases, like language learning, can better leverage the benefits of speech-to-speech (1 model)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;models used (speed and quality, time-to-first token/byte more important than tokens/s)
&lt;ul&gt;
&lt;li&gt;STT:
&lt;ul&gt;
&lt;li&gt;Deepgram&lt;/li&gt;
&lt;li&gt;Whisper (optimized for streaming, original not built for streaming), down at 400&lt;/li&gt;
&lt;li&gt;Gladia (for non-english languages)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;LLM:
&lt;ul&gt;
&lt;li&gt;GPT-4o
&lt;ul&gt;
&lt;li&gt;still the workhorse, still more than 4.1&lt;/li&gt;
&lt;li&gt;big model changes take work and good evals (nobody has good ones),&lt;/li&gt;
&lt;li&gt;models usually not optimized for voice AI,&lt;/li&gt;
&lt;li&gt;not yet better results&lt;/li&gt;
&lt;li&gt;4o-mini was cheaper but slower and worse for tool/function&lt;/li&gt;
&lt;li&gt;note from a fellow student: 4.1-mini might be interesting&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Gemini 2.0 Flash
&lt;ul&gt;
&lt;li&gt;very good, fast, cost efficient&lt;/li&gt;
&lt;li&gt;best audio model for voice AI, today (gemini in audio input mode for voice-to-voice)&lt;/li&gt;
&lt;li&gt;multi-lingual input is ok, output not so much (use English)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;mid-sized open models get better, esp. fine-tuned Llama4 (upcoming OpenPipe finetuning session)&lt;/li&gt;
&lt;li&gt;additional notes about LLM use in voice AI:
&lt;ul&gt;
&lt;li&gt;reliable function calling/tool use is mostly a question of how to prompt 4o against Gemini&lt;/li&gt;
&lt;li&gt;llama can get there&lt;/li&gt;
&lt;li&gt;evals are important as ever&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;TTS:
&lt;ul&gt;
&lt;li&gt;PlayAI, Grok&lt;/li&gt;
&lt;li&gt;OpenAI, Google&lt;/li&gt;
&lt;li&gt;Cartesia&lt;/li&gt;
&lt;li&gt;Rime&lt;/li&gt;
&lt;li&gt;Elevenlabs&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;network transport:
&lt;ul&gt;
&lt;li&gt;why network? We don&amp;rsquo;t get capable enough models on mobile or laptop (yet),&lt;/li&gt;
&lt;li&gt;hybrid architectures might be relevant, already&lt;/li&gt;
&lt;li&gt;telephone is a great transport for voice AI, too
&lt;ul&gt;
&lt;li&gt;PSTN is with a phone number (eg. from Twilio)&lt;/li&gt;
&lt;li&gt;SIP is for interconnectivity with digital telephony infra&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;WebSockets are not good for real time client-to-server audio (and video), but ok for server-to-server audio&lt;/li&gt;
&lt;li&gt;WebRTC is best, complex, but PipeCat supports it ootb, local for testing, in the cloud offering for production,&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Voice AI building blocks in 2025
&lt;ul&gt;
&lt;li&gt;evals - huge topic&lt;/li&gt;
&lt;li&gt;hosting and scaling - very different (if you love k8s, your topic)&lt;/li&gt;
&lt;li&gt;workflow/multi-agent/state machines&lt;/li&gt;
&lt;li&gt;&amp;ldquo;perfect&amp;rdquo; speech
&lt;ul&gt;
&lt;li&gt;LLMs in text mode have passed the turing test&lt;/li&gt;
&lt;li&gt;not quite reached the point for real-time audio recognition and generation&lt;/li&gt;
&lt;li&gt;eg. accurately recording email, postal addess&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;human-like turn detection
&lt;ul&gt;
&lt;li&gt;very important for qualitative experience&lt;/li&gt;
&lt;li&gt;fun and hard ML problem&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;what&amp;rsquo;s next?
&lt;ul&gt;
&lt;li&gt;speech to speech models&lt;/li&gt;
&lt;li&gt;realtime video&lt;/li&gt;
&lt;li&gt;programming with voice&lt;/li&gt;
&lt;li&gt;voice as universal user experience&lt;/li&gt;
&lt;li&gt;LLM as a judge&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;orchestration and flows questions
&lt;ul&gt;
&lt;li&gt;voice in text out can be done with PipeCat aggregator (no great example, yet)&lt;/li&gt;
&lt;li&gt;accents and voice models: Gladia input, PlayAI output&lt;/li&gt;
&lt;li&gt;tip for Gemini: if you want accents, specify your country.&lt;/li&gt;
&lt;li&gt;how to solve cold-start?
&lt;ul&gt;
&lt;li&gt;Daily solved it for us&lt;/li&gt;
&lt;li&gt;own infra: Combine optimized startup and warm capacity&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;faster responses:
&lt;ul&gt;
&lt;li&gt;greedily inference/speculative before turn-detection&lt;/li&gt;
&lt;li&gt;helpful but more costly&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;multi-speaker:
&lt;ul&gt;
&lt;li&gt;hard, no great solution, so far.&lt;/li&gt;
&lt;li&gt;training data doesn&amp;rsquo;t map well to multi-person/agent conversation&lt;/li&gt;
&lt;li&gt;challenge: know when models should or shouldn&amp;rsquo;t respond (emit a no-response-token)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;observability tooling (multiple vendors, eg. Coval, will be in Discord and hold sessions)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Even after the first session, I can already say: If you are interested in Voice AI: &lt;strong&gt;Take the course!&lt;/strong&gt;.&lt;/p&gt;</description></item></channel></rss>