<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>WebRTC on Mitja Martini</title><link>https://mitjamartini.com/en/tags/webrtc/</link><description>Recent content in WebRTC on Mitja Martini</description><generator>Hugo -- gohugo.io</generator><language>en</language><copyright>© 2026 Mitja Martini</copyright><lastBuildDate>Sun, 11 May 2025 00:00:00 +0000</lastBuildDate><atom:link href="https://mitjamartini.com/en/tags/webrtc/index.xml" rel="self" type="application/rss+xml"/><item><title>Pipecat Cloud Latency for EU Users</title><link>https://mitjamartini.com/en/posts/pipecat-cloud-latency-for-eu-users/</link><pubDate>Sun, 11 May 2025 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/pipecat-cloud-latency-for-eu-users/</guid><description>Pipecat Cloud is located in the US. Is its latency ok for voice agents for EU Users.</description></item><item><title>Deploying Voice Agents to Production</title><link>https://mitjamartini.com/en/posts/deploying-voice-agents-to-production/</link><pubDate>Fri, 09 May 2025 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/deploying-voice-agents-to-production/</guid><description>&lt;p&gt;Here are my notes on the session about deploying voice agents to production which is part of the &lt;a
href="https://maven.com/pipecat/voice-ai-and-voice-agents-a-technical-deep-dive"
target="_blank"
&gt;Voice Agents Course&lt;/a&gt;.&lt;/p&gt;
&lt;!-- more --&gt;
&lt;div
class="flex px-4 py-3 rounded-md bg-primary-100 dark:bg-primary-900"
&gt;
&lt;span
class="text-primary-400 ltr:pr-3 rtl:pl-3 flex items-center"
&gt;
&lt;span class="relative block icon"&gt;&lt;svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 384 512"&gt;&lt;path fill="currentColor" d="M112.1 454.3c0 6.297 1.816 12.44 5.284 17.69l17.14 25.69c5.25 7.875 17.17 14.28 26.64 14.28h61.67c9.438 0 21.36-6.401 26.61-14.28l17.08-25.68c2.938-4.438 5.348-12.37 5.348-17.7L272 415.1h-160L112.1 454.3zM191.4 .0132C89.44 .3257 16 82.97 16 175.1c0 44.38 16.44 84.84 43.56 115.8c16.53 18.84 42.34 58.23 52.22 91.45c.0313 .25 .0938 .5166 .125 .7823h160.2c.0313-.2656 .0938-.5166 .125-.7823c9.875-33.22 35.69-72.61 52.22-91.45C351.6 260.8 368 220.4 368 175.1C368 78.61 288.9-.2837 191.4 .0132zM192 96.01c-44.13 0-80 35.89-80 79.1C112 184.8 104.8 192 96 192S80 184.8 80 176c0-61.76 50.25-111.1 112-111.1c8.844 0 16 7.159 16 16S200.8 96.01 192 96.01z"/&gt;&lt;/svg&gt;
&lt;/span&gt;
&lt;/span&gt;
&lt;span
class="dark:text-neutral-300"
&gt;The course if held by kwindla und swyx, the CTO und an investor of Daily.co, a WebRTC und Voice AI infrastructure provider. Some of their recommendations might be predisposed. I still state them as is as I trust them and because I don&amp;rsquo;t have enough experience with voice agents in production.&lt;/span&gt;
&lt;/div&gt;
&lt;h2 class="relative group"&gt;TL;DR
&lt;div id="tldr" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#tldr" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Use a voice AI provider for simple, scalable deployment for production.&lt;/li&gt;
&lt;li&gt;Use a single VM or your homelab for demos.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 class="relative group"&gt;Differences between voice agents and traditional web apps
&lt;div id="differences-between-voice-agents-and-traditional-web-apps" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#differences-between-voice-agents-and-traditional-web-apps" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;are mostly in the transport&lt;/li&gt;
&lt;li&gt;persistent connnection (minutes)&lt;/li&gt;
&lt;li&gt;bidirectional streaming&lt;/li&gt;
&lt;li&gt;stateful sessions&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 class="relative group"&gt;Voice agents in production need
&lt;div id="voice-agents-in-production-need" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#voice-agents-in-production-need" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;A http service for
&lt;ul&gt;
&lt;li&gt;API endpoints,&lt;/li&gt;
&lt;li&gt;a website, and&lt;/li&gt;
&lt;li&gt;webhooks,&lt;/li&gt;
&lt;li&gt;spawning bots.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;A media transport layer or service:
&lt;ul&gt;
&lt;li&gt;WebRTC based for client-to-server (udp), or&lt;/li&gt;
&lt;li&gt;websocket based for server-to-server (tcp).&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Bots (udp or tcp, connect to media transport layer)&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 class="relative group"&gt;Bots
&lt;div id="bots" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#bots" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Are instances of the agents.&lt;/li&gt;
&lt;li&gt;Can be written in Python with PipeCat.&lt;/li&gt;
&lt;li&gt;Use STT, LLM, and TTS providers, which also are the main cost and latency drivers.&lt;/li&gt;
&lt;li&gt;Usually come packaged with small models, eg. for voice activity detection (VAD)&lt;/li&gt;
&lt;li&gt;Each spawned bot serves one session and needs allocated resources during the whole session:
&lt;ul&gt;
&lt;li&gt;0,5 vCPU&lt;/li&gt;
&lt;li&gt;1 GB RAM&lt;/li&gt;
&lt;li&gt;40kbps for WebRTC audio (in 30-60 kbps range)&lt;/li&gt;
&lt;li&gt;video requires more CPU (eg. 1 vCPU), and bandwidth&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Need to be quickly available. Target time-to-first-word:
&lt;ul&gt;
&lt;li&gt;2-3 secs (web),&lt;/li&gt;
&lt;li&gt;3-5 secs (phone)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 class="relative group"&gt;Ways to solve the &amp;ldquo;fast start challenge&amp;rdquo;
&lt;div id="ways-to-solve-the-fast-start-challenge" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#ways-to-solve-the-fast-start-challenge" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;percentage-based warm pool&lt;/li&gt;
&lt;li&gt;fast startup times (caching, pre-loading)&lt;/li&gt;
&lt;li&gt;proactive/predictive scheduling&lt;/li&gt;
&lt;li&gt;fallbacks from reactive world (eg. UX based solutions, not just silent fails)&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 class="relative group"&gt;Infra providers
&lt;div id="infra-providers" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#infra-providers" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;need to support tcp and udp&lt;/li&gt;
&lt;li&gt;voice ai providers are easiest (Pipecat Cloud, Daily, Vapi, Layercode)&lt;/li&gt;
&lt;li&gt;Fly.io (and potentially other container platforms) are good if they support udp (Fly does)&lt;/li&gt;
&lt;li&gt;ML focused provides are good for converged bots with larger models included (gpu clouds)&lt;/li&gt;
&lt;li&gt;hyperscalers are flexible but complex&lt;/li&gt;
&lt;li&gt;BTW: CloudRun does not support udp&lt;/li&gt;
&lt;li&gt;demos can run on single VMs or even be served from a home lab&lt;/li&gt;
&lt;li&gt;by serving everything converged, time-to-first word can get down to 500ms&lt;/li&gt;
&lt;li&gt;otherwise 800-1000 ms is good enough and achievable&lt;/li&gt;
&lt;li&gt;proximity to users matters (Daily plans global regions for PipeCat cloud, currently only us-west)&lt;/li&gt;
&lt;li&gt;conn between servers can be implemented with WebSockets&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 class="relative group"&gt;What&amp;rsquo;s next?
&lt;div id="whats-next" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#whats-next" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;p&gt;After the session I have looked at the PipeCat examples and realized it should be easy enough to run a basic voice agent with PipeCat&amp;rsquo;s &lt;a
href="https://docs.pipecat.ai/server/services/transport/small-webrt"
target="_blank"
&gt;SmallWebRTCTransport&lt;/a&gt; on a virtual server hosted in Europe and then switch the transport and deploy it to production on PipeCat Cloud.&lt;/p&gt;
&lt;p&gt;I will probably try that to see if the latency between US based PipeCat cloud and users in Europe is low enough for a good user experience.&lt;/p&gt;</description></item></channel></rss>