<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>AI Engineering on Mitja Martini</title><link>https://mitjamartini.com/en/categories/ai-engineering/</link><description>Recent content in AI Engineering on Mitja Martini</description><generator>Hugo -- gohugo.io</generator><language>en</language><copyright>© 2026 Mitja Martini</copyright><lastBuildDate>Fri, 23 Jan 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://mitjamartini.com/en/categories/ai-engineering/index.xml" rel="self" type="application/rss+xml"/><item><title>Export ChatGPT Conversations as Markdown</title><link>https://mitjamartini.com/en/posts/2026/01/export-chatgpt-conversations-as-markdown/</link><pubDate>Fri, 23 Jan 2026 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/2026/01/export-chatgpt-conversations-as-markdown/</guid><description>&lt;p&gt;Today I learned about &lt;a
href="https://rashidazarang.com"
target="_blank"
&gt;Rashid&amp;rsquo;s&lt;/a&gt; solution to export ChatGPT conversations to Markdown or even PDF if you like. I prefer Markdown as I can use it to continue the conversation with Claude for example, or change the content and use it in other contexts.&lt;/p&gt;
&lt;p&gt;You don&amp;rsquo;t need to install anything because it&amp;rsquo;s just a JavaScript snippet. You open the conversation in a browser, select anything in the conversation, inspect it with developer tools and use the console to paste and run the JavaScript snippet and then the Markdown gets downloaded. It&amp;rsquo;s really simple and extremely helpful in my view.&lt;/p&gt;
&lt;p&gt;Here&amp;rsquo;s the link: &lt;a
href="https://rashidazarang.com/c/export-your-chatgpt-conversations-to-markdown-pdf"
target="_blank"
&gt;Export Your ChatGPT Conversations to Markdown &amp;amp; PDF&lt;/a&gt;&lt;/p&gt;</description></item><item><title>From 8 Lines with Dokku to 200 with Kubernetes – Why I'm Still Switching</title><link>https://mitjamartini.com/en/posts/2026/01/from-8-lines-with-dokku-to-200-with-k8s/</link><pubDate>Tue, 13 Jan 2026 18:21:46 +0200</pubDate><guid>https://mitjamartini.com/en/posts/2026/01/from-8-lines-with-dokku-to-200-with-k8s/</guid><description>&lt;p&gt;So far, my web apps run on a Dokku server. I haven&amp;rsquo;t tried Vercel or Fly because I didn&amp;rsquo;t want to deal with complex pricing models that incur more costs with every additional project.&lt;/p&gt;
&lt;p&gt;Dokku works like Heroku:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;dokku apps:create myapp
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;dokku postgres:create myapp-db
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;dokku postgres:link myapp-db myapp
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;dokku config:set myapp &lt;span class="nv"&gt;APP_SECRET&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;&lt;/span&gt;&lt;span class="k"&gt;$(&lt;/span&gt;openssl rand -base64 48&lt;span class="k"&gt;)&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;dokku git:from-image myapp myregistry.example.com/image:tag
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;dokku domains:set myapp myapp.mitjas.com
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;dokku ports:set myapp http:80:8000 https:443:8000
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;dokku letsencrypt:enable myapp
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;It doesn&amp;rsquo;t get any simpler. After just 8 lines I have&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;an app on the server,&lt;/li&gt;
&lt;li&gt;a PostgreSQL DB with a user and password,&lt;/li&gt;
&lt;li&gt;the app linked to Postgres,&lt;/li&gt;
&lt;li&gt;a secret configured,&lt;/li&gt;
&lt;li&gt;the app deployed from a Docker image,&lt;/li&gt;
&lt;li&gt;a reverse proxy configured to forward HTTP and HTTPS to the app&amp;rsquo;s container, and&lt;/li&gt;
&lt;li&gt;TLS with certificates signed by Let&amp;rsquo;s Encrypt (including automatic renewal).&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;With Kubernetes I need almost &lt;strong&gt;200 lines of YAML&lt;/strong&gt; for this. Why do I still want to switch to Kubernetes?&lt;/p&gt;
&lt;p&gt;Dokku is great for small projects that can make do with the Dokku plugins and run on a single server. But I believe Kubernetes is a better fit for me in the long run.&lt;/p&gt;
&lt;p&gt;Kubernetes manifests tell me (and LLMs) what&amp;rsquo;s running on the cluster, and I can continuously develop and improve them. Dokku, in contrast, feels like &amp;ldquo;fire, forget, and start from scratch.&amp;rdquo; Other advantages of Kubernetes like better scalability, more choice, and security are nice, too, but for me the most important advantage right now is:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Kubernetes is well-suited for Vibe Coding.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;That&amp;rsquo;s my assumption, anyway. Let&amp;rsquo;s see how it goes.&lt;/p&gt;</description></item><item><title>What Takes Time in Vibe Coding</title><link>https://mitjamartini.com/en/posts/2026/01/what-takes-time-in-vibe-coding/</link><pubDate>Tue, 13 Jan 2026 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/2026/01/what-takes-time-in-vibe-coding/</guid><description>&lt;p&gt;Vibe scripting, for me, is when I develop small tools for myself with the help of coding agents. It works extremely well, especially for command-line tools.&lt;/p&gt;
&lt;p&gt;I recently told a tax advisor about it. He was interested and wanted to see how it works. So I developed a small example on the spot: a VAT calculator. Not the best example, but I couldn&amp;rsquo;t think of anything better on short notice.&lt;/p&gt;
&lt;p&gt;It worked well, and after five minutes the VAT calculator with a Flet/Flutter GUI was up and running.&lt;/p&gt;
&lt;p&gt;But I could also see: there&amp;rsquo;s still room for improvement. For example, the layout wasn&amp;rsquo;t great and the functionality was too limited.&lt;/p&gt;
&lt;p&gt;As i wanted to know how long it would take to turn it into an actually useful app, I later developed it to become a VAT calculator for all EU countries, which uses a small AI pipeline to load rates and descriptions of which product categories are subject to which VAT rate from official EU pages and displays them in an improved GUI.&lt;/p&gt;
&lt;figure&gt;
&lt;img
class="my-0 rounded-md"
srcset="
/en/posts/2026/01/what-takes-time-in-vibe-coding/vat-calculator_hu_51fa9f847f16906c.png 330w,
/en/posts/2026/01/what-takes-time-in-vibe-coding/vat-calculator_hu_547addd196020eb9.png 660w,
/en/posts/2026/01/what-takes-time-in-vibe-coding/vat-calculator_hu_8750735acc10413d.png 1024w,
/en/posts/2026/01/what-takes-time-in-vibe-coding/vat-calculator_hu_290d4f807141e56e.png 2x"
src="https://mitjamartini.com/en/posts/2026/01/what-takes-time-in-vibe-coding/vat-calculator_hu_547addd196020eb9.png"
data-zoom-src="https://mitjamartini.com/en/posts/2026/01/what-takes-time-in-vibe-coding/vat-calculator_hu_290d4f807141e56e.png"
alt="EU VAT Calculator"
/&gt;
&lt;figcaption&gt;The EU VAT Calculator after 2h dev time&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;What started with five minutes for the simple version turned into about two hours. Of course, it&amp;rsquo;s much better now. But it&amp;rsquo;s interesting how big the effort difference is between a solution that serves a specific, narrowly defined purpose and a tool with a broader scope. There&amp;rsquo;s a lot of work involved, and finesse is needed to make a tool that&amp;rsquo;s truly useful. Personally, I need iterations with a human in the loop for that. Maybe there are developers who can perfectly specify everything upfront, but I usually need to see and use something to evaluate and improve it.&lt;/p&gt;
&lt;p&gt;AI accelerates iterations enormously, and you can decide whether to consider it &amp;ldquo;good enough&amp;rdquo; sooner or do a few more iterations. I think this is an important reason why I don&amp;rsquo;t develop faster with AI. I invest the time in more iterations and perhaps less thinking upfront, which then requires more iterations again.&lt;/p&gt;</description></item><item><title>MCP in ChatGPT Developer Mode Beta</title><link>https://mitjamartini.com/en/posts/mcp-in-chatgpt-developer-mode-beta/</link><pubDate>Fri, 17 Oct 2025 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/mcp-in-chatgpt-developer-mode-beta/</guid><description>&lt;p&gt;OpenAI just releasesd MCP connectors and ChatGPT developer mode beta. In this post, I describe the process of connecting MCP servers to ChatGPT, show how they look and feel right now in a chat session and give an overview of their current limitations.&lt;/p&gt;
&lt;!-- more --&gt;
&lt;h2 class="relative group"&gt;Activating developer mode
&lt;div id="activating-developer-mode" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#activating-developer-mode" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;p&gt;MCP connectors can only be created and edited in developer mode which can look a bit scary:&lt;/p&gt;
&lt;p&gt;
&lt;figure&gt;
&lt;img
class="my-0 rounded-md"
loading="lazy"
decoding="async"
fetchpriority="low"
alt="ChatGPT&amp;rsquo;s input in developer mode"
srcset="
/en/posts/mcp-in-chatgpt-developer-mode-beta/developer-mode-input_hu_e2dc8f0400ad0174.webp 330w,
/en/posts/mcp-in-chatgpt-developer-mode-beta/developer-mode-input_hu_a3c1a803a54605c4.webp 660w,
/en/posts/mcp-in-chatgpt-developer-mode-beta/developer-mode-input_hu_f8fe73140609504e.webp 1280w
"
data-zoom-src="https://mitjamartini.com/en/posts/mcp-in-chatgpt-developer-mode-beta/developer-mode-input.webp"
src="https://mitjamartini.com/en/posts/mcp-in-chatgpt-developer-mode-beta/developer-mode-input.webp"&gt;
&lt;/figure&gt;
&lt;/p&gt;
&lt;div
class="flex px-4 py-3 rounded-md bg-primary-100 dark:bg-primary-900"
&gt;
&lt;span
class="text-primary-400 ltr:pr-3 rtl:pl-3 flex items-center"
&gt;
&lt;span class="relative block icon"&gt;&lt;svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 512 512"&gt;&lt;path fill="currentColor" d="M506.3 417l-213.3-364c-16.33-28-57.54-28-73.98 0l-213.2 364C-10.59 444.9 9.849 480 42.74 480h426.6C502.1 480 522.6 445 506.3 417zM232 168c0-13.25 10.75-24 24-24S280 154.8 280 168v128c0 13.25-10.75 24-23.1 24S232 309.3 232 296V168zM256 416c-17.36 0-31.44-14.08-31.44-31.44c0-17.36 14.07-31.44 31.44-31.44s31.44 14.08 31.44 31.44C287.4 401.9 273.4 416 256 416z"/&gt;&lt;/svg&gt;
&lt;/span&gt;
&lt;/span&gt;
&lt;span
class="dark:text-neutral-300"
&gt;&lt;strong&gt;Note:&lt;/strong&gt; This article is written from the perspective of a Pro account user. OpenAI&amp;rsquo;s &lt;a
href="https://help.openai.com/de-de/articles/12584461-developer-mode-and-full-mcp-connectors-in-chatgpt-beta"
target="_blank"
&gt;developer mode and MCP connectors in ChatGPT documentation&lt;/a&gt; describes that admins can publish MCPs in their organization. I cannot test this but I assume that MCP connectors are then also usable in normal mode.&lt;/span&gt;
&lt;/div&gt;
&lt;p&gt;Developer mode can be activated in &lt;code&gt;Settings &amp;gt; Apps &amp;amp; Connectors &amp;gt; Advanced Settings&lt;/code&gt;.&lt;/p&gt;
&lt;h2 class="relative group"&gt;Creating an MCP connection
&lt;div id="creating-an-mcp-connection" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#creating-an-mcp-connection" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;p&gt;Once the developer mode is active, MCP server connections can be added in &lt;code&gt;Settings &amp;gt; Apps &amp;amp; Connectors&lt;/code&gt;:&lt;/p&gt;
&lt;p&gt;
&lt;figure&gt;
&lt;img
class="my-0 rounded-md"
loading="lazy"
decoding="async"
fetchpriority="low"
alt="ChatGPT Apps &amp;amp; Connector settings"
srcset="
/en/posts/mcp-in-chatgpt-developer-mode-beta/chatgpt-apps-and-connectors-settings_hu_5a1042c2b213bb56.webp 330w,
/en/posts/mcp-in-chatgpt-developer-mode-beta/chatgpt-apps-and-connectors-settings_hu_41ba0e8b332c5822.webp 660w,
/en/posts/mcp-in-chatgpt-developer-mode-beta/chatgpt-apps-and-connectors-settings_hu_968767f9b892510d.webp 1280w
"
data-zoom-src="https://mitjamartini.com/en/posts/mcp-in-chatgpt-developer-mode-beta/chatgpt-apps-and-connectors-settings.webp"
src="https://mitjamartini.com/en/posts/mcp-in-chatgpt-developer-mode-beta/chatgpt-apps-and-connectors-settings.webp"&gt;
&lt;/figure&gt;
&lt;/p&gt;
&lt;p&gt;Connectors take an icon, name, description, the MCP server url, and authentication information. Both SSE and streaming transports are supported and OAuth 2.0 can be used for authentication.&lt;/p&gt;
&lt;p&gt;
&lt;figure&gt;
&lt;img
class="my-0 rounded-md"
loading="lazy"
decoding="async"
fetchpriority="low"
alt="ChatGPT Create MCP Connector"
srcset="
/en/posts/mcp-in-chatgpt-developer-mode-beta/chatgpt-create-mcp-connector_hu_fd41aa920e51ecd2.webp 330w,
/en/posts/mcp-in-chatgpt-developer-mode-beta/chatgpt-create-mcp-connector_hu_c75172c4f980f518.webp 660w,
/en/posts/mcp-in-chatgpt-developer-mode-beta/chatgpt-create-mcp-connector_hu_215c530d93e9d155.webp 1280w
"
data-zoom-src="https://mitjamartini.com/en/posts/mcp-in-chatgpt-developer-mode-beta/chatgpt-create-mcp-connector.webp"
src="https://mitjamartini.com/en/posts/mcp-in-chatgpt-developer-mode-beta/chatgpt-create-mcp-connector.webp"&gt;
&lt;/figure&gt;
&lt;/p&gt;
&lt;h2 class="relative group"&gt;Using an MCP connection
&lt;div id="using-an-mcp-connection" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#using-an-mcp-connection" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;p&gt;MCP connections need to be activated for each chat session:&lt;/p&gt;
&lt;p&gt;
&lt;figure&gt;
&lt;img
class="my-0 rounded-md"
loading="lazy"
decoding="async"
fetchpriority="low"
alt="Activating an MCP in a chat"
srcset="
/en/posts/mcp-in-chatgpt-developer-mode-beta/activate-custom-mcps-in-a-chat_hu_939be1a88d553781.webp 330w,
/en/posts/mcp-in-chatgpt-developer-mode-beta/activate-custom-mcps-in-a-chat_hu_f3c95f47370b2bf5.webp 660w,
/en/posts/mcp-in-chatgpt-developer-mode-beta/activate-custom-mcps-in-a-chat_hu_eed1daf625e890c5.webp 1280w
"
data-zoom-src="https://mitjamartini.com/en/posts/mcp-in-chatgpt-developer-mode-beta/activate-custom-mcps-in-a-chat.webp"
src="https://mitjamartini.com/en/posts/mcp-in-chatgpt-developer-mode-beta/activate-custom-mcps-in-a-chat.webp"&gt;
&lt;/figure&gt;
&lt;/p&gt;
&lt;p&gt;In my testing, ChatGPT only uses an MCP connection when it has been instructed to do so. Actions need to be confirmed, but decisions can be stored for the duration of a chat:&lt;/p&gt;
&lt;p&gt;
&lt;figure&gt;
&lt;img
class="my-0 rounded-md"
loading="lazy"
decoding="async"
fetchpriority="low"
alt="Custom MCP in a chat session"
srcset="
/en/posts/mcp-in-chatgpt-developer-mode-beta/custom-mcp-in-a-chat-session_hu_aaa2d002974b56da.webp 330w,
/en/posts/mcp-in-chatgpt-developer-mode-beta/custom-mcp-in-a-chat-session_hu_f01ee0872e2d84f3.webp 660w,
/en/posts/mcp-in-chatgpt-developer-mode-beta/custom-mcp-in-a-chat-session_hu_3a18f1259a533a73.webp 1280w
"
data-zoom-src="https://mitjamartini.com/en/posts/mcp-in-chatgpt-developer-mode-beta/custom-mcp-in-a-chat-session.webp"
src="https://mitjamartini.com/en/posts/mcp-in-chatgpt-developer-mode-beta/custom-mcp-in-a-chat-session.webp"&gt;
&lt;/figure&gt;
&lt;/p&gt;
&lt;div
class="flex px-4 py-3 rounded-md bg-primary-100 dark:bg-primary-900"
&gt;
&lt;span
class="text-primary-400 ltr:pr-3 rtl:pl-3 flex items-center"
&gt;
&lt;span class="relative block icon"&gt;&lt;svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 512 512"&gt;&lt;path fill="currentColor" d="M506.3 417l-213.3-364c-16.33-28-57.54-28-73.98 0l-213.2 364C-10.59 444.9 9.849 480 42.74 480h426.6C502.1 480 522.6 445 506.3 417zM232 168c0-13.25 10.75-24 24-24S280 154.8 280 168v128c0 13.25-10.75 24-23.1 24S232 309.3 232 296V168zM256 416c-17.36 0-31.44-14.08-31.44-31.44c0-17.36 14.07-31.44 31.44-31.44s31.44 14.08 31.44 31.44C287.4 401.9 273.4 416 256 416z"/&gt;&lt;/svg&gt;
&lt;/span&gt;
&lt;/span&gt;
&lt;span
class="dark:text-neutral-300"
&gt;&lt;strong&gt;Note&lt;/strong&gt;: ChatGPT marks this action as a &lt;em&gt;write&lt;/em&gt; action, even though it&amp;rsquo;s really a read action from a user point of view. I don&amp;rsquo;t know if it&amp;rsquo;s a mistake on the side of the DeepWiki MCP, but the interesting part for me is, that Pro accounts seem to support write actions, already, even though this is still documented as a limitation.&lt;/span&gt;
&lt;/div&gt;
&lt;p&gt;MCP connections cannot be added to custom GPTs. Thus, there is no way to preconfigure the custom instructions needed to call an MCP server together with the MCP connection.&lt;/p&gt;
&lt;h2 class="relative group"&gt;Current Limitations
&lt;div id="current-limitations" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#current-limitations" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;p&gt;Currently, there are still quite a few limitations:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Not available on free accounts.&lt;/li&gt;
&lt;li&gt;Pro accounts can only use read/fetch actions (according to the documentation, this might not be true anymore)&lt;/li&gt;
&lt;li&gt;local MCPs are not possible.&lt;/li&gt;
&lt;li&gt;Agent mode does not support custom connectors.&lt;/li&gt;
&lt;li&gt;Deep research mode only supports read/fetch actions.&lt;/li&gt;
&lt;li&gt;Only available on the web, not in the mobile app. In the desktop app, connected MCPs are visible but cannot be used (in Pro accounts), as the desktop app does not support developer mode. I assume MCPs might be usable in Business and Enterprise/Edu accounts.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 class="relative group"&gt;Trying to make sense of it
&lt;div id="trying-to-make-sense-of-it" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#trying-to-make-sense-of-it" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;p&gt;Right now, I think MCP connectors in ChatGPT are not yet ready for day-to-day use - no wonder, it&amp;rsquo;s a beta, after all. Meanwhile, actions in custom GPTs are still usable, more open and not as scary to use.&lt;/p&gt;
&lt;p&gt;OpenAI could have added MCP connectors to custom GPTs and provided an option to pre-confirm actions for certain MCP servers but decided to go a different route with the developer mode.&lt;/p&gt;
&lt;p&gt;For me, MCP is a way to customize chatbots and give them &amp;ldquo;agency&amp;rdquo;. I find use-case-specific MCP servers better than generic ones which tend to bloat the context and distract the LLM.&lt;/p&gt;
&lt;p&gt;I hope, OpenAI will evolve ChatGPT&amp;rsquo;s MCP connectors without introducing an obligatory validation procedure, but I won&amp;rsquo;t bet on it, right now.&lt;/p&gt;</description></item><item><title>Claude Code in Devcontainers</title><link>https://mitjamartini.com/en/posts/claude-code-in-devcontainer/</link><pubDate>Tue, 14 Oct 2025 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/claude-code-in-devcontainer/</guid><description>&lt;p&gt;&lt;a
href="https://containers.dev"
target="_blank"
&gt;Development Containers&lt;/a&gt; or just &amp;ldquo;devcontainers&amp;rdquo; add a layer of security, simplify developer onboarding enables developing in parallel with isolated environments.&lt;/p&gt;
&lt;p&gt;For me, Devcontainers are a great addition to an AI Engineer&amp;rsquo;s toolbox, even though I don&amp;rsquo;t use them day-to-day.&lt;/p&gt;
&lt;p&gt;I have setup Devcontainers with Claude Code using VS Code and the devcontainers cli, adapted it to a Python FastAPI based project, and documented it in this article.&lt;/p&gt;</description></item><item><title>Evals for Voice Agents (Session Notes)</title><link>https://mitjamartini.com/en/posts/evals-for-voice-agents-session-notes/</link><pubDate>Mon, 30 Jun 2025 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/evals-for-voice-agents-session-notes/</guid><description>Notes of a whirlwind intro to evals for voice agents by Kwindla and swyx</description></item><item><title>Coding Agents as Slot Machines</title><link>https://mitjamartini.com/en/posts/coding-agents-as-slot-machines/</link><pubDate>Sun, 01 Jun 2025 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/coding-agents-as-slot-machines/</guid><description>&lt;p&gt;After reading a nice post &lt;a
href="https://news.ycombinator.com/item?id=44147966"
target="_blank"
&gt;via HN&lt;/a&gt; about &lt;a
href="https://rjp.io/blog/2025-05-31-stepping-back"
target="_blank"
&gt;On Stepping Back&lt;/a&gt;, I habitually scrolled through the comments and found this nugget by &lt;a
href="https://news.ycombinator.com/user?id=evrimoztamur"
target="_blank"
&gt;evrimoztamur&lt;/a&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Interacting with LLM coding tools is much like playing a slot machine, it grabs and chokeholds your gambling instincts. You&amp;rsquo;re rolling dice for the perfect result without much thought.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This probably applies to many systems that use Generative AI, especially systems that work on the cutting edge where randomness is more pronounced.&lt;/p&gt;
&lt;p&gt;If you also &amp;ldquo;&lt;a
href="https://www.youtube.com/watch?v=sq6a3WC5_Ns"
target="_blank"
&gt;Can&amp;rsquo;t sleep gud anymore&lt;/a&gt;&amp;rdquo; as Mario Zechner beautifully put it, be aware that coding with agents can be addictive.&lt;/p&gt;</description></item><item><title>Pipecat Cloud Latency for EU Users</title><link>https://mitjamartini.com/en/posts/pipecat-cloud-latency-for-eu-users/</link><pubDate>Sun, 11 May 2025 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/pipecat-cloud-latency-for-eu-users/</guid><description>Pipecat Cloud is located in the US. Is its latency ok for voice agents for EU Users.</description></item><item><title>K/V Cache Quantization in Ollama</title><link>https://mitjamartini.com/en/posts/ollama-kv-cache-quantization/</link><pubDate>Sat, 10 May 2025 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/ollama-kv-cache-quantization/</guid><description>&lt;p&gt;A somewhat hidden feature of Ollama is K/V Cache quantization. This is relevant for local AI as it reduces memory consumption, especially for small LLMs with large context windows.&lt;/p&gt;
&lt;p&gt;This post describes how to activate K/V Cache in Ollama and gives an overview of its benefits, drawbacks and use cases.&lt;/p&gt;
&lt;!-- more --&gt;
&lt;h2 class="relative group"&gt;Activating K/V cache quantization in Ollama
&lt;div id="activating-kv-cache-quantization-in-ollama" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#activating-kv-cache-quantization-in-ollama" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;p&gt;K/V Cache quantization in Ollama is not on by default, so you need to activate it by setting the &lt;code&gt;OLLAMA_KV_CACHE_TYPE&lt;/code&gt; environment variable. Supported values are documented in &lt;a
href="https://github.com/ollama/ollama/blob/main/docs/faq.md#how-can-i-set-the-quantization-type-for-the-kv-cache"
target="_blank"
&gt;How can I set the quantization type for the K/V cache?&lt;/a&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;f16&lt;/li&gt;
&lt;li&gt;q8_0&lt;/li&gt;
&lt;li&gt;q4_0&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 class="relative group"&gt;Benefits and Drawbacks
&lt;div id="benefits-and-drawbacks" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#benefits-and-drawbacks" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;p&gt;K/V Cache quantization can be the difference between being able to run a model on a machine, or not. You can use Sam McLeod&amp;rsquo;s &lt;a
href="https://smcleod.net/vram-estimator/"
target="_blank"
&gt;vram-estimator&lt;/a&gt; to estimate the memory consumption of models with different quantization settings.&lt;/p&gt;
&lt;p&gt;The benefit in brief is lower memory consumption which is most pronounced when using small models with large context windows.&lt;/p&gt;
&lt;p&gt;The main drawback is reduced model accuracy. The lmdeploy team has written a blog post about &lt;a
href="https://lmdeploy.readthedocs.io/en/v0.2.3/quantization/kv_int8.html#accuracy-test"
target="_blank"
&gt;K/V cache quantization accuracy test results&lt;/a&gt; if you want to see quantified impacts of K/V quantization on model accuracy.&lt;/p&gt;
&lt;h2 class="relative group"&gt;Example numbers
&lt;div id="example-numbers" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#example-numbers" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;p&gt;Llama 3.2 8B supports 128.000 tokens context windows.&lt;/p&gt;
&lt;p&gt;When you run it with Q4_K_M quantization and the longest possible context length, it consumes&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;23.3 GB memory without K/V cache qantization,&lt;/li&gt;
&lt;li&gt;17.0 GB with Q8_K_0 K/V cache quantization, and&lt;/li&gt;
&lt;li&gt;13.8 GB with Q4_K_0 K/V cache quantization.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;With Q4_K_0 K/V cache quantization it now fits into 16 GB vRAM.&lt;/p&gt;
&lt;h2 class="relative group"&gt;Use cases
&lt;div id="use-cases" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#use-cases" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Application&lt;/th&gt;
&lt;th&gt;Benefit&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;code generation&lt;/td&gt;
&lt;td&gt;more code in context&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;question answering&lt;/td&gt;
&lt;td&gt;whole docs fit into the context&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;function calling&lt;/td&gt;
&lt;td&gt;more tools&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;chat&lt;/td&gt;
&lt;td&gt;longer conversations&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;multi-modal&lt;/td&gt;
&lt;td&gt;images need many tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 class="relative group"&gt;More Info
&lt;div id="more-info" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#more-info" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;p&gt;To learn more, read Sam McLeod&amp;rsquo;s in-dept blog post about &lt;a
href="https://smcleod.net/2024/12/bringing-k/v-context-quantisation-to-ollama/"
target="_blank"
&gt;Bringing K/V Context Quantisation to Ollama&lt;/a&gt;. Sam helped to implement this in Ollama.&lt;/p&gt;</description></item><item><title>Deploying Voice Agents to Production</title><link>https://mitjamartini.com/en/posts/deploying-voice-agents-to-production/</link><pubDate>Fri, 09 May 2025 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/deploying-voice-agents-to-production/</guid><description>&lt;p&gt;Here are my notes on the session about deploying voice agents to production which is part of the &lt;a
href="https://maven.com/pipecat/voice-ai-and-voice-agents-a-technical-deep-dive"
target="_blank"
&gt;Voice Agents Course&lt;/a&gt;.&lt;/p&gt;
&lt;!-- more --&gt;
&lt;div
class="flex px-4 py-3 rounded-md bg-primary-100 dark:bg-primary-900"
&gt;
&lt;span
class="text-primary-400 ltr:pr-3 rtl:pl-3 flex items-center"
&gt;
&lt;span class="relative block icon"&gt;&lt;svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 384 512"&gt;&lt;path fill="currentColor" d="M112.1 454.3c0 6.297 1.816 12.44 5.284 17.69l17.14 25.69c5.25 7.875 17.17 14.28 26.64 14.28h61.67c9.438 0 21.36-6.401 26.61-14.28l17.08-25.68c2.938-4.438 5.348-12.37 5.348-17.7L272 415.1h-160L112.1 454.3zM191.4 .0132C89.44 .3257 16 82.97 16 175.1c0 44.38 16.44 84.84 43.56 115.8c16.53 18.84 42.34 58.23 52.22 91.45c.0313 .25 .0938 .5166 .125 .7823h160.2c.0313-.2656 .0938-.5166 .125-.7823c9.875-33.22 35.69-72.61 52.22-91.45C351.6 260.8 368 220.4 368 175.1C368 78.61 288.9-.2837 191.4 .0132zM192 96.01c-44.13 0-80 35.89-80 79.1C112 184.8 104.8 192 96 192S80 184.8 80 176c0-61.76 50.25-111.1 112-111.1c8.844 0 16 7.159 16 16S200.8 96.01 192 96.01z"/&gt;&lt;/svg&gt;
&lt;/span&gt;
&lt;/span&gt;
&lt;span
class="dark:text-neutral-300"
&gt;The course if held by kwindla und swyx, the CTO und an investor of Daily.co, a WebRTC und Voice AI infrastructure provider. Some of their recommendations might be predisposed. I still state them as is as I trust them and because I don&amp;rsquo;t have enough experience with voice agents in production.&lt;/span&gt;
&lt;/div&gt;
&lt;h2 class="relative group"&gt;TL;DR
&lt;div id="tldr" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#tldr" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Use a voice AI provider for simple, scalable deployment for production.&lt;/li&gt;
&lt;li&gt;Use a single VM or your homelab for demos.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 class="relative group"&gt;Differences between voice agents and traditional web apps
&lt;div id="differences-between-voice-agents-and-traditional-web-apps" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#differences-between-voice-agents-and-traditional-web-apps" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;are mostly in the transport&lt;/li&gt;
&lt;li&gt;persistent connnection (minutes)&lt;/li&gt;
&lt;li&gt;bidirectional streaming&lt;/li&gt;
&lt;li&gt;stateful sessions&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 class="relative group"&gt;Voice agents in production need
&lt;div id="voice-agents-in-production-need" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#voice-agents-in-production-need" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;A http service for
&lt;ul&gt;
&lt;li&gt;API endpoints,&lt;/li&gt;
&lt;li&gt;a website, and&lt;/li&gt;
&lt;li&gt;webhooks,&lt;/li&gt;
&lt;li&gt;spawning bots.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;A media transport layer or service:
&lt;ul&gt;
&lt;li&gt;WebRTC based for client-to-server (udp), or&lt;/li&gt;
&lt;li&gt;websocket based for server-to-server (tcp).&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Bots (udp or tcp, connect to media transport layer)&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 class="relative group"&gt;Bots
&lt;div id="bots" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#bots" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Are instances of the agents.&lt;/li&gt;
&lt;li&gt;Can be written in Python with PipeCat.&lt;/li&gt;
&lt;li&gt;Use STT, LLM, and TTS providers, which also are the main cost and latency drivers.&lt;/li&gt;
&lt;li&gt;Usually come packaged with small models, eg. for voice activity detection (VAD)&lt;/li&gt;
&lt;li&gt;Each spawned bot serves one session and needs allocated resources during the whole session:
&lt;ul&gt;
&lt;li&gt;0,5 vCPU&lt;/li&gt;
&lt;li&gt;1 GB RAM&lt;/li&gt;
&lt;li&gt;40kbps for WebRTC audio (in 30-60 kbps range)&lt;/li&gt;
&lt;li&gt;video requires more CPU (eg. 1 vCPU), and bandwidth&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Need to be quickly available. Target time-to-first-word:
&lt;ul&gt;
&lt;li&gt;2-3 secs (web),&lt;/li&gt;
&lt;li&gt;3-5 secs (phone)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 class="relative group"&gt;Ways to solve the &amp;ldquo;fast start challenge&amp;rdquo;
&lt;div id="ways-to-solve-the-fast-start-challenge" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#ways-to-solve-the-fast-start-challenge" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;percentage-based warm pool&lt;/li&gt;
&lt;li&gt;fast startup times (caching, pre-loading)&lt;/li&gt;
&lt;li&gt;proactive/predictive scheduling&lt;/li&gt;
&lt;li&gt;fallbacks from reactive world (eg. UX based solutions, not just silent fails)&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 class="relative group"&gt;Infra providers
&lt;div id="infra-providers" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#infra-providers" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;need to support tcp and udp&lt;/li&gt;
&lt;li&gt;voice ai providers are easiest (Pipecat Cloud, Daily, Vapi, Layercode)&lt;/li&gt;
&lt;li&gt;Fly.io (and potentially other container platforms) are good if they support udp (Fly does)&lt;/li&gt;
&lt;li&gt;ML focused provides are good for converged bots with larger models included (gpu clouds)&lt;/li&gt;
&lt;li&gt;hyperscalers are flexible but complex&lt;/li&gt;
&lt;li&gt;BTW: CloudRun does not support udp&lt;/li&gt;
&lt;li&gt;demos can run on single VMs or even be served from a home lab&lt;/li&gt;
&lt;li&gt;by serving everything converged, time-to-first word can get down to 500ms&lt;/li&gt;
&lt;li&gt;otherwise 800-1000 ms is good enough and achievable&lt;/li&gt;
&lt;li&gt;proximity to users matters (Daily plans global regions for PipeCat cloud, currently only us-west)&lt;/li&gt;
&lt;li&gt;conn between servers can be implemented with WebSockets&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 class="relative group"&gt;What&amp;rsquo;s next?
&lt;div id="whats-next" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#whats-next" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;p&gt;After the session I have looked at the PipeCat examples and realized it should be easy enough to run a basic voice agent with PipeCat&amp;rsquo;s &lt;a
href="https://docs.pipecat.ai/server/services/transport/small-webrt"
target="_blank"
&gt;SmallWebRTCTransport&lt;/a&gt; on a virtual server hosted in Europe and then switch the transport and deploy it to production on PipeCat Cloud.&lt;/p&gt;
&lt;p&gt;I will probably try that to see if the latency between US based PipeCat cloud and users in Europe is low enough for a good user experience.&lt;/p&gt;</description></item><item><title>An Overview of the Voice AI Landscape (Session Notes)</title><link>https://mitjamartini.com/en/posts/overview-of-voice-ai-landscape/</link><pubDate>Thu, 08 May 2025 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/overview-of-voice-ai-landscape/</guid><description>&lt;p&gt;I&amp;rsquo;m so happy to be part of the &lt;a
href="https://maven.com/pipecat/voice-ai-and-voice-agents-a-technical-deep-dive"
target="_blank"
&gt;Voice Agents Course&lt;/a&gt; by Kwindla and swyx. Yesterday, Kwindla kicked it off with an overview of the voice AI landscape. The pace, insights, and questions from the audience were just great.&lt;/p&gt;
&lt;p&gt;Here are my personal notes, probably incomplete and maybe not always correct. For a more authoritative overview of the Voice AI landscape, check out their free online book &lt;a
href="https://voiceaiandvoiceagents.com"
target="_blank"
&gt;Voice AI &amp;amp; Voice Agents - An Illustrated Primer&lt;/a&gt;.&lt;/p&gt;
&lt;!-- more --&gt;
&lt;ul&gt;
&lt;li&gt;Voice AI has highly valuable use cases with actual real business value.&lt;/li&gt;
&lt;li&gt;Benefits:
&lt;ul&gt;
&lt;li&gt;Today: lower cost.&lt;/li&gt;
&lt;li&gt;Soon: Better. (peak load response, better answers than most humans can give)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;RAG is still important&lt;/li&gt;
&lt;li&gt;Some challenges:
&lt;ul&gt;
&lt;li&gt;latency
&lt;ul&gt;
&lt;li&gt;measure regularly end-to-end from/to clients,&lt;/li&gt;
&lt;li&gt;record conversation with mic,&lt;/li&gt;
&lt;li&gt;visually look at gaps in waveforms,&lt;/li&gt;
&lt;li&gt;test calls from different regions and cell phone providers&lt;/li&gt;
&lt;li&gt;aim for 800ms, tough but possible to hit with hosted inference,&lt;/li&gt;
&lt;li&gt;very optimized/limited deployments can hit 500ms with quality compromises&lt;/li&gt;
&lt;li&gt;1000ms is not uncommon (still makes users happy)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;turn detection
&lt;ul&gt;
&lt;li&gt;from VAD to semantic&lt;/li&gt;
&lt;li&gt;OpenAI shipped a good text mode TDM&lt;/li&gt;
&lt;li&gt;Gemini flash in audio is ok, but needs to run as parallel flow (&amp;ldquo;greedily&amp;rdquo;)&lt;/li&gt;
&lt;li&gt;LifeKit vs. PipeCat is text vs. audio / end of speech, no audio cues vs. with audio cues&lt;/li&gt;
&lt;li&gt;both a hard ML challenge and important for experience&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;interruption handling&lt;/li&gt;
&lt;li&gt;context management&lt;/li&gt;
&lt;li&gt;function calling, tool use&lt;/li&gt;
&lt;li&gt;sripting, instruction following&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;standard architecture
&lt;ul&gt;
&lt;li&gt;today: 3 models (STT - LLM - STT), easier to achieve robust results with LLMs in text mode&lt;/li&gt;
&lt;li&gt;future probably converged&lt;/li&gt;
&lt;li&gt;3 models, because
&lt;ul&gt;
&lt;li&gt;LLMs&amp;rsquo; text mode is their mode&lt;/li&gt;
&lt;li&gt;we ride on the edge what the best models can do&lt;/li&gt;
&lt;li&gt;today, we need to use their best mode&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;other use cases, like language learning, can better leverage the benefits of speech-to-speech (1 model)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;models used (speed and quality, time-to-first token/byte more important than tokens/s)
&lt;ul&gt;
&lt;li&gt;STT:
&lt;ul&gt;
&lt;li&gt;Deepgram&lt;/li&gt;
&lt;li&gt;Whisper (optimized for streaming, original not built for streaming), down at 400&lt;/li&gt;
&lt;li&gt;Gladia (for non-english languages)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;LLM:
&lt;ul&gt;
&lt;li&gt;GPT-4o
&lt;ul&gt;
&lt;li&gt;still the workhorse, still more than 4.1&lt;/li&gt;
&lt;li&gt;big model changes take work and good evals (nobody has good ones),&lt;/li&gt;
&lt;li&gt;models usually not optimized for voice AI,&lt;/li&gt;
&lt;li&gt;not yet better results&lt;/li&gt;
&lt;li&gt;4o-mini was cheaper but slower and worse for tool/function&lt;/li&gt;
&lt;li&gt;note from a fellow student: 4.1-mini might be interesting&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Gemini 2.0 Flash
&lt;ul&gt;
&lt;li&gt;very good, fast, cost efficient&lt;/li&gt;
&lt;li&gt;best audio model for voice AI, today (gemini in audio input mode for voice-to-voice)&lt;/li&gt;
&lt;li&gt;multi-lingual input is ok, output not so much (use English)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;mid-sized open models get better, esp. fine-tuned Llama4 (upcoming OpenPipe finetuning session)&lt;/li&gt;
&lt;li&gt;additional notes about LLM use in voice AI:
&lt;ul&gt;
&lt;li&gt;reliable function calling/tool use is mostly a question of how to prompt 4o against Gemini&lt;/li&gt;
&lt;li&gt;llama can get there&lt;/li&gt;
&lt;li&gt;evals are important as ever&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;TTS:
&lt;ul&gt;
&lt;li&gt;PlayAI, Grok&lt;/li&gt;
&lt;li&gt;OpenAI, Google&lt;/li&gt;
&lt;li&gt;Cartesia&lt;/li&gt;
&lt;li&gt;Rime&lt;/li&gt;
&lt;li&gt;Elevenlabs&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;network transport:
&lt;ul&gt;
&lt;li&gt;why network? We don&amp;rsquo;t get capable enough models on mobile or laptop (yet),&lt;/li&gt;
&lt;li&gt;hybrid architectures might be relevant, already&lt;/li&gt;
&lt;li&gt;telephone is a great transport for voice AI, too
&lt;ul&gt;
&lt;li&gt;PSTN is with a phone number (eg. from Twilio)&lt;/li&gt;
&lt;li&gt;SIP is for interconnectivity with digital telephony infra&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;WebSockets are not good for real time client-to-server audio (and video), but ok for server-to-server audio&lt;/li&gt;
&lt;li&gt;WebRTC is best, complex, but PipeCat supports it ootb, local for testing, in the cloud offering for production,&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Voice AI building blocks in 2025
&lt;ul&gt;
&lt;li&gt;evals - huge topic&lt;/li&gt;
&lt;li&gt;hosting and scaling - very different (if you love k8s, your topic)&lt;/li&gt;
&lt;li&gt;workflow/multi-agent/state machines&lt;/li&gt;
&lt;li&gt;&amp;ldquo;perfect&amp;rdquo; speech
&lt;ul&gt;
&lt;li&gt;LLMs in text mode have passed the turing test&lt;/li&gt;
&lt;li&gt;not quite reached the point for real-time audio recognition and generation&lt;/li&gt;
&lt;li&gt;eg. accurately recording email, postal addess&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;human-like turn detection
&lt;ul&gt;
&lt;li&gt;very important for qualitative experience&lt;/li&gt;
&lt;li&gt;fun and hard ML problem&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;what&amp;rsquo;s next?
&lt;ul&gt;
&lt;li&gt;speech to speech models&lt;/li&gt;
&lt;li&gt;realtime video&lt;/li&gt;
&lt;li&gt;programming with voice&lt;/li&gt;
&lt;li&gt;voice as universal user experience&lt;/li&gt;
&lt;li&gt;LLM as a judge&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;orchestration and flows questions
&lt;ul&gt;
&lt;li&gt;voice in text out can be done with PipeCat aggregator (no great example, yet)&lt;/li&gt;
&lt;li&gt;accents and voice models: Gladia input, PlayAI output&lt;/li&gt;
&lt;li&gt;tip for Gemini: if you want accents, specify your country.&lt;/li&gt;
&lt;li&gt;how to solve cold-start?
&lt;ul&gt;
&lt;li&gt;Daily solved it for us&lt;/li&gt;
&lt;li&gt;own infra: Combine optimized startup and warm capacity&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;faster responses:
&lt;ul&gt;
&lt;li&gt;greedily inference/speculative before turn-detection&lt;/li&gt;
&lt;li&gt;helpful but more costly&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;multi-speaker:
&lt;ul&gt;
&lt;li&gt;hard, no great solution, so far.&lt;/li&gt;
&lt;li&gt;training data doesn&amp;rsquo;t map well to multi-person/agent conversation&lt;/li&gt;
&lt;li&gt;challenge: know when models should or shouldn&amp;rsquo;t respond (emit a no-response-token)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;observability tooling (multiple vendors, eg. Coval, will be in Discord and hold sessions)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Even after the first session, I can already say: If you are interested in Voice AI: &lt;strong&gt;Take the course!&lt;/strong&gt;.&lt;/p&gt;</description></item><item><title>The Fast Solopreneur</title><link>https://mitjamartini.com/en/posts/the-fast-solopreneur/</link><pubDate>Thu, 08 May 2025 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/the-fast-solopreneur/</guid><description>&lt;p&gt;In &lt;a
href="https://www.deeplearning.ai/the-batch/issue-300/"
target="_blank"
&gt;The Batch 300&lt;/a&gt;, Andrew Ng shared some insights about the importance of speed for startups and how to move fast as a startup. I think they apply well to ideas for a fast solopreneur:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Focus on one idea,&lt;/li&gt;
&lt;li&gt;Code prototypes intuitively,&lt;/li&gt;
&lt;li&gt;Be creative about getting user feedback quickly (it doesn&amp;rsquo;t have to scale),&lt;/li&gt;
&lt;li&gt;Pivot quickly but stay within your domain,&lt;/li&gt;
&lt;li&gt;Use a KISS stack you know, and&lt;/li&gt;
&lt;li&gt;Learn AI technology well.&lt;/li&gt;
&lt;/ul&gt;
&lt;!-- more --&gt;
&lt;p&gt;Here is a summary of Andrew&amp;rsquo;s original perspective:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Focus on a concrete idea (don&amp;rsquo;t get distracted)&lt;/li&gt;
&lt;li&gt;Switch quickly to another hypothesis when data shows that the original hypothesis is flawed&lt;/li&gt;
&lt;li&gt;Trust a domain expert&amp;rsquo;s gut instinct&lt;/li&gt;
&lt;li&gt;Build and test prototypes quickly with AI-assisted coding&lt;/li&gt;
&lt;li&gt;Be fast at getting user feedback (this becomes the bottleneck and the competitive advantage)&lt;/li&gt;
&lt;li&gt;Know the technology well (a deep understanding of what AI is good at and what it isn&amp;rsquo;t saves time and avoids dead ends)&lt;/li&gt;
&lt;/ul&gt;</description></item><item><title>A note on the hidden complexities of WebSockets</title><link>https://mitjamartini.com/en/posts/a-note-about-the-hidden-complexities-of-websockets/</link><pubDate>Sat, 25 Jan 2025 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/a-note-about-the-hidden-complexities-of-websockets/</guid><description>&lt;p&gt;AI Apps are often expected to be realtime. On the web, realtime communication can be implemented with WebSockets. I&amp;rsquo;ve started with WebSockets to create chatbots and other live-updated interfaces, but then switched to SSE and now mostly follow these rules of thumbs:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Use WebSockets for server-to-server communication.&lt;/li&gt;
&lt;li&gt;Use SSE for server-to-client communication.&lt;/li&gt;
&lt;li&gt;If you still want or need to use WebSockets for server-to-client communication, add a fallback to SSE.&lt;/li&gt;
&lt;li&gt;If you need realtime voice/video server-to-client communication, use WebRTC.&lt;/li&gt;
&lt;/ul&gt;
&lt;!-- more --&gt;
&lt;p&gt;If you still want or need to build a websocket service, please read Atul Jalan&amp;rsquo;s blog post about &lt;a
href="https://composehq.com/blog/scaling-websockets-1-23-25"
target="_blank"
&gt;The Hidden Complexity of Scaling WebSockets&lt;/a&gt;. It&amp;rsquo;s a quick read and captures important lessons to keep in mind when working with WebSockets.&lt;/p&gt;
&lt;p&gt;All of his lessons are important, even when working not at scale. Here is a quick summary:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Downtimeless deployments are much more involved than in HTTP services.&lt;/li&gt;
&lt;li&gt;Establish a good message schema, eg. 2 byte prefixes and single character field delimiters.&lt;/li&gt;
&lt;li&gt;Use heartbeats to detect dead connections - both ways.&lt;/li&gt;
&lt;li&gt;Have an http fallback as WebSockets are often blocked. Usually: Server sent events (SSE) for server to client communication and simple requests for client to server.&lt;/li&gt;
&lt;li&gt;More, like standard tooling (rate limiting, validation, error handling), no caching at the edge, per-message authentication.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Standard framworks like Django, FastHTML, and Quart (an async Flask clone) have good basic support for WebSockets, but don&amp;rsquo;t really help dealing with their hidden complexities. I hope, frameworks will level up a bit, as this is mostly &amp;ldquo;undifferentiated heavy lifting&amp;rdquo;.&lt;/p&gt;</description></item><item><title>RTX 5090 for Local AI</title><link>https://mitjamartini.com/en/posts/rtx-5090-for-local-ai/</link><pubDate>Tue, 26 Nov 2024 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/rtx-5090-for-local-ai/</guid><description>A look at the NVIDIA RTX 5090 specs for local LLM inference.</description></item></channel></rss>