<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Posts on Mitja Martini</title><link>https://mitjamartini.com/en/posts/</link><description>Recent content in Posts on Mitja Martini</description><generator>Hugo -- gohugo.io</generator><language>en</language><copyright>© 2026 Mitja Martini</copyright><lastBuildDate>Sun, 07 Sep 2025 00:00:00 +0000</lastBuildDate><atom:link href="https://mitjamartini.com/en/posts/index.xml" rel="self" type="application/rss+xml"/><item><title>Senior means simpler</title><link>https://mitjamartini.com/en/posts/2026/01/senior-means-simpler/</link><pubDate>Tue, 27 Jan 2026 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/2026/01/senior-means-simpler/</guid><description>&lt;p&gt;When I was young, I thought of senior engineers, consultants, etc. as experts that are more effective. My assumption was that seniors work harder to get there and be senior.&lt;/p&gt;
&lt;p&gt;Now that I&amp;rsquo;m older, I realize: Being efficient is as important and probably more distintctive than being effective: I need to be calm, save time, and work efficiently to achieve the goals in a relaxed way, as I don&amp;rsquo;t have as much energy anymore and because time feels much more precious.&lt;/p&gt;
&lt;p&gt;Being senior is all about simplifying.&lt;/p&gt;</description></item><item><title>Export ChatGPT Conversations as Markdown</title><link>https://mitjamartini.com/en/posts/2026/01/export-chatgpt-conversations-as-markdown/</link><pubDate>Fri, 23 Jan 2026 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/2026/01/export-chatgpt-conversations-as-markdown/</guid><description>&lt;p&gt;Today I learned about &lt;a
href="https://rashidazarang.com"
target="_blank"
&gt;Rashid&amp;rsquo;s&lt;/a&gt; solution to export ChatGPT conversations to Markdown or even PDF if you like. I prefer Markdown as I can use it to continue the conversation with Claude for example, or change the content and use it in other contexts.&lt;/p&gt;
&lt;p&gt;You don&amp;rsquo;t need to install anything because it&amp;rsquo;s just a JavaScript snippet. You open the conversation in a browser, select anything in the conversation, inspect it with developer tools and use the console to paste and run the JavaScript snippet and then the Markdown gets downloaded. It&amp;rsquo;s really simple and extremely helpful in my view.&lt;/p&gt;
&lt;p&gt;Here&amp;rsquo;s the link: &lt;a
href="https://rashidazarang.com/c/export-your-chatgpt-conversations-to-markdown-pdf"
target="_blank"
&gt;Export Your ChatGPT Conversations to Markdown &amp;amp; PDF&lt;/a&gt;&lt;/p&gt;</description></item><item><title>Hardware for local coding models is still affordable. For how long?</title><link>https://mitjamartini.com/en/posts/2026/01/hardware-for-local-coding-models-still-affordable/</link><pubDate>Sat, 17 Jan 2026 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/2026/01/hardware-for-local-coding-models-still-affordable/</guid><description>&lt;p&gt;The recent RAM price hikes have pushed GPU prices up as well. The only systems that have not yet been affected to the same extent are Macs and high-end GPUs (RTX 6000 Pro and above). However, I would already classify GPUs as out of reach: running coding models with sufficiently large context windows would require one or two RTX 6000 Pro cards, or three to six RTX 5090s.&lt;/p&gt;
&lt;p&gt;Let’s take a look at Macs instead.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Price&lt;/th&gt;
&lt;th&gt;Mem Bandwidth&lt;/th&gt;
&lt;th&gt;CPU&lt;/th&gt;
&lt;th&gt;RAM&lt;/th&gt;
&lt;th&gt;NVMe&lt;/th&gt;
&lt;th&gt;CPU&lt;/th&gt;
&lt;th&gt;GPU&lt;/th&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Price/Bandwidth&lt;/th&gt;
&lt;th&gt;Price/Memory&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;2.500 EUR&lt;/td&gt;
&lt;td&gt;273 GB/s&lt;/td&gt;
&lt;td&gt;M4 Pro&lt;/td&gt;
&lt;td&gt;64 GB&lt;/td&gt;
&lt;td&gt;1 TB&lt;/td&gt;
&lt;td&gt;12C&lt;/td&gt;
&lt;td&gt;16C&lt;/td&gt;
&lt;td&gt;Mac Mini&lt;/td&gt;
&lt;td&gt;9,15 EUR/GB/s&lt;/td&gt;
&lt;td&gt;39,06 EUR/GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4.027 EUR&lt;/td&gt;
&lt;td&gt;546 GB/s&lt;/td&gt;
&lt;td&gt;M4 Max&lt;/td&gt;
&lt;td&gt;128 GB&lt;/td&gt;
&lt;td&gt;1 TB&lt;/td&gt;
&lt;td&gt;16C&lt;/td&gt;
&lt;td&gt;40C&lt;/td&gt;
&lt;td&gt;Mac Studio&lt;/td&gt;
&lt;td&gt;7,14 EUR/GB/s&lt;/td&gt;
&lt;td&gt;31,43 EUR/GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4.200 EUR&lt;/td&gt;
&lt;td&gt;800 GB/s&lt;/td&gt;
&lt;td&gt;M3 Ultra&lt;/td&gt;
&lt;td&gt;96 GB&lt;/td&gt;
&lt;td&gt;1 TB&lt;/td&gt;
&lt;td&gt;28C&lt;/td&gt;
&lt;td&gt;60C&lt;/td&gt;
&lt;td&gt;Mac Studio&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;5,52 EUR/GB/s&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;43,75 EUR/GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6.720 EUR&lt;/td&gt;
&lt;td&gt;800 GB/s&lt;/td&gt;
&lt;td&gt;M3 Ultra&lt;/td&gt;
&lt;td&gt;256 GB&lt;/td&gt;
&lt;td&gt;2 TB&lt;/td&gt;
&lt;td&gt;28C&lt;/td&gt;
&lt;td&gt;60C&lt;/td&gt;
&lt;td&gt;Mac Studio&lt;/td&gt;
&lt;td&gt;8,40 EUR/GB/s&lt;/td&gt;
&lt;td&gt;26,25 EUR/GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;11.900 EUR&lt;/td&gt;
&lt;td&gt;800 GB/s&lt;/td&gt;
&lt;td&gt;M3 Ultra&lt;/td&gt;
&lt;td&gt;512 GB&lt;/td&gt;
&lt;td&gt;4 TB&lt;/td&gt;
&lt;td&gt;32C&lt;/td&gt;
&lt;td&gt;80C&lt;/td&gt;
&lt;td&gt;Mac Studio&lt;/td&gt;
&lt;td&gt;14,87 EUR/GB/s&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;23,24 EUR/GB&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Mac Studios with the M3 Ultra currently offer the best overall value. The 96 GB version is fast, and its RAM capacity is sufficient to run many general-purpose models using quantization. The 256 GB version can handle coding models at reasonable quantization levels with 64k context, which is just about sufficient for running coding agents &lt;sup id="fnref:1"&gt;&lt;a href="#fn:1" class="footnote-ref" role="doc-noteref"&gt;1&lt;/a&gt;&lt;/sup&gt;. The 512 GB model, however, exceeds the €10k, which I consider a hard upper limit for what can reasonably be called an “affordable” system.&lt;/p&gt;
&lt;p&gt;The price difference between the 96 GB and 256 GB Mac Studio models is surprisingly close to current RAM market prices. For comparison, 256 GB of DDR5-6000 RAM currently costs around €3,430, or €13.39 per GB. The price delta between the 96 GB and 256 GB Mac Studio, and between the 256 GB and 512 GB versions, is approximately €12.88 per GB and €18.50 per GB, respectively. If Apple were to return to its usual memory pricing premiums, the 256 GB configuration would likely end up much closer to €10,000.&lt;/p&gt;
&lt;p&gt;Hosted inference is both cheaper and more capable, but it cannot be used in all scenarios. In particular, freelancers and small companies are often required to rely on local systems when working for larger clients. The M3 Ultra, launched in March 2025, together with improved models, quantization techniques, and REAP, has effectively kicked off an era of affordable local coding models. That era may now enter a pause: RAM prices are expected to rise through 2026 and are likely to remain elevated until 2027 or 2028 &lt;sup id="fnref:2"&gt;&lt;a href="#fn:2" class="footnote-ref" role="doc-noteref"&gt;2&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;div class="footnotes" role="doc-endnotes"&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id="fn:1"&gt;
&lt;p&gt;Based on benchmark results shared on &lt;a
href="https://www.reddit.com/r/LocalLLaMA/comments/1pw8h6w/glm476bit_mlx_vs_minimaxm216bit_mlx_benchmark/"
target="_blank"
&gt;r/localllama&lt;/a&gt;&amp;#160;&lt;a href="#fnref:1" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:2"&gt;
&lt;p&gt;Based on an analysis from December 2025 on &lt;a
href="https://wccftech.com/memory-ddr5-ddr4-shortages-last-till-q4-2027-higher-prices-throughout-2026/"
target="_blank"
&gt;wccftech&lt;/a&gt;&amp;#160;&lt;a href="#fnref:2" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</description></item><item><title>MCP and A2A Attack Vectors for AI Agents</title><link>https://mitjamartini.com/en/posts/2026/01/mcp-a2a-attack-vectors/</link><pubDate>Wed, 14 Jan 2026 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/2026/01/mcp-a2a-attack-vectors/</guid><description>&lt;p&gt;Christian Posta from Solo.io has written an interesting &lt;a
href="https://www.solo.io/blog/deep-dive-mcp-and-a2a-attack-vectors-for-ai-agents"
target="_blank"
&gt;Deep Dive into MCP and A2A Attack Vectors for AI Agents&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Here is a shortlist. There are certainly more attack vectors, and the mitigations are a start, but certainly not perfect.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Attack Vector&lt;/th&gt;
&lt;th&gt;How it works&lt;/th&gt;
&lt;th&gt;Mitigation&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Naming Attacks&lt;/td&gt;
&lt;td&gt;Attackers create look-alike names or typosquatted services that trick the AI into picking a malicious resource instead of the legitimate one&lt;/td&gt;
&lt;td&gt;Enforce unique identities, cryptographically verify server/agent identities, and use a trusted registry rather than blind name matching&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Context Poisoning / Indirect Prompt Injection&lt;/td&gt;
&lt;td&gt;An attacker embeds hidden instructions in context that steer the model to do harmful actions or leak data; in A2A this can happen via malicious task states or malformed skill descriptions&lt;/td&gt;
&lt;td&gt;Sanitize and vet descriptions, use strict schema constraints, and limit or filter natural-language metadata that models see&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Shadowing Attacks&lt;/td&gt;
&lt;td&gt;A malicious service shadows a legitimate tool or agent by registering something that alters how other trusted components behave (e.g., injecting hidden guidance that changes billing logic or influences other agents&amp;rsquo; outputs)&lt;/td&gt;
&lt;td&gt;Require authentication and authorization for every component, use whitelists, and ensure that tools/agents don&amp;rsquo;t get used solely based on unverified contextual presence&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rug Pulls&lt;/td&gt;
&lt;td&gt;An attacker initially builds trust by providing a useful tool or agent, but once widely adopted, subtly changes its behavior to perform harmful operations, manipulate outputs, or exfiltrate data&lt;/td&gt;
&lt;td&gt;Continuous monitoring, evaluate behavioral changes over time, use policy controls and version gating so tools can&amp;rsquo;t suddenly change semantics without review&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;AI agents use protocols like MCP (Model Context Protocol) and A2A (Agent-to-Agent) and decide when and how to call tools or other agents. They do this dynamic discovery and usage based on natural language context which makes them susceptible to semantic manipulation.&lt;/p&gt;</description></item><item><title>Try vibe coding (again)</title><link>https://mitjamartini.com/en/posts/2026/01/try-vibe-coding-again/</link><pubDate>Tue, 13 Jan 2026 20:39:52 +0100</pubDate><guid>https://mitjamartini.com/en/posts/2026/01/try-vibe-coding-again/</guid><description>&lt;p&gt;In case you haven&amp;rsquo;t tried vibe coding, recently, you should probably try it (again). Their performance has increased quite a bit thanks to recent models like Opus 4.5. Here are some recent takes on vibe coding:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Everyone&amp;rsquo;s talking about vibe coding without looking at code. I was skeptical.
I decided to give it a shot on a challenging problem and was blown away by what
I could accomplish in 8 hours.
I&amp;rsquo;m not skeptical anymore, but I also do NOT think it kills SaaS.&lt;/p&gt;
&lt;p&gt;&amp;ndash; Hamel Husain on &lt;a
href="https://www.youtube.com/watch?v=97hYbaVCJ74"
target="_blank"
&gt;Youtube&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;I&amp;rsquo;m not joking and this isn&amp;rsquo;t funny. We have been trying to build distributed
agent orchestrators at Google since last year. There are various options, not
everyone is aligned&amp;hellip; I gave Claude Code a description of the problem, it
generated what we built last year in an hour.&lt;/p&gt;
&lt;p&gt;&amp;ndash; Jaana Dogan, Principal Engineer at Google on &lt;a
href="https://x.com/rakyll/status/2007239758158975130"
target="_blank"
&gt;X&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;I&amp;rsquo;ve never felt this much behind as a programmer. The profession is being
dramatically refactored as the bits contributed by the programmer are
increasingly sparse and between. I have a sense that I could be 10X more
powerful if I just properly string together what has become available over
the last ~year and a failure to claim the boost feels decidedly like skill
issue. There&amp;rsquo;s a new programmable layer of abstraction to master [&amp;hellip;] Clearly
some powerful alien tool was handed around except it comes with no manual and
everyone has to figure out how to hold it and operate it, while the resulting
magnitude 9 earthquake is rocking the profession.&lt;/p&gt;
&lt;p&gt;Roll up your sleeves to not fall behind.&lt;/p&gt;
&lt;p&gt;&amp;ndash; Andrej Karpathy on &lt;a
href="https://x.com/karpathy/status/2004607146781278521"
target="_blank"
&gt;X&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;</description></item><item><title>The Van Halen Test can check if an agent knows it's context</title><link>https://mitjamartini.com/en/posts/2026/01/van-halen-test-agent-context/</link><pubDate>Tue, 13 Jan 2026 19:43:45 +0100</pubDate><guid>https://mitjamartini.com/en/posts/2026/01/van-halen-test-agent-context/</guid><description>&lt;p&gt;Here&amp;rsquo;s a trick from the 80s that&amp;rsquo;s still useful, today: Van Halen&amp;rsquo;s live shows were potentially dangerous. To make sure the local crew read the rooster, they asked for a candy bowl with M&amp;amp;Ms and added:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Warning: Absolutely no brown ones.&lt;/strong&gt;&lt;/p&gt;
&lt;figure&gt;
&lt;img
class="my-0 rounded-md"
srcset="
/en/posts/2026/01/van-halen-test-agent-context/van-halen-test_hu_6e962dc59d7a873c.jpg 330w,
/en/posts/2026/01/van-halen-test-agent-context/van-halen-test_hu_c038fd99310b0364.jpg 660w,
/en/posts/2026/01/van-halen-test-agent-context/van-halen-test_hu_5f80de5536d19535.jpg 1024w,
/en/posts/2026/01/van-halen-test-agent-context/van-halen-test_hu_df5918a5d31f3f91.jpg 2x"
src="https://mitjamartini.com/en/posts/2026/01/van-halen-test-agent-context/van-halen-test_hu_c038fd99310b0364.jpg"
data-zoom-src="https://mitjamartini.com/en/posts/2026/01/van-halen-test-agent-context/van-halen-test_hu_df5918a5d31f3f91.jpg"
alt="The Van Halen Test"
/&gt;
&lt;figcaption&gt;The Van Halen Test via &lt;a
href="http://lefthandwrites.weebly.com/reveries/van-halen-brown-mms-and-english-assignments"
target="_blank"
&gt;Left hand writes&lt;/a&gt;&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;One day in 1982, they trashed their backstage area because David Lee Roth found a brown M&amp;amp;M.&lt;/p&gt;
&lt;p&gt;With coding agents, it&amp;rsquo;s also not always clear, if they still know all the context. Adding an easy-to-check piece of information helps finding out. I hope that won&amp;rsquo;t make you trash your office, though.&lt;/p&gt;</description></item><item><title>From 8 Lines with Dokku to 200 with Kubernetes – Why I'm Still Switching</title><link>https://mitjamartini.com/en/posts/2026/01/from-8-lines-with-dokku-to-200-with-k8s/</link><pubDate>Tue, 13 Jan 2026 18:21:46 +0200</pubDate><guid>https://mitjamartini.com/en/posts/2026/01/from-8-lines-with-dokku-to-200-with-k8s/</guid><description>&lt;p&gt;So far, my web apps run on a Dokku server. I haven&amp;rsquo;t tried Vercel or Fly because I didn&amp;rsquo;t want to deal with complex pricing models that incur more costs with every additional project.&lt;/p&gt;
&lt;p&gt;Dokku works like Heroku:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;dokku apps:create myapp
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;dokku postgres:create myapp-db
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;dokku postgres:link myapp-db myapp
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;dokku config:set myapp &lt;span class="nv"&gt;APP_SECRET&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;&lt;/span&gt;&lt;span class="k"&gt;$(&lt;/span&gt;openssl rand -base64 48&lt;span class="k"&gt;)&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;dokku git:from-image myapp myregistry.example.com/image:tag
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;dokku domains:set myapp myapp.mitjas.com
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;dokku ports:set myapp http:80:8000 https:443:8000
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;dokku letsencrypt:enable myapp
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;It doesn&amp;rsquo;t get any simpler. After just 8 lines I have&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;an app on the server,&lt;/li&gt;
&lt;li&gt;a PostgreSQL DB with a user and password,&lt;/li&gt;
&lt;li&gt;the app linked to Postgres,&lt;/li&gt;
&lt;li&gt;a secret configured,&lt;/li&gt;
&lt;li&gt;the app deployed from a Docker image,&lt;/li&gt;
&lt;li&gt;a reverse proxy configured to forward HTTP and HTTPS to the app&amp;rsquo;s container, and&lt;/li&gt;
&lt;li&gt;TLS with certificates signed by Let&amp;rsquo;s Encrypt (including automatic renewal).&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;With Kubernetes I need almost &lt;strong&gt;200 lines of YAML&lt;/strong&gt; for this. Why do I still want to switch to Kubernetes?&lt;/p&gt;
&lt;p&gt;Dokku is great for small projects that can make do with the Dokku plugins and run on a single server. But I believe Kubernetes is a better fit for me in the long run.&lt;/p&gt;
&lt;p&gt;Kubernetes manifests tell me (and LLMs) what&amp;rsquo;s running on the cluster, and I can continuously develop and improve them. Dokku, in contrast, feels like &amp;ldquo;fire, forget, and start from scratch.&amp;rdquo; Other advantages of Kubernetes like better scalability, more choice, and security are nice, too, but for me the most important advantage right now is:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Kubernetes is well-suited for Vibe Coding.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;That&amp;rsquo;s my assumption, anyway. Let&amp;rsquo;s see how it goes.&lt;/p&gt;</description></item><item><title>What Takes Time in Vibe Coding</title><link>https://mitjamartini.com/en/posts/2026/01/what-takes-time-in-vibe-coding/</link><pubDate>Tue, 13 Jan 2026 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/2026/01/what-takes-time-in-vibe-coding/</guid><description>&lt;p&gt;Vibe scripting, for me, is when I develop small tools for myself with the help of coding agents. It works extremely well, especially for command-line tools.&lt;/p&gt;
&lt;p&gt;I recently told a tax advisor about it. He was interested and wanted to see how it works. So I developed a small example on the spot: a VAT calculator. Not the best example, but I couldn&amp;rsquo;t think of anything better on short notice.&lt;/p&gt;
&lt;p&gt;It worked well, and after five minutes the VAT calculator with a Flet/Flutter GUI was up and running.&lt;/p&gt;
&lt;p&gt;But I could also see: there&amp;rsquo;s still room for improvement. For example, the layout wasn&amp;rsquo;t great and the functionality was too limited.&lt;/p&gt;
&lt;p&gt;As i wanted to know how long it would take to turn it into an actually useful app, I later developed it to become a VAT calculator for all EU countries, which uses a small AI pipeline to load rates and descriptions of which product categories are subject to which VAT rate from official EU pages and displays them in an improved GUI.&lt;/p&gt;
&lt;figure&gt;
&lt;img
class="my-0 rounded-md"
srcset="
/en/posts/2026/01/what-takes-time-in-vibe-coding/vat-calculator_hu_51fa9f847f16906c.png 330w,
/en/posts/2026/01/what-takes-time-in-vibe-coding/vat-calculator_hu_547addd196020eb9.png 660w,
/en/posts/2026/01/what-takes-time-in-vibe-coding/vat-calculator_hu_8750735acc10413d.png 1024w,
/en/posts/2026/01/what-takes-time-in-vibe-coding/vat-calculator_hu_290d4f807141e56e.png 2x"
src="https://mitjamartini.com/en/posts/2026/01/what-takes-time-in-vibe-coding/vat-calculator_hu_547addd196020eb9.png"
data-zoom-src="https://mitjamartini.com/en/posts/2026/01/what-takes-time-in-vibe-coding/vat-calculator_hu_290d4f807141e56e.png"
alt="EU VAT Calculator"
/&gt;
&lt;figcaption&gt;The EU VAT Calculator after 2h dev time&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;What started with five minutes for the simple version turned into about two hours. Of course, it&amp;rsquo;s much better now. But it&amp;rsquo;s interesting how big the effort difference is between a solution that serves a specific, narrowly defined purpose and a tool with a broader scope. There&amp;rsquo;s a lot of work involved, and finesse is needed to make a tool that&amp;rsquo;s truly useful. Personally, I need iterations with a human in the loop for that. Maybe there are developers who can perfectly specify everything upfront, but I usually need to see and use something to evaluate and improve it.&lt;/p&gt;
&lt;p&gt;AI accelerates iterations enormously, and you can decide whether to consider it &amp;ldquo;good enough&amp;rdquo; sooner or do a few more iterations. I think this is an important reason why I don&amp;rsquo;t develop faster with AI. I invest the time in more iterations and perhaps less thinking upfront, which then requires more iterations again.&lt;/p&gt;</description></item><item><title>MCP in ChatGPT Developer Mode Beta</title><link>https://mitjamartini.com/en/posts/mcp-in-chatgpt-developer-mode-beta/</link><pubDate>Fri, 17 Oct 2025 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/mcp-in-chatgpt-developer-mode-beta/</guid><description>&lt;p&gt;OpenAI just releasesd MCP connectors and ChatGPT developer mode beta. In this post, I describe the process of connecting MCP servers to ChatGPT, show how they look and feel right now in a chat session and give an overview of their current limitations.&lt;/p&gt;
&lt;!-- more --&gt;
&lt;h2 class="relative group"&gt;Activating developer mode
&lt;div id="activating-developer-mode" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#activating-developer-mode" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;p&gt;MCP connectors can only be created and edited in developer mode which can look a bit scary:&lt;/p&gt;
&lt;p&gt;
&lt;figure&gt;
&lt;img
class="my-0 rounded-md"
loading="lazy"
decoding="async"
fetchpriority="low"
alt="ChatGPT&amp;rsquo;s input in developer mode"
srcset="
/en/posts/mcp-in-chatgpt-developer-mode-beta/developer-mode-input_hu_e2dc8f0400ad0174.webp 330w,
/en/posts/mcp-in-chatgpt-developer-mode-beta/developer-mode-input_hu_a3c1a803a54605c4.webp 660w,
/en/posts/mcp-in-chatgpt-developer-mode-beta/developer-mode-input_hu_f8fe73140609504e.webp 1280w
"
data-zoom-src="https://mitjamartini.com/en/posts/mcp-in-chatgpt-developer-mode-beta/developer-mode-input.webp"
src="https://mitjamartini.com/en/posts/mcp-in-chatgpt-developer-mode-beta/developer-mode-input.webp"&gt;
&lt;/figure&gt;
&lt;/p&gt;
&lt;div
class="flex px-4 py-3 rounded-md bg-primary-100 dark:bg-primary-900"
&gt;
&lt;span
class="text-primary-400 ltr:pr-3 rtl:pl-3 flex items-center"
&gt;
&lt;span class="relative block icon"&gt;&lt;svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 512 512"&gt;&lt;path fill="currentColor" d="M506.3 417l-213.3-364c-16.33-28-57.54-28-73.98 0l-213.2 364C-10.59 444.9 9.849 480 42.74 480h426.6C502.1 480 522.6 445 506.3 417zM232 168c0-13.25 10.75-24 24-24S280 154.8 280 168v128c0 13.25-10.75 24-23.1 24S232 309.3 232 296V168zM256 416c-17.36 0-31.44-14.08-31.44-31.44c0-17.36 14.07-31.44 31.44-31.44s31.44 14.08 31.44 31.44C287.4 401.9 273.4 416 256 416z"/&gt;&lt;/svg&gt;
&lt;/span&gt;
&lt;/span&gt;
&lt;span
class="dark:text-neutral-300"
&gt;&lt;strong&gt;Note:&lt;/strong&gt; This article is written from the perspective of a Pro account user. OpenAI&amp;rsquo;s &lt;a
href="https://help.openai.com/de-de/articles/12584461-developer-mode-and-full-mcp-connectors-in-chatgpt-beta"
target="_blank"
&gt;developer mode and MCP connectors in ChatGPT documentation&lt;/a&gt; describes that admins can publish MCPs in their organization. I cannot test this but I assume that MCP connectors are then also usable in normal mode.&lt;/span&gt;
&lt;/div&gt;
&lt;p&gt;Developer mode can be activated in &lt;code&gt;Settings &amp;gt; Apps &amp;amp; Connectors &amp;gt; Advanced Settings&lt;/code&gt;.&lt;/p&gt;
&lt;h2 class="relative group"&gt;Creating an MCP connection
&lt;div id="creating-an-mcp-connection" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#creating-an-mcp-connection" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;p&gt;Once the developer mode is active, MCP server connections can be added in &lt;code&gt;Settings &amp;gt; Apps &amp;amp; Connectors&lt;/code&gt;:&lt;/p&gt;
&lt;p&gt;
&lt;figure&gt;
&lt;img
class="my-0 rounded-md"
loading="lazy"
decoding="async"
fetchpriority="low"
alt="ChatGPT Apps &amp;amp; Connector settings"
srcset="
/en/posts/mcp-in-chatgpt-developer-mode-beta/chatgpt-apps-and-connectors-settings_hu_5a1042c2b213bb56.webp 330w,
/en/posts/mcp-in-chatgpt-developer-mode-beta/chatgpt-apps-and-connectors-settings_hu_41ba0e8b332c5822.webp 660w,
/en/posts/mcp-in-chatgpt-developer-mode-beta/chatgpt-apps-and-connectors-settings_hu_968767f9b892510d.webp 1280w
"
data-zoom-src="https://mitjamartini.com/en/posts/mcp-in-chatgpt-developer-mode-beta/chatgpt-apps-and-connectors-settings.webp"
src="https://mitjamartini.com/en/posts/mcp-in-chatgpt-developer-mode-beta/chatgpt-apps-and-connectors-settings.webp"&gt;
&lt;/figure&gt;
&lt;/p&gt;
&lt;p&gt;Connectors take an icon, name, description, the MCP server url, and authentication information. Both SSE and streaming transports are supported and OAuth 2.0 can be used for authentication.&lt;/p&gt;
&lt;p&gt;
&lt;figure&gt;
&lt;img
class="my-0 rounded-md"
loading="lazy"
decoding="async"
fetchpriority="low"
alt="ChatGPT Create MCP Connector"
srcset="
/en/posts/mcp-in-chatgpt-developer-mode-beta/chatgpt-create-mcp-connector_hu_fd41aa920e51ecd2.webp 330w,
/en/posts/mcp-in-chatgpt-developer-mode-beta/chatgpt-create-mcp-connector_hu_c75172c4f980f518.webp 660w,
/en/posts/mcp-in-chatgpt-developer-mode-beta/chatgpt-create-mcp-connector_hu_215c530d93e9d155.webp 1280w
"
data-zoom-src="https://mitjamartini.com/en/posts/mcp-in-chatgpt-developer-mode-beta/chatgpt-create-mcp-connector.webp"
src="https://mitjamartini.com/en/posts/mcp-in-chatgpt-developer-mode-beta/chatgpt-create-mcp-connector.webp"&gt;
&lt;/figure&gt;
&lt;/p&gt;
&lt;h2 class="relative group"&gt;Using an MCP connection
&lt;div id="using-an-mcp-connection" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#using-an-mcp-connection" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;p&gt;MCP connections need to be activated for each chat session:&lt;/p&gt;
&lt;p&gt;
&lt;figure&gt;
&lt;img
class="my-0 rounded-md"
loading="lazy"
decoding="async"
fetchpriority="low"
alt="Activating an MCP in a chat"
srcset="
/en/posts/mcp-in-chatgpt-developer-mode-beta/activate-custom-mcps-in-a-chat_hu_939be1a88d553781.webp 330w,
/en/posts/mcp-in-chatgpt-developer-mode-beta/activate-custom-mcps-in-a-chat_hu_f3c95f47370b2bf5.webp 660w,
/en/posts/mcp-in-chatgpt-developer-mode-beta/activate-custom-mcps-in-a-chat_hu_eed1daf625e890c5.webp 1280w
"
data-zoom-src="https://mitjamartini.com/en/posts/mcp-in-chatgpt-developer-mode-beta/activate-custom-mcps-in-a-chat.webp"
src="https://mitjamartini.com/en/posts/mcp-in-chatgpt-developer-mode-beta/activate-custom-mcps-in-a-chat.webp"&gt;
&lt;/figure&gt;
&lt;/p&gt;
&lt;p&gt;In my testing, ChatGPT only uses an MCP connection when it has been instructed to do so. Actions need to be confirmed, but decisions can be stored for the duration of a chat:&lt;/p&gt;
&lt;p&gt;
&lt;figure&gt;
&lt;img
class="my-0 rounded-md"
loading="lazy"
decoding="async"
fetchpriority="low"
alt="Custom MCP in a chat session"
srcset="
/en/posts/mcp-in-chatgpt-developer-mode-beta/custom-mcp-in-a-chat-session_hu_aaa2d002974b56da.webp 330w,
/en/posts/mcp-in-chatgpt-developer-mode-beta/custom-mcp-in-a-chat-session_hu_f01ee0872e2d84f3.webp 660w,
/en/posts/mcp-in-chatgpt-developer-mode-beta/custom-mcp-in-a-chat-session_hu_3a18f1259a533a73.webp 1280w
"
data-zoom-src="https://mitjamartini.com/en/posts/mcp-in-chatgpt-developer-mode-beta/custom-mcp-in-a-chat-session.webp"
src="https://mitjamartini.com/en/posts/mcp-in-chatgpt-developer-mode-beta/custom-mcp-in-a-chat-session.webp"&gt;
&lt;/figure&gt;
&lt;/p&gt;
&lt;div
class="flex px-4 py-3 rounded-md bg-primary-100 dark:bg-primary-900"
&gt;
&lt;span
class="text-primary-400 ltr:pr-3 rtl:pl-3 flex items-center"
&gt;
&lt;span class="relative block icon"&gt;&lt;svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 512 512"&gt;&lt;path fill="currentColor" d="M506.3 417l-213.3-364c-16.33-28-57.54-28-73.98 0l-213.2 364C-10.59 444.9 9.849 480 42.74 480h426.6C502.1 480 522.6 445 506.3 417zM232 168c0-13.25 10.75-24 24-24S280 154.8 280 168v128c0 13.25-10.75 24-23.1 24S232 309.3 232 296V168zM256 416c-17.36 0-31.44-14.08-31.44-31.44c0-17.36 14.07-31.44 31.44-31.44s31.44 14.08 31.44 31.44C287.4 401.9 273.4 416 256 416z"/&gt;&lt;/svg&gt;
&lt;/span&gt;
&lt;/span&gt;
&lt;span
class="dark:text-neutral-300"
&gt;&lt;strong&gt;Note&lt;/strong&gt;: ChatGPT marks this action as a &lt;em&gt;write&lt;/em&gt; action, even though it&amp;rsquo;s really a read action from a user point of view. I don&amp;rsquo;t know if it&amp;rsquo;s a mistake on the side of the DeepWiki MCP, but the interesting part for me is, that Pro accounts seem to support write actions, already, even though this is still documented as a limitation.&lt;/span&gt;
&lt;/div&gt;
&lt;p&gt;MCP connections cannot be added to custom GPTs. Thus, there is no way to preconfigure the custom instructions needed to call an MCP server together with the MCP connection.&lt;/p&gt;
&lt;h2 class="relative group"&gt;Current Limitations
&lt;div id="current-limitations" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#current-limitations" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;p&gt;Currently, there are still quite a few limitations:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Not available on free accounts.&lt;/li&gt;
&lt;li&gt;Pro accounts can only use read/fetch actions (according to the documentation, this might not be true anymore)&lt;/li&gt;
&lt;li&gt;local MCPs are not possible.&lt;/li&gt;
&lt;li&gt;Agent mode does not support custom connectors.&lt;/li&gt;
&lt;li&gt;Deep research mode only supports read/fetch actions.&lt;/li&gt;
&lt;li&gt;Only available on the web, not in the mobile app. In the desktop app, connected MCPs are visible but cannot be used (in Pro accounts), as the desktop app does not support developer mode. I assume MCPs might be usable in Business and Enterprise/Edu accounts.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 class="relative group"&gt;Trying to make sense of it
&lt;div id="trying-to-make-sense-of-it" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#trying-to-make-sense-of-it" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;p&gt;Right now, I think MCP connectors in ChatGPT are not yet ready for day-to-day use - no wonder, it&amp;rsquo;s a beta, after all. Meanwhile, actions in custom GPTs are still usable, more open and not as scary to use.&lt;/p&gt;
&lt;p&gt;OpenAI could have added MCP connectors to custom GPTs and provided an option to pre-confirm actions for certain MCP servers but decided to go a different route with the developer mode.&lt;/p&gt;
&lt;p&gt;For me, MCP is a way to customize chatbots and give them &amp;ldquo;agency&amp;rdquo;. I find use-case-specific MCP servers better than generic ones which tend to bloat the context and distract the LLM.&lt;/p&gt;
&lt;p&gt;I hope, OpenAI will evolve ChatGPT&amp;rsquo;s MCP connectors without introducing an obligatory validation procedure, but I won&amp;rsquo;t bet on it, right now.&lt;/p&gt;</description></item><item><title>Claude Code in Devcontainers</title><link>https://mitjamartini.com/en/posts/claude-code-in-devcontainer/</link><pubDate>Tue, 14 Oct 2025 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/claude-code-in-devcontainer/</guid><description>&lt;p&gt;&lt;a
href="https://containers.dev"
target="_blank"
&gt;Development Containers&lt;/a&gt; or just &amp;ldquo;devcontainers&amp;rdquo; add a layer of security, simplify developer onboarding enables developing in parallel with isolated environments.&lt;/p&gt;
&lt;p&gt;For me, Devcontainers are a great addition to an AI Engineer&amp;rsquo;s toolbox, even though I don&amp;rsquo;t use them day-to-day.&lt;/p&gt;
&lt;p&gt;I have setup Devcontainers with Claude Code using VS Code and the devcontainers cli, adapted it to a Python FastAPI based project, and documented it in this article.&lt;/p&gt;</description></item><item><title>Tipps for Migrating to Hugo</title><link>https://mitjamartini.com/en/posts/tipps-for-migrating-to-hugo/</link><pubDate>Sat, 06 Sep 2025 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/tipps-for-migrating-to-hugo/</guid><description>&lt;p&gt;In this article, I document some non-technical things I&amp;rsquo;ve learned migrating my blog from Jekyll to &lt;a
href="https://gohugo.io"
target="_blank"
&gt;Hugo&lt;/a&gt;. The gist is to &lt;strong&gt;keep it as simple as possible&lt;/strong&gt;:&lt;/p&gt;</description></item><item><title>Evals for Voice Agents (Session Notes)</title><link>https://mitjamartini.com/en/posts/evals-for-voice-agents-session-notes/</link><pubDate>Mon, 30 Jun 2025 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/evals-for-voice-agents-session-notes/</guid><description>Notes of a whirlwind intro to evals for voice agents by Kwindla and swyx</description></item><item><title>Coding Agents as Slot Machines</title><link>https://mitjamartini.com/en/posts/coding-agents-as-slot-machines/</link><pubDate>Sun, 01 Jun 2025 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/coding-agents-as-slot-machines/</guid><description>&lt;p&gt;After reading a nice post &lt;a
href="https://news.ycombinator.com/item?id=44147966"
target="_blank"
&gt;via HN&lt;/a&gt; about &lt;a
href="https://rjp.io/blog/2025-05-31-stepping-back"
target="_blank"
&gt;On Stepping Back&lt;/a&gt;, I habitually scrolled through the comments and found this nugget by &lt;a
href="https://news.ycombinator.com/user?id=evrimoztamur"
target="_blank"
&gt;evrimoztamur&lt;/a&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Interacting with LLM coding tools is much like playing a slot machine, it grabs and chokeholds your gambling instincts. You&amp;rsquo;re rolling dice for the perfect result without much thought.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This probably applies to many systems that use Generative AI, especially systems that work on the cutting edge where randomness is more pronounced.&lt;/p&gt;
&lt;p&gt;If you also &amp;ldquo;&lt;a
href="https://www.youtube.com/watch?v=sq6a3WC5_Ns"
target="_blank"
&gt;Can&amp;rsquo;t sleep gud anymore&lt;/a&gt;&amp;rdquo; as Mario Zechner beautifully put it, be aware that coding with agents can be addictive.&lt;/p&gt;</description></item><item><title>Pipecat Cloud Latency for EU Users</title><link>https://mitjamartini.com/en/posts/pipecat-cloud-latency-for-eu-users/</link><pubDate>Sun, 11 May 2025 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/pipecat-cloud-latency-for-eu-users/</guid><description>Pipecat Cloud is located in the US. Is its latency ok for voice agents for EU Users.</description></item><item><title>K/V Cache Quantization in Ollama</title><link>https://mitjamartini.com/en/posts/ollama-kv-cache-quantization/</link><pubDate>Sat, 10 May 2025 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/ollama-kv-cache-quantization/</guid><description>&lt;p&gt;A somewhat hidden feature of Ollama is K/V Cache quantization. This is relevant for local AI as it reduces memory consumption, especially for small LLMs with large context windows.&lt;/p&gt;
&lt;p&gt;This post describes how to activate K/V Cache in Ollama and gives an overview of its benefits, drawbacks and use cases.&lt;/p&gt;
&lt;!-- more --&gt;
&lt;h2 class="relative group"&gt;Activating K/V cache quantization in Ollama
&lt;div id="activating-kv-cache-quantization-in-ollama" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#activating-kv-cache-quantization-in-ollama" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;p&gt;K/V Cache quantization in Ollama is not on by default, so you need to activate it by setting the &lt;code&gt;OLLAMA_KV_CACHE_TYPE&lt;/code&gt; environment variable. Supported values are documented in &lt;a
href="https://github.com/ollama/ollama/blob/main/docs/faq.md#how-can-i-set-the-quantization-type-for-the-kv-cache"
target="_blank"
&gt;How can I set the quantization type for the K/V cache?&lt;/a&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;f16&lt;/li&gt;
&lt;li&gt;q8_0&lt;/li&gt;
&lt;li&gt;q4_0&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 class="relative group"&gt;Benefits and Drawbacks
&lt;div id="benefits-and-drawbacks" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#benefits-and-drawbacks" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;p&gt;K/V Cache quantization can be the difference between being able to run a model on a machine, or not. You can use Sam McLeod&amp;rsquo;s &lt;a
href="https://smcleod.net/vram-estimator/"
target="_blank"
&gt;vram-estimator&lt;/a&gt; to estimate the memory consumption of models with different quantization settings.&lt;/p&gt;
&lt;p&gt;The benefit in brief is lower memory consumption which is most pronounced when using small models with large context windows.&lt;/p&gt;
&lt;p&gt;The main drawback is reduced model accuracy. The lmdeploy team has written a blog post about &lt;a
href="https://lmdeploy.readthedocs.io/en/v0.2.3/quantization/kv_int8.html#accuracy-test"
target="_blank"
&gt;K/V cache quantization accuracy test results&lt;/a&gt; if you want to see quantified impacts of K/V quantization on model accuracy.&lt;/p&gt;
&lt;h2 class="relative group"&gt;Example numbers
&lt;div id="example-numbers" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#example-numbers" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;p&gt;Llama 3.2 8B supports 128.000 tokens context windows.&lt;/p&gt;
&lt;p&gt;When you run it with Q4_K_M quantization and the longest possible context length, it consumes&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;23.3 GB memory without K/V cache qantization,&lt;/li&gt;
&lt;li&gt;17.0 GB with Q8_K_0 K/V cache quantization, and&lt;/li&gt;
&lt;li&gt;13.8 GB with Q4_K_0 K/V cache quantization.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;With Q4_K_0 K/V cache quantization it now fits into 16 GB vRAM.&lt;/p&gt;
&lt;h2 class="relative group"&gt;Use cases
&lt;div id="use-cases" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#use-cases" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Application&lt;/th&gt;
&lt;th&gt;Benefit&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;code generation&lt;/td&gt;
&lt;td&gt;more code in context&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;question answering&lt;/td&gt;
&lt;td&gt;whole docs fit into the context&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;function calling&lt;/td&gt;
&lt;td&gt;more tools&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;chat&lt;/td&gt;
&lt;td&gt;longer conversations&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;multi-modal&lt;/td&gt;
&lt;td&gt;images need many tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 class="relative group"&gt;More Info
&lt;div id="more-info" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#more-info" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;p&gt;To learn more, read Sam McLeod&amp;rsquo;s in-dept blog post about &lt;a
href="https://smcleod.net/2024/12/bringing-k/v-context-quantisation-to-ollama/"
target="_blank"
&gt;Bringing K/V Context Quantisation to Ollama&lt;/a&gt;. Sam helped to implement this in Ollama.&lt;/p&gt;</description></item><item><title>Deploying Voice Agents to Production</title><link>https://mitjamartini.com/en/posts/deploying-voice-agents-to-production/</link><pubDate>Fri, 09 May 2025 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/deploying-voice-agents-to-production/</guid><description>&lt;p&gt;Here are my notes on the session about deploying voice agents to production which is part of the &lt;a
href="https://maven.com/pipecat/voice-ai-and-voice-agents-a-technical-deep-dive"
target="_blank"
&gt;Voice Agents Course&lt;/a&gt;.&lt;/p&gt;
&lt;!-- more --&gt;
&lt;div
class="flex px-4 py-3 rounded-md bg-primary-100 dark:bg-primary-900"
&gt;
&lt;span
class="text-primary-400 ltr:pr-3 rtl:pl-3 flex items-center"
&gt;
&lt;span class="relative block icon"&gt;&lt;svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 384 512"&gt;&lt;path fill="currentColor" d="M112.1 454.3c0 6.297 1.816 12.44 5.284 17.69l17.14 25.69c5.25 7.875 17.17 14.28 26.64 14.28h61.67c9.438 0 21.36-6.401 26.61-14.28l17.08-25.68c2.938-4.438 5.348-12.37 5.348-17.7L272 415.1h-160L112.1 454.3zM191.4 .0132C89.44 .3257 16 82.97 16 175.1c0 44.38 16.44 84.84 43.56 115.8c16.53 18.84 42.34 58.23 52.22 91.45c.0313 .25 .0938 .5166 .125 .7823h160.2c.0313-.2656 .0938-.5166 .125-.7823c9.875-33.22 35.69-72.61 52.22-91.45C351.6 260.8 368 220.4 368 175.1C368 78.61 288.9-.2837 191.4 .0132zM192 96.01c-44.13 0-80 35.89-80 79.1C112 184.8 104.8 192 96 192S80 184.8 80 176c0-61.76 50.25-111.1 112-111.1c8.844 0 16 7.159 16 16S200.8 96.01 192 96.01z"/&gt;&lt;/svg&gt;
&lt;/span&gt;
&lt;/span&gt;
&lt;span
class="dark:text-neutral-300"
&gt;The course if held by kwindla und swyx, the CTO und an investor of Daily.co, a WebRTC und Voice AI infrastructure provider. Some of their recommendations might be predisposed. I still state them as is as I trust them and because I don&amp;rsquo;t have enough experience with voice agents in production.&lt;/span&gt;
&lt;/div&gt;
&lt;h2 class="relative group"&gt;TL;DR
&lt;div id="tldr" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#tldr" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Use a voice AI provider for simple, scalable deployment for production.&lt;/li&gt;
&lt;li&gt;Use a single VM or your homelab for demos.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 class="relative group"&gt;Differences between voice agents and traditional web apps
&lt;div id="differences-between-voice-agents-and-traditional-web-apps" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#differences-between-voice-agents-and-traditional-web-apps" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;are mostly in the transport&lt;/li&gt;
&lt;li&gt;persistent connnection (minutes)&lt;/li&gt;
&lt;li&gt;bidirectional streaming&lt;/li&gt;
&lt;li&gt;stateful sessions&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 class="relative group"&gt;Voice agents in production need
&lt;div id="voice-agents-in-production-need" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#voice-agents-in-production-need" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;A http service for
&lt;ul&gt;
&lt;li&gt;API endpoints,&lt;/li&gt;
&lt;li&gt;a website, and&lt;/li&gt;
&lt;li&gt;webhooks,&lt;/li&gt;
&lt;li&gt;spawning bots.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;A media transport layer or service:
&lt;ul&gt;
&lt;li&gt;WebRTC based for client-to-server (udp), or&lt;/li&gt;
&lt;li&gt;websocket based for server-to-server (tcp).&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Bots (udp or tcp, connect to media transport layer)&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 class="relative group"&gt;Bots
&lt;div id="bots" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#bots" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Are instances of the agents.&lt;/li&gt;
&lt;li&gt;Can be written in Python with PipeCat.&lt;/li&gt;
&lt;li&gt;Use STT, LLM, and TTS providers, which also are the main cost and latency drivers.&lt;/li&gt;
&lt;li&gt;Usually come packaged with small models, eg. for voice activity detection (VAD)&lt;/li&gt;
&lt;li&gt;Each spawned bot serves one session and needs allocated resources during the whole session:
&lt;ul&gt;
&lt;li&gt;0,5 vCPU&lt;/li&gt;
&lt;li&gt;1 GB RAM&lt;/li&gt;
&lt;li&gt;40kbps for WebRTC audio (in 30-60 kbps range)&lt;/li&gt;
&lt;li&gt;video requires more CPU (eg. 1 vCPU), and bandwidth&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Need to be quickly available. Target time-to-first-word:
&lt;ul&gt;
&lt;li&gt;2-3 secs (web),&lt;/li&gt;
&lt;li&gt;3-5 secs (phone)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 class="relative group"&gt;Ways to solve the &amp;ldquo;fast start challenge&amp;rdquo;
&lt;div id="ways-to-solve-the-fast-start-challenge" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#ways-to-solve-the-fast-start-challenge" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;percentage-based warm pool&lt;/li&gt;
&lt;li&gt;fast startup times (caching, pre-loading)&lt;/li&gt;
&lt;li&gt;proactive/predictive scheduling&lt;/li&gt;
&lt;li&gt;fallbacks from reactive world (eg. UX based solutions, not just silent fails)&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 class="relative group"&gt;Infra providers
&lt;div id="infra-providers" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#infra-providers" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;need to support tcp and udp&lt;/li&gt;
&lt;li&gt;voice ai providers are easiest (Pipecat Cloud, Daily, Vapi, Layercode)&lt;/li&gt;
&lt;li&gt;Fly.io (and potentially other container platforms) are good if they support udp (Fly does)&lt;/li&gt;
&lt;li&gt;ML focused provides are good for converged bots with larger models included (gpu clouds)&lt;/li&gt;
&lt;li&gt;hyperscalers are flexible but complex&lt;/li&gt;
&lt;li&gt;BTW: CloudRun does not support udp&lt;/li&gt;
&lt;li&gt;demos can run on single VMs or even be served from a home lab&lt;/li&gt;
&lt;li&gt;by serving everything converged, time-to-first word can get down to 500ms&lt;/li&gt;
&lt;li&gt;otherwise 800-1000 ms is good enough and achievable&lt;/li&gt;
&lt;li&gt;proximity to users matters (Daily plans global regions for PipeCat cloud, currently only us-west)&lt;/li&gt;
&lt;li&gt;conn between servers can be implemented with WebSockets&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 class="relative group"&gt;What&amp;rsquo;s next?
&lt;div id="whats-next" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#whats-next" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;p&gt;After the session I have looked at the PipeCat examples and realized it should be easy enough to run a basic voice agent with PipeCat&amp;rsquo;s &lt;a
href="https://docs.pipecat.ai/server/services/transport/small-webrt"
target="_blank"
&gt;SmallWebRTCTransport&lt;/a&gt; on a virtual server hosted in Europe and then switch the transport and deploy it to production on PipeCat Cloud.&lt;/p&gt;
&lt;p&gt;I will probably try that to see if the latency between US based PipeCat cloud and users in Europe is low enough for a good user experience.&lt;/p&gt;</description></item><item><title>An Overview of the Voice AI Landscape (Session Notes)</title><link>https://mitjamartini.com/en/posts/overview-of-voice-ai-landscape/</link><pubDate>Thu, 08 May 2025 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/overview-of-voice-ai-landscape/</guid><description>&lt;p&gt;I&amp;rsquo;m so happy to be part of the &lt;a
href="https://maven.com/pipecat/voice-ai-and-voice-agents-a-technical-deep-dive"
target="_blank"
&gt;Voice Agents Course&lt;/a&gt; by Kwindla and swyx. Yesterday, Kwindla kicked it off with an overview of the voice AI landscape. The pace, insights, and questions from the audience were just great.&lt;/p&gt;
&lt;p&gt;Here are my personal notes, probably incomplete and maybe not always correct. For a more authoritative overview of the Voice AI landscape, check out their free online book &lt;a
href="https://voiceaiandvoiceagents.com"
target="_blank"
&gt;Voice AI &amp;amp; Voice Agents - An Illustrated Primer&lt;/a&gt;.&lt;/p&gt;
&lt;!-- more --&gt;
&lt;ul&gt;
&lt;li&gt;Voice AI has highly valuable use cases with actual real business value.&lt;/li&gt;
&lt;li&gt;Benefits:
&lt;ul&gt;
&lt;li&gt;Today: lower cost.&lt;/li&gt;
&lt;li&gt;Soon: Better. (peak load response, better answers than most humans can give)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;RAG is still important&lt;/li&gt;
&lt;li&gt;Some challenges:
&lt;ul&gt;
&lt;li&gt;latency
&lt;ul&gt;
&lt;li&gt;measure regularly end-to-end from/to clients,&lt;/li&gt;
&lt;li&gt;record conversation with mic,&lt;/li&gt;
&lt;li&gt;visually look at gaps in waveforms,&lt;/li&gt;
&lt;li&gt;test calls from different regions and cell phone providers&lt;/li&gt;
&lt;li&gt;aim for 800ms, tough but possible to hit with hosted inference,&lt;/li&gt;
&lt;li&gt;very optimized/limited deployments can hit 500ms with quality compromises&lt;/li&gt;
&lt;li&gt;1000ms is not uncommon (still makes users happy)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;turn detection
&lt;ul&gt;
&lt;li&gt;from VAD to semantic&lt;/li&gt;
&lt;li&gt;OpenAI shipped a good text mode TDM&lt;/li&gt;
&lt;li&gt;Gemini flash in audio is ok, but needs to run as parallel flow (&amp;ldquo;greedily&amp;rdquo;)&lt;/li&gt;
&lt;li&gt;LifeKit vs. PipeCat is text vs. audio / end of speech, no audio cues vs. with audio cues&lt;/li&gt;
&lt;li&gt;both a hard ML challenge and important for experience&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;interruption handling&lt;/li&gt;
&lt;li&gt;context management&lt;/li&gt;
&lt;li&gt;function calling, tool use&lt;/li&gt;
&lt;li&gt;sripting, instruction following&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;standard architecture
&lt;ul&gt;
&lt;li&gt;today: 3 models (STT - LLM - STT), easier to achieve robust results with LLMs in text mode&lt;/li&gt;
&lt;li&gt;future probably converged&lt;/li&gt;
&lt;li&gt;3 models, because
&lt;ul&gt;
&lt;li&gt;LLMs&amp;rsquo; text mode is their mode&lt;/li&gt;
&lt;li&gt;we ride on the edge what the best models can do&lt;/li&gt;
&lt;li&gt;today, we need to use their best mode&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;other use cases, like language learning, can better leverage the benefits of speech-to-speech (1 model)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;models used (speed and quality, time-to-first token/byte more important than tokens/s)
&lt;ul&gt;
&lt;li&gt;STT:
&lt;ul&gt;
&lt;li&gt;Deepgram&lt;/li&gt;
&lt;li&gt;Whisper (optimized for streaming, original not built for streaming), down at 400&lt;/li&gt;
&lt;li&gt;Gladia (for non-english languages)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;LLM:
&lt;ul&gt;
&lt;li&gt;GPT-4o
&lt;ul&gt;
&lt;li&gt;still the workhorse, still more than 4.1&lt;/li&gt;
&lt;li&gt;big model changes take work and good evals (nobody has good ones),&lt;/li&gt;
&lt;li&gt;models usually not optimized for voice AI,&lt;/li&gt;
&lt;li&gt;not yet better results&lt;/li&gt;
&lt;li&gt;4o-mini was cheaper but slower and worse for tool/function&lt;/li&gt;
&lt;li&gt;note from a fellow student: 4.1-mini might be interesting&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Gemini 2.0 Flash
&lt;ul&gt;
&lt;li&gt;very good, fast, cost efficient&lt;/li&gt;
&lt;li&gt;best audio model for voice AI, today (gemini in audio input mode for voice-to-voice)&lt;/li&gt;
&lt;li&gt;multi-lingual input is ok, output not so much (use English)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;mid-sized open models get better, esp. fine-tuned Llama4 (upcoming OpenPipe finetuning session)&lt;/li&gt;
&lt;li&gt;additional notes about LLM use in voice AI:
&lt;ul&gt;
&lt;li&gt;reliable function calling/tool use is mostly a question of how to prompt 4o against Gemini&lt;/li&gt;
&lt;li&gt;llama can get there&lt;/li&gt;
&lt;li&gt;evals are important as ever&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;TTS:
&lt;ul&gt;
&lt;li&gt;PlayAI, Grok&lt;/li&gt;
&lt;li&gt;OpenAI, Google&lt;/li&gt;
&lt;li&gt;Cartesia&lt;/li&gt;
&lt;li&gt;Rime&lt;/li&gt;
&lt;li&gt;Elevenlabs&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;network transport:
&lt;ul&gt;
&lt;li&gt;why network? We don&amp;rsquo;t get capable enough models on mobile or laptop (yet),&lt;/li&gt;
&lt;li&gt;hybrid architectures might be relevant, already&lt;/li&gt;
&lt;li&gt;telephone is a great transport for voice AI, too
&lt;ul&gt;
&lt;li&gt;PSTN is with a phone number (eg. from Twilio)&lt;/li&gt;
&lt;li&gt;SIP is for interconnectivity with digital telephony infra&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;WebSockets are not good for real time client-to-server audio (and video), but ok for server-to-server audio&lt;/li&gt;
&lt;li&gt;WebRTC is best, complex, but PipeCat supports it ootb, local for testing, in the cloud offering for production,&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Voice AI building blocks in 2025
&lt;ul&gt;
&lt;li&gt;evals - huge topic&lt;/li&gt;
&lt;li&gt;hosting and scaling - very different (if you love k8s, your topic)&lt;/li&gt;
&lt;li&gt;workflow/multi-agent/state machines&lt;/li&gt;
&lt;li&gt;&amp;ldquo;perfect&amp;rdquo; speech
&lt;ul&gt;
&lt;li&gt;LLMs in text mode have passed the turing test&lt;/li&gt;
&lt;li&gt;not quite reached the point for real-time audio recognition and generation&lt;/li&gt;
&lt;li&gt;eg. accurately recording email, postal addess&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;human-like turn detection
&lt;ul&gt;
&lt;li&gt;very important for qualitative experience&lt;/li&gt;
&lt;li&gt;fun and hard ML problem&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;what&amp;rsquo;s next?
&lt;ul&gt;
&lt;li&gt;speech to speech models&lt;/li&gt;
&lt;li&gt;realtime video&lt;/li&gt;
&lt;li&gt;programming with voice&lt;/li&gt;
&lt;li&gt;voice as universal user experience&lt;/li&gt;
&lt;li&gt;LLM as a judge&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;orchestration and flows questions
&lt;ul&gt;
&lt;li&gt;voice in text out can be done with PipeCat aggregator (no great example, yet)&lt;/li&gt;
&lt;li&gt;accents and voice models: Gladia input, PlayAI output&lt;/li&gt;
&lt;li&gt;tip for Gemini: if you want accents, specify your country.&lt;/li&gt;
&lt;li&gt;how to solve cold-start?
&lt;ul&gt;
&lt;li&gt;Daily solved it for us&lt;/li&gt;
&lt;li&gt;own infra: Combine optimized startup and warm capacity&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;faster responses:
&lt;ul&gt;
&lt;li&gt;greedily inference/speculative before turn-detection&lt;/li&gt;
&lt;li&gt;helpful but more costly&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;multi-speaker:
&lt;ul&gt;
&lt;li&gt;hard, no great solution, so far.&lt;/li&gt;
&lt;li&gt;training data doesn&amp;rsquo;t map well to multi-person/agent conversation&lt;/li&gt;
&lt;li&gt;challenge: know when models should or shouldn&amp;rsquo;t respond (emit a no-response-token)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;observability tooling (multiple vendors, eg. Coval, will be in Discord and hold sessions)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Even after the first session, I can already say: If you are interested in Voice AI: &lt;strong&gt;Take the course!&lt;/strong&gt;.&lt;/p&gt;</description></item><item><title>The Fast Solopreneur</title><link>https://mitjamartini.com/en/posts/the-fast-solopreneur/</link><pubDate>Thu, 08 May 2025 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/the-fast-solopreneur/</guid><description>&lt;p&gt;In &lt;a
href="https://www.deeplearning.ai/the-batch/issue-300/"
target="_blank"
&gt;The Batch 300&lt;/a&gt;, Andrew Ng shared some insights about the importance of speed for startups and how to move fast as a startup. I think they apply well to ideas for a fast solopreneur:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Focus on one idea,&lt;/li&gt;
&lt;li&gt;Code prototypes intuitively,&lt;/li&gt;
&lt;li&gt;Be creative about getting user feedback quickly (it doesn&amp;rsquo;t have to scale),&lt;/li&gt;
&lt;li&gt;Pivot quickly but stay within your domain,&lt;/li&gt;
&lt;li&gt;Use a KISS stack you know, and&lt;/li&gt;
&lt;li&gt;Learn AI technology well.&lt;/li&gt;
&lt;/ul&gt;
&lt;!-- more --&gt;
&lt;p&gt;Here is a summary of Andrew&amp;rsquo;s original perspective:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Focus on a concrete idea (don&amp;rsquo;t get distracted)&lt;/li&gt;
&lt;li&gt;Switch quickly to another hypothesis when data shows that the original hypothesis is flawed&lt;/li&gt;
&lt;li&gt;Trust a domain expert&amp;rsquo;s gut instinct&lt;/li&gt;
&lt;li&gt;Build and test prototypes quickly with AI-assisted coding&lt;/li&gt;
&lt;li&gt;Be fast at getting user feedback (this becomes the bottleneck and the competitive advantage)&lt;/li&gt;
&lt;li&gt;Know the technology well (a deep understanding of what AI is good at and what it isn&amp;rsquo;t saves time and avoids dead ends)&lt;/li&gt;
&lt;/ul&gt;</description></item><item><title>Installing Docker on Raspberry PI</title><link>https://mitjamartini.com/en/posts/docker-on-raspberrypi/</link><pubDate>Mon, 05 May 2025 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/docker-on-raspberrypi/</guid><description>How to Install Docker and Docker Compose on Raspberry Pi OS (rootful and rootless).</description></item><item><title>A note on the hidden complexities of WebSockets</title><link>https://mitjamartini.com/en/posts/a-note-about-the-hidden-complexities-of-websockets/</link><pubDate>Sat, 25 Jan 2025 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/a-note-about-the-hidden-complexities-of-websockets/</guid><description>&lt;p&gt;AI Apps are often expected to be realtime. On the web, realtime communication can be implemented with WebSockets. I&amp;rsquo;ve started with WebSockets to create chatbots and other live-updated interfaces, but then switched to SSE and now mostly follow these rules of thumbs:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Use WebSockets for server-to-server communication.&lt;/li&gt;
&lt;li&gt;Use SSE for server-to-client communication.&lt;/li&gt;
&lt;li&gt;If you still want or need to use WebSockets for server-to-client communication, add a fallback to SSE.&lt;/li&gt;
&lt;li&gt;If you need realtime voice/video server-to-client communication, use WebRTC.&lt;/li&gt;
&lt;/ul&gt;
&lt;!-- more --&gt;
&lt;p&gt;If you still want or need to build a websocket service, please read Atul Jalan&amp;rsquo;s blog post about &lt;a
href="https://composehq.com/blog/scaling-websockets-1-23-25"
target="_blank"
&gt;The Hidden Complexity of Scaling WebSockets&lt;/a&gt;. It&amp;rsquo;s a quick read and captures important lessons to keep in mind when working with WebSockets.&lt;/p&gt;
&lt;p&gt;All of his lessons are important, even when working not at scale. Here is a quick summary:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Downtimeless deployments are much more involved than in HTTP services.&lt;/li&gt;
&lt;li&gt;Establish a good message schema, eg. 2 byte prefixes and single character field delimiters.&lt;/li&gt;
&lt;li&gt;Use heartbeats to detect dead connections - both ways.&lt;/li&gt;
&lt;li&gt;Have an http fallback as WebSockets are often blocked. Usually: Server sent events (SSE) for server to client communication and simple requests for client to server.&lt;/li&gt;
&lt;li&gt;More, like standard tooling (rate limiting, validation, error handling), no caching at the edge, per-message authentication.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Standard framworks like Django, FastHTML, and Quart (an async Flask clone) have good basic support for WebSockets, but don&amp;rsquo;t really help dealing with their hidden complexities. I hope, frameworks will level up a bit, as this is mostly &amp;ldquo;undifferentiated heavy lifting&amp;rdquo;.&lt;/p&gt;</description></item><item><title>RTX 5090 for Local AI</title><link>https://mitjamartini.com/en/posts/rtx-5090-for-local-ai/</link><pubDate>Tue, 26 Nov 2024 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/rtx-5090-for-local-ai/</guid><description>A look at the NVIDIA RTX 5090 specs for local LLM inference.</description></item><item><title>ChatGPT can use information from internal sites</title><link>https://mitjamartini.com/en/posts/chatgpt-can-use-information-from-internal-sites/</link><pubDate>Tue, 05 Nov 2024 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/chatgpt-can-use-information-from-internal-sites/</guid><description>The ChatGPT search feature can also search and use information from intranet websites.</description></item><item><title>AI as Repository of Human Knowledge</title><link>https://mitjamartini.com/en/posts/ai-as-repository-of-human-knowledge/</link><pubDate>Wed, 16 Oct 2024 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/ai-as-repository-of-human-knowledge/</guid><description>An interesting statement by Yann LeCun.</description></item><item><title>Deploy a Static Site on Dokku</title><link>https://mitjamartini.com/en/posts/deploy-static-site-on-dokku/</link><pubDate>Thu, 03 Oct 2024 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/deploy-static-site-on-dokku/</guid><description>&lt;p&gt;This article describes how to deploy a static site on Dokku, including activating Let&amp;rsquo;s encrypt signed certificates for the domain.&lt;/p&gt;
&lt;!-- more --&gt;
&lt;h2 class="relative group"&gt;Configure DNS
&lt;div id="configure-dns" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#configure-dns" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;p&gt;Register the domain and point it to the Dokku server. I usually set these records in the domain&amp;rsquo;s DNS zone:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;A @ &lt;span class="m"&gt;7200&lt;/span&gt; &lt;span class="nv"&gt;$DOKKU_SERVER_IPv4&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;A * &lt;span class="m"&gt;7200&lt;/span&gt; &lt;span class="nv"&gt;$DOKKU_SERVER_IPv4&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;CNAME www &lt;span class="m"&gt;7200&lt;/span&gt; &lt;span class="nv"&gt;$APP_DOMAIN&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;AAAA @ &lt;span class="m"&gt;7200&lt;/span&gt; &lt;span class="nv"&gt;$DOKKU_SERVER_IPv6&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;AAAA * &lt;span class="m"&gt;7200&lt;/span&gt; &lt;span class="nv"&gt;$DOKKU_SERVER_IPv6&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;h2 class="relative group"&gt;Create a git repository
&lt;div id="create-a-git-repository" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#create-a-git-repository" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;p&gt;Create a folder and turn it into a git repository:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;mkdir static-dokku
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nb"&gt;cd&lt;/span&gt; static-dokku
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;git init .
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Add an &lt;code&gt;index.html&lt;/code&gt; file with the desired content:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;hi&amp;#34;&lt;/span&gt; &amp;gt; index.html
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Create an empty &lt;code&gt;.static&lt;/code&gt; file. This tells Dokku that the repository contains a static site. Dokku then uses the Nginx builder to serve it.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;touch .static
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Commit the changes to git:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;git add .
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;git commit -m &lt;span class="s2"&gt;&amp;#34;initial commit&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;h2 class="relative group"&gt;Create an app on the Dokku host and git push to deploy the app
&lt;div id="create-an-app-on-the-dokku-host-and-git-push-to-deploy-the-app" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#create-an-app-on-the-dokku-host-and-git-push-to-deploy-the-app" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;p&gt;Use these commands to create, configure, and deploy the site:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nb"&gt;export&lt;/span&gt; &lt;span class="nv"&gt;DOKKU_HOST&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&amp;lt;your dokku server&amp;gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nb"&gt;export&lt;/span&gt; &lt;span class="nv"&gt;APP_DOMAIN&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&amp;lt;your app domain&amp;gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nb"&gt;export&lt;/span&gt; &lt;span class="nv"&gt;APP_NAME&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&amp;lt;your app name&amp;gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;ssh dokku@&lt;span class="nv"&gt;$DOKKU_HOST&lt;/span&gt; apps:create &lt;span class="nv"&gt;$APP_NAME&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;git remote add dokku dokku@&lt;span class="nv"&gt;$DOKKU_HOST&lt;/span&gt;:&lt;span class="nv"&gt;$APP_NAME&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;git push dokku main
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;ssh dokku@&lt;span class="nv"&gt;$DOKKU_HOST&lt;/span&gt; domains:add &lt;span class="nv"&gt;$APP_NAME&lt;/span&gt; &lt;span class="nv"&gt;$APP_DOMAIN&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;ssh dokku@&lt;span class="nv"&gt;$DOKKU_HOST&lt;/span&gt; letsencrypt:enable &lt;span class="nv"&gt;$APP_NAME&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;That&amp;rsquo;s it. The site is now live at &lt;a
href=""&gt;your app domain&lt;/a&gt;.&lt;/p&gt;
&lt;h2 class="relative group"&gt;Deploy updates
&lt;div id="deploy-updates" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#deploy-updates" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;p&gt;Updates can be deployed by checking them into git and running this command:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;git push dokku main
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;h2 class="relative group"&gt;Serving from sub directory
&lt;div id="serving-from-sub-directory" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#serving-from-sub-directory" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;p&gt;Serving files from a subdirectory (eg. &lt;code&gt;public&lt;/code&gt;) is possible by setting the &lt;code&gt;NGINX_ROOT&lt;/code&gt; environment variable of the app on the Dokku host:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;ssh dokku@&lt;span class="nv"&gt;$DOKKU_HOST&lt;/span&gt; config:set &lt;span class="nv"&gt;$APP_NAME&lt;/span&gt; &lt;span class="nv"&gt;NGINX_ROOT&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;/app/www/public
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;h2 class="relative group"&gt;Building the site during deployment
&lt;div id="building-the-site-during-deployment" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#building-the-site-during-deployment" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;p&gt;If you want to build the site during deployment, you can use multiple &lt;a
href="https://dokku.com/docs/deployment/builders/herokuish-buildpacks/"
target="_blank"
&gt;buildpacks&lt;/a&gt;, for example by adding a &lt;code&gt;.buildpacks&lt;/code&gt; file to the root directory of the repository.&lt;/p&gt;
&lt;p&gt;This way, you can use the first buildpack for building and the second fro serving.&lt;/p&gt;
&lt;p&gt;Another, even more flexible option is to use &lt;a
href="https://dokku.com/docs/deployment/builders/dockerfiles/"
target="_blank"
&gt;Dockerfile deployments&lt;/a&gt; with &lt;a
href="https://docs.docker.com/build/building/multi-stage/"
target="_blank"
&gt;multi-stage builds&lt;/a&gt;.&lt;/p&gt;</description></item><item><title>Dokku SSH Alias</title><link>https://mitjamartini.com/en/posts/dokku-ssh-alias/</link><pubDate>Thu, 03 Oct 2024 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/dokku-ssh-alias/</guid><description>A quick tip to make working a bit easier.</description></item><item><title>Launch VS Code from the Command Line</title><link>https://mitjamartini.com/en/posts/launch-vscode-from-the-cli/</link><pubDate>Sat, 28 Sep 2024 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/launch-vscode-from-the-cli/</guid><description>A quick tip to make your life a bit easier with VSCode and Python.</description></item><item><title>Deploying a SaaS Pegasus Based Django App on Dokku</title><link>https://mitjamartini.com/en/posts/deploying-a-django-app-on-dokku/</link><pubDate>Sun, 22 Sep 2024 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/deploying-a-django-app-on-dokku/</guid><description>In this post I&amp;rsquo;ll share the steps I did to deploy a SaaS Pegasus bootstrapped Django app (Scriv from the SaaS Pegasus marketplace to be precise) with Celery, Redis and Postgres on a Dokku host. There were some hickups along the way, but I believe, when you go step-by-step, it is quite straightforward.</description></item><item><title>The Furo Sphinx Theme Looks Good with Small Documentations</title><link>https://mitjamartini.com/en/posts/furo-sphinx-theme/</link><pubDate>Sat, 10 Aug 2024 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/furo-sphinx-theme/</guid><description>A quick note about the Furo theme for the Sphinx documentation system.</description></item><item><title>A Script to Export Models from Ollama</title><link>https://mitjamartini.com/en/posts/export-models-from-ollama/</link><pubDate>Tue, 28 May 2024 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/export-models-from-ollama/</guid><description>A workaround for transferring models to air-gapped Ollama instances.</description></item><item><title>Ollama on Windows</title><link>https://mitjamartini.com/en/posts/ollama-on-windows/</link><pubDate>Sun, 18 Feb 2024 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/ollama-on-windows/</guid><description>A tutorial and video about installing and using Ollama and OpenWebUI on Windows.</description></item><item><title>Summarizing Large Texts with LLMs and a Tree of Summaries</title><link>https://mitjamartini.com/en/posts/tree-of-summaries/</link><pubDate>Sun, 04 Feb 2024 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/tree-of-summaries/</guid><description>How to summarize texts with LLMs that are too large for their context window.</description></item><item><title>Hands-on ChatGPT in Excel (Book)</title><link>https://mitjamartini.com/en/posts/hands-on-chatgpt-in-excel/</link><pubDate>Mon, 18 Dec 2023 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/hands-on-chatgpt-in-excel/</guid><description>Example Excel workbooks for a book about using OpenAI models in Excel Formulas.</description></item><item><title>ChatGPT Image Inputs Use Cases</title><link>https://mitjamartini.com/en/posts/chatgpt-image-input-use-cases/</link><pubDate>Wed, 25 Oct 2023 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/chatgpt-image-input-use-cases/</guid><description>Practical use cases of ChatGPT&amp;rsquo;s / GPT-4&amp;rsquo;s image input feature.</description></item><item><title>Prompting Small LLMs.</title><link>https://mitjamartini.com/en/posts/prompting-small-llms/</link><pubDate>Tue, 10 Oct 2023 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/prompting-small-llms/</guid><description>What are small LLMs good at and how to get good results from small LLMs.</description></item><item><title>Fetching Data from REST APIs with Python and httpx</title><link>https://mitjamartini.com/en/posts/python-httpx-get/</link><pubDate>Sun, 15 Jan 2023 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/python-httpx-get/</guid><description>How to install and use httpx.</description></item><item><title>What is new in Excel in Q1 2022</title><link>https://mitjamartini.com/en/posts/new-in-excel-q1-2022/</link><pubDate>Sun, 24 Apr 2022 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/new-in-excel-q1-2022/</guid><description>&lt;p&gt;Microsoft keeps adding new exciting features to Excel all the time. Here is an overview of new features added in the first quarter of 2022.&lt;/p&gt;
&lt;p&gt;For me, &lt;code&gt;TEXTSPLIT&lt;/code&gt; is the best new Excel feature of this quarter. With it you can split text that is eg. formatted as CSV into multiple rows and columns. This is extremely helpful if you want to use ChatGPT in Excel.&lt;/p&gt;
&lt;!-- more --&gt;
&lt;p&gt;I base this on the Excel Blog. Please drop me a mail if I missed something. As Microsoft rolls out new features gradually, not everything might be available for you, right now.&lt;/p&gt;
&lt;h2 class="relative group"&gt;3 New Excel Functions for Text Extraction
&lt;div id="3-new-excel-functions-for-text-extraction" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#3-new-excel-functions-for-text-extraction" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;p&gt;Three new functions make it easier to extract text from a cell:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;TEXTBEFORE&lt;/strong&gt;:
Returns text that’s before delimiting characters&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;TEXTAFTER&lt;/strong&gt;:
Returns text that’s after delimiting characters&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;TEXTSPLIT&lt;/strong&gt;:
Splits text into rows or columns using delimiters&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 class="relative group"&gt;11 New Excel Functions for Dynamic Array Manipulation
&lt;div id="11-new-excel-functions-for-dynamic-array-manipulation" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#11-new-excel-functions-for-dynamic-array-manipulation" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;p&gt;Dynamic arrays have been the break trough of array formulas in Excel. Eleven new functions make working with dynamic arrays even more fun:&lt;/p&gt;
&lt;p&gt;Functions to &lt;strong&gt;combine arrays&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;VSTACK&lt;/strong&gt; - Stacks arrays vertically&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;HSTACK&lt;/strong&gt; - Stacks arrays horizontally&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Functions to &lt;strong&gt;shape arrays&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;WRAPROWS&lt;/strong&gt; - Wraps a row array into a 2D array&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;WRAPCOLS&lt;/strong&gt; - Wraps a column array into a 2D array&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;TOROW&lt;/strong&gt; - Returns an array as one row&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;TOCOL&lt;/strong&gt; - Returns an array as one column&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Functions to &lt;strong&gt;resize arrays&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;TAKE&lt;/strong&gt; - Returns rows or columns from an array start to end&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;DROP&lt;/strong&gt; - Drops rows or columns from an array start to end&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;CHOOSEROWS&lt;/strong&gt; - Returns the specified rows from an array&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;CHOOSECOLS&lt;/strong&gt; - Returns the specified columns from an array&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;EXPAND&lt;/strong&gt; - Expands an array to the specified dimensions&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 class="relative group"&gt;Improvements in Excel for the Web
&lt;div id="improvements-in-excel-for-the-web" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#improvements-in-excel-for-the-web" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Conditional Formating Pane&lt;/strong&gt; - Microsoft enhanced the user experience for creating and managing conditional formatting rules in Excel for the Web. Now you can use a pane to inspect and manage the conditional formatting rules.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Function Library&lt;/strong&gt; - Microsoft added the function library to the formula menu. Now it is more consistent with the desktop experience.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;New filter menu&lt;/strong&gt; - The filter menu has been enhanced to make filtering data more convenient.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Insert slicer&lt;/strong&gt; - Now you can insert slicers for PivotTables and PivotCharts in Excel for the Web, too.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Insert online pictures&lt;/strong&gt; - With this feature, you can add images from Microsoft Stock images and from Bing search results to your worksheets.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 class="relative group"&gt;Improved Name Manager in Excel for Mac
&lt;div id="improved-name-manager-in-excel-for-mac" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#improved-name-manager-in-excel-for-mac" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;p&gt;The Name Manager in Excel for Mac was not on par with Excel for Windows. Not anymore. With this update you get all the capabilities you might have been missing.&lt;/p&gt;
&lt;h2 class="relative group"&gt;LAMBDA and LAMBDA helper functions are Generally Available for Production
&lt;div id="lambda-and-lambda-helper-functions-are-generally-available-for-production" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#lambda-and-lambda-helper-functions-are-generally-available-for-production" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;p&gt;In February 2022, Microsoft announced general availability for &lt;strong&gt;LAMBDA&lt;/strong&gt; and the &lt;strong&gt;LAMBDA helper functions&lt;/strong&gt;. With LAMBDA you can turn Excel formulas into custom functions. This revolutionizes how you can build formulas in Excel. Along with the GA Microsoft added an &lt;strong&gt;Avanced Formula Environment&lt;/strong&gt;. This gives you a pane to edit more complicated Excel formulas in a proper editor. In addition, you can now use &lt;strong&gt;function tool tips&lt;/strong&gt; and &lt;strong&gt;auto-completion&lt;/strong&gt; with self-defined functions. In practice, your LAMBDA functions now look and feel like an native Excel function. Microsoft also increased the &lt;strong&gt;recursion limit&lt;/strong&gt; by 16x.&lt;/p&gt;
&lt;h2 class="relative group"&gt;Charting on the Web Feature Updates
&lt;div id="charting-on-the-web-feature-updates" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#charting-on-the-web-feature-updates" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;p&gt;Charting is a powerful tool in Excel to explore, analyze and present data. Charting in Excel for the Web was more like a viewer with only few editing capabilities. This has now changed. Microsoft has added a &lt;strong&gt;Chart Format Pane&lt;/strong&gt; and &lt;strong&gt;On-Chart Interactions&lt;/strong&gt; to Excel for the Web. You an also now &lt;strong&gt;add and remove chart elements&lt;/strong&gt; and &lt;strong&gt;create charts from non-adjacent cell ranges&lt;/strong&gt;.&lt;/p&gt;
&lt;h2 class="relative group"&gt;AutoComplete for Dropdown Lists in Excel for Windows
&lt;div id="autocomplete-for-dropdown-lists-in-excel-for-windows" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#autocomplete-for-dropdown-lists-in-excel-for-windows" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;p&gt;AutoComplete for dropdown lists matches the string you type in a dropdown list cell with words from items in the dropdown list. It then shows only the matching items. As you type more characters, the list gets smaller and smaller. It matches words anywhere in the list item string.&lt;/p&gt;
&lt;p&gt;This enhancement greatly speeds up data entry in fields with large lists. It also enables new use cases for this kind of data entry.&lt;/p&gt;
&lt;p&gt;At time of writing (April 2022) it is only available in the Beta Channel. As an upcoming enhancement before releasing for production, Microsoft plans to add excluding duplicates from the drop-down.&lt;/p&gt;
&lt;h2 class="relative group"&gt;Excel 4.0 (XLM) Macros are now Restricted by Default
&lt;div id="excel-40-xlm-macros-are-now-restricted-by-default" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#excel-40-xlm-macros-are-now-restricted-by-default" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;p&gt;You can now &lt;strong&gt;enable or disable Excel 4.0 (XLM) macros in the Excel Trust Center&lt;/strong&gt; for enhanced security.&lt;/p&gt;</description></item></channel></rss>