<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Ollama on Mitja Martini</title><link>https://mitjamartini.com/en/tags/ollama/</link><description>Recent content in Ollama on Mitja Martini</description><generator>Hugo -- gohugo.io</generator><language>en</language><copyright>© 2026 Mitja Martini</copyright><lastBuildDate>Sat, 10 May 2025 00:00:00 +0000</lastBuildDate><atom:link href="https://mitjamartini.com/en/tags/ollama/index.xml" rel="self" type="application/rss+xml"/><item><title>K/V Cache Quantization in Ollama</title><link>https://mitjamartini.com/en/posts/ollama-kv-cache-quantization/</link><pubDate>Sat, 10 May 2025 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/ollama-kv-cache-quantization/</guid><description>&lt;p&gt;A somewhat hidden feature of Ollama is K/V Cache quantization. This is relevant for local AI as it reduces memory consumption, especially for small LLMs with large context windows.&lt;/p&gt;
&lt;p&gt;This post describes how to activate K/V Cache in Ollama and gives an overview of its benefits, drawbacks and use cases.&lt;/p&gt;
&lt;!-- more --&gt;
&lt;h2 class="relative group"&gt;Activating K/V cache quantization in Ollama
&lt;div id="activating-kv-cache-quantization-in-ollama" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#activating-kv-cache-quantization-in-ollama" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;p&gt;K/V Cache quantization in Ollama is not on by default, so you need to activate it by setting the &lt;code&gt;OLLAMA_KV_CACHE_TYPE&lt;/code&gt; environment variable. Supported values are documented in &lt;a
href="https://github.com/ollama/ollama/blob/main/docs/faq.md#how-can-i-set-the-quantization-type-for-the-kv-cache"
target="_blank"
&gt;How can I set the quantization type for the K/V cache?&lt;/a&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;f16&lt;/li&gt;
&lt;li&gt;q8_0&lt;/li&gt;
&lt;li&gt;q4_0&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 class="relative group"&gt;Benefits and Drawbacks
&lt;div id="benefits-and-drawbacks" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#benefits-and-drawbacks" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;p&gt;K/V Cache quantization can be the difference between being able to run a model on a machine, or not. You can use Sam McLeod&amp;rsquo;s &lt;a
href="https://smcleod.net/vram-estimator/"
target="_blank"
&gt;vram-estimator&lt;/a&gt; to estimate the memory consumption of models with different quantization settings.&lt;/p&gt;
&lt;p&gt;The benefit in brief is lower memory consumption which is most pronounced when using small models with large context windows.&lt;/p&gt;
&lt;p&gt;The main drawback is reduced model accuracy. The lmdeploy team has written a blog post about &lt;a
href="https://lmdeploy.readthedocs.io/en/v0.2.3/quantization/kv_int8.html#accuracy-test"
target="_blank"
&gt;K/V cache quantization accuracy test results&lt;/a&gt; if you want to see quantified impacts of K/V quantization on model accuracy.&lt;/p&gt;
&lt;h2 class="relative group"&gt;Example numbers
&lt;div id="example-numbers" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#example-numbers" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;p&gt;Llama 3.2 8B supports 128.000 tokens context windows.&lt;/p&gt;
&lt;p&gt;When you run it with Q4_K_M quantization and the longest possible context length, it consumes&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;23.3 GB memory without K/V cache qantization,&lt;/li&gt;
&lt;li&gt;17.0 GB with Q8_K_0 K/V cache quantization, and&lt;/li&gt;
&lt;li&gt;13.8 GB with Q4_K_0 K/V cache quantization.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;With Q4_K_0 K/V cache quantization it now fits into 16 GB vRAM.&lt;/p&gt;
&lt;h2 class="relative group"&gt;Use cases
&lt;div id="use-cases" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#use-cases" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Application&lt;/th&gt;
&lt;th&gt;Benefit&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;code generation&lt;/td&gt;
&lt;td&gt;more code in context&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;question answering&lt;/td&gt;
&lt;td&gt;whole docs fit into the context&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;function calling&lt;/td&gt;
&lt;td&gt;more tools&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;chat&lt;/td&gt;
&lt;td&gt;longer conversations&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;multi-modal&lt;/td&gt;
&lt;td&gt;images need many tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 class="relative group"&gt;More Info
&lt;div id="more-info" class="anchor"&gt;&lt;/div&gt;
&lt;span
class="absolute top-0 w-6 transition-opacity opacity-0 ltr:-left-6 rtl:-right-6 not-prose group-hover:opacity-100 select-none"&gt;
&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700 !no-underline" href="#more-info" aria-label="Anchor"&gt;#&lt;/a&gt;
&lt;/span&gt;
&lt;/h2&gt;
&lt;p&gt;To learn more, read Sam McLeod&amp;rsquo;s in-dept blog post about &lt;a
href="https://smcleod.net/2024/12/bringing-k/v-context-quantisation-to-ollama/"
target="_blank"
&gt;Bringing K/V Context Quantisation to Ollama&lt;/a&gt;. Sam helped to implement this in Ollama.&lt;/p&gt;</description></item><item><title>A Script to Export Models from Ollama</title><link>https://mitjamartini.com/en/posts/export-models-from-ollama/</link><pubDate>Tue, 28 May 2024 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/export-models-from-ollama/</guid><description>A workaround for transferring models to air-gapped Ollama instances.</description></item><item><title>Ollama on Windows</title><link>https://mitjamartini.com/en/posts/ollama-on-windows/</link><pubDate>Sun, 18 Feb 2024 00:00:00 +0000</pubDate><guid>https://mitjamartini.com/en/posts/ollama-on-windows/</guid><description>A tutorial and video about installing and using Ollama and OpenWebUI on Windows.</description></item></channel></rss>