Ollama reserves RAM for far more context than you actually use. If a model is loaded with a 32K context window while your prompts and responses rarely cross a few thousand tokens, most of that capacity sits unused. And that unused capacity still has a memory cost. Ollama needs to allocate a KV cache for the configured context length, and the cache can add several GBs on top of the downloaded model. It's also not just an Ollama problem. The same fundamental problem applies to llama.cpp and GGUF models running through LM Studio’s llama.cpp engine. On a machine with less RAM, you can just lower the context window to something closer to your actual usage and get a lot more memory available for everything else.
Ollama reserves memory for context you may never use
And that slows everything down
The context length is a maximum capacity rather than a measurement of what your current prompt is using. It includes the system prompt, conversation history, current input, and generated response. As I mentioned above, you may configure a model for 32K tokens and use only 3K during a normal conversation, but Ollama still prepares the model to handle the larger limit when it loads.
Most of this extra memory goes into the KV cache. The cache stores calculations from tokens the model has already processed so it doesn’t have to process the entire conversation again. A larger context window needs space for more of these calculations. In Ollama and other llama.cpp-based runtimes, the size of that allocation is tied to the configured capacity rather than the length of the short prompt you happen to send afterward.
The downloaded model size doesn’t account for this clearly. For example, a 5 GB Q4 model tells you how much space its quantized weights occupy, but Ollama still needs compute buffers and the KV cache. You can therefore reduce the size of the model weights and still lose several gigabytes to an oversized context setting.
On Apple Silicon, that allocation comes from the same unified-memory pool used by macOS and every other application. On a system with a dedicated GPU, it normally occupies VRAM and can spill into system or shared memory when VRAM runs out. The unused context is then taking memory from other applications or pushing part of the model onto slower memory without contributing anything to the requests you normally make. Ollama’s defaults can make this worse on systems with more VRAM. It currently assigns 4K context below 24 GiB, 32K between 24 and 48 GiB, and 256K at 48 GiB or above.
Choose a context limit from your actual requests
You don't need as much as you think
Ollama’s API gives you useful numbers for choosing a context window. Its responses include prompt_eval_count, which reports input tokens, and eval_count, which reports generated tokens. You can check these across conversations to get a better starting point than the model’s advertised maximum. Include your longer exchanges in that comparison so you aren’t sizing everything around a single short question.
For the few-thousand-token conversations I’m describing here, I’d start by testing an 8K window. That leaves room for follow-up questions and an unusually long response without immediately jumping to 32K. You'll need to treat reasoning models differently because they generate thinking tokens before the final answer. Ollama exposes that thinking separately, so the answer you see in the chat interface doesn’t necessarily represent everything the model generates. You need to budget for that output too.
You can change the context slider in Ollama’s app or enter /set parameter num_ctx 8192 during a terminal session. For a persistent model-specific setting, you can use a Modelfile, which supports PARAMETER num_ctx 8192. Test the smaller window with a conversation that needs information from earlier messages. When truncation is enabled, Ollama removes older messages to meet the limit, preserving system messages and the latest message.
Cache precision and parallel requests need attention too
These can also help reduce memory use
Ollama gives you another way to reduce memory use — storing the conversation cache with slightly less numerical precision. This is called cache quantization. The 8-bit option, commonly labelled Q8, uses roughly half the memory of the default cache. You keep the same conversation limit, although the numbers stored in the cache are slightly less exact. That saving applies only to the cache, so it won’t cut the model’s entire memory requirement in half.
Q8 is the option I’d try first because its effect on response quality is usually negligible. The 4-bit option, labelled Q4, saves more memory, bringing the cache down to approximately a quarter of its original size. But as you'd guess, it also loses more precision, which can affect answers, especially in longer conversations. I’d compare answers to a few familiar questions before keeping either setting enabled.
You should also check how many requests Ollama is configured to handle simultaneously. If you are allowing four conversations to generate answers at once, you are wasting your cache space. You can have this enabled if several people or apps share your local server. For my use, sending a question and waiting for the answer, there’s little reason to reserve that extra capacity. Ollama currently defaults to one simultaneous request, so you only need to revisit this if you’ve increased it.
Give other inferences a try
While Ollama does a very good job of running your models, you might want to explore other options. For me, LM Studio has turned out to be a lot better than Ollama. I've also had good experiences with Docker Model Runner and BaseRT.



















