Back to blog
12 min read

Best Ollama Models for CPU Servers by RAM Size (Gemma 4, Qwen 3.5)

The best Ollama model for CPU servers depends on RAM. Picks for 2 to 32 GB, Gemma 4 and Qwen 3.5 tags, Q4_K_M vs Q8_0, and measured Qwen tokens per second.

TL;DR:

  • The best Ollama model for CPU servers depends on RAM: budget the download size plus 1 to 4 GB, then the OS.
  • Picks by RAM: 4 GB qwen3.5:2b-q4_K_M, 8 GB qwen3.5:4b, 16 GB qwen3.5:9b or gemma4:12b-it-q4_K_M, 32 GB qwen3.6:35b-a3b-q4_K_M.
  • Mixture-of-experts (MoE) models are the fast large models on CPU: Qwen3.6-35B-A3B generated 14.24 tokens/s after an 8K prompt on 16 vCPU, against 2.41 tokens/s for dense Qwen3.5-27B.
  • Use Q4_K_M for models of 4B parameters and up. qwen3.5:9b is a 6.6 GB download at Q4_K_M and 10 GB at Q8_0.
  • Ollama defaults to a 4,096-token context on CPU. Its docs recommend at least 64,000 tokens for agents and coding tools.

Applies to: Ubuntu 24.04 LTS and 26.04 LTS · Ollama 0.35.1 · Ollama library tags checked October 2026

Ollama is an open-source runtime, released under the MIT license, that downloads open-weight language models and serves them through a command line and a local API on port 11434. It runs without a GPU: on a CPU-only server it loads the whole model into system RAM. This guide ranks current Ollama library models by server RAM, with picks for chatbots, extraction, coding, agents with tool calling, and embeddings.

Can Ollama run on a CPU?

Yes, Ollama runs on CPU-only servers. When a model sits entirely in system memory, ollama ps shows 100% CPU in the PROCESSOR column. RAM decides which models load. vCPU count and memory bandwidth decide how fast they answer.

A request has two phases: prompt processing, then generation, measured in tokens per second. On a CPU, the first phase dominates long prompts. In the CPU benchmark of self-hosted LLMs, Qwen3.5-4B generated 12.23 tokens/s after an 8K-token prompt, and the full request still took 43.07 seconds.

Speeds and memory figures in this guide come from that benchmark: an Arct Cloud High Performance hvm.xlarge plan with 16 vCPU (the AMD Ryzen 9 option) and 32 GB RAM, running Ollama 0.32.15 with one request at a time in August 2026. Plans with fewer vCPUs generate more slowly.

How much RAM do Ollama models need?

An Ollama model needs its download size in RAM plus memory for the context cache and the runtime. Per Ollama's FAQ, required RAM scales with OLLAMA_NUM_PARALLEL times OLLAMA_CONTEXT_LENGTH, and a loaded model stays in memory for 5 minutes after the last request.

Peak memory measured at an 8K context:

ModelDownloadPeak memory, 8K context
Qwen3.5-4B3.4 GB4.18 GiB
Qwen3.5-9B6.6 GB7.28 GiB
Gemma 3 12B8.1 GB9.71 GiB
Qwen3.5-27B17 GB18.28 GiB
Qwen3.6-35B-A3B24 GB22.90 GiB

One GiB is 1.07 GB, so these models used 0.6 to 2.6 GB more than their download size. Plan for 1 to 4 GB of overhead plus the OS. Swap lets an oversized model load, but generation becomes too slow to use.

Best Ollama models by RAM size

The plan column lists the plans in each tier with that much RAM; Cost Optimized stops at 16 GB. Speeds are short-prompt generation rates from the 16-vCPU benchmark plan, and the 2B model was not part of that benchmark.

RAMPlans (vCPU)PickDownloadSpeed on 16 vCPU
2 GBcvm.nano, vm.nano, hvm.nano (1)embeddinggemma (embeddings only)622 MBNot applicable
4 GBcvm.micro, vm.micro, hvm.micro (2)qwen3.5:2b-q4_K_M1.9 GBNot measured
8 GBcvm.small, vm.tiny, hvm.tiny (4)qwen3.5:4b3.4 GB13.43 tokens/s
16 GBcvm.xlarge, vm.medium, hvm.medium (8)qwen3.5:9b6.6 GB7.78 tokens/s
32 GBvm.xlarge, hvm.xlarge (16)qwen3.6:35b-a3b-q4_K_M24 GB15.67 tokens/s
  • 2 GB: a chat model leaves little room here. In the benchmark, Qwen3-0.6B, a 523 MB download, peaked at 2.03 GiB at 8K. Run an embedding model such as embeddinggemma. To test qwen3.5:0.8b (1.0 GB), keep the default 4,096-token context and check free memory with free -h after it loads.
  • 4 GB: the default qwen3.5:2b tag is 2.7 GB. The q4_K_M tag saves 0.8 GB for the OS and the context cache.
  • 8 GB: qwen3.5:4b peaked at 4.18 GiB at 8K and 5.01 GiB at 32K. gemma4:e2b-it-q4_K_M (4.6 GB) is the Gemma 4 option at this size.
  • 16 GB: qwen3.5:9b peaked at 8.09 GiB on a 32K prompt, which took 254 seconds. gemma4:12b-it-q4_K_M (8.0 GB) is the Gemma 4 option at this size.
  • 32 GB: qwen3.6:35b-a3b-q4_K_M activates about 3B of its 35B parameters per token, so it generates faster than dense 9B models.

These picks assume Ollama's default 4,096-token context. A longer context adds cache memory, so check the loaded size with ollama ps after you raise it.

Best Ollama model for CPU by task

Chatbots and general use

Use qwen3.5:4b on 8 GB and qwen3.5:9b or gemma4:12b-it-q4_K_M on 16 GB. Qwen3.5 lists 201 languages and dialects. Reasoning tokens add generation time, so turn thinking off for short replies with --think=false.

Extraction and classification

Every Qwen3.5 package in the benchmark passed both output checks: a strict JSON extraction and a Python task with five assertions. Ollama's structured outputs enforce a JSON schema through the format field. For labels and routing, tev1:0.8b (812 MB), an experimental decision model fine-tuned from Qwen3.5, returns a choice with probabilities through /v1/systemone on Ollama 0.35 or later.

Coding

On 32 GB, qwen3.6:35b-a3b-q4_K_M generated 14.24 tokens/s after an 8K prompt, the fastest large model in the benchmark at that length. The library also lists a qwen3.6:35b-a3b-coding tag (23 to 24 GB), which the benchmark did not test, and qwen3-coder:30b (19 GB, MoE) is the coding model from the Qwen3 generation. On 16 GB, use qwen3.5:9b.

Ollama recommends at least 64,000 tokens of context for coding agents, and a 32K prompt took 183 to 254 seconds on the 4B and 9B models. Run agentic coding on an API model and keep the local model for short questions, as in OpenCode with API or Ollama models.

Agents with tool calling

Gemma 4 and Qwen3.5 both support tool calling in Ollama. The 64,000-token context sets the RAM floor: qwen3.5:9b peaked at 8.09 GiB at 32K and 11.24 GiB on a 128K attempt, so start with it on 16 GB and time prompts with your own tool definitions. Gemma 4 memory at 64K was not measured. The n8n AI agent tutorial connects either an API model or Ollama.

Embeddings

Embedding models are the smallest downloads in the library. Ollama's docs recommend embeddinggemma (622 MB, 2K context) and qwen3-embedding, whose 0.6b tag is 639 MB with a 32K context. nomic-embed-text is 274 MB. Ollama keeps up to 3 models loaded on CPU when they fit, so a RAG setup can hold a chat model and an embedding model in RAM.

Gemma 4 on Ollama without a GPU

Gemma 4 is Google DeepMind's open model family with text and image input, configurable thinking and native function calling. The CPU benchmark did not include Gemma 4, so the plan RAM column applies the download-plus-overhead rule above.

TagParametersDownload (Q4_K_M)ContextPlan RAM
gemma4:e2b-it-q4_K_M2.3B effective4.6 GB128K8 GB
gemma4:e4b-it-q4_K_M4.5B effective6.6 GB128K12 GB
gemma4:12b-it-q4_K_M12B8.0 GB256K16 GB
gemma4:26b-a4b-it-q4_K_M25.2B, 3.8B active18 GB256K24 GB
gemma4:31b-it-q4_K_M30.7B dense20 GB256K32 GB

The previous generation, Gemma 3 12B (8.1 GB), generated 5.48 tokens/s with a short prompt and peaked at 9.71 GiB at 8K. Dense 27B models generated about 2.4 tokens/s on the same plan, and the dense 31B build is a 20 GB file against 17 GB for Qwen3.5-27B.

Default tags show a size range (6.6 to 9.5 GB for gemma4:e4b), and the upper figure matches the MLX build. On a Linux server, pull the explicit -it-q4_K_M tag or the smaller -it-qat tag (gemma4:12b-it-qat is 7.2 GB).

Q4_K_M or Q8_0 on a CPU server

Quantization stores model weights at lower precision: Q4_K_M is a 4-bit format, Q8_0 an 8-bit format. CPU generation reads the weights from RAM for every token, so a larger file generates more slowly.

ModelQ4_K_MQ8_0
Qwen3.5-2B1.9 GB2.7 GB
Qwen3.5-9B6.6 GB10 GB
Gemma 4 12B8.0 GB13 GB
  • Use Q4_K_M for models of 4B parameters and up, and for the Gemma 4 E2B and E4B builds.
  • The default qwen3.5:0.8b and qwen3.5:2b tags match their q8_0 sizes. On 4 GB RAM, the 2B q4_K_M tag saves 0.8 GB.
  • Skip mlx, nvfp4 and mxfp8 tags, which the tags page labels MLX, and pull the -q4_K_M or -qat builds on a CPU server. Skip bf16 (unquantized) and mtp tags (extra tensors for speculative decoding).
  • For long contexts, OLLAMA_KV_CACHE_TYPE=q8_0 cuts the context cache to about half when Flash Attention is active.

Test a model on your server

Install Ollama with the Ollama on Ubuntu 24.04 tutorial, then time a candidate on a prompt from your own workload:

ollama pull qwen3.5:4b
ollama run qwen3.5:4b --verbose --think=false "Summarize in two sentences: YOUR_TEXT"
ollama ps

Replace YOUR_TEXT with your own text. --verbose prints prompt eval rate and eval rate in tokens/s after the answer. ollama ps shows the loaded size, 100% CPU and the context length.

To change the context length, open the systemd override:

sudo systemctl edit ollama.service

In the editor, add these lines below the line ### Anything between here and the comment below will become the contents of the drop-in file and above ### Edits below this comment will be discarded. If a [Service] section is already there from the install tutorial, change its OLLAMA_CONTEXT_LENGTH line instead. Save and exit:

[Service]
Environment="OLLAMA_CONTEXT_LENGTH=16384"

Apply the change:

sudo systemctl daemon-reload
sudo systemctl restart ollama

A longer context raises memory use for every loaded model. Run ollama ps and free -h after the next request, and lower the value if available memory runs short.

Ollama listens on 127.0.0.1:11434 by default. Keep it there and reach it over an SSH tunnel. From your own computer, run:

ssh -N -L 11434:127.0.0.1:11434 YOUR_USER@YOUR_SERVER_IP

Replace YOUR_USER with your sudo user and YOUR_SERVER_IP with the server's IP address. While the tunnel is open, send requests to http://127.0.0.1:11434 on your computer.

Run local models on demand: a request keeps the CPU busy while it runs, then the server idles. Sustained full-CPU load is not permitted on Arct plans, so continuous or batch generation belongs on an API model.

FAQ

Can Ollama work on a CPU?

Yes. Ollama loads the whole model into system RAM when a server has no GPU, and ollama ps reports 100% CPU. Small and mixture-of-experts models answer at interactive speed; dense 27B models generated about 2.4 tokens per second on 16 vCPU.

How much RAM is needed to run Ollama models?

Plan for the download size plus 1 to 4 GB. Qwen3.5 4B fits 8 GB RAM, Qwen3.5 9B or Gemma 4 12B fit 16 GB, and Qwen3.6 35B-A3B fits 32 GB. On 2 GB, run an embedding model such as EmbeddingGemma. Longer contexts need more RAM.

Which is the best Ollama model for CPU servers?

Qwen3.5 is the best starting family on CPU: 4B on 8 GB RAM and 9B on 16 GB. On 32 GB, Qwen3.6 35B-A3B generated 14.24 tokens per second after an 8K prompt, the fastest result of any model above 1B in the same test.

What is the best Ollama model for coding in 2026?

On a 32 GB CPU server, start with Qwen3.6 35B-A3B, which generated 14.24 tokens per second after an 8K prompt. Qwen3-Coder 30B is the coding-tuned option. Ollama recommends a 64,000-token context for coding agents, and long prompts take minutes on CPU, so use an API model for agentic coding.

Does Gemma 4 run on Ollama without a GPU?

Yes. Gemma 4 E2B is a 4.6 GB download at Q4_K_M, sized for 8 GB RAM. The 12B build is 8.0 GB, sized for 16 GB. The 26B mixture-of-experts build activates 3.8B parameters per token and is sized for 24 GB.

Are Ollama models free?

Yes. Ollama is open-source software under the MIT license, and library models download at no cost. Each model has its own license from its publisher. Tags ending in cloud run on Ollama's hosted service, need an Ollama account and count toward its cloud usage plans.

Should I use Q4_K_M or Q8_0 on a CPU?

Use Q4_K_M for most models on a CPU server. Q8_0 files are 1.4 to 1.8 times larger, so they need more RAM and generate more slowly on the same CPU. The default Qwen3.5 0.8B and 2B tags are already Q8_0-sized.

Next steps

High Performance plans run on AMD Ryzen 9 up to 5.7 GHz or AMD EPYC Genoa 9004 Series up to 4.3 GHz, with DDR5 RAM. See High Performance VPS plans or compare all plans.