All tutorials
11 min read

How to Install Ollama on Ubuntu 24.04 with Open WebUI (CPU Only)

Install Ollama on Ubuntu 24.04 without a GPU, keep its API on 127.0.0.1:11434, and run Open WebUI in Docker behind Nginx, HTTPS and UFW in 8 steps.

TL;DR:

  • Install Ollama 0.35.1 on Ubuntu 24.04 with the official script, pin its API to 127.0.0.1:11434 in a systemd override, then run Open WebUI v0.11.4 in Docker behind Nginx with a Let's Encrypt certificate.
  • No GPU is needed: the installer reports CPU-only mode and loads models into system RAM. Plan for 8 GB RAM to run qwen3.5:4b and 16 GB for qwen3.5:9b.
  • Ollama's local API has no authentication. Keep port 11434 on loopback and open only 22, 80 and 443 in UFW.
  • Open WebUI uses host networking on 127.0.0.1:8080, so Docker publishes no ports. The first account becomes the admin and sign-ups switch off.
  • Disk: Ollama takes 2.3 GB, the Open WebUI image is a 1.66 GB download that takes more once extracted (check with docker image ls), and each model adds its own download size.

Applies to: Ubuntu 24.04 LTS · Ollama 0.35.1 · Open WebUI v0.11.4 · checked October 2026

Ollama is an open-source runtime, released under the MIT license, that downloads open-weight language models and serves them through a command line and a local API on port 11434. Open WebUI is a self-hosted chat interface for Ollama and API model providers. This tutorial installs both on a CPU-only Ubuntu 24.04 server: Ollama as a systemd service on the loopback address, Open WebUI in Docker, and Nginx with HTTPS in front, behind the Uncomplicated Firewall (UFW).

How much RAM does Ollama need on a CPU VPS?

Without a GPU, Ollama loads the whole model into system RAM: budget the download size plus 1 to 4 GB, then the OS and Open WebUI. In the CPU benchmarks for self-hosted LLMs, qwen3.5:4b, a 3.4 GB download, peaked at 4.18 GiB at an 8K-token context, so it runs on 8 GB of server RAM. qwen3.5:9b fits 16 GB, and qwen3.6:35b-a3b-q4_K_M fits 32 GB.

On 4 GB RAM, use a model under 2 GB, such as qwen3.5:2b-q4_K_M (1.9 GB), at the default context. The RAM table for each model and the picks by task and plan size are in best Ollama models for CPU servers.

Why port 11434 stays on 127.0.0.1

Ollama's documentation states that the local API on port 11434 does not require authentication. Any client that reaches it can run prompts, pull models until the disk fills, and delete models. The installer binds it to 127.0.0.1:11434; setting OLLAMA_HOST=0.0.0.0 puts that open API on the internet.

Here, only SSH and Nginx accept outside connections, and Open WebUI asks for a login before it forwards anything to Ollama. Open WebUI's hardening guide describes the app as built for private networks. For a single user, the SSH tunnel in Step 5 or a WireGuard VPN keeps it off the public internet.

Prerequisites

  • A VPS running Ubuntu 24.04 LTS with SSH access and enough RAM for your model (see the RAM section above). Ubuntu 26.04 LTS uses the same commands. The CPU figures here were measured on an Arct Cloud High Performance hvm.xlarge plan (16 vCPU, 32 GB RAM).
  • A non-root user with sudo privileges and SSH key login. See How to Generate an SSH Key.
  • Docker Engine and the Docker Compose plugin. See How to Install Docker and Docker Compose on Ubuntu 24.04 and 26.04.
  • A domain with a DNS A record, such as chat.example.com, pointing to the server's IP address (Step 7 only).

Commands run as the sudo user. Replace values written in capitals, such as YOUR_DOMAIN, with your own.

Step 1: Install Ollama on Ubuntu with the official script

Ollama's Linux package is a .tar.zst archive, and the installer stops with an error if zstd is missing. Ubuntu's cloud image includes it; this command adds it where it is missing:

sudo apt update
sudo apt install -y curl zstd

Download the script from ollama.com, read it, then run it:

curl -fsSL https://ollama.com/install.sh -o ollama-install.sh
less ollama-install.sh
sh ollama-install.sh

The script asks for your sudo password. The output looks similar to this:

>>> Installing ollama to /usr/local
>>> Downloading ollama-linux-amd64.tar.zst
...
>>> Creating ollama user...
...
>>> Creating ollama systemd service...
>>> Enabling and starting ollama service...
>>> The Ollama API is now available at 127.0.0.1:11434.
>>> Install complete. Run "ollama" from the command line.
WARNING: No NVIDIA/AMD GPU detected. Ollama will run in CPU-only mode.

The warning confirms CPU mode. The script creates an ollama system user and the ollama systemd service, and stores models in /usr/share/ollama/.ollama/models. Check the version:

ollama -v
ollama version is 0.35.1

At the time of writing, the script installed 0.35.1. Your output may show a newer version.

Step 2: Pin the Ollama API to 127.0.0.1

Settings for the service go in a systemd drop-in file, which survives updates because the installer rewrites only ollama.service:

sudo mkdir -p /etc/systemd/system/ollama.service.d
sudo tee /etc/systemd/system/ollama.service.d/override.conf > /dev/null <<'EOF'
[Service]
Environment="OLLAMA_HOST=127.0.0.1:11434"
Environment="OLLAMA_CONTEXT_LENGTH=8192"
Environment="OLLAMA_NO_CLOUD=1"
EOF
  • OLLAMA_HOST sets the loopback address explicitly, so the bind address is recorded in one file.
  • OLLAMA_CONTEXT_LENGTH raises the context from 4,096 tokens, the default without a GPU, to 8,192. RAM use grows with it; on 4 GB RAM, delete this line.
  • OLLAMA_NO_CLOUD turns off Ollama's cloud models and web search, so prompts stay on the server.

Apply the change and test the API:

sudo systemctl daemon-reload
sudo systemctl restart ollama
curl http://127.0.0.1:11434
Ollama is running

Service logs are in journalctl -u ollama, which shows Ollama cloud disabled: true after the restart.

Step 3: Pull and test a first model

Download the model sized for your RAM. qwen3.5:4b is a 3.4 GB download:

ollama pull qwen3.5:4b
pulling manifest
...
verifying sha256 digest
writing manifest
success

Run one prompt with timings, and turn off the model's thinking step for a short reply:

ollama run qwen3.5:4b --verbose --think=false "Explain a reverse proxy in two sentences."
ollama ps

After the answer, --verbose prints prompt eval rate and eval rate in tokens per second. ollama ps shows 100% CPU in the PROCESSOR column and 8192 under CONTEXT. In the benchmark, qwen3.5:4b generated 13.43 tokens/s on 16 vCPU with a short prompt; plans with fewer vCPU generate more slowly.

Use local models on demand: a request keeps the CPU busy while the reply generates, then the server idles, and Ollama unloads the model after 5 idle minutes. Sustained full-CPU load is not permitted on Arct plans, so continuous or batch generation belongs on an API model.

Step 4: Run Open WebUI with Docker Compose

Open WebUI runs with host networking so it can reach Ollama on 127.0.0.1:11434. Docker publishes no ports in this mode, and the HOST variable binds the app to loopback. Create the project directory and an .env file with a fixed secret key, which keeps users logged in when the container is recreated. Replace YOUR_DOMAIN in the command below with your domain, such as chat.example.com, before you run it. If you will use only the SSH tunnel and skip Step 7, enter the domain you plan to use later:

mkdir -p ~/open-webui
cd ~/open-webui
printf 'DOMAIN=%s\nWEBUI_SECRET_KEY=%s\n' YOUR_DOMAIN "$(openssl rand -hex 32)" > .env
chmod 600 .env

Open WebUI saves WEBUI_URL to its database on first start, so change it later under Settings > Admin > General > WebUI URL, not in .env. Create compose.yaml with nano compose.yaml, paste this file, then save with Ctrl+O and exit with Ctrl+X:

services:
  open-webui:
    image: ghcr.io/open-webui/open-webui:v0.11.4
    container_name: open-webui
    network_mode: host
    environment:
      HOST: "127.0.0.1"
      PORT: "8080"
      OLLAMA_BASE_URL: "http://127.0.0.1:11434"
      WEBUI_SECRET_KEY: "${WEBUI_SECRET_KEY}"
      WEBUI_URL: "https://${DOMAIN}"
      CORS_ALLOW_ORIGIN: "https://${DOMAIN};http://localhost:8080"
      WEBUI_SESSION_COOKIE_SECURE: "true"
      FORWARDED_ALLOW_IPS: "127.0.0.1"
    volumes:
      - open-webui:/app/backend/data
    restart: unless-stopped

volumes:
  open-webui:
    name: open-webui

CORS_ALLOW_ORIGIN lists the HTTPS and SSH tunnel addresses, which Open WebUI's docs require behind a reverse proxy. FORWARDED_ALLOW_IPS trusts forwarded headers only from Nginx on the same host. Chats, users and settings live in the open-webui volume. Start the container:

docker compose up -d

The first start takes about a minute. Check the health endpoint:

curl -s http://127.0.0.1:8080/health
{"status":true}

Step 5: Create the admin account over an SSH tunnel

Create the admin account before the site is public. On your own computer, open a tunnel to the server and leave it running:

ssh -N -L 8080:127.0.0.1:8080 YOUR_USER@YOUR_SERVER_IP

Replace YOUR_USER with your sudo user and YOUR_SERVER_IP with the server's IP address, such as 203.0.113.10. Open http://localhost:8080 in a browser and select Create Admin Account. Use Chrome or Firefox for the tunnel address; both accept secure cookies on http://localhost.

The first account becomes the administrator, and Open WebUI switches sign-ups off once it exists. To add users later, turn on New Sign Ups under your avatar > Settings > Admin > Authentication; new accounts wait as Pending until you approve them.

Start a new chat, select qwen3.5:4b and send a message. For a private setup, keep using the tunnel: allow only OpenSSH in Step 6 and skip Step 7.

Step 6: Open the firewall for SSH and HTTPS

UFW blocks every incoming port you do not allow. Allow SSH before you enable it, or you lose SSH access to the server:

sudo ufw allow OpenSSH
sudo ufw allow 80/tcp
sudo ufw allow 443/tcp
sudo ufw enable
sudo ufw status

Type y when UFW asks to confirm. The output looks similar to this:

Status: active

To                         Action      From
--                         ------      ----
OpenSSH                    ALLOW       Anywhere
80/tcp                     ALLOW       Anywhere
443/tcp                    ALLOW       Anywhere
OpenSSH (v6)               ALLOW       Anywhere (v6)
80/tcp (v6)                ALLOW       Anywhere (v6)
443/tcp (v6)               ALLOW       Anywhere (v6)

Ports 11434 and 8080 stay closed, and both services listen only on 127.0.0.1.

Step 7: Put Open WebUI behind Nginx with HTTPS

Install Nginx and Certbot with its Nginx plugin from Ubuntu's repositories:

sudo apt install -y nginx certbot python3-certbot-nginx

Create the site file with sudo nano /etc/nginx/sites-available/open-webui and paste this configuration:

server {
    listen 80;
    listen [::]:80;
    server_name YOUR_DOMAIN;

    client_max_body_size 20M;

    location / {
        proxy_pass http://127.0.0.1:8080;
        proxy_http_version 1.1;
        proxy_set_header Upgrade $http_upgrade;
        proxy_set_header Connection "upgrade";
        proxy_set_header Host $host;
        proxy_set_header X-Real-IP $remote_addr;
        proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
        proxy_set_header X-Forwarded-Proto $scheme;
        proxy_buffering off;
        proxy_cache off;
        proxy_read_timeout 1800;
        proxy_send_timeout 1800;
    }
}

Replace YOUR_DOMAIN. The Upgrade headers carry Open WebUI's WebSocket connections, proxy_buffering off streams replies token by token, and the 1,800-second timeouts cover slow CPU replies. Enable the site and test the configuration:

sudo ln -s /etc/nginx/sites-available/open-webui /etc/nginx/sites-enabled/
sudo nginx -t
sudo systemctl reload nginx
nginx: the configuration file /etc/nginx/nginx.conf syntax is ok
nginx: configuration file /etc/nginx/nginx.conf test is successful

Request a TLS certificate, often called an SSL certificate, from Let's Encrypt:

sudo certbot --nginx -d YOUR_DOMAIN -m YOUR_EMAIL --agree-tos --no-eff-email --redirect

Replace YOUR_EMAIL with your email address; --no-eff-email skips the prompt to share it with the EFF. Certbot adds the certificate and port 443 listeners to the site file and redirects HTTP to HTTPS. Let's Encrypt stopped sending expiry emails in June 2025, so renewal depends on the certbot.timer systemd timer. Test renewal with sudo certbot renew --dry-run, then open https://YOUR_DOMAIN and sign in.

To limit the site to your own network, add allow YOUR_IP; and deny all; inside the location / block and reload Nginx.

Step 8: Verify the ports and the firewall

List the listening TCP sockets on the server with ss -tln and check the Local Address column. A fresh Ubuntu 24.04 server shows these:

Local addressServiceReachable from outside
127.0.0.1:11434Ollama APINo
127.0.0.1:8080Open WebUINo
127.0.0.53%lo:53, 127.0.0.54:53systemd-resolved (DNS)No
0.0.0.0:22, [::]:22SSHYes
0.0.0.0:80, 0.0.0.0:443, [::]:80, [::]:443NginxYes

A line with 0.0.0.0:11434 or *:11434 means OLLAMA_HOST changed; repeat Step 2. From your own computer, test the API port:

curl --max-time 5 http://YOUR_SERVER_IP:11434

The request times out after 5 seconds. A reply of Ollama is running means the API is public.

Update Ollama and Open WebUI

Download, read and run the install script again to update Ollama; the Step 2 drop-in stays in place. OLLAMA_VERSION installs a specific release; leave it out for the latest:

curl -fsSL https://ollama.com/install.sh -o ollama-install.sh
less ollama-install.sh
OLLAMA_VERSION=0.35.1 sh ollama-install.sh

Open WebUI updates can run one-way database migrations, so back up the volume first:

cd ~/open-webui
docker compose stop
sudo tar -czf ~/open-webui-$(date +%F).tar.gz -C /var/lib/docker/volumes/open-webui/_data .

Change the image: tag in compose.yaml to the new version from the Open WebUI releases page, then pull and restart:

docker compose pull
docker compose up -d

For server-level copies, see VPS backups.

Troubleshoot common Ollama and Open WebUI errors

SymptomCauseFix
ERROR: This version requires zstd for extractionzstd is missingsudo apt install zstd, then rerun the script
Open WebUI lists no modelsNo model pulled, or Ollama is downollama list, curl http://127.0.0.1:11434, journalctl -u ollama -n 50
502 Bad Gateway from NginxOpen WebUI is stopped or still startingdocker compose ps, then docker compose logs open-webui
Replies show raw ** or ##, or stop mid-streamNginx buffers the streamKeep proxy_buffering off, then sudo systemctl reload nginx
is not an accepted origin in Open WebUI logsThe browser address is missing from CORS_ALLOW_ORIGINAdd it in compose.yaml, then docker compose up -d
Model fails to load or stops partwayThe model and context exceed free RAMCheck free -h, use a smaller tag or lower OLLAMA_CONTEXT_LENGTH

FAQ

How much RAM does Ollama need?

Ollama needs the model's download size in RAM plus 1 to 4 GB. Qwen3.5 4B, a 3.4 GB download, peaked at 4.18 GiB with an 8K context, so plan 8 GB of server RAM for it. Qwen3.5 9B fits 16 GB. Longer contexts and parallel requests need more.

Is Ollama safe to use?

Yes, as long as its API stays on 127.0.0.1, the default address. The API on port 11434 has no authentication, so anyone who reaches it can run prompts, pull models and delete them. Keep it on loopback and reach it through Open WebUI, an SSH tunnel or a VPN.

Can I install Ollama on Linux without sudo?

Yes, without a system service. Extract Ollama's Linux tarball into a folder in your home directory and run bin/ollama serve from that folder. The install script needs sudo because it writes to /usr/local and creates the ollama systemd service.

Should I install Ollama with Docker on Ubuntu?

You can. The official ollama/ollama image runs on CPU, but its documented command publishes port 11434 on every interface, and Docker-published ports bypass UFW. Publish it as 127.0.0.1:11434:11434 instead. This tutorial uses the script for a systemd service on the host.

How do I install an older version of Ollama?

Set the OLLAMA_VERSION variable when you run the install script, for example OLLAMA_VERSION=0.34.4. Version numbers are on Ollama's GitHub releases page, and the same variable installs pre-releases.

Next steps

High Performance plans run on AMD Ryzen 9 up to 5.7 GHz or AMD EPYC Genoa 9004 Series up to 4.3 GHz, with DDR5 RAM. Compare plans.

This work is licensed under CC BY-NC-SA 4.0.