TL;DR:
- Install Ollama 0.35.1 on Ubuntu 24.04 with the official script, pin its API to
127.0.0.1:11434in a systemd override, then run Open WebUI v0.11.4 in Docker behind Nginx with a Let's Encrypt certificate. - No GPU is needed: the installer reports CPU-only mode and loads models into system RAM. Plan for 8 GB RAM to run
qwen3.5:4band 16 GB forqwen3.5:9b. - Ollama's local API has no authentication. Keep port 11434 on loopback and open only 22, 80 and 443 in UFW.
- Open WebUI uses host networking on
127.0.0.1:8080, so Docker publishes no ports. The first account becomes the admin and sign-ups switch off. - Disk: Ollama takes 2.3 GB, the Open WebUI image is a 1.66 GB download that takes more once extracted (check with
docker image ls), and each model adds its own download size.
Applies to: Ubuntu 24.04 LTS · Ollama 0.35.1 · Open WebUI v0.11.4 · checked October 2026
Ollama is an open-source runtime, released under the MIT license, that downloads open-weight language models and serves them through a command line and a local API on port 11434. Open WebUI is a self-hosted chat interface for Ollama and API model providers. This tutorial installs both on a CPU-only Ubuntu 24.04 server: Ollama as a systemd service on the loopback address, Open WebUI in Docker, and Nginx with HTTPS in front, behind the Uncomplicated Firewall (UFW).
How much RAM does Ollama need on a CPU VPS?
Without a GPU, Ollama loads the whole model into system RAM: budget the download size plus 1 to 4 GB, then the OS and Open WebUI. In the CPU benchmarks for self-hosted LLMs, qwen3.5:4b, a 3.4 GB download, peaked at 4.18 GiB at an 8K-token context, so it runs on 8 GB of server RAM. qwen3.5:9b fits 16 GB, and qwen3.6:35b-a3b-q4_K_M fits 32 GB.
On 4 GB RAM, use a model under 2 GB, such as qwen3.5:2b-q4_K_M (1.9 GB), at the default context. The RAM table for each model and the picks by task and plan size are in best Ollama models for CPU servers.
Why port 11434 stays on 127.0.0.1
Ollama's documentation states that the local API on port 11434 does not require authentication. Any client that reaches it can run prompts, pull models until the disk fills, and delete models. The installer binds it to 127.0.0.1:11434; setting OLLAMA_HOST=0.0.0.0 puts that open API on the internet.
Here, only SSH and Nginx accept outside connections, and Open WebUI asks for a login before it forwards anything to Ollama. Open WebUI's hardening guide describes the app as built for private networks. For a single user, the SSH tunnel in Step 5 or a WireGuard VPN keeps it off the public internet.
Prerequisites
- A VPS running Ubuntu 24.04 LTS with SSH access and enough RAM for your model (see the RAM section above). Ubuntu 26.04 LTS uses the same commands. The CPU figures here were measured on an Arct Cloud High Performance
hvm.xlargeplan (16 vCPU, 32 GB RAM). - A non-root user with sudo privileges and SSH key login. See How to Generate an SSH Key.
- Docker Engine and the Docker Compose plugin. See How to Install Docker and Docker Compose on Ubuntu 24.04 and 26.04.
- A domain with a DNS A record, such as
chat.example.com, pointing to the server's IP address (Step 7 only).
Commands run as the sudo user. Replace values written in capitals, such as YOUR_DOMAIN, with your own.
Step 1: Install Ollama on Ubuntu with the official script
Ollama's Linux package is a .tar.zst archive, and the installer stops with an error if zstd is missing. Ubuntu's cloud image includes it; this command adds it where it is missing:
sudo apt update
sudo apt install -y curl zstd
Download the script from ollama.com, read it, then run it:
curl -fsSL https://ollama.com/install.sh -o ollama-install.sh
less ollama-install.sh
sh ollama-install.sh
The script asks for your sudo password. The output looks similar to this:
>>> Installing ollama to /usr/local
>>> Downloading ollama-linux-amd64.tar.zst
...
>>> Creating ollama user...
...
>>> Creating ollama systemd service...
>>> Enabling and starting ollama service...
>>> The Ollama API is now available at 127.0.0.1:11434.
>>> Install complete. Run "ollama" from the command line.
WARNING: No NVIDIA/AMD GPU detected. Ollama will run in CPU-only mode.
The warning confirms CPU mode. The script creates an ollama system user and the ollama systemd service, and stores models in /usr/share/ollama/.ollama/models. Check the version:
ollama -v
ollama version is 0.35.1
At the time of writing, the script installed 0.35.1. Your output may show a newer version.
Step 2: Pin the Ollama API to 127.0.0.1
Settings for the service go in a systemd drop-in file, which survives updates because the installer rewrites only ollama.service:
sudo mkdir -p /etc/systemd/system/ollama.service.d
sudo tee /etc/systemd/system/ollama.service.d/override.conf > /dev/null <<'EOF'
[Service]
Environment="OLLAMA_HOST=127.0.0.1:11434"
Environment="OLLAMA_CONTEXT_LENGTH=8192"
Environment="OLLAMA_NO_CLOUD=1"
EOF
OLLAMA_HOSTsets the loopback address explicitly, so the bind address is recorded in one file.OLLAMA_CONTEXT_LENGTHraises the context from 4,096 tokens, the default without a GPU, to 8,192. RAM use grows with it; on 4 GB RAM, delete this line.OLLAMA_NO_CLOUDturns off Ollama's cloud models and web search, so prompts stay on the server.
Apply the change and test the API:
sudo systemctl daemon-reload
sudo systemctl restart ollama
curl http://127.0.0.1:11434
Ollama is running
Service logs are in journalctl -u ollama, which shows Ollama cloud disabled: true after the restart.
Step 3: Pull and test a first model
Download the model sized for your RAM. qwen3.5:4b is a 3.4 GB download:
ollama pull qwen3.5:4b
pulling manifest
...
verifying sha256 digest
writing manifest
success
Run one prompt with timings, and turn off the model's thinking step for a short reply:
ollama run qwen3.5:4b --verbose --think=false "Explain a reverse proxy in two sentences."
ollama ps
After the answer, --verbose prints prompt eval rate and eval rate in tokens per second. ollama ps shows 100% CPU in the PROCESSOR column and 8192 under CONTEXT. In the benchmark, qwen3.5:4b generated 13.43 tokens/s on 16 vCPU with a short prompt; plans with fewer vCPU generate more slowly.
Use local models on demand: a request keeps the CPU busy while the reply generates, then the server idles, and Ollama unloads the model after 5 idle minutes. Sustained full-CPU load is not permitted on Arct plans, so continuous or batch generation belongs on an API model.
Step 4: Run Open WebUI with Docker Compose
Open WebUI runs with host networking so it can reach Ollama on 127.0.0.1:11434. Docker publishes no ports in this mode, and the HOST variable binds the app to loopback. Create the project directory and an .env file with a fixed secret key, which keeps users logged in when the container is recreated. Replace YOUR_DOMAIN in the command below with your domain, such as chat.example.com, before you run it. If you will use only the SSH tunnel and skip Step 7, enter the domain you plan to use later:
mkdir -p ~/open-webui
cd ~/open-webui
printf 'DOMAIN=%s\nWEBUI_SECRET_KEY=%s\n' YOUR_DOMAIN "$(openssl rand -hex 32)" > .env
chmod 600 .env
Open WebUI saves WEBUI_URL to its database on first start, so change it later under Settings > Admin > General > WebUI URL, not in .env. Create compose.yaml with nano compose.yaml, paste this file, then save with Ctrl+O and exit with Ctrl+X:
services:
open-webui:
image: ghcr.io/open-webui/open-webui:v0.11.4
container_name: open-webui
network_mode: host
environment:
HOST: "127.0.0.1"
PORT: "8080"
OLLAMA_BASE_URL: "http://127.0.0.1:11434"
WEBUI_SECRET_KEY: "${WEBUI_SECRET_KEY}"
WEBUI_URL: "https://${DOMAIN}"
CORS_ALLOW_ORIGIN: "https://${DOMAIN};http://localhost:8080"
WEBUI_SESSION_COOKIE_SECURE: "true"
FORWARDED_ALLOW_IPS: "127.0.0.1"
volumes:
- open-webui:/app/backend/data
restart: unless-stopped
volumes:
open-webui:
name: open-webui
CORS_ALLOW_ORIGIN lists the HTTPS and SSH tunnel addresses, which Open WebUI's docs require behind a reverse proxy. FORWARDED_ALLOW_IPS trusts forwarded headers only from Nginx on the same host. Chats, users and settings live in the open-webui volume. Start the container:
docker compose up -d
The first start takes about a minute. Check the health endpoint:
curl -s http://127.0.0.1:8080/health
{"status":true}
Step 5: Create the admin account over an SSH tunnel
Create the admin account before the site is public. On your own computer, open a tunnel to the server and leave it running:
ssh -N -L 8080:127.0.0.1:8080 YOUR_USER@YOUR_SERVER_IP
Replace YOUR_USER with your sudo user and YOUR_SERVER_IP with the server's IP address, such as 203.0.113.10. Open http://localhost:8080 in a browser and select Create Admin Account. Use Chrome or Firefox for the tunnel address; both accept secure cookies on http://localhost.
The first account becomes the administrator, and Open WebUI switches sign-ups off once it exists. To add users later, turn on New Sign Ups under your avatar > Settings > Admin > Authentication; new accounts wait as Pending until you approve them.
Start a new chat, select qwen3.5:4b and send a message. For a private setup, keep using the tunnel: allow only OpenSSH in Step 6 and skip Step 7.
Step 6: Open the firewall for SSH and HTTPS
UFW blocks every incoming port you do not allow. Allow SSH before you enable it, or you lose SSH access to the server:
sudo ufw allow OpenSSH
sudo ufw allow 80/tcp
sudo ufw allow 443/tcp
sudo ufw enable
sudo ufw status
Type y when UFW asks to confirm. The output looks similar to this:
Status: active
To Action From
-- ------ ----
OpenSSH ALLOW Anywhere
80/tcp ALLOW Anywhere
443/tcp ALLOW Anywhere
OpenSSH (v6) ALLOW Anywhere (v6)
80/tcp (v6) ALLOW Anywhere (v6)
443/tcp (v6) ALLOW Anywhere (v6)
Ports 11434 and 8080 stay closed, and both services listen only on 127.0.0.1.
Step 7: Put Open WebUI behind Nginx with HTTPS
Install Nginx and Certbot with its Nginx plugin from Ubuntu's repositories:
sudo apt install -y nginx certbot python3-certbot-nginx
Create the site file with sudo nano /etc/nginx/sites-available/open-webui and paste this configuration:
server {
listen 80;
listen [::]:80;
server_name YOUR_DOMAIN;
client_max_body_size 20M;
location / {
proxy_pass http://127.0.0.1:8080;
proxy_http_version 1.1;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection "upgrade";
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
proxy_buffering off;
proxy_cache off;
proxy_read_timeout 1800;
proxy_send_timeout 1800;
}
}
Replace YOUR_DOMAIN. The Upgrade headers carry Open WebUI's WebSocket connections, proxy_buffering off streams replies token by token, and the 1,800-second timeouts cover slow CPU replies. Enable the site and test the configuration:
sudo ln -s /etc/nginx/sites-available/open-webui /etc/nginx/sites-enabled/
sudo nginx -t
sudo systemctl reload nginx
nginx: the configuration file /etc/nginx/nginx.conf syntax is ok
nginx: configuration file /etc/nginx/nginx.conf test is successful
Request a TLS certificate, often called an SSL certificate, from Let's Encrypt:
sudo certbot --nginx -d YOUR_DOMAIN -m YOUR_EMAIL --agree-tos --no-eff-email --redirect
Replace YOUR_EMAIL with your email address; --no-eff-email skips the prompt to share it with the EFF. Certbot adds the certificate and port 443 listeners to the site file and redirects HTTP to HTTPS. Let's Encrypt stopped sending expiry emails in June 2025, so renewal depends on the certbot.timer systemd timer. Test renewal with sudo certbot renew --dry-run, then open https://YOUR_DOMAIN and sign in.
To limit the site to your own network, add allow YOUR_IP; and deny all; inside the location / block and reload Nginx.
Step 8: Verify the ports and the firewall
List the listening TCP sockets on the server with ss -tln and check the Local Address column. A fresh Ubuntu 24.04 server shows these:
| Local address | Service | Reachable from outside |
|---|---|---|
127.0.0.1:11434 | Ollama API | No |
127.0.0.1:8080 | Open WebUI | No |
127.0.0.53%lo:53, 127.0.0.54:53 | systemd-resolved (DNS) | No |
0.0.0.0:22, [::]:22 | SSH | Yes |
0.0.0.0:80, 0.0.0.0:443, [::]:80, [::]:443 | Nginx | Yes |
A line with 0.0.0.0:11434 or *:11434 means OLLAMA_HOST changed; repeat Step 2. From your own computer, test the API port:
curl --max-time 5 http://YOUR_SERVER_IP:11434
The request times out after 5 seconds. A reply of Ollama is running means the API is public.
Update Ollama and Open WebUI
Download, read and run the install script again to update Ollama; the Step 2 drop-in stays in place. OLLAMA_VERSION installs a specific release; leave it out for the latest:
curl -fsSL https://ollama.com/install.sh -o ollama-install.sh
less ollama-install.sh
OLLAMA_VERSION=0.35.1 sh ollama-install.sh
Open WebUI updates can run one-way database migrations, so back up the volume first:
cd ~/open-webui
docker compose stop
sudo tar -czf ~/open-webui-$(date +%F).tar.gz -C /var/lib/docker/volumes/open-webui/_data .
Change the image: tag in compose.yaml to the new version from the Open WebUI releases page, then pull and restart:
docker compose pull
docker compose up -d
For server-level copies, see VPS backups.
Troubleshoot common Ollama and Open WebUI errors
| Symptom | Cause | Fix |
|---|---|---|
ERROR: This version requires zstd for extraction | zstd is missing | sudo apt install zstd, then rerun the script |
| Open WebUI lists no models | No model pulled, or Ollama is down | ollama list, curl http://127.0.0.1:11434, journalctl -u ollama -n 50 |
502 Bad Gateway from Nginx | Open WebUI is stopped or still starting | docker compose ps, then docker compose logs open-webui |
Replies show raw ** or ##, or stop mid-stream | Nginx buffers the stream | Keep proxy_buffering off, then sudo systemctl reload nginx |
is not an accepted origin in Open WebUI logs | The browser address is missing from CORS_ALLOW_ORIGIN | Add it in compose.yaml, then docker compose up -d |
| Model fails to load or stops partway | The model and context exceed free RAM | Check free -h, use a smaller tag or lower OLLAMA_CONTEXT_LENGTH |
FAQ
How much RAM does Ollama need?
Ollama needs the model's download size in RAM plus 1 to 4 GB. Qwen3.5 4B, a 3.4 GB download, peaked at 4.18 GiB with an 8K context, so plan 8 GB of server RAM for it. Qwen3.5 9B fits 16 GB. Longer contexts and parallel requests need more.
Is Ollama safe to use?
Yes, as long as its API stays on 127.0.0.1, the default address. The API on port 11434 has no authentication, so anyone who reaches it can run prompts, pull models and delete them. Keep it on loopback and reach it through Open WebUI, an SSH tunnel or a VPN.
Can I install Ollama on Linux without sudo?
Yes, without a system service. Extract Ollama's Linux tarball into a folder in your home directory and run bin/ollama serve from that folder. The install script needs sudo because it writes to /usr/local and creates the ollama systemd service.
Should I install Ollama with Docker on Ubuntu?
You can. The official ollama/ollama image runs on CPU, but its documented command publishes port 11434 on every interface, and Docker-published ports bypass UFW. Publish it as 127.0.0.1:11434:11434 instead. This tutorial uses the script for a systemd service on the host.
How do I install an older version of Ollama?
Set the OLLAMA_VERSION variable when you run the install script, for example OLLAMA_VERSION=0.34.4. Version numbers are on Ollama's GitHub releases page, and the same variable installs pre-releases.
Next steps
- Best Ollama Models for CPU Servers by RAM Size
- How to Install OpenCode on Ubuntu 24.04 with API or Ollama Models or How to Build an n8n AI Agent with an API Model or Ollama
- Ollama Linux documentation and Open WebUI HTTPS with Nginx
High Performance plans run on AMD Ryzen 9 up to 5.7 GHz or AMD EPYC Genoa 9004 Series up to 4.3 GHz, with DDR5 RAM. Compare plans.
This work is licensed under CC BY-NC-SA 4.0.