All tutorials
13 min read

How to Build an n8n AI Agent with an API Model or Ollama

Build an n8n AI agent in 9 steps: a Telegram bot with the AI Agent node, an API chat model, Postgres memory, two tools, and an optional Ollama model.

TL;DR:

  • Build a Telegram assistant from a Telegram Trigger, the AI Agent node, an API chat model, Postgres Chat Memory, two tools (Calculator and an HTTP Request tool for Open-Meteo weather) and a Telegram reply node.
  • Since n8n 1.82.0 every AI Agent node runs as a Tools Agent and needs at least one tool connected; the steps start from a self-hosted n8n 2.x instance on an HTTPS domain.
  • Telegram delivers updates only to an HTTPS webhook on port 443, 80, 88 or 8443, so n8n needs N8N_WEBHOOK_URL set to its public HTTPS address.
  • Postgres Chat Memory sends the last 5 exchanges per Telegram chat to the model and keeps them in a table that survives restarts.
  • The optional local model, qwen3.5:4b, is a 3.4 GB download that runs in an Ollama container with no published port and peaked at 4.18 GiB of memory in a CPU benchmark.

Applies to: Ubuntu 24.04 LTS · n8n 2.41.6 · Ollama 0.35.1 · checked October 2026

An n8n AI agent is a workflow built around n8n's AI Agent node: the node passes each message to a chat model, runs the tools the model picks, and reads earlier messages from a memory node. n8n, short for "nodemation", is a fair-code workflow automation platform distributed under the Sustainable Use License. This n8n AI agent tutorial builds one working agent, a Telegram assistant with memory and two tools, on a self-hosted n8n 2.x instance with an API model, then shows how to run the same agent on a local Ollama model.

What an n8n AI agent is made of

The AI Agent node is a root node. It runs with sub-nodes attached to its Chat Model, Memory and Tool connectors, and n8n's documentation requires at least one tool.

PartNode in this tutorialJob
TriggerTelegram TriggerStarts the workflow when you message the bot
AgentAI AgentSends the prompt to the model and runs the tools it calls
Chat modelOpenAI Chat Model or Ollama Chat ModelWrites the reply and decides on tool calls
MemoryPostgres Chat MemoryStores the conversation per Telegram chat
ToolsCalculator, HTTP Request ToolDoes arithmetic and fetches current weather
OutputTelegramSends the reply to your chat

Since n8n 1.82.0, every AI Agent node works as a Tools Agent, so there is no agent type to choose. The node steps also work on n8n Cloud. The database and Ollama steps assume self-hosted n8n on your own server.

API model or Ollama: which to use

An API model runs inference on the provider's servers and adds no RAM load to your server. A local Ollama model keeps prompts on your server, but it needs RAM for the model and generates more slowly on a CPU.

API modelOllama on the same server
Where inference runsProvider: OpenAI, Anthropic, OpenRouter or GoogleYour server's CPU
Extra server RAMNone4.18 GiB peak for qwen3.5:4b with an 8K prompt
Response timeSet by the provider43 s for a 128-token reply to an 8K prompt with qwen3.5:4b on 16 vCPU (12.23 tokens/s generation)
Use it forContinuous, batch or high-volume workPrivate, on-demand chats

The CPU figures come from the Self-Hosted LLM on a VPS benchmark, measured with Ollama 0.32.15 on a 16 vCPU, 32 GB High Performance plan with no GPU. On an 8 GB server, the 4.18 GiB (4.5 GB) peak of qwen3.5:4b leaves about 3.5 GB for Ubuntu, n8n, its task runner, PostgreSQL and the reverse proxy, so 16 GB RAM gives more headroom.

Build the agent with an API model first. Step 8 swaps in Ollama.

Prerequisites

Step 1: Create a Telegram bot

In Telegram, open a chat with @BotFather and send /newbot. Enter a display name, then a username that ends in bot. BotFather replies with an access token. Anyone with the token controls the bot, so keep it out of chats and repositories.

Step 2: Create a database for agent memory

Postgres Chat Memory creates its table on first use. A separate role and database keep chat history apart from n8n's own tables. Generate a password for the new role:

openssl rand -hex 24

Change to the directory that holds your n8n compose file, then create the role and the database:

cd ~/n8n
docker compose exec postgres sh -c 'createuser -U "$POSTGRES_USER" --pwprompt agent_memory'
docker compose exec postgres sh -c 'createdb -U "$POSTGRES_USER" --owner agent_memory agent_memory'

Paste the generated password twice when createuser asks. $POSTGRES_USER expands inside the container to the database superuser from your .env file. Replace ~/n8n and the service name postgres if your compose file uses others.

The backup commands in the self-host tutorial dump only the n8n database. To keep chat history, add this line to them, after the line that sets STAMP:

docker compose exec -T postgres pg_dump -U agent_memory -d agent_memory -Fc > ~/n8n-backups/agent-memory-$STAMP.dump

Step 3: Add the Telegram Trigger

Create a workflow in n8n and add a Telegram Trigger node. Under the credential field, select Create new credential and paste the bot token as the Access Token. Set Trigger On to Message.

Select Execute step, then send any message to your bot. The node output shows the Telegram update. Copy the number at message.from.id: this is your Telegram user ID.

Under Additional Fields, add Restrict to User IDs and paste your user ID. The trigger now ignores messages from other accounts. Telegram allows one webhook per bot, so each bot gets one Telegram Trigger. After you publish in Step 9, do not test the trigger again: a test run moves the bot's webhook to the test URL, and the published workflow receives nothing until you unpublish and publish it again.

Step 4: Add the AI Agent and an API chat model

Add an AI Agent node after the trigger and set:

  • Source for Prompt (User Message): Define below
  • Prompt (User Message): {{ $json.message.text }}

The default source, Connected Chat Trigger Node, expects a chatInput field that the Telegram Trigger does not send. Under Options, add System Message and replace its text:

You are a personal assistant in a Telegram chat.
Reply in plain text without Markdown or HTML, in 5 sentences or fewer.
Use the calculator tool for arithmetic.
Use the weather tool for current weather. Pass the latitude and longitude of the place the user names.
If a tool fails, say so instead of guessing.

The plain-text rule matters because the Telegram node parses text as Markdown (Legacy) unless you set Parse Mode, and Telegram rejects a reply with an unmatched * or _.

Select the plus sign under Chat Model and pick OpenAI Chat Model. Anthropic Chat Model, OpenRouter Chat Model and Google Gemini Chat Model work the same way. Create a credential with your API key and pick a model from the Model list. A small model is enough for two tools.

Step 5: Connect Postgres Chat Memory

Select the plus sign under Memory and pick Postgres Chat Memory. Create a Postgres credential with these values:

FieldValue
Hostpostgres
Databaseagent_memory
Useragent_memory
PasswordThe password from Step 2
Port5432

Leave SSL on Disable. The connection stays on the Docker Compose network. Then set the node:

  • Session ID: Define below
  • Key: {{ $('Telegram Trigger').item.json.message.chat.id }}
  • Table Name: n8n_chat_histories (default)
  • Context Window Length: 5 (default)

The key gives each Telegram chat its own history. Context Window Length sets how many past exchanges the model receives with each new message.

Step 6: Add the Calculator and HTTP Request tools

Select the plus sign under Tool and add Calculator. It has no settings.

Add a second tool, HTTP Request Tool, and rename the node to Get weather. The agent sees the node name as the tool name. Set:

FieldValue
Tool DescriptionSet Manually
DescriptionGets the current temperature, wind speed and precipitation for a latitude and longitude.
MethodGET
URLhttps://api.open-meteo.com/v1/forecast

Turn on Send Query Parameters, keep Using Fields Below, and add four parameters:

NameValue
latitude{{ $fromAI('latitude', 'Latitude in decimal degrees', 'number') }}
longitude{{ $fromAI('longitude', 'Longitude in decimal degrees', 'number') }}
currenttemperature_2m,wind_speed_10m,precipitation
timezoneauto

$fromAI() lets the model fill a parameter each time it calls the tool. Open-Meteo needs no API key. Its terms allow non-commercial use, up to 10,000 calls per day.

Step 7: Send the reply to Telegram

Add a Telegram node after the AI Agent and choose Send a text message. Set:

  • Chat ID: {{ $('Telegram Trigger').item.json.message.chat.id }}
  • Text: {{ $json.output }}

The AI Agent returns its reply in the output field. Under Additional Fields, add Append n8n Attribution and turn it off, or every reply ends with "This message was sent automatically with n8n".

Step 8: Connect n8n to Ollama for a local model (optional)

Ollama serves open-weight models from the server's CPU. Here it runs as a container in the same Compose project, so n8n reaches it by service name and no port opens to the internet. How to Install Ollama on Ubuntu 24.04 covers a standalone install.

Add this service under services: in your compose file:

  ollama:
    image: ollama/ollama:0.35.1
    restart: unless-stopped
    environment:
      - OLLAMA_CONTEXT_LENGTH=8192
    volumes:
      - ollama_data:/root/.ollama

Add the volume under the top-level volumes: key, next to your existing volumes:

  ollama_data:

Start the container and pull a model with tool support:

docker compose up -d ollama
docker compose exec ollama ollama pull qwen3.5:4b

Without a GPU, Ollama gives models a 4,096-token context. OLLAMA_CONTEXT_LENGTH=8192 doubles it for the system message, tool definitions and chat history, and Ollama drops the oldest messages when the input runs longer. Ollama's documentation recommends at least 64,000 tokens for agents; raise the value only if the server has RAM to spare, because a longer context needs more memory. Check that n8n reaches Ollama:

docker compose exec n8n node -e "fetch('http://ollama:11434/api/version').then(r => r.text()).then(console.log)"

The output looks similar to this:

{"version":"0.35.1"}

In n8n, disconnect the API chat model, select the plus sign under Chat Model, and pick Ollama Chat Model. Create an Ollama credential with the Base URL http://ollama:11434 and leave API Key empty. Pick qwen3.5:4b from the Model list. Under Options, add Enable Thinking and turn it off: Qwen3.5 otherwise writes reasoning tokens before each reply, which adds CPU time. To keep the API model as a backup, turn on Enable Fallback Model in the AI Agent and connect the API model to the Fallback Model connector.

On 4 vCPU, a reply that fills the 8K context takes longer than the 43 seconds measured on 16 vCPU, and each tool call adds a second model round. For larger models by RAM size, see Best Ollama Models for CPU Servers.

Use the local model for on-demand chats: each message loads the CPU for the length of one reply, then the server idles, and Ollama unloads the model after 5 idle minutes. Sustained full-CPU load is not permitted on Arct plans, so continuous or batch generation belongs on an API model.

Step 9: Publish and test the agent

Select Publish in the canvas header, then Publish in the dialog. Publishing registers the production webhook with Telegram. Send your bot two messages:

What is 18% of 2,340?
What is the weather in Lisbon right now?

Open the workflow's Executions tab to see which tool each run called. Then ask "What did I ask first?". A reply that names the percentage question confirms the memory.

Check the webhook from the server:

read -rs BOT_TOKEN
printf 'url = "https://api.telegram.org/bot%s/getWebhookInfo"\n' "$BOT_TOKEN" | curl -s -K -
unset BOT_TOKEN

Paste the token after the first command and press Enter. read -rs keeps the token out of your shell history, and curl -K - reads the URL from standard input, so the token does not appear in the process list. The output looks similar to this:

{"ok":true,"result":{"url":"https://n8n.example.com/webhook/YOUR_WEBHOOK_ID/webhook","has_custom_certificate":false,"pending_update_count":0,
...

A url under /webhook/, not /webhook-test/, shows the published workflow holds the webhook. When delivery fails, a last_error_message field names the cause. Check the stored history:

docker compose exec postgres psql -U agent_memory -d agent_memory -c "SELECT session_id, count(*) FROM n8n_chat_histories GROUP BY session_id;"

The output lists your Telegram chat ID with a message count above zero.

Troubleshoot common n8n AI agent errors

SymptomCauseFix
Bad Request: bad webhook: An HTTPS URL must be provided for webhookN8N_WEBHOOK_URL is unset or uses http://Set N8N_WEBHOOK_URL to the HTTPS address in compose.yaml, then run docker compose up -d
The "text" parameter is empty.The message was a photo or sticker without textSend text, or add an If node that checks message.text
Bad Request: can't parse entitiesThe reply has an unmatched * or _, and the Telegram node parses Markdown (Legacy)Keep the plain-text rule in the system message
connect ECONNREFUSED on port 11434The Ollama credential still has the default Base URL http://localhost:11434, which inside the n8n container points at n8nSet the Base URL to http://ollama:11434
does not support toolsThe Ollama model has no tool callingPull a model with the tools tag, such as qwen3.5:4b

FAQ

What are n8n AI agents?

An n8n AI agent is a workflow built on the AI Agent node, which connects a chat model, optional memory and at least one tool. The model reads each message and decides which tools to call, and the node passes its final reply to the next step.

Can I use n8n for free?

Yes. The self-hosted Community edition is free to use under the Sustainable Use License for internal business, personal and non-commercial use. The source code is public on GitHub as fair-code.

How does memory work in an n8n AI agent?

A memory sub-node stores past messages under a session key and sends the most recent exchanges to the model, 5 by default. Simple Memory keeps them inside the n8n process and drops a session after 1 hour idle. Postgres Chat Memory stores them in a table that survives restarts.

Can n8n use Ollama for AI agents?

Yes. Connect the Ollama Chat Model node to the AI Agent and set the credential's Base URL to your Ollama server, such as http://ollama:11434 when both run in one Docker Compose project. The model must support tool calling.

Which Ollama model works with the n8n AI Agent node on a CPU?

Qwen3.5 4B is a 3.4 GB download with tool support that peaked at 4.18 GiB of memory. On a 16 vCPU benchmark server it answered an 8K prompt in 43 seconds, generating 12.23 tokens/s. Qwen3.5 9B peaked at 7.28 GiB and generated 7.35 tokens/s on the same server.

Why does my n8n Telegram trigger not receive messages?

Telegram delivers updates only to an HTTPS webhook, so n8n needs N8N_WEBHOOK_URL set to its public HTTPS address. Telegram also keeps one webhook per bot, so a test run and a published workflow cannot both receive messages.

Next steps

Arct Cloud plans with 4 vCPU and 8 GB RAM are Cost Optimized cvm.small, General Purpose vm.tiny and High Performance hvm.tiny, each with NVMe storage, and the 16 GB sizes add headroom for the Ollama option. Compare plans.

This work is licensed under CC BY-NC-SA 4.0.