For the complete documentation index, see llms.txt. This page is also available as Markdown.

How to Run Local AI Models with Hermes Agent

Guide on using open LLMs with Hermes Agent locally.

This guide enables you to run open LLMs locally with Hermes Agent via Unsloth. Hermes Agent by Nous Research is an open-source autonomous AI agent that connects to a model endpoint, executes tasks, and improves over time through memory and learned skills.

Hermes will work with any local model exposed through Unsloth’s OpenAI-compatible API, including: DeepSeek, Qwen, Gemma, and more. Hermes acts as the agent client, while Unsloth loads and serves models via the local API entirely offline.

After setup, every prompt sent through Hermes will run using your local model on your device.

Qwen3.5 running locally in Hermes via Unsloth.

Setup Hermes🦥 Connect your local model

In this tutorial, you’ll install Hermes and configure it to use unsloth/Qwen3.6-27B-GGUF served from Unsloth. Prefer a different model? Swap in any other model by loading it in Unsloth and updating the configuration.

Setup Hermes Agent

Prerequisites:

The Hermes command-line installer supports Linux, macOS, and WSL2. Make sure Git is installed; on Linux, also install curl and xz-utils. The installer automatically provisions uv, Python 3.11, Node.js 22, ripgrep, and ffmpeg.

1. Run the installer

curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash

The installer:

  • Detects your platform and checks dependencies.

  • Clones Hermes to ~/.hermes/hermes-agent/.

  • Creates a Python virtual environment and installs the Python dependencies.

  • Installs the browser-tool dependencies and Playwright’s Chromium engine.

  • Adds the hermes command and launches the setup wizard.

Playwright may request sudo to install Chromium’s shared system libraries. Hermes itself does not require root access.

2. Reload your shell so the hermes command is on your PATH:

3. Verify the install:

If the command resolves, Hermes is installed. Everything lives under ~/.hermes/:

Path
What it is

~/.hermes/config.yaml

Main settings (model, provider, tools, TTS, …)

~/.hermes/.env

API keys and other secrets

~/.hermes/hermes-agent/

The Hermes source + virtualenv

~/.hermes/cron/, sessions/, logs/

Runtime data

~/.hermes/skills/

Installed skills (synced from the Skills Hub)

Full install reference: hermes-agent.nousresearch.com/docs/getting-started/installation. If the installer reports a missing prerequisite, install it and re-run the one-liner. The installer is idempotent.

⚡ Quickstart

After installing Hermes, we'll need to install Unsloth Desktop to enable Hermes to serve and run inference of local models.

1

Download Unsloth

The easiest way to get started is by installing the Unsloth Desktop app. It supports MacOS, Linux, Windows, NVIDIA, AMD, Intel and CPU setups.

Download Unsloth

Or, if you prefer manual installation:

MacOS, Linux, WSL:

Windows PowerShell:

2

Install

  1. Open the Unsloth installer (.dmg, .exe files)

  2. Drag Unsloth to Applications for Mac or complete setup for Windows.

  3. Launch the app and wait for installation to complete

3

Choose a model

Open 'Select model' dropdown ontop or 'Model hub' tab, choose a model and a quantization that fits your device, then download it. Once it finishes, start chatting - no setup required.

4

Unsloth is now ready

To start chatting, type a message and press Enter.

Connect Hermes. Run unsloth start hermes in your terminal while Unsloth is open. It mints an API key, writes the config, and launches Hermes against your loaded model.

⚡ Run Hermes Agent with unsloth start

To launch Hermes directly with a model, run:

Unsloth automatically selects the correct parameters to use but you can still change it.

With a model loaded in Unsloth Studio, run:

Hermes Agent connected to a local model through Unsloth Studio
Hermes Agent running through its Unsloth Studio provider.

Unsloth launches Hermes from a separate managed home with the Unsloth provider, model, and context settings already configured. Your existing Hermes setup is left unchanged.

This managed home is temporary by default. To keep your sessions and state, add --persist from your first launch:

To return to your latest session later, run:

To reopen a specific session, use --resume <session-id-or-title> instead.

See the complete unsloth start reference for model selection, remote connections, and advanced options.

The setup wizard below remains available if you prefer to manage the Hermes provider yourself.

🔑 Creating an API key

  1. Open the sidebar, click your Unsloth avatar at the bottom-left.

  2. Go to SettingsAPI.

  3. Enter a friendly name (e.g. hermes-agent-macbook).

  4. (Optional) Set an expiry.

  5. Click Create.

  6. Copy the key immediately. Unsloth stores only a hash and you won't be able to view it again.

All keys start with the sk-unsloth- prefix. Revoke a key from the same page at any time. Requests made with a revoked key will fail with 401 Unauthorized.

🦥 Integrate Hermes with Unsloth API

Hermes sends each chat turn to a configured inference provider and connects to OpenAI-compatible endpoints. Configure the provider during installation or later in the setup wizard.

1. Open the setup wizard:

Pick Model & Provider from the "What would you like to do?" menu to configure only the inference endpoint, or Full Setup to walk through everything (TTS, tools, messaging gateway, agent settings).

2. Select the Custom OpenAI-compatible endpoint when Hermes prompts you for an inference provider.

3. Fill in the prompts as Hermes walks through them:

Prompt
Value

API base URL

http://localhost:8888/v1 (your Unsloth port + /v1)

API key

Your sk-unsloth-… key

Detected model: … Use this model?

Y (Hermes auto-detects the model via GET /v1/models)

Context length in tokens

(leave blank for auto-detect)

Display name

Anything you like, e.g. unsloth-api

Hermes verifies the endpoint against /v1/models and confirms the detected model before continuing.

4. Accept defaults for the remaining prompts (TTS, tools, messaging gateway, agent settings) you can reconfigure any of them later. Hermes writes everything to ~/.hermes/config.yaml and ~/.hermes/.env.

5. Launch Hermes:

The startup banner shows your Unsloth model name in the status bar (e.g. unsloth/Qwen3.6-27B-GGUF), and the prompt is ready for input.

To reconfigure just the model later, run hermes setup model. To edit the config file directly, hermes config edit opens ~/.hermes/config.yaml in your $EDITOR.

Optional: tune the Unsloth server

unsloth run starts the local API server and loads a model for your app to connect to. You can also customize how the server behaves when starting it.

Use --reasoning off to turn thinking off, or --reasoning on to turn it on for models that support reasoning.

This starts the server on 0.0.0.0:8888, allowing other devices on your local network to connect. -p changes which port the server runs on. If you want phones, laptops, or other devices on your network to connect to the API server, start it with -H 0.0.0.0.

Some apps may still override generation settings for individual requests. For more advanced runtime configuration, see the main API tuning section.

Last updated

Was this helpful?