> For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://unsloth.ai/docs/zh/ji-chu/claude-code.md).

# 如何使用 Claude Code 运行本地 LLM

这份分步指南向你展示如何将开源 LLM 和 API 完全在本地连接到 Claude Code，并配有截图。可使用任何开源模型运行，例如 Qwen3.6、DeepSeek 和 Gemma。

{% columns %}
{% column width="58.333333333333336%" %}
在本教程中，我们将使用以下开源模型： [Gemma 4](/docs/zh/mo-xing/gemma-4.md) 和 [Qwen3.5](/docs/zh/mo-xing/qwen3.5.md) 它们是很强的智能体和代码模型（可在 24GB RAM/统一内存设备上运行）。

用于推理，我们将使用 [Unsloth Desktop](https://github.com/unslothai/unsloth) 和 [`llama.cpp`](https://github.com/ggml-org/llama.cpp) 它可让你在 macOS、Linux 和 Windows 上运行/提供 LLM 服务。你也可以使用任何其他模型。对于模型量化版本，我们使用 Unsloth [动态 GGUF](/docs/zh/ji-chu/dynamic-3.0-ggufs.md) ，以在保持准确性的同时运行任何量化 LLM。
{% endcolumn %}

{% column width="41.666666666666664%" %}

<figure><img src="/files/86e035acbaafcca2b57e4bb3f4191ad68bf30547" alt=""><figcaption><p>Claude Code 在本地使用 Qwen3.5 运行。</p></figcaption></figure>
{% endcolumn %}
{% endcolumns %}

<a href="/pages/1a707991086189a8e5cd8374f3ce1b81915bc159#claude-code-setup" class="button primary" data-icon="claude">Claude Code 设置</a><a href="/pages/1a707991086189a8e5cd8374f3ce1b81915bc159#quickstart-tutorials" class="button primary">📖 本地模型设置教程</a>

## <i class="fa-claude">:claude:</i> Claude Code 设置

在设置本地 LLM 之前，我们需要先安装 Claude Code。Claude Code 是一个基于终端的编程智能体，它能够理解你的代码库，并使用自然语言处理复杂的 Git 工作流。

{% tabs %}
{% tab title="macOS、Linux、WSL" %}

#### **安装 Claude Code：**

将以下内容粘贴到你的终端中以安装 Claude Code：

```bash
curl -fsSL https://claude.ai/install.sh | bash
```

安装完成后，进入你的项目文件夹。然后输入 `claude` 到 `shell` 开始。

```bash
cd ~/projects/my-project 
claude
```

{% endtab %}

{% tab title="Windows" %}

#### **安装 Claude Code：**

在 `PowerShell` 中安装 Claude Code：

```powershell
irm https://claude.ai/install.ps1 | iex
```

安装完成后，进入你的项目文件夹。然后输入 `claude` 到 `PowerShell` 开始。

<pre class="language-powershell"><code class="lang-powershell"><strong>cd /path/to/your/project
</strong>claude
</code></pre>

<div data-with-frame="true"><figure><img src="/files/7fbec167ff74cf9abb427048194317d761229e04" alt="" width="563"><figcaption></figcaption></figure></div>
{% endtab %}
{% endtabs %}

### :detective:修复 Claude Code 中 90% 更慢的推理

{% hint style="warning" %}
Claude Code 最近会在前面添加一个 Claude Code Attribution 标头，这会 **使 KV 缓存失效，导致本地模型的推理速度变慢 90%**.
{% endhint %}

归因信息是一行前置到 **系统提示词开头的内容** (`x-anthropic-billing-header: cc_version=...; cch=...;`) 其值会在每次请求时变化，因此整个提示前缀在每一轮都会错过 KV 缓存。

最简单的修复方式是在启动 Claude Code 时直接在命令行中禁用它，这样就不需要编辑文件：

{% code overflow="wrap" %}

```bash
claude --settings '{"env":{"CLAUDE_CODE_ATTRIBUTION_HEADER":"0","CLAUDE_CODE_ENABLE_TELEMETRY":"0"}}' --model unsloth/gemma-4-26B-A4B-it-GGUF
```

{% endcode %}

{% hint style="info" %}
较新的 Claude Code 版本也会遵循 `export CLAUDE_CODE_ATTRIBUTION_HEADER=0`；旧版本会忽略 shell 变量，因此上面的 `--settings` 形式（或下面的设置文件）是更可靠的选择。
{% endhint %}

要将其永久生效，请在 `CLAUDE_CODE_ATTRIBUTION_HEADER` 设为 0，并放入 `"env"` 中设置 `~/.claude/settings.json`中。例如，执行： `cat > ~/.claude/settings.json` 然后添加下面的内容（粘贴后按回车，再按 CTRL+D 保存）。如果你已有一个之前的 `~/.claude/settings.json` 文件，只需将 `"CLAUDE_CODE_ATTRIBUTION_HEADER" : "0"` 添加到 "env" 部分，并保持设置文件其余内容不变。

<pre class="language-json"><code class="lang-json">{
  "promptSuggestionEnabled": false,
  "env": {
    "CLAUDE_CODE_ENABLE_TELEMETRY": "0",
    "CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC": "1",
    <a data-footnote-ref href="#user-content-fn-1">"CLAUDE_CODE_ATTRIBUTION_HEADER" : "0"</a>
  },
  "attribution": {
    "commit": "",
    "pr": ""
  },
  "plansDirectory" : "./plans",
  "prefersReducedMotion" : true,
  "terminalProgressBarEnabled" : false,
  "effortLevel" : "high"
}
</code></pre>

## 📖 快速入门教程

{% columns %}
{% column %}
在开始之前，我们首先需要完成你将要使用的特定模型的设置。我们使用 [Unsloth](/docs/zh/xin/studio.md) （一个网页 UI）和 llama.cpp，它们是用于在你的 Mac、Linux、Windows 设备上运行和提供 LLM 服务的开源框架。

Unsloth 还具有独特的自我修复 [工具调用](/docs/zh/xin/studio/chat.md#auto-healing-tool-calling) 和 [网页搜索](/docs/zh/xin/studio/chat.md#code-execution) 能力。右侧可见连接到 Unsloth 的 Claude Code：
{% endcolumn %}

{% column %}

<div data-with-frame="true"><figure><img src="/files/c633f6e5a61522d2d7fa76b1c6c3376b956d223d" alt=""><figcaption></figcaption></figure></div>
{% endcolumn %}
{% endcolumns %}

<a href="/pages/1a707991086189a8e5cd8374f3ce1b81915bc159#connect-claude-code" class="button primary" data-icon="claude">连接 Claude Code</a><a href="/pages/1a707991086189a8e5cd8374f3ce1b81915bc159#unsloth-tutorial" class="button primary">🦥 Unsloth 教程</a><a href="https://unsloth.ai/docs/basics/claude-code#llama.cpp-tutorial" class="button primary"> llama.cpp 教程</a>

## 🦥 Unsloth 教程

在本教程中，我们将通过 UI 使用 [Unsloth](https://github.com/unslothai/unsloth)将本地模型提供/连接到 Claude Code。Unsloth 可在 Windows、WSL、Linux 和 MacOS 上运行。

{% columns %}
{% column %}

* 搜索、下载、 [运行 GGUF](/docs/zh/xin/studio.md#run-models-locally) 和 safetensor 模型
* [**自我修复** 工具调用](/docs/zh/xin/studio.md#execute-code--heal-tool-calling) + **网页搜索**
* [**代码执行**](/docs/zh/xin/studio.md#run-models-locally) （Python、Bash）
* [自动推理](https://unsloth.ai/docs/desktop#feature-deep-dive) 参数选择（temp、top-p 等）
* 通过 llama.cpp 实现快速 CPU + GPU 推理
* [训练 LLM](/docs/zh/xin/studio.md#no-code-training) 速度提升 2 倍，VRAM 减少 70%

安装说明见下：
{% endcolumn %}

{% column %}

<div data-with-frame="true"><figure><img src="/files/f8801eb887a6a996de4f745df66b4617f383d216" alt=""><figcaption><p>Unsloth 中运行的 Qwen3.6 2-bit 示例。</p></figcaption></figure></div>
{% endcolumn %}
{% endcolumns %}

{% stepper %}
{% step %}

#### 下载 Unsloth

最简单的入门方式是安装 [Unsloth Desktop](/docs/zh/desktop.md) 应用。它支持 [MacOS](/docs/zh/kuai-su-kai-shi/install/mac.md)、Linux、 [Windows](/docs/zh/kuai-su-kai-shi/install/windows-installation.md), [NVIDIA](/docs/zh/kuai-su-kai-shi/install/pip-install.md), [AMD](/docs/zh/kuai-su-kai-shi/install/amd.md)、Intel 和 CPU 配置。

<a href="https://unsloth.ai/download" class="button primary" data-icon="down-to-bracket">下载 Unsloth</a>

* <i class="fa-apple">:apple:</i> [下载 macOS 版本](https://unsloth.ai/download/mac)
* <i class="fa-windows">:windows:</i> [下载 Windows 版本](https://unsloth.ai/download/windows)
* <i class="fa-linux">:linux:</i> [下载 Linux 版本](https://unsloth.ai/download/linux)

或者，如果你更喜欢手动安装：

**MacOS、Linux、WSL：**

```bash
curl -fsSL https://unsloth.ai/install.sh | sh
```

**Windows PowerShell：**

```bash
irm https://unsloth.ai/install.ps1 | iex
```

{% endstep %}

{% step %}

#### 安装

1. 打开 Unsloth 安装程序（`.dmg`, `.exe` 文件）
2. 在 Mac 上将 Unsloth 拖到 Applications，或在 Windows 上完成安装。
3. 启动应用并等待安装完成
   {% endstep %}

{% step %}

#### 选择模型

打开顶部的“选择模型”下拉菜单或“模型中心”标签页，选择适合你设备的模型和量化方式，然后下载它。完成后，直接开始聊天——无需额外设置。

<figure><img src="/files/ee716919ab359455030e8acd4d55b5ec13ca328d" alt="" width="563"><figcaption></figcaption></figure>
{% endstep %}

{% step %}

#### Unsloth 已准备就绪

要开始聊天，输入消息并按 Enter。

* **连接工具：** [Claude Code](/docs/zh/ji-chu/claude-code.md), [Codex](/docs/zh/ji-chu/codex.md)、网页搜索、 [MCP](/docs/zh/ji-chu/mcp.md) 等等
* **训练模型：** 微调文本、扩散、嵌入等更多内容
* **生成媒体：** 在本地创建并训练图像、视频、TTS

<figure><img src="/files/b0dfb852351654c29812a5aa23885913432eaa10" alt="" width="563"><figcaption></figcaption></figure>
{% endstep %}
{% endstepper %}

### 模型加载 + API 指南

{% stepper %}
{% step %}

#### 选择模型

在使用 API 之前，请从 Chat 页面左上角的 **选择模型** 下拉菜单中加载一个模型。

<figure><img src="/files/e7c1a267b0c4f58689066eddfc57a2c2211f1e13" alt=""><figcaption></figcaption></figure>

在本指南中，我们将使用： `unsloth/gemma-4-26B-A4B-it-GGUF` 以及推荐的 `UD-Q4_K_XL` 量化。
{% endstep %}

{% step %}

#### 测试模型

在使用客户端之前，先发送一条简短消息：

<div data-with-frame="true"><figure><img src="/files/d1ef1d199c3aee2da86cc3da46a133801d2683ad" alt="" width="563"><figcaption></figcaption></figure></div>

{% hint style="info" %}
这可以确认模型已正确加载并已准备好响应。
{% endhint %}
{% endstep %}

{% step %}

#### **Unsloth API 密钥**

在 Unsloth 中打开 **设置 → API** 以查看或创建你的 API 密钥。

<figure><img src="/files/7d43e8763d1ae72290485151822b1ea2e4fce42a" alt=""><figcaption></figcaption></figure>

请像对待密码一样对待你的 API 密钥，不要在截图或仓库中暴露它。
{% endstep %}
{% endstepper %}

## ⚙️ 连接 Claude Code

现在我们已经为 Claude Code 设置好了本地 LLM，接下来配置 Claude Code 以与你的工具协同工作。你可以通过以下方式轻松连接 `unsloth start` 轻松连接，或者也可以 [手动](#connect-manually).

### ⚡ 使用 `unsloth start`

要直接用某个模型启动 Claude，请运行：

```bash
unsloth start claude \
    --model unsloth/qwen3.8-27B-GGUF-GGUF:UD-Q4_K_XL
    --temp 1.0 \
    --top-p 0.95 \
    --top-k 20 \
    --min-p 0.0 \
    --chat-template-kwargs '{"reasoning_effort":"medium"}'
```

Unsloth 会自动选择正确的参数，但你仍然可以更改。

在 Unsloth Studio 中加载好模型后，打开你的项目文件夹并运行：

```bash
unsloth start claude
```

<figure><img src="/files/23549bd18ccaa6adc6bd5383f5cc9fe161b45e92" alt="Claude Code connected to a local model through Unsloth Studio"><figcaption><p>Claude Code 正在使用 Unsloth Studio 中加载的模型运行。</p></figcaption></figure>

Unsloth 会为这次启动设置本地端点、API 密钥、模型和上下文长度。你原有的 Claude Code 配置不会受到影响。

Claude Code 已经会将对话保存在其正常会话存储中，因此 `--persist` 不需要。继续你最近的会话，使用：

```bash
unsloth start claude --continue
```

查看完整的 `unsloth start` 从命令行加载模型、远程 Unsloth 服务器和高级选项的参考。

本指南的其余部分涵盖手动 `llama.cpp` 设置。

#### 🔌 手动连接

如果你更喜欢手动设置，可以先设置以下环境变量。默认情况下，这些变量不会在会话之间保留。

{% tabs %}
{% tab title="macOS、Linux、WSL" %}
**配置：** 设置本地 API URL：

```bash
export ANTHROPIC_BASE_URL="http://localhost:8888"
```

从 Unsloth Studio → Settings → API 复制你的密钥（或者在你用 `unsloth run`启动它时从控制台获取，其中显示为 `sk-unsloth-...`），然后设置它。

{% code overflow="wrap" %}

```bash
export ANTHROPIC_AUTH_TOKEN="sk-unsloth-xxxxxxxxxxxx"
```

{% endcode %}

还要设置一个空的 `ANTHROPIC_API_KEY` ，这样 Claude Code 就不会提示你输入云端密钥：

```bash
export ANTHROPIC_API_KEY=""
```

可选：使用当前在 Unsloth 中加载的模型名称作为默认值。

```bash
export ANTHROPIC_MODEL="unsloth/gemma-4-26B-A4B-it-GGUF"
```

请使用它在 `GET http://localhost:8888/v1/models` 中显示的完整模型 ID `（与传递给`).
{% endtab %}

{% tab title="Windows" %}
**配置：** 在 Powershell 中设置本地 API URL：

```powershell
$env:ANTHROPIC_BASE_URL = "http://localhost:8888"
```

复制你的密钥，从 **Unsloth Studio → Settings → API**，然后设置：

```powershell
$env:ANTHROPIC_AUTH_TOKEN = "sk-unsloth-xxxxxxxxxxxx"
```

**可选：** 使用当前在 Unsloth 中加载的模型名称作为默认值。

```powershell
$env:ANTHROPIC_MODEL = "gemma-4-26B-A4B-it-GGUF"
```

{% hint style="info" %}
模型名称应是当前在 Unsloth Studio 中加载的模型。
{% endhint %}
{% endtab %}
{% endtabs %}

### 启动 Claude Code

使用当前在 Unsloth 中加载的模型启动 Claude Code。

我们将使用 `gemma-4-26B-A4B-it-GGUF`，但你可以使用任何与 Unsloth 兼容的模型。

```shellscript
claude --model unsloth/gemma-4-26B-A4B-it-GGUF
```

{% hint style="info" %}
为了进一步提升本地模型速度，你还可以使用 `--bare --exclude-dynamic-system-prompt-sections`。参见下面的“可选：缩减系统提示词”。
{% endhint %}

Claude Code 应该会打开并显示所选模型。

<figure><img src="/files/6d4bc21d09fd5681680ae40843b06b44cfd15323" alt=""><figcaption></figcaption></figure>

{% hint style="warning" %}
参见 [#fixing-90-slower-inference-in-claude-code](#fixing-90-slower-inference-in-claude-code "mention") 先修复由于 KV 缓存失效导致开源模型慢 90% 的问题。
{% endhint %}

尝试使用这个提示词来研究并排名高质量的 SFT 数据集：

{% code overflow="wrap" %}

```
你只能在 project/ 中工作。不要搜索 CLAUDE.md——这就是它。使用网页搜索在 Hugging Face 上找出 10 个真实的 instruction/chat/SFT 数据集，在研究过程中简要总结你的发现并解释每个数据集为何与 SFT 相关，然后创建 sft_report.md 作为一份精美的 Markdown 报告，包含排名、数据集名称、创建者、3–5 个相关标签、简短的通俗英文摘要，以及它对 SFT 有用的原因。保持内容简洁易读，不要有巨大的元数据倾倒、粘贴原始描述、过长的标签列表或无关数据集。任务在 sft_report.md 包含 10 条干净、写得很好的数据集条目时完成，并以：“Successfully finetuned a model with Unsloth!",
```

{% endcode %}

在你提交提示词后，智能体将搜索网页、评估结果并撰写最终报告。这可能需要几分钟。

某些工作流可能需要你批准操作或回复后续提示。

<figure><img src="/files/b62a6c3b23315b9c625654765c0c4e78269d7198" alt="" width="563"><figcaption></figcaption></figure>

{% hint style="info" %}
某些工作流可能需要你批准操作或回答后续提示。
{% endhint %}

完成后，生成的 `sft_report.md` 会看起来类似这样。

<figure><img src="/files/f00c03605b12f627ff1970d0d9d0721b7be2039c" alt="" width="375"><figcaption></figcaption></figure>

{% hint style="warning" %}
如果你看到 `Unable to connect to API (ConnectionRefused)` ，请记得取消设置 `ANTHROPIC_BASE_URL` 通过 `unset ANTHROPIC_BASE_URL`

如果你发现开源模型慢了 90%， [先看这里](#fixing-90-slower-inference-in-claude-code) 来修复 KV 缓存被失效的问题。
{% endhint %}

### 可选：缩减系统提示词

Claude Code 是为 Anthropic 的托管模型构建的，因此其默认系统提示词很大。在本地模型上，你可以在启动时添加两个标志来缩减它，以获得更快的响应并更好地复用 KV 缓存：

{% code overflow="wrap" %}

```shellscript
claude --model unsloth/gemma-4-26B-A4B-it-GGUF --bare --exclude-dynamic-system-prompt-sections
```

{% endcode %}

{% hint style="info" %}
`--bare` 会跳过对 hooks、skills、plugins、MCP 服务器和 CLAUDE.md 的自动发现（Claude 仍保留 Bash 和文件读写/编辑），并且 `--exclude-dynamic-system-prompt-sections` 会将每台机器的部分移出提示前缀。这两个选项都会缩短提示词并提升 KV 缓存复用，使本地模型明显更快。它们是可选的，不会改变上面的连接设置。
{% endhint %}

### 可选：调整 Unsloth 服务器

Claude Code 使用在 Unsloth 中运行的模型。你可以在启动服务器时自定义其行为。

```bash
# 面向编程智能体的服务：--disable-tools 会透传智能体自己的工具
unsloth run \\
  --model unsloth/gemma-4-26B-A4B-it-GGUF \\
  --disable-tools \\
  --reasoning off \\
  -p 8888
```

{% hint style="warning" %}
使用 `--disable-tools` 当驱动 Claude Code（或任何外部编程智能体）时。默认情况下，Unsloth Studio 运行自己的服务端工具，这会吞掉智能体的工具调用，所以 Claude Code 会回答，但从不编辑文件。 `--disable-tools` 切换为透传，因此会使用 Claude Code 自己的 Write/Edit/Bash 工具。
{% endhint %}

使用 `--reasoning off` 用于关闭思考，或者使用 `--reasoning on` 为支持推理的模型开启它。

```bash
# 在本地网络上暴露 API
unsloth run \\
  --model unsloth/gemma-4-26B-A4B-it-GGUF \\
  -H 0.0.0.0 \\
  -p 8888
```

这会在 `0.0.0.0:8888`上启动服务器，从而允许本地网络中的其他设备连接。

使用 `-p` 以更改服务器运行的端口。若你希望网络中的手机、笔记本或其他设备连接，请使用 `-H 0.0.0.0` 。

如需更高级的运行时配置，请参阅主 [API 调优](https://unsloth.ai/docs/basics/api#unsloth-run-command) 部分。

## 🦙 Llama.cpp 教程

在开始之前，我们首先需要完成你将要使用的特定模型的设置。我们使用 `llama.cpp` 它是一个开源框架，可在你的 Mac、Linux、Windows 等设备上运行 LLM。Llama.cpp 包含 `llama-server` 这使你能够高效地提供和部署 LLM。模型将通过 8001 端口提供服务，所有智能体工具都通过单一的 OpenAI 兼容端点路由。

#### Qwen3.5 教程

我们将使用 [Qwen3.5](/docs/zh/mo-xing/qwen3.5.md)-35B-A3B 和特定设置，以实现快速且准确的编码任务。如果你的 VRAM 不足，并且想要一个 **更聪明的** 模型， **Qwen3.5-27B** 是个不错的选择，但它大约会慢 2 倍；或者你也可以使用其他 Qwen3.5 变体，例如 9B、4B 或 2B。

{% hint style="info" %}
如果你想要一个 **更聪明的** 模型，或者你的 VRAM 不足，可以使用 Qwen3.5-27B。不过，它会比 35B-A3B 慢大约 2 倍。或者你也可以使用 [**Qwen3-Coder-Next**](/docs/zh/mo-xing/qwen3-coder-next.md) 如果你有足够的 VRAM，它会非常棒。
{% endhint %}

{% stepper %}
{% step %}

#### 安装 llama.cpp

我们需要安装 `llama.cpp` 来部署/提供本地 LLM 服务，以便在 Claude Code 等中使用。我们遵循官方构建说明，以获得正确的 GPU 绑定和最高性能。更改 `-DGGML_CUDA=ON` 更改为 `-DGGML_CUDA=OFF` 如果你没有 GPU，或者只想进行 CPU 推理。 **对于 Apple Mac / Metal 设备**，设置 `-DGGML_CUDA=OFF` 然后按常规继续——Metal 支持默认开启。

```bash
apt-get update
apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev git-all -y
git clone https://github.com/ggml-org/llama.cpp
cmake llama.cpp -B llama.cpp/build \
    -DBUILD_SHARED_LIBS=OFF -DGGML_CUDA=ON
cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-mtmd-cli llama-server llama-gguf-split
cp llama.cpp/build/bin/llama-* llama.cpp
```

<figure><img src="/files/675408caceb845def8b58ffd8fabf4e741a6a196" alt="" width="563"><figcaption></figcaption></figure>
{% endstep %}

{% step %}

#### 下载并在本地使用模型

通过以下方式下载模型： `huggingface_hub` 在 Python 中（安装后通过 `pip install huggingface_hub hf_transfer`）。我们使用 **UD-Q4\_K\_XL** 量化版本，以获得最佳的体积/准确率平衡。你可以在我们的 [此处合集](/docs/zh/kuai-su-kai-shi/unsloth-model-catalog.md)。如果下载卡住，请参阅 [Hugging Face Hub，XET 调试](/docs/zh/ji-chu/troubleshooting-and-faqs/hugging-face-hub-xet-debugging.md)

```bash
hf download unsloth/Qwen3.5-35B-A3B-GGUF \
    --local-dir unsloth/Qwen3.5-35B-A3B-GGUF \
    --include "*UD-Q4_K_XL*" # 对于动态 2bit 使用 "*UD-Q2_K_XL*"
```

<figure><img src="/files/9d9a42a31d027b31aaf6635b3710f3a53d4a6caa" alt=""><figcaption></figcaption></figure>

{% hint style="success" %}
我们使用了 `unsloth/Qwen3.5-35B-A3B-GGUF` ，但你也可以使用其他变体，如 27B，或其他模型，例如 `unsloth/`[`Qwen3-Coder-Next`](/docs/zh/mo-xing/qwen3-coder-next.md)`-GGUF`.
{% endhint %}

<figure><img src="/files/56f8dbb1752bda89c4b66bc6ba2989c83811280a" alt="" width="563"><figcaption></figcaption></figure>
{% endstep %}

{% step %}

#### 启动 Llama-server

要为智能体工作负载部署 Qwen3.5，我们使用 `llama-server`。我们应用 [Qwen 推荐的采样参数](/docs/zh/mo-xing/qwen3.5.md#recommended-settings) 用于思考模式： `temp 0.6`, `top_p 0.95` , `top-k 20`。请注意，如果你使用非思考模式或其他任务，这些数值会变化。

在新的终端中运行此命令（使用 `tmux` ，或者打开一个新终端）。下面的配置应该 **可完美装入一块 24GB GPU（RTX 4090）（占用 23GB）** `--fit on` 上自动卸载，但如果你看到性能不佳，请降低 `--ctx-size` .

{% hint style="info" %}
我们使用了 `--cache-type-k q8_0 --cache-type-v q8_0` 用于 KV 缓存量化，以减少 VRAM 占用。对于全精度，使用 `--cache-type-k bf16 --cache-type-v bf16` 。注意，在某些机器上 bf16 KV Cache 可能会稍慢一些。
{% endhint %}

```bash
./llama.cpp/llama-server \
    --model unsloth/Qwen3.5-35B-A3B-GGUF/Qwen3.5-35B-A3B-UD-Q4_K_XL.gguf \
    --alias "unsloth/Qwen3.5-35B-A3B" \
    --temp 0.6 \\
    --top-p 0.95 \
    --top-k 20 \
    --min-p 0.00 \
    --port 8001 \
    --kv-unified \
    --cache-type-k q8_0 --cache-type-v q8_0
```

{% hint style="success" %}
你也可以为 Qwen3.5 关闭思考，这可以提高智能体编程任务的性能。要在 llama.cpp 中关闭思考，请将以下内容添加到 llama-server 命令：

`--chat-template-kwargs "{\"enable_thinking\": false}"`

<img src="/files/122fbec0dd1c1450b0f18c87633b0cabb192a69b" alt="" data-size="original">
{% endhint %}
{% endstep %}
{% endstepper %}

### 使用 llama-server 启动 Claude Code

{% hint style="success" %}
我们使用了 `unsloth/GLM-4.7-Flash-GGUF` ，但你可以使用任何类似的内容，例如 `unsloth/Qwen3.6-27B-GGUF`.
{% endhint %}

{% hint style="warning" %}
参见 [#fixing-90-slower-inference-in-claude-code](#fixing-90-slower-inference-in-claude-code "mention") 先修复由于 KV 缓存失效导致开源模型慢 90% 的问题。
{% endhint %}

进入你的项目文件夹（`mkdir project ; cd project`）并运行：

```bash
claude --model unsloth/GLM-4.7-Flash
```

要使用 Qwen3.6-35B-A3B，只需改成：

```bash
claude --model unsloth/Qwen3.6-35B-A3B
```

<div data-with-frame="true"><figure><img src="/files/9895a3039484599f93e699644a0bf20622f32921" alt="" width="563"><figcaption></figcaption></figure></div>

要让 Claude Code 在不经过任何批准的情况下执行命令，请执行 **（注意：这会让 Claude Code 在未经任何批准的情况下随意执行和运行代码！）**

{% code overflow="wrap" %}

```bash
claude --model unsloth/GLM-4.7-Flash --dangerously-skip-permissions
```

{% endcode %}

试试这个提示，以安装并运行一个简单的 Unsloth 微调：

{% code overflow="wrap" %}

```
你只能在当前工作目录 project/ 中工作。不要搜索 CLAUDE.md——这就是它。通过 uv 使用虚拟环境安装 Unsloth。若可能，请先使用 `python -m venv unsloth_env`，然后执行 `source unsloth_env/bin/activate`。请查看 https://unsloth.ai/docs/get-started/install/pip-install 了解方法（获取并阅读）。然后按照 https://github.com/unslothai/unsloth 中描述的方式进行一次简单的 Unsloth 微调运行。你可以使用 1 块 GPU。
```

{% endcode %}

<div data-with-frame="true"><figure><img src="/files/a3431c9353badb6bb35de75882b82d4a13ca2b49" alt="" width="563"><figcaption></figcaption></figure></div>

稍等片刻后，Unsloth 将通过 uv 安装到虚拟环境中，并完成加载：

<div data-with-frame="true"><figure><img src="/files/6ba0cfee745ccd165528bf1db06ec5ad1ab35122" alt="" width="563"><figcaption></figcaption></figure></div>

最后你将看到一个使用 Unsloth 成功微调的模型！

<div data-with-frame="true"><figure><img src="/files/4e98c753b70b30afb88ce8604503bcf46e1229ba" alt="" width="563"><figcaption></figcaption></figure></div>

{% hint style="warning" %}
如果你看到 `Unable to connect to API (ConnectionRefused)` ，请记得取消设置 `ANTHROPIC_BASE_URL` 通过 `unset ANTHROPIC_BASE_URL`

如果你发现开源模型慢了 90%， [先看这里](#fixing-90-slower-inference-in-claude-code) 来修复 KV 缓存被失效的问题。
{% endhint %}

[^1]: 必须使用这个！


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://unsloth.ai/docs/zh/ji-chu/claude-code.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
