> For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://unsloth.ai/docs/jp/moderu/glm-5.3-flash.md).

# GLM-5.3-Flash: ローカルでの実行方法

Z.aiによる新しいGLM-5.3-Flash、別名ox-alphaモデルを実行します。

GLM-5.3-Flashは、別名 **`ox-alpha`**、Z.ai の新しい 320B パラメータ（18B アクティブ）のマルチモーダル・オープンモデルで、 **上回ります** [GLM-5.2](/docs/jp/moderu/glm-5.2.md)。GLM-5.3-Flash は、次のより小さい版です [GLM-5.3](/docs/jp/moderu/glm-5.3.md) であり、次と競合します **Claude Opus 4.8** のコーディングおよびエージェント系ベンチマークにおいて。今なら 1-bit モデルを 102GB の RAM/VRAM、または 3-bit を 128GB 構成で llama.cpp または [Unsloth](https://github.com/unslothai/unsloth)を使ってローカルで実行できます。初日アクセスを提供してくれた Z.ai に感謝します。

Unsloth の動的 **1-bit** （93GB）GGUF は **トップ1%精度の 71% を維持しながら** 、 **85% 小さく** なっています。BF16（642GB）と比べて。動的 3-bit は 76% 小さく、87% の精度を維持します。

​​<a href="/pages/1e2ff4c2c3ab2b8bf7abf958679105509f3070d3#run-glm-5.3-flash-ox-alpha-locally" class="button primary">GLM-5.3-Flash 実行ガイド</a><a href="https://unsloth.ai/download" class="button secondary">Unsloth をダウンロード</a>

{% hint style="success" %}
**9月4日：** GLM-5.3-Flash は現在、 [**3.3倍高速な推論で**](#faster-inference-and-mtp-support)**!**
{% endhint %}

{% columns %}
{% column width="50%" %}
GLM-5.3-Flash は 30T トークンで学習され、新たに学習されたベースモデル上に構築されています。ハイブリッドな疎注意機構と線形注意機構により、精度を損なうことなく長文コンテキストの提供コストを下げています。

今ならモデルを直接 [Unsloth Desktop](#run-glm-5.3-flash-in-unsloth).
{% endcolumn %}

{% column width="50%" %}

<figure><img src="https://735611837-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxhOjnexMCB3dmuQFQ2Zq%2Fuploads%2FqPVJA6BHcCJIkQQrDvpD%2FScreenshot%202026-08-27%20at%207.38.11%E2%80%AFAM.png?alt=media&amp;token=576a6810-3f32-4e4c-a0d8-8a84cb733a52" alt=""><figcaption></figcaption></figure>
{% endcolumn %}
{% endcolumns %}

### :gear: 使用ガイド

#### GLM-5.3-Flash の要件：

最小の 1-bit 量子化は 100GB RAM で動作し、3-bit は Mac や NVIDIA DGX Spark のような 128GB デバイスで動作します。\
**表：ハードウェア要件** （単位 = 総メモリ：RAM + VRAM、または統合メモリ）

| 1-bit  | 2-bit  | 3-bit      | 4-bit      | 8-bit  | BF16   |
| ------ | ------ | ---------- | ---------- | ------ | ------ |
| 100 GB | 115 GB | 128-150 GB | 162-210 GB | 350 GB | 650 GB |

### 推奨設定

GLM-5.3-Flash には **3つの思考モード**があります：Low、High、Max です。複雑なタスクには Max Thinking を使ってください。 [Unsloth](#run-glm-5.2-in-unsloth-studio)では、チャット領域のトグルで Low、High、Max Thinking を簡単に選択できます。

ほとんどのユースケースでは、次の設定を使ってください：

| デフォルト設定（ほとんどのタスク）   | DeepSWE              |
| ------------------- | -------------------- |
| `temperature` = 1.0 | `temperature` = 0.95 |
| `top_p` = 0.95      | `top_p` = 1.0        |

* **最大コンテキストウィンドウ：** `1,048,576`.

#### 推論強度の変更

GLM-5.3-Flash はデフォルトで Max 推論を使用します。また、 `reasoning_effort` が「low」「high」「max」のいずれかになる推論強度もサポートしています。

### 高速化された推論と MTP サポート

9月4日時点で、私たちは初日版にいくつかの改善と最適化を追加しました [llama.cpp の PR](https://github.com/ggml-org/llama.cpp/pull/27754)。より高速なデコードパスと追加の MTP サポートを実装し、最大で **3.3倍高速な推論** を長いコンテキスト長でも実現しました！

すべてが [Unsloth Desktop](#run-glm-5.3-flash-in-unsloth)でそのまま使えます。必要に応じて最新バージョンに更新するだけです。追加のモジュールや MTP ファイルは不要です。あるいは、次の [llama.cpp](#run-glm-5.3-flash-in-llama.cpp) ガイドに従うこともできます。

<div><figure><img src="https://735611837-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxhOjnexMCB3dmuQFQ2Zq%2Fuploads%2FRKSzhmcluWnoNwUxFS40%2Fimage.png?alt=media&amp;token=872cd9b0-e968-4a0d-977c-5315d0ca3e49" alt=""><figcaption></figcaption></figure> <figure><img src="https://735611837-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxhOjnexMCB3dmuQFQ2Zq%2Fuploads%2F2tUYioUQhnyCv1pJTfld%2Fimage.png?alt=media&amp;token=b8bfd0de-22c8-4f30-8bfd-36656ae3c16e" alt=""><figcaption></figcaption></figure></div>

1xB200 上で GLM-5.3-Flash UD-IQ1\_S を使用し、まず MTP を無視すると、次の結果になります：

| テスト          | ベースライン tok/s | 最適化後 tok/s |
| ------------ | -----------: | ---------: |
| pp512        |      1121.80 |     1122.0 |
| tg32         |        62.79 |      63.10 |
| tg32 @ 4096  |        53.52 |      59.50 |
| tg32 @ 16384 |        41.02 |      57.99 |
| tg32 @ 65536 |        20.66 |      48.99 |

その後 MTP を追加すると、特に長いコンテキストでさらに大きな効果が得られます。ただし、ドラフトトークンが増えるほど推論が遅くなるため、n=2 付近で止めるべきです。

| プロンプト | MTP オフ |  n=2 |  n=3 |  n=5 |
| ----- | -----: | ---: | ---: | ---: |
| 4096  |   58.6 | 86.5 | 80.2 | 63.7 |
| 16K   |   55.0 |      | 77.2 |      |

短いコンテキスト長でも、最大 1.6倍の高速化ながら速度向上が見られます。

<div><figure><img src="https://735611837-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxhOjnexMCB3dmuQFQ2Zq%2Fuploads%2FmX36IJYj83OgfeQbCjh5%2Fimage.png?alt=media&amp;token=56e0ea43-df0b-47c4-8551-849921ef3913" alt=""><figcaption></figcaption></figure> <figure><img src="https://735611837-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxhOjnexMCB3dmuQFQ2Zq%2Fuploads%2FCebn6JiIDQpQMSOz8nIX%2Fimage.png?alt=media&amp;token=1f76ea17-4970-45ae-b92e-3abce29628d7" alt=""><figcaption></figcaption></figure></div>

### 📈 量子化分析

GLM-5.3-Flash を UD-IQ1\_S 1bit（93.09GB）まで量子化すると、BF16（641.64GB）と比べて 85% 小さくなりながら、トップ1%精度の 71% を維持します

動的 2-bit の UD-Q2\_K\_XL は 109GB で、83% 小さく、精度の 78% を維持します。\
動的 3-bit の UD-IQ3\_XXS は 120GB で、81% 小さく、精度の 82% を維持します。\
動的 4-bit の UD-Q4\_K\_XL は 200GB で、69% 小さく、精度の 93% を維持します。

<div><figure><img src="https://735611837-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxhOjnexMCB3dmuQFQ2Zq%2Fuploads%2FHB3bFxI8Cm3IOw3CRzvR%2Fglm53_flash_dynamic_ggufs_top1_accuracy_new_data.png?alt=media&amp;token=49882cca-1643-4e81-ad64-9d53c75ab93a" alt=""><figcaption></figcaption></figure> <figure><img src="https://735611837-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxhOjnexMCB3dmuQFQ2Zq%2Fuploads%2FthGU2pMvyHLSsnOal90W%2Fglm53_flash_dynamic_ggufs_kld_benchmarks_new_data.png?alt=media&amp;token=204f70ee-cae2-4a63-b958-d4ec515f4ce3" alt=""><figcaption></figcaption></figure></div>

| 量子化          | サイズ    | top-1 精度 | 平均 KLD   | KLD 99.9% |
| ------------ | ------ | -------- | -------- | --------- |
| UD-IQ1\_S    | 93.09  | 70.89%   | 0.669714 | 9.1658    |
| UD-IQ1\_M    | 97.58  | 73.06%   | 0.572413 | 8.5069    |
| UD-IQ2\_XXS  | 101.84 | 76.30%   | 0.450148 | 7.5764    |
| UD-Q2\_K\_XL | 108.72 | 78.34%   | 0.380134 | 6.8412    |
| UD-IQ3\_XXS  | 120.37 | 81.63%   | 0.283772 | 5.9611    |
| UD-Q3\_K\_XL | 147.54 | 86.25%   | 0.159697 | 4.0281    |
| UD-IQ4\_XS   | 156.82 | 88.18%   | 0.116652 | 3.1014    |
| UD-Q4\_K\_XL | 199.71 | 92.22%   | 0.049294 | 1.4894    |
| UD-Q5\_K\_XL | 240.31 | 94.35%   | 0.027052 | 0.8696    |
| UD-Q6\_K\_XL | 291.83 | 95.23%   | 0.019007 | 0.6267    |

## GLM-5.3-Flash (Ox-Alpha) をローカルで実行

今なら、GLM-5.3-Flash (Ox-Alpha) を Unsloth Desktop と llama.cpp で、私たちの [専用 PR](https://github.com/ggml-org/llama.cpp/pull/27754)を使って実行できます。 `UD-IQ3_XXS` デモでは 128GB デバイスに収まるため 3-bit を使用しています。量子化タイプは自由に変更してください。

* Hugging Face： [GLM-5.3-Flash-GGUF](https://huggingface.co/unsloth/GLM-5.3-Flash-GGUF)

<a href="/pages/1e2ff4c2c3ab2b8bf7abf958679105509f3070d3#run-glm-5.3-flash-in-unsloth" class="button primary">Unsloth Desktop で実行</a><a href="/pages/1e2ff4c2c3ab2b8bf7abf958679105509f3070d3#run-glm-5.3-flash-in-llama.cpp" class="button secondary">llama.cpp で実行</a>

### 🦥 Unsloth で GLM-5.3-Flash を実行

GLM-5.3-Flash は現在、 [Unsloth Desktop](#run-qwen3.8-in-unsloth-desktop)ローカル AI 向けのオープンソース UI アプリです。 **Unsloth は RAM へ自動的にオフロードし、マルチ GPU 構成を検出します**。Unsloth Desktop を使えば、モデルをローカルで **MacOS、Windows**、Linux で実行でき、さらに：

{% columns %}
{% column %}

* 検索、ダウンロード、 [GGUF を実行](/docs/jp/xin-zhe/studio.md#run-models-locally)、MLX および safetensor モデル
* [**自己修復** ツール呼び出し](/docs/jp/xin-zhe/studio/chat.md#auto-healing-tool-calling) + **ウェブ検索**
* [**コード実行**](/docs/jp/desktop.md#code-execution) （Python、Bash）
* [自動推論](https://unsloth.ai/docs/desktop#feature-deep-dive) パラメータ調整（temp、top-p など）
* MLX と llama.cpp による高速な CPU + GPU 推論
* [LLM を学習](/docs/jp/xin-zhe/studio.md#no-code-training) VRAM を 70% 削減し、2倍高速
  {% endcolumn %}

{% column %}

<figure><img src="https://735611837-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxhOjnexMCB3dmuQFQ2Zq%2Fuploads%2F6IXaXdTVyvbrnjehlxys%2Fkimik3.gif?alt=media&amp;token=31e1213b-d7da-46e9-bc7f-3a8c402513fc" alt=""><figcaption></figcaption></figure>
{% endcolumn %}
{% endcolumns %}

{% stepper %}
{% step %}

#### Unsloth をインストール

始める最も簡単な方法は、 [Unsloth Desktop アプリ](/docs/jp/desktop.md)をダウンロードすることです。対応環境： [macOS](/docs/jp/meru/install/mac.md), [Windows](/docs/jp/meru/install/windows-installation.md)、 [Linux](/docs/jp/meru/install/linux.md).

<a href="https://unsloth.ai/download" class="button primary" data-icon="down-to-bracket">Unsloth をダウンロード</a>

* <i class="fa-apple">:apple:</i> [macOS 用をダウンロード](https://unsloth.ai/download/mac)
* <i class="fa-windows">:windows:</i> [Windows 用をダウンロード](https://unsloth.ai/download/windows)
* <i class="fa-linux">:linux:</i> [Linux 用をダウンロード](https://unsloth.ai/download/linux)

または、手動でインストールしたい場合は：

MacOS、Linux、WSL：

```bash
curl -fsSL https://unsloth.ai/install.sh | sh
```

Windows PowerShell：

```bash
irm https://unsloth.ai/install.ps1 | iex
```

{% endstep %}

{% step %}

#### GLM-5.3-Flash を検索してダウンロード

移動先： [Unsloth Chat](/docs/jp/xin-zhe/studio/chat.md) または Model hub で検索バーに GLM-5.3-Flash を入力し、目的のモデルと量子化をダウンロードしてください。

<figure><img src="https://735611837-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxhOjnexMCB3dmuQFQ2Zq%2Fuploads%2FvOQkVli16S26FygKi2gw%2FScreenshot%202026-08-27%20at%204.18.11%E2%80%AFAM.png?alt=media&amp;token=df7914de-597e-43b9-8147-d568ccca0a51" alt=""><figcaption></figcaption></figure>
{% endstep %}

{% step %}

#### GLM-5.3-Flash を実行

Unsloth を使うと推論パラメータは自動設定されるはずですが、手動で変更することもできます。コンテキスト長、チャットテンプレート、その他の設定も編集できます。

詳細は、以下をご覧ください： [Unsloth 推論ガイド](/docs/jp/xin-zhe/studio/chat.md)。以下は 1-bit 実行例：

<figure><img src="https://735611837-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxhOjnexMCB3dmuQFQ2Zq%2Fuploads%2FwmOvZRaD61XCmMDBlHOD%2FScreenshot%202026-08-27%20at%207.08.52%E2%80%AFAM.png?alt=media&amp;token=4980f172-4323-43a0-9338-fe0e2ca749b6" alt=""><figcaption></figcaption></figure>
{% endstep %}

{% step %}

#### Unsloth API で GLM-5.3-Flash を提供

次を使用できます `unsloth run` コマンドを使い、次を利用して API 経由で GLM-5.3-Flash を提供できます `llama-server` コンテキストサイズ、GPU レイヤー、スレッディング、サンプリング、ネットワーキング、ツール設定などのランタイムフラグ。詳細は以下をご覧ください： [API ドキュメント](/docs/jp/ji-ben/api.md) または [unsloth start](/docs/jp/lian-xie/unsloth-start.md).

{% code overflow="wrap" %}

```bash
unsloth run --model unsloth/GLM-5.3-Flash-GGUF:UD-IQ3_XXS
```

{% endcode %}
{% endstep %}

{% step %}

#### Unsloth の準備ができました

Unsloth Desktop 経由で GLM-5.3-Flash を使って、他にも次のようなことができます：

* **ツール連携：** [Claude Code](/docs/jp/ji-ben/claude-code.md), [Codex](/docs/jp/ji-ben/codex.md), [ウェブ検索](/docs/jp/xin-zhe/studio/chat.md#advanced-web-search), [MCP](/docs/jp/ji-ben/mcp.md) など
* **モデルの学習：** テキスト、拡散モデル、 [埋め込み](/docs/jp/ji-ben/embedding-finetuning.md)など
* **メディア生成：** 作成して学習 [画像](/docs/jp/ji-ben/diffusion-image.md)、動画、 [TTS](/docs/jp/ji-ben/text-to-speech-tts-fine-tuning.md) をローカルで

<figure><img src="https://735611837-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxhOjnexMCB3dmuQFQ2Zq%2Fuploads%2FNYiWxq5OX7NdPoqxp2Hu%2FScreenshot%202026-08-27%20at%2011.59.01%E2%80%AFPM.png?alt=media&amp;token=d9311caa-6935-47b3-a767-5c48e92b7c25" alt=""><figcaption></figcaption></figure>
{% endstep %}
{% endstepper %}

### :llama: llama.cpp で GLM-5.3-Flash を実行

{% stepper %}
{% step %}
専用の llama.cpp PR を使用する必要があります [ここ](https://github.com/unslothai/llama.cpp/pull/61)。以下のビルド手順に従うこともできます。次を変更してください `-DGGML_CUDA=ON` を `-DGGML_CUDA=OFF` GPU がない場合、または CPU 推論のみを使いたい場合は。 **Apple Mac / Metal デバイスでは**、次に `-DGGML_CUDA=OFF` と設定し、その後は通常どおり続行してください。Metal サポートはデフォルトで有効です。

```bash
apt-get update
apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y
git clone --branch glm5next/upstream https://github.com/unslothai/llama.cpp
cmake llama.cpp -B llama.cpp/build \
    -DBUILD_SHARED_LIBS=OFF -DGGML_CUDA=ON
cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-mtmd-cli llama-server llama-gguf-split
cp llama.cpp/build/bin/llama-* llama.cpp
```

{% endstep %}

{% step %}
モデルを実行するには、次のようにします：

{% code overflow="wrap" %}

```bash
pip install -U "huggingface_hub[cli]"
hf download unsloth/GLM-5.3-Flash-GGUF \
    --local-dir unsloth/GLM-5.3-Flash-GGUF \
    --include "*UD-IQ3_XXS*" # 2-bit には "*IQ2_XXS*" を使用
```

{% endcode %}
{% endstep %}

{% step %}
その後、実行します：

{% code overflow="wrap" %}

```bash
./llama.cpp/llama-cli \
    --model unsloth/GLM-5.3-Flash-GGUF/UD-IQ3_XXS/GLM-5.3-Flash-UD-IQ3_XXS-00001-of-00004.gguf \
    --temp 1.0 \
    --top-p 0.95 \
    --chat-template-kwargs '{"reasoning_effort":"max"}'
```

{% endcode %}

次に置き換えてください `UD-IQ3_XXS` お好みの量子化に、例えば `IQ2_XXS` を、アップロード後は 2-bit 用に。
{% endstep %}
{% endstepper %}

## 📊 ベンチマーク

<div><figure><img src="https://735611837-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxhOjnexMCB3dmuQFQ2Zq%2Fuploads%2FfxF1H1wdQe3vKBepuhdA%2Fimage.png?alt=media&amp;token=b6e590cc-fafb-43b3-8e8b-0b318200cbcb" alt=""><figcaption></figcaption></figure> <figure><img src="https://735611837-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxhOjnexMCB3dmuQFQ2Zq%2Fuploads%2FUy73TYPneV9NarmiUgJZ%2Fimage.png?alt=media&amp;token=8b344785-cb9b-4c61-bf55-4fdc77be1d0a" alt=""><figcaption></figcaption></figure></div>

| ベンチマーク                              | GLM-5.3-Flash | GLM-5.2 | DeepSeek-V4-Vision-Exp | Opus 4.8 | GPT-5.6 Terra | Gemini 3.7 Flash |
| ----------------------------------- | ------------- | ------- | ---------------------- | -------- | ------------- | ---------------- |
| コーディング                              |               |         |                        |          |               |                  |
| Terminal Bench 2.1                  | 84.3          | 81.0    | 83.9                   | 85.0     | 87.4          | 85.8             |
| <p>DeepSWE</p><p>v1.1</p>           | 63.4          | 46.2    | 59.3                   | 58.0     | 69.6          | 65.3             |
| NL2Repo                             | 56.3          | 48.9    | 57.7                   | 69.7     | -             | -                |
| エージェント                              |               |         |                        |          |               |                  |
| Toolathlon 検証済み                     | 78.4          | 59.9    | 75.9                   | 76.2     | 74.9          | -                |
| <p>AutomationBench</p><p>v1.0.6</p> | 48.8          | 26.2    | 38.8                   | 41.0     | 37.2          | 52.3             |
| Agents' Last Exam                   | 26.3          | 20.4    | 27.3                   | 27.0     | 28.0          | -                |
| ツール付き HLE                           | 55.3          | 54.7    | 55.1                   | 57.9     | -             | -                |
| GDPval-AA v2                        | 1773          | 1504    | 1675                   | 1582     | 1571          | 1527             |
| ビジョン                                |               |         |                        |          |               |                  |
| OfficeQA Pro                        | 62.4          | -       | 57.9                   | 48.9     | -             | -                |
| <p>CharXiv 推論</p><p>ツール付き</p>       | 89.4          | -       | 80.4                   | 89.9     | 88.0          | 88.7             |
| <p>Chartography</p><p>ツール付き</p>     | 78.0          | -       | 64.3                   | 75.0     | 68.0          | 65.0             |
| BabyVision                          | 53.4          | -       | 35.1                   | 46.8     | 61.6          | 70.9             |
| MVbench                             | 77.8          | -       | 69.4                   | 67.1     | 75.0          | 82.2             |
| MMVU                                | 80.5          | -       | 72.7                   | 67.4     | 75.8          | 82.3             |


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://unsloth.ai/docs/jp/moderu/glm-5.3-flash.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
