> For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://unsloth.ai/docs/de/modelle/qwen3.8-next.md).

# Qwen3.8-Flash-Next: So führst du es lokal aus

Leitfaden zum lokalen Ausführen von Qwen3.8-Flash-Next.

Qwen3.8-Flash-Next ist ein neues Open-Weight-Modell, **125B-Parameter** multimodales MoE-Modell von Qwen. Auf der neuen Qwen4-Architektur aufgebaut, unterstützt es ein Kontextfenster von 262K und fortgeschrittenes Schlussfolgern. [Qwen3.8](/docs/de/modelle/qwen3.8.md)-Flash-Next übertrifft Claude-4.6-Opus (Max) und kann lokal auf Geräten mit **75 GB RAM**/Unified Memory ohne erforderlichen GPU-VRAM ausgeführt werden. Um das Modell auszuführen, verwenden Sie unsere [GGUFs](https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF) über llama.cpp oder [Unsloth Desktop](/docs/de/desktop.md). Vielen Dank an Qwen für den Zugriff ab Tag 0.

{% columns %}
{% column %}
**1-Bit sind 75 GB** und verwendet 4-Bit für den Ngram / PLE. Das ist **79 % kleiner** als BF16 (355 GB) und behält eine **Top-1%-Genauigkeit von 80 %**.

<a href="/pages/46b3f9c112a1132791a9a655fa0ca5625d2b54cf#run-qwen3.8-flash-next-in-unsloth" class="button primary">Qwen3.8-Flash ausführen</a><a href="https://unsloth.ai/download" class="button secondary">Unsloth herunterladen</a>

{% hint style="success" %}
[**MTP**](#mtp-guide) ist da! Führen Sie Qwen3.8-Flash 1,3- bis 1,7-mal schneller aus in [Unsloth Desktop](#run-qwen3.8-flash-next-in-unsloth)!
{% endhint %}
{% endcolumn %}

{% column %}

<figure><img src="https://797013937-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxhOjnexMCB3dmuQFQ2Zq%2Fuploads%2F1YneA51oWzaIiwep9H5I%2F1000024423.gif?alt=media&amp;token=40a169cc-12f1-403a-898c-b310f81ff52a" alt=""><figcaption><p>4-Bit Qwen3.8-Flash, ausgeführt in Unsloth</p></figcaption></figure>
{% endcolumn %}
{% endcolumns %}

### :gear: Nutzungsanleitung

Egal, ob Sie **Qwen3.8-Flash-Next** auf einer CPU mit Systemspeicher oder auf einer GPU mit VRAM ausführen, macht relativ wenig Unterschied. Seine einzigartige Architektur ermöglicht Inferenz mit RAM oder Unified Memory und erreicht eine Leistung, die näher an GPU-VRAM liegt als bei anderen Modellen üblich ist. Dadurch eignet es sich besonders gut für Macs, NVIDIA-DGX-Spark-Systeme und andere Geräte mit großen Speicherkapazitäten.

Sie benötigen mindestens **75 GB RAM oder Unified Memory** um das Modell auszuführen. Seine kleinste 1-Bit-quantisierte Version ist größer als üblich, da es neue Ngram-Schichten bzw. pro Schicht Embeddings gibt, die wie eine Nachschlagetabelle funktionieren. Das bedeutet jedoch auch, dass die Quantisierung weniger aggressiv ist, sodass das Modell mehr von seiner ursprünglichen Genauigkeit behält als stärker quantisierte Modelle. Sie können die PLE-/Ngram-Schicht auch auf die SSD auslagern und mmap verwenden, wodurch weniger CPU- und GPU-VRAM genutzt wird.

#### Anforderungen für Qwen3.8-Flash-Next:

Die kleinste Quantisierung funktioniert mit 75 GB RAM, daher ist ein Gerät mit 96 GB RAM/Unified Memory am besten.\
**Tabelle: Hardwareanforderungen** (Einheiten = Gesamtspeicher: RAM + VRAM oder Unified Memory)

<table><thead><tr><th>1-Bit</th><th>2-Bit</th><th>3-Bit</th><th>4-Bit</th><th width="128">5-Bit</th><th>8-Bit</th><th>BF16</th></tr></thead><tbody><tr><td>75 GB</td><td>79 GB</td><td>90 GB</td><td>96-114 GB</td><td>163 GB</td><td>200 GB</td><td>355 GB</td></tr></tbody></table>

{% hint style="info" %}
Wenn Sie [MTP](/docs/de/modelle/mtp.md) für schnellere Inferenz verwenden möchten, sollten Sie 1–2 GB zusätzlichen Spielraum einplanen.
{% endhint %}

### Empfohlene Einstellungen

Qwen3.8-Flash-Next ist ein **hybrider Denk-** Modell mit unterschiedlichen Standardeinstellungen für Denk- und Nicht-Denk-Modi. Extra High ist standardmäßig aktiviert. Wenn Sie also kürzere Denkspuren möchten, können Sie [den Denkaufwand anpassen](#thinking--preserve-thinking):

| Parameter            | Denkmodus | Instruct- (Nicht-Denk-) Modus |
| -------------------- | --------- | ----------------------------- |
| `Temperatur`         | 1.0       | 0.7                           |
| `top_p`              | 0.95      | 0.80                          |
| `top_k`              | 20        | 20                            |
| `min_p`              | 0.0       | 0.0                           |
| `presence_penalty`   | 0.0       | 1.5                           |
| `repetition_penalty` | 1.0       | 1.0                           |

* Kontextlänge = bis zu `262,144`
* Denkmodus: `temperature=1.0`, `top_p=0.95`, `top_k=20`, `min_p=0.0`, `presence_penalty=0.0`, `repetition_penalty=1.0`
* Instruct- (oder Nicht-Denk-) Modus: `temperature=0.7`, `top_p=0.80`, `top_k=20`, `min_p=0.0`, `presence_penalty=1.5`, `repetition_penalty=1.0`

### 💡 Denken + Denken beibehalten

{% columns %}
{% column %}
Qwen3.8-Flash-Next hat **Denken beibehalten** wodurch die Denkspur aus dem vorherigen Gespräch erhalten bleibt. Das erhöht die Anzahl der verwendeten Tokens, kann aber die Genauigkeit in fortgesetzten Gesprächen verbessern. [Unsloth](#run-qwen3.8-in-unsloth-desktop) hat Schalter für 'Think' und beibehaltenes Denken für Qwen3.8 (siehe rechts):
{% endcolumn %}

{% column %}

<figure><img src="https://797013937-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxhOjnexMCB3dmuQFQ2Zq%2Fuploads%2FLdgmjRrb5qhpbY9PwYe8%2FScreenshot%202026-08-14%20at%2011.26.15%E2%80%AFAM.png?alt=media&amp;token=6333f5ca-196d-46ae-9efd-2e522014e6db" alt=""><figcaption></figcaption></figure>
{% endcolumn %}
{% endcolumns %}

Qwen3.8-Flash-Next wird mit Unterstützung für `reasoning_effort`geliefert, mit dem sich die Tiefe des Schlussfolgerns anpassen und die Kosten steuern lassen. Diese Schalter werden in Unsloth automatisch aktiviert:

* `xhigh` (Standard): für komplexe Aufgaben, die gründliche Analyse erfordern
* `mittel`: Ausgleich zwischen Genauigkeit und Geschwindigkeit
* `niedrig`: effizientes Schlussfolgern, optimiert auf Geschwindigkeit und Kosten
* none

{% hint style="warning" %}
Um[ Denk-/Schlussfolgerungs-](#how-to-enable-or-disable-reasoning-and-thinking) Aufwand in `unsloth run` oder `llama-server`zu ändern, verwenden Sie `--chat-template-kwargs '{"reasoning_effort":"medium"}'`

Wenn Sie unter **Windows** Powershell sind, verwenden Sie: `--chat-template-kwargs "{\"reasoning_effort\":\"medium\"}"`

Ändern Sie `mittel` auf das gewünschte Schlussfolgerungsniveau.
{% endhint %}

### Quantisierungsanalyse

Wir haben KLD für Qwen3.8-Flash-Quants ausgeführt und zeigen, dass eine Wiederherstellung der Top-1%-Genauigkeit von 80 % mit 79 % weniger Festplattenspeicher möglich ist. Die neue Architektur verwendet PLE / Ngrams, und diese werden nicht so stark quantisiert (mindestens 4-Bit), da sie ein Zufriffsmuster mit zufälligem Zugriff haben, und eine starke Quantisierung würde dem Modell schaden.

<div><figure><img src="https://797013937-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxhOjnexMCB3dmuQFQ2Zq%2Fuploads%2FPemX6jyt4OWohwHfcjqo%2Fqwen38_flash_unsloth_top1_accuracy_new_data.png?alt=media&amp;token=3fd7713b-e9d9-43d4-bbca-96d56df43a80" alt=""><figcaption></figcaption></figure> <figure><img src="https://797013937-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxhOjnexMCB3dmuQFQ2Zq%2Fuploads%2F8bGa4A4jbGsn0qrNY6Gc%2Fqwen38_flash_unsloth_kld_new_data.png?alt=media&amp;token=45ac0e9a-c9b3-4530-90f6-ed8a2328ef08" alt=""><figcaption></figcaption></figure></div>

| Quant        | GB    | Top-1 % | mittlere KLD | 99,9 % KLD |
| ------------ | ----- | ------- | ------------ | ---------- |
| UD-IQ1\_S    | 72.5  | 77.325  | 0.396070     | 7.2126     |
| UD-IQ1\_M    | 74.5  | 79.691  | 0.314739     | 6.1965     |
| UD-Q2\_K\_XL | 78.9  | 82.715  | 0.224607     | 4.9121     |
| UD-IQ3\_XXS  | 82.0  | 85.414  | 0.165120     | 4.0375     |
| UD-Q3\_K\_XL | 90.0  | 88.315  | 0.106504     | 3.0538     |
| UD-IQ4\_XS   | 93.7  | 89.554  | 0.083630     | 2.3677     |
| UD-Q4\_K\_XL | 111.3 | 92.255  | 0.046893     | 1.5468     |
| UD-Q5\_K\_XL | 158.3 | 93.680  | 0.030415     | 1.0036     |
| UD-Q6\_K\_XL | 169.2 | 94.089  | 0.027091     | 0.8416     |
| Q8\_0        | 188.2 | 94.122  | 0.026574     | 0.8118     |

## Anleitung zum Ausführen von Qwen3.8-Flash-Next

Sie können Qwen3.8-Flash-Next jetzt in Unsloth Desktop und llama.cpp ausführen. Sie können den Quantisierungstyp frei ändern.

* Hugging Face: [Qwen3.8-Flash-Next-**GGUF**](https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF)
* ModelScope: [Qwen3.8-Flash-Next-GGUF](https://www.modelscope.cn/models/unsloth/Qwen3.8-Flash-Next-GGUF)

<a href="/pages/5695d6b6dec2df015f4427f80047e860ea228616#run-qwen3.8-in-unsloth-desktop" class="button primary">In Unsloth Desktop ausführen</a><a href="/pages/5695d6b6dec2df015f4427f80047e860ea228616#run-qwen3.8-in-llama.cpp" class="button secondary">In llama.cpp ausführen</a><a href="/docs/de/modelle/qwen3.8-next.md#mtp-guide" class="button primary">MTP-Anleitung</a>

{% hint style="success" %}
Qwen3.8-Flash-Next kann jetzt lokal ausgeführt werden in [Unsloth Desktop](#run-qwen3.8-flash-next-in-unsloth)!
{% endhint %}

### 🦥 Qwen3.8-Flash-Next in Unsloth ausführen

Qwen3.8-Flash-Next kann jetzt ausgeführt werden in [Unsloth Desktop](#run-qwen3.8-in-unsloth-desktop), einer Open-Source-UI-App für lokale KI. **Unsloth lagert automatisch in den RAM aus und erkennt Multi-GPU-Setups**. Mit Unsloth Desktop können Sie Modelle lokal ausführen auf **MacOS, Windows**, Linux und:

{% columns %}
{% column %}

* Suchen, herunterladen, [GGUFs ausführen](/docs/de/neu/studio.md#run-models-locally), MLX- und Safetensor-Modelle
* [**Selbstheilende** Tool-Aufrufe](/docs/de/neu/studio/chat.md#auto-healing-tool-calling) + **Websuche**
* [**Codeausführung**](/docs/de/desktop.md#code-execution) (Python, Bash)
* [Automatische Inferenz](https://unsloth.ai/docs/desktop#feature-deep-dive) Parameter-Tuning (Temp, Top-p usw.)
* Schnelle CPU- + GPU-Inferenz über MLX und llama.cpp
* [LLMs trainieren](/docs/de/neu/studio.md#no-code-training) 2x schneller mit 70 % weniger VRAM
  {% endcolumn %}

{% column %}

<figure><img src="https://797013937-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxhOjnexMCB3dmuQFQ2Zq%2Fuploads%2F6IXaXdTVyvbrnjehlxys%2Fkimik3.gif?alt=media&amp;token=31e1213b-d7da-46e9-bc7f-3a8c402513fc" alt=""><figcaption></figcaption></figure>
{% endcolumn %}
{% endcolumns %}

{% stepper %}
{% step %}

#### Unsloth installieren

Am einfachsten starten Sie, indem Sie die [Unsloth Desktop-App](/docs/de/desktop.md)herunterladen. Funktioniert auf [macOS](/docs/de/erste-schritte/install/mac.md), [Windows](/docs/de/erste-schritte/install/windows-installation.md), und [Linux](/docs/de/erste-schritte/install/linux.md).

<a href="https://unsloth.ai/download" class="button primary" data-icon="down-to-bracket">Unsloth herunterladen</a>

* <i class="fa-apple">:apple:</i> [Download für macOS](https://unsloth.ai/download/mac)
* <i class="fa-windows">:windows:</i> [Download für Windows](https://unsloth.ai/download/windows)
* <i class="fa-linux">:linux:</i> [Download für Linux](https://unsloth.ai/download/linux)

Oder, wenn Sie lieber manuell installieren möchten:

MacOS, Linux, WSL:

```bash
curl -fsSL https://unsloth.ai/install.sh | sh
```

Windows PowerShell:

```bash
irm https://unsloth.ai/install.ps1 | iex
```

{% endstep %}

{% step %}

#### Qwen3.8-Flash-Next suchen und herunterladen

Gehen Sie zu [Unsloth Chat](/docs/de/neu/studio/chat.md) oder zum Model Hub und suchen Sie in der Suchleiste nach Qwen3.8-Flash und laden Sie das gewünschte Modell und die gewünschte Quantisierung herunter.

<figure><img src="https://797013937-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxhOjnexMCB3dmuQFQ2Zq%2Fuploads%2FiEMEgtWrGc0DMZ4FLRez%2FScreenshot%202026-08-27%20at%204.21.05%E2%80%AFAM.png?alt=media&amp;token=94ad9ccf-f882-48b0-8aaa-83e90bfc2630" alt=""><figcaption></figcaption></figure>
{% endstep %}

{% step %}

#### Qwen3.8-Flash-Next ausführen

MTP ist automatisch aktiviert, aber Sie können es deaktivieren. Die Inferenzparameter sollten bei Verwendung von Unsloth automatisch gesetzt werden, Sie können sie jedoch weiterhin manuell ändern. Sie können auch die Kontextlänge, das Chat-Template und andere Einstellungen bearbeiten.

Für weitere Informationen können Sie unsere [Unsloth-Inferenzanleitung](/docs/de/neu/studio/chat.md).

Zum Beispiel ermöglicht Ihnen die Verwendung von Unsloth Desktop mit dem 397 GB großen Qwen3.8 (-91 % kleiner) das Umschalten von Denkmodi, Inline-Canvas, Websuche und Codeausführung und vieles mehr.

<figure><img src="https://797013937-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxhOjnexMCB3dmuQFQ2Zq%2Fuploads%2F1YneA51oWzaIiwep9H5I%2F1000024423.gif?alt=media&amp;token=40a169cc-12f1-403a-898c-b310f81ff52a" alt=""><figcaption></figcaption></figure>
{% endstep %}

{% step %}

#### Qwen3.8-Flash-Next über die Unsloth-API bereitstellen

Sie können `unsloth run` Befehl verwenden und Qwen3.8 über eine API bereitstellen mit `llama-server` Laufzeit-Flags, einschließlich Kontextgröße, GPU-Layer, Threading, Sampling, Netzwerk und Werkzeugkonfiguration. Weitere Informationen finden Sie in unseren [API-Dokumentation](/docs/de/grundlagen/api.md) oder [unsloth start](/docs/de/integrationen/unsloth-start.md).

{% code overflow="wrap" %}

```bash
unsloth run --model unsloth/Qwen3.8-Flash-Next-GGUF:UD-Q4_K_XL
```

{% endcode %}
{% endstep %}

{% step %}

#### Unsloth ist jetzt bereit

Sie können mit Qwen3.8-Flash-Next über Unsloth Desktop auch viele andere Dinge tun, wie:

* **Werkzeuge verbinden:** [Claude Code](/docs/de/grundlagen/claude-code.md), [Codex](/docs/de/grundlagen/codex.md), [Websuche](/docs/de/neu/studio/chat.md#advanced-web-search), [MCP](/docs/de/grundlagen/mcp.md) und mehr
* **Modelle trainieren:** Text, Diffusion, [Embedding](/docs/de/grundlagen/embedding-finetuning.md)und mehr feinabstimmen
* **Medien generieren:** Erstellen und trainieren [Bilder](/docs/de/grundlagen/diffusion-image.md), Video, [TTS](/docs/de/grundlagen/text-to-speech-tts-fine-tuning.md) lokal

<figure><img src="https://797013937-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxhOjnexMCB3dmuQFQ2Zq%2Fuploads%2FRnU2breyPzalRzIHyq8U%2FScreenshot%202026-08-28%20at%2012.12.20%E2%80%AFAM.png?alt=media&amp;token=da925810-e1d3-4c06-bf27-7caad15a2330" alt=""><figcaption></figcaption></figure>
{% endstep %}
{% endstepper %}

### :llama: Qwen3.8-Flash-Next in llama.cpp ausführen

{% stepper %}
{% step %}
Installieren Sie die neueste Version von llama.cpp. Sie können auch den Build-Anweisungen unten folgen. Ändern Sie `-DGGML_CUDA=ON` zu `-DGGML_CUDA=OFF` wenn Sie keine GPU haben oder nur CPU-Inferenz möchten. **Für Apple-Mac- / Metal-Geräte**, setzen Sie `-DGGML_CUDA=OFF` und fahren Sie dann wie gewohnt fort – Metal-Unterstützung ist standardmäßig aktiviert.

```bash
apt-get update
apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y
git clone https://github.com/ggml-org/llama.cpp
cmake llama.cpp -B llama.cpp/build \\
    -DBUILD_SHARED_LIBS=OFF -DGGML_CUDA=ON
cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-mtmd-cli llama-server llama-gguf-split
cp llama.cpp/build/bin/llama-* llama.cpp
```

{% endstep %}

{% step %}
Um das Modell auszuführen, können Sie Folgendes tun:

{% code overflow="wrap" %}

```bash
pip install -U "huggingface_hub[cli]"
hf download unsloth/Qwen3.8-Flash-Next-GGUF \\
    --local-dir unsloth/Qwen3.8-Flash-Next-GGUF \\
    --include "*UD-Q4_K_XL*" # Verwenden Sie "*IQ2_XXS*" für 2-Bit
```

{% endcode %}
{% endstep %}

{% step %}
Dann zum Ausführen:

{% code overflow="wrap" %}

```bash
./llama.cpp/llama-cli \\
    --model unsloth/Qwen3.8-Flash-Next-GGUF/UD-IQ1_S/Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf \\
    --temp 1.0 \\
    --top-p 0.95 \\
    --top-k 20 \\
    --min-p 0.0
```

{% endcode %}
{% endstep %}
{% endstepper %}

### MTP-Anleitung

Qwen3.8-Flash kann mit 1,3 bis **1,7x schnellerer Inferenz** über [MTP](/docs/de/modelle/mtp.md) (Multi-Token Prediction) ohne Genauigkeitsverlust ausgeführt werden! MTP ermöglicht es Qwen3.8-Flash, **170 Token/s** auf 1x RTX 6000 PRO GPU zu erreichen, verglichen mit der 100-Token-Basislinie. MTP beschleunigt die Inferenz, indem ein Modell mehrere kommende Tokens auf einmal vorhersagt, anstatt pro Schritt nur einen Token zu erzeugen, und ist besonders auf GPUs wirksam.

Um Qwen3.8-Flash mit MTP auszuführen, ist MTP in [Unsloth Desktop](#run-qwen3.8-flash-next-in-unsloth) standardmäßig aktiviert oder Sie können unseren benutzerdefinierten llama.cpp-PR verwenden.

<div><figure><img src="https://797013937-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxhOjnexMCB3dmuQFQ2Zq%2Fuploads%2F39BmPdGOL8IlwrdQFsGh%2Fqwen38_flash_next_unsloth_ggufs_mtp_speedup_no_mtp.png?alt=media&amp;token=77c52179-821a-406b-a5ab-fd027ebe8d30" alt=""><figcaption></figcaption></figure> <figure><img src="https://797013937-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxhOjnexMCB3dmuQFQ2Zq%2Fuploads%2FF9IF8IWG2DAgfLBXzhyz%2Fqwen38_flash_next_unsloth_ggufs_mtp_decode_tokens_per_s.png?alt=media&amp;token=def0f8fb-5268-4e33-9e85-0a1fed4a1156" alt=""><figcaption></figcaption></figure></div>

Die Gewinne sind auf Geräten mit geringerer Speicherbandbreite, wie älteren Macs, kleiner. Wir haben sowohl gemeinsame MTP-Module erstellt (schließt die embed\_tokens aus und teilt sie mit dem Hauptmodell), um Speicherplatz sowie RAM- und VRAM-Nutzung um etwa 1 bis 2 GB zu sparen.

| MTP-Typ  | Allgemeines MTP | Geteiltes MTP | Einsparungen |
| -------- | --------------- | ------------- | ------------ |
| BF16     | 7,77 GB         | 5,23 GB       | 2,54 GB      |
| Q8\_0    | 4,14 GB         | 2,79 GB       | 1,35 GB      |
| Q4\_K\_M | 2,79 GB         | 1,91 GB       | 880 MB       |

Die 3-Bit-MTP-Quantisierung funktioniert mit 91 GB RAM, daher ist ein Gerät mit 96 GB RAM/Unified Memory am besten.\
**Tabelle: MTP-Hardwareanforderungen** (Einheiten = Gesamtspeicher: RAM + VRAM oder Unified Memory)

<table><thead><tr><th>1-Bit</th><th>2-Bit</th><th>3-Bit</th><th>4-Bit</th><th width="128">5-Bit</th><th>8-Bit</th><th>BF16</th></tr></thead><tbody><tr><td>76 GB</td><td>80 GB</td><td>91 GB</td><td>97-115 GB</td><td>164 GB</td><td>200 GB</td><td>355 GB</td></tr></tbody></table>

#### MTP Qwen3.8-Flash ausführen

Um Qwen3.8-Flash mit MTP auszuführen, müssen Sie nur [**Unsloth Desktop installieren**](#run-qwen3.8-flash-next-in-unsloth) oder auf die neueste Version von Unsloth aktualisieren und dann das Modell erneut herunterladen oder die MTP-Datei herunterladen. Weitere Anweisungen für llama.cpp finden Sie weiter unten.

{% columns %}
{% column %}
Unsloth Desktop funktioniert auf [macOS](/docs/de/erste-schritte/install/mac.md), [Windows](/docs/de/erste-schritte/install/windows-installation.md), und [Linux](/docs/de/erste-schritte/install/linux.md).

<a href="https://unsloth.ai/download" class="button primary" data-icon="down-to-bracket">Unsloth herunterladen</a>

* <i class="fa-apple">:apple:</i> [Download für macOS](https://unsloth.ai/download/mac)
* <i class="fa-windows">:windows:</i> [Download für Windows](https://unsloth.ai/download/windows)
* <i class="fa-linux">:linux:</i> [Download für Linux](https://unsloth.ai/download/linux)

In Unsloth Desktop können Sie auch die Anzahl der Draft-Tokens ändern oder MTP anpassen. Verwenden Sie die erweiterten Einstellungen in der rechten Seitenleiste, aktivieren Sie "Erweiterte Einstellungen" und Sie können MTP-/Ngram-Speculative-Decoding, die Anzahl der Draft-Tokens und mehr auswählen:
{% endcolumn %}

{% column %}

<figure><img src="https://797013937-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxhOjnexMCB3dmuQFQ2Zq%2Fuploads%2FqKy281quNOdn5toIr5am%2Fimage.png?alt=media&amp;token=bf0f6f00-1192-4494-977c-1bf4fa346fa0" alt="" width="305"><figcaption></figcaption></figure>
{% endcolumn %}
{% endcolumns %}

#### MTP Llama.cpp-Anleitung

{% code overflow="wrap" %}

```bash
apt-get update
apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y
git clone --branch qwen4exp/mtp https://github.com/danielhanchen/llama.cpp
cmake llama.cpp -B llama.cpp/build \\
    -DBUILD_SHARED_LIBS=OFF -DGGML_CUDA=ON
cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-mtmd-cli llama-server llama-gguf-split
cp llama.cpp/build/bin/llama-* llama.cpp
```

{% endcode %}

Dann um das geteilte MTP-Modul herunterzuladen:

{% code overflow="wrap" %}

```bash
pip install -U "huggingface_hub[cli]"
hf download unsloth/Qwen3.8-Flash-Next-GGUF \\
    --local-dir unsloth/Qwen3.8-Flash-Next-GGUF \\
    --include "*mtp-Qwen3.8-Flash-Next-shared-Q8_0.gguf*"
```

{% endcode %}

Und um llama-server damit zu verwenden:

{% code overflow="wrap" %}

```bash
llama.cpp/llama-server \\
    -hf unsloth/Qwen3.8-Flash-Next-GGUF:UD-Q4_K_XL \\
    -md unsloth/Qwen3.8-Flash-Next-GGUF/MTP/mtp-Qwen3.8-Flash-Next-shared-Q8_0.gguf \\
    --spec-type draft-mtp --spec-draft-n-max 5
```

{% endcode %}

### 📊 Benchmarks

Für GGUF-Quantisierungs-Benchmarks können Sie oben unsere [Quantisierungsanalyse](#quantization-analysis) oder [Dynamic V3.0-Artikel](/docs/de/grundlagen/dynamic-3.0-ggufs.md).

<div><figure><img src="https://797013937-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxhOjnexMCB3dmuQFQ2Zq%2Fuploads%2FJFGDUmmWUvJMD0eCUbiE%2Fqwennextbe.jpg?alt=media&amp;token=0be96d30-9f51-41f3-8e3a-3366a4fdfb93" alt=""><figcaption></figcaption></figure> <figure><img src="https://797013937-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxhOjnexMCB3dmuQFQ2Zq%2Fuploads%2FpQBzoQNziZtrFDHCyN3t%2Fbench2max.jpg?alt=media&amp;token=3d6adbe3-6c8b-4cb4-937f-f434ebd7f106" alt=""><figcaption></figcaption></figure></div>


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://unsloth.ai/docs/de/modelle/qwen3.8-next.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
