> For the complete documentation index, see [llms.txt](https://unsloth.ai/docs/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://unsloth.ai/docs/de/grundlagen.md).

# Grundlagen

- [Wie man Unsloth als API-Endpunkt verwendet](https://unsloth.ai/docs/de/grundlagen/api.md)
- [Inferenz & Bereitstellung](https://unsloth.ai/docs/de/grundlagen/inference-and-deployment.md): Lerne, wie du dein fine-tuned Modell speicherst, damit du es in deiner bevorzugten Inferenz-Engine ausführen kannst.
- [Als GGUF speichern](https://unsloth.ai/docs/de/grundlagen/inference-and-deployment/saving-to-gguf.md)
- [Spekulatives Dekodieren](https://unsloth.ai/docs/de/grundlagen/inference-and-deployment/saving-to-gguf/speculative-decoding.md): Spekulatives Dekodieren mit llama-server, llama.cpp, vLLM und mehr für 2x schnellere Inferenz
- [Leitfaden für vLLM-Bereitstellung & Inferenz](https://unsloth.ai/docs/de/grundlagen/inference-and-deployment/vllm-guide.md): Anleitung zum Speichern und Bereitstellen von LLMs in vLLM zum Einsatz von LLMs in der Produktion
- [vLLM-Engine-Argumente](https://unsloth.ai/docs/de/grundlagen/inference-and-deployment/vllm-guide/vllm-engine-arguments.md)
- [Leitfaden zum Hot-Swapping von LoRA](https://unsloth.ai/docs/de/grundlagen/inference-and-deployment/vllm-guide/lora-hot-swapping-guide.md)
- [Modelle in Ollama speichern](https://unsloth.ai/docs/de/grundlagen/inference-and-deployment/saving-to-ollama.md)
- [Modelle in LM Studio bereitstellen](https://unsloth.ai/docs/de/grundlagen/inference-and-deployment/lm-studio.md): Modelle als GGUF speichern, damit du sie in LM Studio ausführen und bereitstellen kannst
- [Wie man das LM Studio CLI im Linux-Terminal installiert](https://unsloth.ai/docs/de/grundlagen/inference-and-deployment/lm-studio/how-to-install-lm-studio-cli-in-linux-terminal.md): Anleitung zur Installation der LM Studio CLI ohne UI in einer Terminal-Instanz.
- [Leitfaden für SGLang-Bereitstellung & Inferenz](https://unsloth.ai/docs/de/grundlagen/inference-and-deployment/sglang-guide.md): Anleitung zum Speichern und Bereitstellen von LLMs in SGLang zum Einsatz von LLMs in der Produktion
- [Unsloth-Inferenz](https://unsloth.ai/docs/de/grundlagen/inference-and-deployment/unsloth-inference.md): Lerne, wie du dein fine-tuned Modell mit der schnelleren Inferenz von Unsloth ausführen kannst.
- [Leitfaden für die Bereitstellung von llama-server & OpenAI-Endpunkten](https://unsloth.ai/docs/de/grundlagen/inference-and-deployment/llama-server-and-openai-endpoint.md): Bereitstellung über llama-server mit einem OpenAI-kompatiblen Endpunkt
- [Wie man LLMs auf deinem iOS- oder Android-Telefon ausführt und bereitstellt](https://unsloth.ai/docs/de/grundlagen/inference-and-deployment/deploy-llms-phone.md): Tutorial zum Fine-Tuning deines eigenen LLMs und zur Bereitstellung auf deinem Android- oder iPhone-Gerät mit ExecuTorch.
- [Fehlerbehebung bei der Inferenz](https://unsloth.ai/docs/de/grundlagen/inference-and-deployment/troubleshooting-inference.md): Wenn du Probleme beim Ausführen oder Speichern deines Modells hast.
- [LLMs mit Hugging Face Jobs bereitstellen](https://unsloth.ai/docs/de/grundlagen/inference-and-deployment/deploying-llms-with-hugging-face-jobs.md): Hugging Face Jobs und Skills verwenden, um LFM mit Codex / Claude Code mit einem SKILL zu fine-tunen.
- [Wie man lokale LLMs mit Claude Code ausführt](https://unsloth.ai/docs/de/grundlagen/claude-code.md): Anleitung zur Verwendung offener Modelle mit Claude Code auf deinem lokalen Gerät.
- [Wie man lokale LLMs mit OpenAI Codex ausführt](https://unsloth.ai/docs/de/grundlagen/codex.md): Verwende offene Modelle mit OpenAI Codex lokal auf deinem Gerät.
- [Unsloth Dynamic NVFP4-Anleitung ausführen](https://unsloth.ai/docs/de/grundlagen/nvfp4.md): Erfahren Sie, wie Unsloth Dynamic NVFP4 schnelle, genaue 4-Bit-Inferenz auf NVIDIA-Blackwell-GPUs ermöglicht.
- [Modelle auf AMD-GPUs mit Unsloth trainieren und ausführen](https://unsloth.ai/docs/de/grundlagen/amd.md)
- [Wie man MCP-Server mit lokalen LLMs verwendet](https://unsloth.ai/docs/de/grundlagen/mcp.md): Lerne mit Screenshots, wie man MCP-Server mit offenen KI-Modellen verbindet.
- [Multi-GPU-Fine-Tuning mit Unsloth](https://unsloth.ai/docs/de/grundlagen/multi-gpu-training-with-unsloth.md): Lerne, wie du LLMs auf mehreren GPUs und mit Parallelisierung mit Unsloth fine-tunest.
- [Multi-GPU-Fine-Tuning mit Distributed Data Parallel (DDP)](https://unsloth.ai/docs/de/grundlagen/multi-gpu-training-with-unsloth/ddp.md): Lerne, wie du die Unsloth-CLI verwendest, um auf mehreren GPUs mit Distributed Data Parallel (DDP) zu trainieren!
- [Leitfaden zum Fine-Tuning von Embedding-Modellen mit Unsloth](https://unsloth.ai/docs/de/grundlagen/embedding-finetuning.md): Lerne, wie du Embedding-Modelle mit Unsloth ganz einfach fine-tunest.
- [Fine-Tune MoE-Modelle 12x schneller mit Unsloth](https://unsloth.ai/docs/de/grundlagen/faster-moe.md): Trainiere MoE-LLMs lokal mit dem Unsloth-Leitfaden.
- [Leitfaden zum Fine-Tuning von Text-to-Speech (TTS)](https://unsloth.ai/docs/de/grundlagen/text-to-speech-tts-fine-tuning.md): Lerne, wie du TTS- und STT-Sprachmodelle mit Unsloth fine-tunest.
- [Unsloth Dynamic 2.0 GGUFs](https://unsloth.ai/docs/de/grundlagen/unsloth-dynamic-2.0-ggufs.md): Ein großes neues Upgrade für unsere Dynamic Quants!
- [Unsloth Dynamic GGUFs auf Aider Polyglot](https://unsloth.ai/docs/de/grundlagen/unsloth-dynamic-2.0-ggufs/unsloth-dynamic-ggufs-on-aider-polyglot.md): Leistung von Unsloth Dynamic GGUFs auf Aider-Polyglot-Benchmarks
- [Anleitung zur Tool-Calling für lokale LLMs](https://unsloth.ai/docs/de/grundlagen/tool-calling-guide-for-local-llms.md)
- [Vision-Fine-Tuning](https://unsloth.ai/docs/de/grundlagen/vision-fine-tuning.md): Lerne, wie du Vision-/multimodale LLMs mit Unsloth fine-tunest
- [Fehlerbehebung & FAQs](https://unsloth.ai/docs/de/grundlagen/troubleshooting-and-faqs.md): Tipps zur Lösung von Problemen und häufig gestellte Fragen.
- [Hugging Face Hub, XET-Debugging](https://unsloth.ai/docs/de/grundlagen/troubleshooting-and-faqs/hugging-face-hub-xet-debugging.md): Debugging, Fehlerbehebung bei hängenden, festgefahrenen Downloads und langsamen Downloads
- [Chat-Vorlagen](https://unsloth.ai/docs/de/grundlagen/chat-templates.md): Lerne die Grundlagen und Anpassungsoptionen von Chat-Vorlagen kennen, einschließlich Conversational-, ChatML-, ShareGPT-, Alpaca-Formate und mehr!
- [Unsloth-Umgebungs-Flags](https://unsloth.ai/docs/de/grundlagen/unsloth-environment-flags.md): Erweiterte Flags, die nützlich sein können, wenn du fehlschlagende Fine-Tunes siehst oder Dinge abschalten möchtest.
- [Fortgesetztes Pretraining](https://unsloth.ai/docs/de/grundlagen/continued-pretraining.md): Auch bekannt als kontinuierliches Fine-Tuning. Unsloth ermöglicht es dir, kontinuierlich vorzutrainieren, damit ein Modell eine neue Sprache lernen kann.
- [Fine-Tuning vom letzten Checkpoint aus](https://unsloth.ai/docs/de/grundlagen/finetuning-from-last-checkpoint.md): Checkpointing ermöglicht es dir, den Fortschritt deines Fine-Tunings zu speichern, damit du es pausieren und später fortsetzen kannst.
- [Unsloth-Benchmarks](https://unsloth.ai/docs/de/grundlagen/unsloth-benchmarks.md): Von Unsloth aufgezeichnete Benchmarks auf NVIDIA-GPUs.


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://unsloth.ai/docs/de/grundlagen.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
