About Unsloth Desktop
Unsloth Desktop is an open-source desktop app for running, training, and deploying AI models locally on macOS, Windows, and Linux.It supports LLMs, diffusion image/video models, audio models (text-to-speech and speech-to-text), and embedding workflows for on-device inference and fine-tuning.
The platform offers LoRA and full fine-tuning, quantization-aware training, multi-GPU and distributed training, and reduced-VRAM inference options.A built-in model hub and model-swapping UI facilitate downloading and managing models and gguf formats, including Qwen, Gemma, Meta Muse, Minimax, Kimi K3, and GLM families.
Developer features include sandboxed code execution, tool-calling with automatic retry/healing, parallel chat sessions, and OpenAI/Anthropic-compatible API integrations.Deployment and remote access options include local HTTPS hosting and a Cloudflare tunnel for remote inference and model serving.
The tool targets ML engineers, researchers, and developers working on local model training, diffusion image/video generation, fine-tuning, and on-device model deployment.
Key Features
Use Cases
Who is it for?
The platform offers LoRA and full fine-tuning, quantization-aware training, multi-GPU and distributed training, and reduced-VRAM inference options.A built-in model hub and model-swapping UI facilitate downloading and managing models and gguf formats, including Qwen, Gemma, Meta Muse, Minimax, Kimi K3, and GLM families.
Developer features include sandboxed code execution, tool-calling with automatic retry/healing, parallel chat sessions, and OpenAI/Anthropic-compatible API integrations.Deployment and remote access options include local HTTPS hosting and a Cloudflare tunnel for remote inference and model serving.
The tool targets ML engineers, researchers, and developers working on local model training, diffusion image/video generation, fine-tuning, and on-device model deployment.
Key Features
- Local desktop app for running, training, and deploying AI models on macOS, Windows, and Linux
- Support for LLMs, diffusion image/video models, audio models (text-to-speech and speech-to-text), and embedding workflows for on-device inference and fine-tuning
- Training and inference tooling: LoRA and full fine-tuning, quantization-aware training, multi-GPU and distributed training, and reduced-VRAM inference options
- Built-in model hub and model-swapping UI with gguf format support for downloading and managing models
- Developer and deployment features: sandboxed code execution, tool-calling with automatic retry/healing, parallel chat sessions, OpenAI/Anthropic-compatible API integrations, local HTTPS hosting and Cloudflare tunnel for remote inference and model serving
Use Cases
- Fine-tune and deploy a privacy-preserving, production-ready LLM with Unsloth Desktop using LoRA and quantization-aware training across multiple GPUs, manage versions in the built-in model manager, and remotely host a secure model endpoint for internal apps without sending data to the cloud
- Create high-fidelity images and videos locally by training and running diffusion models in Unsloth Desktop with multi-GPU acceleration, sandboxed code execution for safe experimentation, and export quantized models for fast edge inference or integration into creative pipelines
- Build an offline voice assistant or speech-analysis pipeline by training and optimizing audio models on-device with Unsloth Desktop, apply quantization and pruning for low-latency on-device inference, and securely share or serve models via its remote hosting for cross-team testing
Who is it for?
- Machine learning engineers
- Machine learning researchers
- Application developers
- Data scientists
- Multimodal researchers
