
Your local AI,
finally under control
Llama Control is a desktop control panel for llama.cpp. Manage GGUF models, chat with your AI, benchmark performance, auto-tune server flags, and browse HuggingFace, all from one beautiful interface.
Up and running in 4 steps
No command line. No config files. No hassle. From download to your first chat in under 5 minutes.
Download & Install
Download the latest installer from GitHub Releases. About 110 MB, a single download, no account required.
Point to Your Models
Open Settings and point Llama Control to your GGUF model folder. It auto-detects LM Studio and common paths. The app scans all .gguf files and displays their metadata instantly.
Install llama.cpp
Use the built-in llama.cpp installer in Settings to download the right prebuilt for your GPU (CUDA, Vulkan, or CPU). One click, no manual extraction needed.
Run & Chat
Select a model, optionally use Auto-Tune for optimized flags, and hit start. Chat with your local AI instantly, no terminal, no config files, no command line.
Six powerful tools, one interface
Your GGUF models, beautifully organized
Browse local GGUF files with rich metadata parsing. Edit names, descriptions, and tags. Start or stop the server per-model with custom launch flags, and organize your collection with favorites and custom folders.
- One-click llama-server start with per-model flags
- GGUF metadata parser, architecture, quantization, parameters
- Server profiles for quick flag switching
- Real-time GPU VRAM, CPU, and RAM usage stats
- Duplicate, export, and delete models
Local Models
4 modelsChat with your local AI
Streamed token responses from your running llama-server. Manage multiple chat sessions, edit system prompts, and switch to terminal mode with xterm.js for advanced workflows.
- Streamed real-time responses from local model
- Multiple sessions with full history
- Customizable system prompts
- Terminal mode with multi-session xterm.js
- Reasoning trace display for CoT models
Browse & download from HuggingFace
Search GGUF models directly from HuggingFace. Get hardware-aware VRAM-fit hints, check disk space before downloading, and stream files straight into your models folder with resumable downloads.
- Search by query, category, or popularity
- VRAM-fit hints based on your GPU
- Disk space check before downloads
- Resumable multi-GB downloads
- GGUF vendor detection badges
Benchmark Results
Measure. Compare. Optimize.
Run llama-bench speed tests to measure tokens/second. Evaluate quality with 18 deterministic tasks covering math, logic, code, and knowledge. Grade everything with AI judges.
- llama-bench speed benchmark (tokens/s)
- 18 quality tasks: math, logic, code, format, knowledge, traps
- Perplexity test (wikitext-2)
- AI quality judging via Claude/OpenAI/Gemini
- Multi-model comparison slots
Let AI tune your server
Forget manual flag tweaking. The AI research phase suggests optimal launch flags, then automatically benchmarks each candidate through a trial loop. KV cache quantization, chunk sizes, fit-to-VRAM, all auto-optimized.
- AI research phase suggests optimal flag candidates
- Automatic benchmarking trial loop
- KV cache quantization ladder testing
- Context and chunk size tuning
- Fit-to-VRAM auto-calibration
Auto-Tune Pipeline
Usage Analytics
Know your AI habits
Track your AI usage patterns with a comprehensive dashboard. Monitor total chats, messages, response times, and character counts. Identify your most active days and top conversations.
- Total chats and message counts
- User vs. assistant message breakdown
- Average and median response times
- Characters generated vs. sent ratio
- Most active day tracking
Common questions
Everything you need to know about Llama Control.
Ready to take control?
Download Llama Control for free and start managing your local AI experience today. Windows 10+, no account needed.
v0.1.60, Windows 10/11 (64-bit), ~110 MB installer