LM Studio: Run and Serve Local LLMs With a GUI and API Server
LM Studio gives you a polished desktop GUI for downloading GGUF models, tuning GPU layers, and serving an OpenAI-compatible local API - no command line required.
Mahmudul Haque Qudrati
CEO & ML Engineer
One AI engineering post, weekly
LLM benchmarks, prompt techniques, and token-cost breakdowns — not another AI news roundup.
Most local LLM tooling assumes you are comfortable in a terminal. LM Studio removes that barrier: it is a desktop application (Mac, Windows, Linux) that handles model discovery, download, hardware tuning, and serving through a graphical interface. It is the fastest path for non-terminal users to run a model locally - and it is equally useful for developers who want a visual config layer on top of llama.cpp.
Installing LM Studio
Download the installer from lmstudio.ai. The app bundles a pre-compiled llama.cpp backend, so there is nothing else to install. On first launch, LM Studio detects your GPU (NVIDIA/AMD/Apple Silicon) and configures defaults accordingly.
Team workspace
Ship faster with chat, meetings, and projects in one place — Zlyqor.
Downloading Models
Use the built-in search to browse HuggingFace Hub directly. LM Studio surfaces GGUF-format models from trusted publishers like TheBloke and Bartowski. Select a quantization level - Q4_K_M is a good default - and click Download. Models are stored in ~/LM Studio/.
The supported formats page lists GGUF, MLX (Apple Silicon native), and GPTQ.
GPU Layer Offloading
In the model settings panel, the GPU Layers slider controls how many transformer layers are offloaded to VRAM. More layers = faster inference. Set it to the maximum your VRAM can hold - LM Studio shows a live VRAM meter. A rule of thumb:
| Hardware | Q4_K_M Llama 3.1 8B GPU Layers |
|---|---|
| 8 GB VRAM | ~22 layers (partial offload) |
| 16 GB VRAM | All 32 layers |
| Apple M2 24 GB unified | All 32 layers |
Chat Playground
The Chat tab gives you a full conversation interface with message history, system prompt editor, and generation parameters (temperature, top-p, repeat penalty). You can save and load presets - useful for comparing prompting strategies across models.
Local API Server
Enable the server from the Local Server tab. LM Studio starts an OpenAI-compatible HTTP server on localhost:1234. Drop-in replacement in any app:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:1234/v1", api_key="lm-studio")
response = client.chat.completions.create(
model="local-model",
messages=[{"role": "user", "content": "What is GGUF?"}],
)
print(response.choices[0].message.content)
The model name "local-model" is a placeholder - LM Studio routes all /v1/chat/completions requests to whichever model is loaded.
Hardware Requirements by Model Size
| Model Size | Min RAM (Q4_K_M) | Recommended |
|---|---|---|
| 7B | 6 GB | 8 GB VRAM |
| 13B | 10 GB | 16 GB VRAM |
| 34B | 24 GB | 24 GB VRAM |
| 70B | 40 GB | 2x24 GB or A100 |
For most laptops, a 7B or 8B model at Q4_K_M runs at 15 - 40 tokens/sec - fast enough for interactive use. Apple M-series chips with unified memory outperform discrete NVIDIA cards of equivalent memory size because there is no PCIe bandwidth bottleneck.
LM Studio is updated frequently - check the changelog for new features like multi-model serving and MLX acceleration on Apple Silicon.

Mahmudul Haque Qudrati
CEO & ML Engineer
Visionary technologist, software engineer, and machine learning specialist. Founder and CEO of Pristren, directing engineering teams that ship production-grade AI/ML pipelines, mission-critical full-stack applications, and developer tooling. Creator of Zlyqor, the unified team workspace platform. Author of 540+ technical guides and benchmark research reports on large language models, agentic workflows, Model Context Protocol (MCP), and modern web stacks.
More from Mahmudul
Related Articles
How to Use Claude to Make Videos Like Vox and Others
Claude can help you make Vox-style videos by generating scripts, editing with code, and automating animation. Here's a practical guide with real workflows and costs.
OpenAI Ends Cursor Model Access on Nov 12, 2026: What Developers Need to Know
OpenAI will terminate Cursor's access to its models on November 12, 2026, following SpaceX's acquisition. This guide explains the timeline, why it happened, and practical steps to migrate your workflow.
I Used Claude Code to Get a Second Opinion on My MRI: A Practical Overview
A developer fed his MRI scan to Claude Code and got a second opinion. Here's how the experiment worked, what it cost, and why you shouldn't rely on it for medical decisions.
// discussion
Comments