Introduction
A 16 GB laptop GPU is enough to run a genuinely capable coding model entirely on your own machine โ no API keys, no data leaving the box. The model I settled on is Qwen3-Coder-30B-A3B-Instruct [1]: a Mixture-of-Experts model with 30 billion total parameters but only ~3 billion active per token, so it reasons like a big model while running at the speed of a small one.
This post is the settings sheet I wish I'd had: the exact LM Studio load configuration, the one VRAM rule that decides whether the model flies or crawls, how to wire it into pi (a terminal coding agent) [2], and honest benchmarks from driving it against a real codebase. Everything here was measured on a laptop RTX 5080 (16 GB VRAM).
โ ๏ธ The one rule: never fill VRAM to the top
This is the single most important thing, so it goes first. The Q4_K_M quantization of this model is about 17.4 GB of weights โ which is larger than a 16 GB card. That means some layers must be offloaded to the CPU no matter what; the full model physically cannot sit on the GPU.
The trap is the "max GPU offload" setting. Turn it up and LM Studio dutifully tries to place everything on the GPU. On Windows, instead of erroring, the NVIDIA driver silently spills the overflow into shared GPU memory โ system RAM accessed over the PCIe bus. Every token then shuttles data across PCIe, and throughput collapses:
Same model, same GPU: 0.6 tokens/sec at "max" offload (VRAM 97% full, spilling) versus ~24 tokens/sec with the offload dialed back to leave headroom. A ~40ร difference from one setting.
The fix is counterintuitive: put less on the GPU, not more. Leave roughly 1.5โ2 GB of VRAM free for the KV cache and driver overhead, and verify the real number with nvidia-smi rather than trusting the loader's estimate. For a MoE model that overflows VRAM, partial offload beats "max" every single time.
๐๏ธ LM Studio load settings
Set these in the model's load panel before loading it (in LM Studio: My Models โ the model's load configuration). The values below are the tested 16 GB baseline.
| Setting | Value | Why |
|---|---|---|
| GPU Offload | 0.60 |
The critical dial. Offload ~60% of layers on 16 GB; target ~1.5โ2 GB free VRAM. Verify with nvidia-smi. |
| Context Length | 32768 |
Agent harnesses need room. An 8K window silently truncates the system prompt and file context, causing "forgotten instructions." 32K is the working floor. |
| Flash Attention | ON |
Large KV-cache savings โ this is the headroom that makes 32K context affordable without spilling. |
| KV Cache Quantization | Q8_0 |
Roughly halves per-token context memory. Enables 48โ64K context, or spend the freed VRAM on a higher offload ratio. Negligible quality loss at Q8. |
For sampling, start deterministic for coding โ temperature 0.2โ0.3. For more exploratory generation, Qwen's recommended defaults are temp 0.7 ยท top_p 0.8 ยท top_k 20 ยท repeat_penalty 1.05 [1].
Scale the offload to your card
The principle is identical at every tier: leave headroom. Only the 16 GB row is measured; the others are sensible starting points to verify and adjust with nvidia-smi.
| VRAM | GPU offload | Context | ~Speed | Feel |
|---|---|---|---|---|
| 12 GB | ~0.35 | 16โ24K | 10โ15 tok/s | Usable, slow |
| 16 GB (tested) | 0.60 | 32K | ~24 tok/s | Sweet spot |
| 24 GB | ~0.90 | 32โ64K | 50โ70 tok/s | Comfortable |
| 32 GB+ | max | 64K+ | Fast | Fully on-GPU |
Below 24 GB the full weights can't sit entirely on the GPU, so throughput is capped by the CPU-offloaded layers. That's the price of running a 30B model on a small card โ and it's still worth paying.
๐ Start the server
LM Studio exposes an OpenAI-compatible endpoint that pi (and other tools like Continue or Cline) talk to. Its lms CLI [3] makes this scriptable:
# start the local server โ http://localhost:1234/v1
lms server start
# load the model with the tuned settings (16 GB baseline)
lms load qwen3-coder-30b-a3b-instruct --gpu 0.60 --context-length 32768 --identifier coder -y
# confirm it's resident and see the real context window
lms ps
Note that Flash Attention and KV-cache quantization are set in the GUI load panel, not on the CLI โ enable them there once, and this lms load command reproduces the rest.
๐ Wire up pi
pi is a terminal coding agent with read, bash, edit, and write tools [2]. The community pi-lmstudio extension [4] bridges it to LM Studio and auto-discovers loaded models.
1. Install the extension โ it connects to http://127.0.0.1:1234 by default:
pi install npm:pi-lmstudio
2. Make it the default model by editing ~/.pi/agent/settings.json:
{
"packages": ["npm:pi-lmstudio"],
"defaultProvider": "lmstudio",
"defaultModel": "qwen3-coder-30b-a3b-instruct"
}
3. Or pick it live inside pi with the /model command or Ctrl+P, searching for models prefixed lmstudio. Then just run pi in your project. For a remote or custom endpoint, add { "url": "โฆ" } to ~/.pi/agent/lmstudio.json.
๐ What to expect
I measured this driving pi against a real 121-file C# codebase. The encouraging finding: quality held up โ the model stayed grounded and did not hallucinate. Latency, not intelligence, is the ceiling, and it scales with how far a task fans out across files.
| Task type | Time | Result |
|---|---|---|
| Search / locate (ripgrep + read a couple files) | 27โ55 sec | Accurate; every cited file verified to exist |
| Single-file edit or explanation | ~1 min | The daily-driver sweet spot |
| Multi-file synthesis (trace a pipeline across 5 projects) | ~5 min | Correct โ all 9 cited classes were real, zero hallucination |
Rule of thumb: it's excellent for scoped work โ find X, explain Y, edit this file โ and tedious for broad autonomous runs that fan out across a whole repo. Keep tasks focused, and reach for a cloud model on the big fan-out jobs.
๐งฎ MoE beats dense on a 16 GB card
The 30B-A3B isn't the only model worth running here, and comparing it against a dense model of similar size exposes the most useful lesson for a VRAM-constrained machine. I tested two more on the same RTX 5080: Qwen3.6-35B-A3B (a Mixture-of-Experts model, ~3 billion active per token) and Qwen3.8-27B (a dense model, all 27 billion active). Neither fits fully in 16 GB, so both must offload to the CPU โ but the results diverge sharply.
| Model | Type | Disk (Q4) | Best config | VRAM | Speed |
|---|---|---|---|---|---|
| Qwen3.6-35B-A3B | MoE (~3B active) | 19.7 GB | 32K @ 0.62 | 14.4 GB | ~23 tok/s |
| Qwen3.8-27B | dense (27B active) | 15.7 GB | 16K @ 0.74 | 15.3 GB | ~8 tok/s |
The MoE model is larger on disk yet runs roughly 3ร faster and leaves more VRAM headroom. The reason is the active-parameter count: the MoE routes to only ~3B parameters per token, so the layers stranded on the CPU barely slow it down โ whereas the dense 27B pushes all 27B parameters through those CPU-bottlenecked offloaded layers on every single token. The dense model is also VRAM-tight (~15.3 GB, little safety margin), so it's the harder one to run safely.
The lesson: when you're VRAM-constrained and offloading is unavoidable, choose your model by active parameters, not total size. On a 16 GB card, a bigger MoE will usually beat a smaller dense model on both speed and headroom.
To see how that plays out in practice, I drove the 35B-A3B through pi on the same 121-file C# codebase and ran three tasks, timing each against the 30B-Coder for reference:
| pi task (real codebase) | Qwen3-Coder-30B | Qwen3.6-35B-A3B |
|---|---|---|
| Comprehension โ read a project, summarize it | 55 s | 28 s |
| Code search โ locate the logic that does X | 27 s | 63 s |
| Multi-file synthesis โ trace a pipeline across 5 projects | 301 s (9 steps) | 492 s (11 steps, with code) |
The two models decode at a similar per-token rate (~24 vs ~23 tok/s), so the wall-clock differences come from how much each model chooses to read and write. The 35B-A3B tends to produce more thorough, code-grounded answers โ so it's quicker when a task is bounded (comprehension) but takes longer when it digs deeper (search and synthesis). Its output was consistently richer; the one caveat is that on the longest synthesis it invented a single class name, where the more terse 30B stayed fully grounded. On balance the 35B-A3B is the one I'd run day to day โ just don't read wall-clock time as raw speed.
๐งช An 11 GB desktop: RTX 2080 Ti + 64 GB RAM
I repeated the exercise on a much older desktop: a Dell Precision 5820 with an 8-core / 16-thread Xeon W-2145, 64 GB of system RAM, and an RTX 2080 Ti with 11 GB of VRAM. This is a useful second data point because the GPU is several generations older and has 5 GB less VRAM than the laptop above, while the extra system memory gives large partially-offloaded models room to breathe.
The Windows desktop was already using roughly 4 GB of VRAM before loading a model, so copying the 16 GB settings was not an option. I kept the same 32K context, Flash Attention and Q8 KV cache, then measured real VRAM with nvidia-smi after each load rather than trusting LM Studio's estimate.
| Model | Load strategy | VRAM free | Decode speed |
|---|---|---|---|
| Qwen3.6-35B-A3B | 20% ordinary layer offload | 1.68 GB | 12.0 tok/s |
| Qwen3.6-35B-A3B | all shared layers on GPU; all expert layers on CPU | 4.14 GB | 24.2 tok/s |
| Qwen3.6-35B-A3B | all shared layers on GPU; 85% of expert layers on CPU | 1.54 GB | 28.6 tok/s |
| Qwen3.8-27B | 23% ordinary layer offload | 1.31 GB | 3.3 tok/s |
| Muse Glimmer 30B | 21% ordinary layer offload | 1.84 GB | 2.9 tok/s |
The surprising result is that the 2080 Ti can decode the 35B MoE faster than the newer 16 GB laptop result above. The winning layout is not ordinary partial layer offload. It puts the model's shared attention and other dense work on the GPU, keeps most expert blocks in the roomy 64 GB system RAM, and lets only the active experts cross the CPU path. LM Studio exposes that split as a separate numCpuExpertLayersRatio setting through its SDK:
{
contextLength: 32768,
maxParallelPredictions: 1,
gpu: {
ratio: "max",
mainGpu: 0,
numCpuExpertLayersRatio: 0.85
},
flashAttention: true,
llamaKCacheQuantizationType: "q8_0",
llamaVCacheQuantizationType: "q8_0",
offloadKVCacheToGpu: true,
keepModelInMemory: true
}
On this machine that became --n-gpu-layers 999999 --n-cpu-moe 35 in the underlying llama.cpp process. Moving every expert layer to the CPU was already twice as fast as naive 20% offload; allowing the remaining 15% of expert layers onto the GPU raised throughput again while preserving the target ~1.5 GB headroom. This is specific to MoE models: the dense Qwen3.8 and Muse models had no equivalent expert split and remained CPU-bound at roughly 3 tok/s.
pi results on the same 121-file C# codebase
I then ran the same three styles of pi task against the same 121-file C# codebase used for the earlier measurements. These are wall-clock agent times, not pure decode benchmarks: they include every file search, read, reasoning pass and generated token.
| pi task | Time | Accuracy |
|---|---|---|
| Comprehension โ summarize the CLI project | 5m 26s | Good; detailed and largely grounded |
| Code search โ locate one-line goal expansion | 7m 43s | Mixed; found the core API but invented two outer call-chain symbols |
| Multi-project synthesis โ trace the execution pipeline | 32m 52s | Poor; fluent but contained substantial invented architecture |
The runtime held close to 29 tok/s, so the long wall-clock times were not a decode problem. The agent chose to perform a great deal of exploration and then generated increasingly long answers. More importantly, answer confidence did not track correctness: the broadest trace looked the most authoritative while hallucinating nonexistent entry points, factory APIs, options and directory layouts. On this setup the 35B-A3B is excellent for bounded questions, but broad repository synthesis still needs aggressive scoping and independent verification.
๐ญ The next step: let the system tune it for you
Everything above is manual โ you pick the offload ratio and leave the headroom yourself. A recent research system, FreeToken [5], automates exactly that decision. Instead of a fixed GPU-offload ratio, it measures your machine's PCIe and CPU-memory bandwidth and then decides, per decode step, how many of the missing experts to stream to the GPU over PCIe versus compute directly on the CPU. It keeps the full expert pool in system RAM as the authoritative copy and treats VRAM purely as a cache โ in its framing, GPU memory "affects only performance, never correctness."
I tried it on this same laptop, and it's genuinely clever to watch: it benchmarked my bandwidth, auto-selected a hybrid backend, and chose to fetch ~34% of each decode step's expert misses over PCIe with the rest on the CPU โ automatically splitting the 35B across GPU and CPU layers. Two honest caveats for a Windows machine like mine, though. First, it flips the bottleneck from VRAM to system RAM: the whole model has to live in your 32 GB, so the 35B sits right at the edge of what fits. Second, on Windows it falls back to a slower router kernel (the optimized one ships only for Linux), so you won't see the paper's headline throughput here. It's early research software and it wobbled at that size โ but the direction is unmistakable: the offload tuning this post does by hand is exactly what these systems are starting to do for you automatically.
๐งญ Conclusion
A 30B-class MoE coding model on a 16 GB laptop is a real, private, offline pair-programmer โ provided you respect the one rule: leave VRAM headroom instead of maxing the GPU offload. Get that right, set a 32K context so agent harnesses don't choke, and wire it into pi, and you have a capable local assistant for scoped coding tasks that never sends a byte to the cloud.
If you can spare the download, Qwen3.6-35B-A3B is the one I'd actually run: on the same card it was faster and more detailed than the 30B-Coder on real tasks, while a similarly-sized dense model (Qwen3.8-27B) crawled at a third of the speed. That's the lasting takeaway โ on a VRAM-constrained machine, a bigger Mixture-of-Experts model beats a smaller dense one, because only its handful of active parameters ever touch the slow CPU-offloaded path. Pick by active parameters, leave your headroom, and a 16 GB laptop is a genuinely useful local coding box.
And the ceiling keeps rising. Even in the short time I spent on this, systems like FreeToken [5] are already turning the manual dial-twiddling above into something automatic โ and coaxing frontier-sized models onto plain consumer hardware. If a 30B feels this capable on a laptop today, I'm genuinely excited to see where local coding lands a year from now. This space is moving fast, and it's well worth watching.
References
- Qwen Team. "Qwen3-Coder-30B-A3B-Instruct." Hugging Face, 2025. Available: https://huggingface.co/Qwen/Qwen3-Coder-30B-A3B-Instruct
- Zechner, Mario. "pi โ a coding agent for the terminal." GitHub Repository, 2025. Available: https://github.com/mario-zechner/pi-coding-agent
- LM Studio. "lms โ LM Studio's CLI for models, server, and runtime." LM Studio Developer Docs, 2025. Available: https://lmstudio.ai/docs/developer
- pi-lmstudio. "LM Studio integration extension for the pi coding agent." npm Registry, 2025. Available: https://www.npmjs.com/package/pi-lmstudio
- Yang, S., et al. "FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution." arXiv preprint arXiv:2608.16157, 2026. Available: https://arxiv.org/abs/2608.16157