Skip to content

Topics

lm-studio

1 post

  1. Evaluating local models on one 12 GB GPU

    Five local models tested on one RTX 4070 (12 GB) with LM Studio and llama.cpp: Qwen3-VL-8B, Gemma 4 12B QAT, Gemma 4 26B A4B, Qwen3.8-27B and Qwen3.6-35B-A3B. Accuracy against 40 hand-labelled documents, time to first token, tokens per second, MoE expert offload tuning (1.1 to 52.8 tok/s), prompt-size limits and vision, thinking and tool-calling checks, as a method to repeat on your own card.

    8 min read