Local AI: The LM Studio Surprise

A researcher observes a brain in a cage

Part 4: the conclusion of the “Local AI” series — The LM Studio surprise: same model, same hardware, 40-200x faster In Part 3, we pushed gemma4:e4b through four demanding tests. The results were impressive — excellent clinical reasoning, comprehensive summaries, working code. But the wait times were brutal: 5 to 16 minutes per response on … Read more

Local AI: Choosing the Right Model for Your Hardware

A researcher observes a brain in a cage

Part 2 of the “Local AI” series — Benchmarks and recommendations for CPU-only inference In Part 1, we installed Ollama and ran our first local model. Now the question becomes: which model should you actually use? The answer depends on your hardware and your patience. On CPU-only systems, model size directly impacts speed — and … Read more

Local AI: Running LLMs Without Dedicated Graphics

A researcher observes a brain in a cage

Part 1 of the “Local AI” series — A practical guide to running local LLMs on CPU-only consumer hardware You’ve probably heard that running AI models locally requires expensive hardware: a Mac with Apple Silicon, or a gaming PC with a high-end NVIDIA GPU. This isn’t entirely wrong — those setups are faster. But they’re … Read more