Blog
Guides on running models on your own hardware: how they are trained, which machine they fit, and how to set them up.
How models are trained
How LLMs Are Trained: Pretraining, Fine-Tuning, and Alignment
A plain-English map of how large language models are trained. First they read and guess the next word. Then people teach them to answer, behave, and sometimes think first.
LLM Pretraining Explained: Base Models and Continued Pretraining
Pretraining in plain English: the model reads the web and guesses the next word, which makes a base model. Continued pretraining is the same practice on one field.
Supervised Fine-Tuning Explained: Instructions, LoRA, QLoRA, and DoRA
Supervised fine-tuning in plain English: show the model the answer you wanted. Covers instructions, chat, tools, and LoRA, QLoRA, and DoRA.
LLM Alignment Explained: RLHF, DPO, GRPO, and RLAIF
Alignment in plain English. RLHF, DPO, GRPO, and RLAIF all push a model toward better answers. They differ in who gives the grade.
Reinforcement Fine-Tuning Explained: Why Reasoning Models Think First
Reinforcement fine-tuning in plain English: grade the final answer, and the model learns to think before it replies. Covers DeepSeek-R1 and OpenAI's o-series.
Hardware and model picks
Best Local LLM for Coding in 2026
A practical 2026 guide to the best open-weight models for coding on your own hardware. Compares Qwen3 Coder, DeepSeek Coder V3, Codestral 2, and Gemma 3, with VRAM tiers and quantization picks.
Open-Weight Reasoning Models in 2026
DeepSeek-R1, Qwen3 in thinking mode, and QwQ 32B on math, coding, and logic, with VRAM requirements and when reasoning mode is overkill.
AMD Radeon for Local LLMs in 2026: Where ROCm Stands
Running local LLMs on AMD Radeon. Compares the RX 9070 XT, 7900 XTX, and Strix Halo against NVIDIA, and where ROCm still trails CUDA.
Choosing the Right Quantization for Local LLMs in 2026
Q4_K_M, Q5_K_M, Q6_K, Q8_0, and the newer IQ formats on quality, VRAM, and speed, with recommendations by hardware tier.
Best Local AI Models by Use Case (2026 Guide)
A curated 2026 guide to the best local AI models for chat, coding, reasoning, vision, lightweight use, and embeddings.
How to Run Gemma on Your Phone in 2026
Which Gemma size fits which phone, the best apps for Android and iOS, real-world tokens-per-second numbers, and setup steps.
Apple Silicon vs RTX 4090 for Local LLMs
M3/M4 Max and M-Ultra versus the RTX 4090 on throughput, capacity, ecosystem, power, noise, and cost.
VRAM vs Unified Memory: When Each Wins for Local AI
Bandwidth, capacity, and sustained throughput, and which wins for chat models, 70B models, and long-context workloads.
Setup and tools
llama.cpp vs Ollama vs LM Studio vs Hugging Face vs MLX
What llama.cpp, Ollama, LM Studio, Hugging Face, and Apple's MLX each do, how they relate, and which to use when.
How to Run AI Models Locally on Windows, Mac, and Linux
A beginner guide for installing Ollama, downloading a model, verifying it runs locally, and using LM Studio.
Ollama vs LM Studio vs RunLocalModel
How the main local runtimes compare, and why a hardware check comes before you download a model.
How It Works: Methodology and Data Sources
How RunLocalModel estimates VRAM requirements and inference speeds, and where the numbers come from.
Jev
What Is Jev? A Beginner's Guide
Jev is TypeSafe AI's System One model: it returns typed choices, scores, and probabilities instead of text.
How to Use Jev: A Practical Support Ticket Example
A practical tutorial that routes a support ticket, checks uncertainty, and hands the reply to a local or hosted language model.
Using Jev with ChatGPT, Claude, and Local Models
Put Jev in front of a language model: pick the model, drop dead context, screen the turn, and send only the uncertain cases onward.