Tag
#gpu
Every story tagged gpu, newest first.

AWS's new G7 instances won its own benchmark with half the GPUs
AWS benchmarked its Blackwell-based G7 instances against G5, G6 and G6e on 30B mixture-of-experts models. The two-GPU box beat the four-GPU boxes, the cheapest configuration and the fastest one turned out to be different machines, and the same model cost 2.5 times more per token on retrieval traffic than on chat.
Ahmad J · Sep 8, 2026 · 6 min read

Which local LLMs fit in 8, 12, 16, 24 or 32GB of VRAM
At Q4_K_M and 8K context, 8GB comfortably holds Gemma 2 9B (5.8 GB) and 24GB holds Gemma 3 27B (17.5 GB). Full tables for every common VRAM tier, computed the same way our GPU checker computes them.
Ahmad J · Sep 6, 2026 · 8 min read

How to Choose Hardware for Local AI
Ask about memory before compute: a model has to fit to run at all, and bandwidth sets the speed. CPU, discrete GPU, or unified memory. A short buying guide, plus an interactive picker that points you to the right path.
Ahmad J · Sep 1, 2026 · 3 min read

Small AI Models Are Quietly Winning in Production
Frontier models get the headlines, but inside real companies smaller, cheaper, faster models do the actual work. Here's how they win, where they don't, and what it costs to ignore them.
Ahmad J · Aug 14, 2026 · 4 min read

How to Build an AI Workstation on a Tight Budget
A working AI machine doesn't need a flagship build. The gate is one number, the memory on your GPU, so spend there first, buy the rest just good enough, and leave room to grow. Includes a simple rule for sizing VRAM to the models you actually want to run.
Ahmad J · Aug 13, 2026 · 3 min read
More stories
Tutorial · softwareSet up GPU drivers and the toolkit for local AI workAug 11, 2026 · 3 min read
Article · securityThe Local Illusion: The Real Security Risks of Running a Local LLMAug 9, 2026 · 6 min read
Article · businessWhat a Model Actually Costs to Run in Production: A Back-of-Envelope Framework for TeamsAug 9, 2026 · 8 min read
Article · aiWhy GPUs Beat CPUs for AI InferenceAug 8, 2026 · 7 min read
Article · hardwarePCIe 4.0 vs 5.0 vs Thunderbolt for AI Workloads: Where the Generational Upgrade Actually MattersAug 7, 2026 · 9 min read
Article · hardwareCUDA Lock-In Is Real: A Precise Cost Accounting of What Switching GPU Vendors Actually BreaksAug 6, 2026 · 9 min read
Article · aiHow Much VRAM You Actually Need to Run a Local LLMAug 5, 2026 · 7 min read