Tag
#quantization
Every story tagged quantization, newest first.

AWS's new G7 instances won its own benchmark with half the GPUs
AWS benchmarked its Blackwell-based G7 instances against G5, G6 and G6e on 30B mixture-of-experts models. The two-GPU box beat the four-GPU boxes, the cheapest configuration and the fastest one turned out to be different machines, and the same model cost 2.5 times more per token on retrieval traffic than on chat.
Ahmad J · Sep 8, 2026 · 6 min read

Which local LLMs fit in 8, 12, 16, 24 or 32GB of VRAM
At Q4_K_M and 8K context, 8GB comfortably holds Gemma 2 9B (5.8 GB) and 24GB holds Gemma 3 27B (17.5 GB). Full tables for every common VRAM tier, computed the same way our GPU checker computes them.
Ahmad J · Sep 6, 2026 · 8 min read

DeepSeek's first V4 vision model is MIT-licensed, and 307 GB
DeepSeek shipped an open-weight multimodal model under MIT this weekend. Its own files carry three things the coverage left out: the download size, the fact that every benchmark is self-reported, and a footnote saying two of the wins are over a model that cannot see.
Ahmad J · Sep 1, 2026 · 5 min read

A cheap model solved open math problems, once Google put a team around it
Google's Antigravity agents cleared seven open problems in mathematics and theoretical computer science. The expensive model produced them and the cheap one reproduced three. What that took is the more interesting number.
Ahmad J · Sep 1, 2026 · 5 min read

Local vs. Cloud AI Processing: The Real Trade-Offs
Where you run AI, on the device or in a datacenter, decides its latency, cost, privacy, and control. A precise breakdown of the real trade-offs, why memory (not compute) limits local models, and why most 2026 systems route between both.
Ahmad J · Aug 12, 2026 · 4 min read
More stories
Article · aiMemory Bandwidth, Not Compute, Limits Local LLM SpeedAug 11, 2026 · 5 min read
Article · aiLoRA and QLoRA Fine-Tuning: How to Customize LLMs Without Burning Your BudgetAug 8, 2026 · 8 min read
Article · softwareVector Databases: How They Actually Work, and When You Don't Need OneAug 7, 2026 · 4 min read
Article · aiQuantization Explained: What Q4, Q8, and FP16 Actually Do to a Local ModelAug 6, 2026 · 8 min read
Article · aiHow Much VRAM You Actually Need to Run a Local LLMAug 5, 2026 · 7 min read
Article · aiHow to Evaluate a Local LLM for a Real Task: A Repeatable Testing FrameworkAug 5, 2026 · 11 min read
Guide · aiThe best vector databases for RAG in 2026Aug 1, 2026 · 12 min read