Tag
#llm
Every story tagged llm, newest first.

Which local LLMs fit in 8, 12, 16, 24 or 32GB of VRAM
At Q4_K_M and 8K context, 8GB comfortably holds Gemma 2 9B (5.8 GB) and 24GB holds Gemma 3 27B (17.5 GB). Full tables for every common VRAM tier, computed the same way our GPU checker computes them.
Ahmad J · Sep 6, 2026 · 8 min read

Evaluate whether a model is good enough for your task
Stop guessing from vibes. A repeatable way to decide if a model clears the bar for your specific job, using your own data.
Ahmad J · Sep 4, 2026 · 3 min read

Choosing the right model size for your task
Bigger is not automatically better. A decision framework for matching model size to the job, the latency budget, and the hardware you actually have.
Ahmad J · Sep 4, 2026 · 4 min read

Claude Fable 5.1 costs 25% less, and its per-token price did not move
Anthropic's new model is cheaper for the reason nobody put in the headline: input and output rates are unchanged, and the whole discount is a 75 per cent cut to cache reads. That makes the saving conditional on how much context your workload re-reads, and the published benchmark gains are just as uneven.
Ahmad J · Sep 2, 2026 · 5 min read

DeepSeek's first V4 vision model is MIT-licensed, and 307 GB
DeepSeek shipped an open-weight multimodal model under MIT this weekend. Its own files carry three things the coverage left out: the download size, the fact that every benchmark is self-reported, and a footnote saying two of the wins are over a model that cannot see.
Ahmad J · Sep 1, 2026 · 5 min read
More stories
News · aiA cheap model solved open math problems, once Google put a team around itSep 1, 2026 · 5 min read
Article · aiSmall AI Models Are Quietly Winning in ProductionAug 14, 2026 · 4 min read
Tutorial · softwareHow to structure prompts for reliable, parseable LLM outputAug 11, 2026 · 4 min read
Article · aiThe Real Cost of AI Is Inference, Not TrainingAug 10, 2026 · 9 min read
Article · securityThe Local Illusion: The Real Security Risks of Running a Local LLMAug 9, 2026 · 6 min read
Article · aiTokens, Explained: How Language Models Read Your Text and How You're BilledAug 7, 2026 · 8 min read
Article · aiMixture-of-Experts Models: How They Work and Why They Cut Inference CostsAug 6, 2026 · 10 min read