explore
Pick a thread.
Every story we have published, by topic. Tap a tile to filter, with no reloads, just the thread you want to pull.
ai
54 stories
Which local LLMs fit in 8, 12, 16, 24 or 32GB of VRAM
At Q4_K_M and 8K context, 8GB comfortably holds Gemma 2 9B (5.8 GB) and 24GB holds Gemma 3 27B (17.5 GB). Full tables for every common VRAM tier, computed the same way our GPU checker computes them.
Ahmad J · Sep 6, 2026 · 8 min read

How to Choose a Laptop for Local AI
Memory decides whether a model runs at all, bandwidth decides how fast, and thermals decide whether that speed lasts. A short buying guide, plus an interactive picker that turns your priority into a recommendation.
Ahmad J · Sep 5, 2026 · 4 min read

Choosing the right model size for your task
Bigger is not automatically better. A decision framework for matching model size to the job, the latency budget, and the hardware you actually have.
Ahmad J · Sep 4, 2026 · 4 min read

AI Agents Are Moving From Demos to Narrow Jobs
The viral agent demos promised software that does everything. What actually ships are agents scoped to one job with tight guardrails, and that narrowing is the point.
Ahmad J · Sep 3, 2026 · 3 min read

Gemini can now skim a video instead of watching every frame
Google replaced fixed-rate video ingest with an agentic loop that fetches only the moments it needs, and reports up to 88 percent fewer tokens. Every headline figure is an upper bound, the gains concentrate on long video, and the cost saving is just the token saving passed through.
Ahmad J · Sep 2, 2026 · 4 min read
More stories
News · aiWorld Labs' Atlas wins six of seven 3D benchmarks, and you cannot run itSep 2, 2026 · 5 min read
News · aiClaude Fable 5.1 costs 25% less, and its per-token price did not moveSep 2, 2026 · 5 min read
News · aiSmall Models Are Quietly Taking Over the Easy WorkSep 2, 2026 · 3 min read
Guide · aiWhich AI subscription should you actually pay for?Sep 2, 2026 · 7 min read
News · aiDeepSeek's first V4 vision model is MIT-licensed, and 307 GBSep 1, 2026 · 5 min read
News · aiA cheap model solved open math problems, once Google put a team around itSep 1, 2026 · 5 min read
Article · aiSmall AI Models Are Quietly Winning in ProductionAug 14, 2026 · 4 min read
Article · aiAI agents in production: the honest 2026 state of playAug 13, 2026 · 8 min read
Article · aiLocal vs. Cloud AI Processing: The Real Trade-OffsAug 12, 2026 · 4 min read
News · aiOpen-Weight Models Changed Who Controls the AI StackAug 12, 2026 · 3 min read
Tutorial · aiHow to fine-tune a small language model with LoRAAug 12, 2026 · 4 min read
Article · aiMemory Bandwidth, Not Compute, Limits Local LLM SpeedAug 11, 2026 · 5 min read
Article · aiWhat an AI Benchmark Actually Measures (And Why Leaderboards Mislead You)Aug 10, 2026 · 5 min read
Article · aiThe Real Cost of AI Is Inference, Not TrainingAug 10, 2026 · 9 min read
Tutorial · aiHow to build a basic RAG pipeline for a local LLMAug 10, 2026 · 4 min read
Article · aiHow LLM Context Windows Actually Work (and Why Bigger Isn't Always Better)Aug 9, 2026 · 4 min read
Article · aiPrompt Caching: How It Actually Cuts Your LLM API BillAug 8, 2026 · 3 min read
Article · aiWhy GPUs Beat CPUs for AI InferenceAug 8, 2026 · 7 min read
Article · aiLoRA and QLoRA Fine-Tuning: How to Customize LLMs Without Burning Your BudgetAug 8, 2026 · 8 min read
Article · aiTokens, Explained: How Language Models Read Your Text and How You're BilledAug 7, 2026 · 8 min read
Article · aiHow AI Coding Assistants Actually Work Under the HoodAug 7, 2026 · 6 min read
Article · aiQuantization Explained: What Q4, Q8, and FP16 Actually Do to a Local ModelAug 6, 2026 · 8 min read
Article · aiMixture-of-Experts Models: How They Work and Why They Cut Inference CostsAug 6, 2026 · 10 min read
Article · aiThe Real Difference Between MCP, Function Calling, and Agent LoopsAug 5, 2026 · 11 min read
Article · aiHow Much VRAM You Actually Need to Run a Local LLMAug 5, 2026 · 7 min read
Article · aiHow to Evaluate a Local LLM for a Real Task: A Repeatable Testing FrameworkAug 5, 2026 · 11 min read
Guide · aiThe best local LLM runners in 2026: Ollama, LM Studio, vLLM, and moreAug 4, 2026 · 9 min read
Guide · aiThe best home-server hardware for self-hosting AI in 2026Aug 3, 2026 · 13 min read
Guide · aiThe best cloud GPU providers for AI training in 2026Aug 3, 2026 · 11 min read
Guide · aiThe best AI note-taking and writing tools in 2026Aug 2, 2026 · 11 min read
Guide · aiThe best Macs for local AI and machine learning in 2026Aug 2, 2026 · 11 min read
Guide · aiThe best cloud hosting for running AI models in 2026Aug 2, 2026 · 11 min read
Guide · aiThe best vector databases for RAG in 2026Aug 1, 2026 · 12 min read
Guide · aiThe best CPUs for AI development workstations in 2026Aug 1, 2026 · 13 min read
Guide · aiThe best AI image generators in 2026Aug 1, 2026 · 13 min read
Article · aiModel Distillation: How Small Models Learn to Punch Above Their WeightJul 28, 2026 · 9 min read
Article · aiRAG vs Fine-Tuning: Which One Your Use Case Actually NeedsJul 27, 2026 · 10 min read
Article · aiWhat an AI Agent Really Is: Stripping Away the HypeJul 27, 2026 · 14 min read
Guide · aiThe best budget laptops for programming and AI work in 2026Jun 20, 2026 · 13 min read
Guide · aiThe best laptops for running local AI models in 2026Jun 20, 2026 · 12 min read
Guide · aiThe best GPUs for running large language models locally in 2026Jun 20, 2026 · 10 min read
Guide · aiThe best mini PCs for local AI inference in 2026Jun 20, 2026 · 10 min read
Article · aiOpenAI Acquires Ona to Give Codex Agents a Persistent Home in Enterprise CloudsJun 19, 2026 · 3 min read
Guide · aiThe best AI coding assistants in 2026Jun 19, 2026 · 10 min read
Tutorial · aiHow to choose the right quantization for a local LLMMay 24, 2026 · 4 min read
Article · aiHow a transformer model actually worksMay 13, 2026 · 4 min read
Article · aiThe real difference between training and inferenceMay 12, 2026 · 4 min read
Article · aiWhat a context window actually isMay 11, 2026 · 4 min read
Article · aiWhat RAG actually is and is notMay 10, 2026 · 4 min read