Running LLMs on a Single GPU — The Practical Playbook
You don't need a data centre to run capable language models. Here's the exact stack — Ollama, llama.cpp, GGUF quantisation and VRAM maths — for running 7B to 70B models on one consumer GPU.
AI-curated · Hand-verified · Updated daily
Discover is a fast, beautiful technology magazine covering AI, open source, home labs, local LLMs, science and New Zealand — hand-curated and continuously updated.
Editor's picks
You don't need a data centre to run capable language models. Here's the exact stack — Ollama, llama.cpp, GGUF quantisation and VRAM maths — for running 7B to 70B models on one consumer GPU.
Astro + Cloudflare Pages is the cheapest, fastest stack on the web: free hosting, global edge CDN, automatic HTTPS, and deploys that take under a minute. Here's the complete setup, including custom domains and DNS.
Single agents hit walls. Multi-agent systems hit each other. Here are the orchestration patterns that survive real workloads — supervisor, pipeline, swarm, and the evaluation loops that keep them honest.
You don't need a seminary degree to study the Bible seriously. Free tools put the Hebrew and Greek originals, interlinear texts and manuscript history within reach of anyone — here's how scholars actually work, and how you can too.
No account? No problem. Ship a live serverless API on Cloudflare's free tier in 20 minutes — install Wrangler, write a Worker, deploy, and test it with curl.
Fresh off the press
Browse by topic
Curated reads
Everything you need to go from zero to running your own AI stack — models, tools and workflows.
Take your data back. Docker, NAS, dashboards and private services that run forever.
Static sites, edge deployment and modern frontend engineering — fast by default.
Hydroponics, home growing and tech from Aotearoa New Zealand.
Signal, not noise