Single agents hit walls. Multi-agent systems hit each other. Here are the orchestration patterns that survive real workloads — supervisor, pipeline, swarm, and the evaluation loops that keep them honest.
You don't need a data centre to run capable language models. Here's the exact stack — Ollama, llama.cpp, GGUF quantisation and VRAM maths — for running 7B to 70B models on one consumer GPU.