September was mostly about bringing AI systems under control: running them on your own hardware, keeping their memory honest, and building development workflows that hold up against agent output.
The dominant thread was self-hosting and the tradeoffs that come with it — how much context fits on a 16 GB card, which AMD backend to build, which runtime to standardize on, and where the open-model cost curve bends. Around it ran a quieter but important theme: agent memory, how it goes bad, and how to store and govern it locally.
1 Running local LLMs: performance, hardware, and cost
A big chunk of the month was about making models actually run on consumer and mid-tier hardware, with the numbers laid out rather than hand-waved. The recurring finding is that the bottleneck is rarely the model file — it is the KV cache, the backend, and how you size the context window.
- KV Cache on 16 GB GPUs: Making Long Context Actually Fit
- ROCm vs Vulkan for AMD Local LLM Hosting: 2026 Guide
- llama.cpp vs Ollama in 2026: Which Runtime Should You Run?
- The Efficient Frontier of Open Models: Finding the Sweet Spot in 2026
2 AI agents: memory, skills, and migration
The second cluster was about the moving parts that make agents usable in practice: what you choose to remember, how to stop memory from feeding on the model’s own inferences, which extension model fits your setup, and how to carry an agent between frameworks without losing state or secrets.
- Agent Skills vs MCP Servers: Decision Framework
- Self-Reinforcing Memory Loops in AI Agents: Causes and Fixes
- Mnemosyne for Hermes Agent: Local Memory Quickstart
- How to Migrate from OpenClaw to Hermes Agent Safely
3 Spec-driven development and the AI engineering stack
Several articles focused on the tooling that structures how agents write code. OpenSpec came up twice — the setup workflow and the problem of proposals that get rejected — and there was a look at a more opinionated stack built on top of Claude Code and how it composes with the rest.
- OpenSpec Quickstart: Install, Workflow, and Common Pitfalls
- OpenSpec Rejected Proposals: A Decision Memory Convention
- gstack: AI Software Engineering Stack
4 Model frontiers and deep research tooling
A smaller pair of pieces looked past the present: where architectures are heading once transformers hit their compute and data limits, and the landscape of self-hosted systems you can point a question at and get a sourced answer back.
- What Comes After LLMs? Mamba, Diffusion & World Models
- Self-Hosted Deep Research Systems: 12 Tools Compared
If one of these articles is useful to someone building AI systems, backend systems, infrastructure, or technical knowledge workflows, please forward this email or share the link with them.
Thanks for reading, Rost