Skip to main content

3 posts tagged with "optimization"

View All Tags

My Model Usage, July 2026: Receipts for the Routing Theory

· 12 min read
Gergely Sipos
Frontend Architect

This is my token telemetry for July 2026 — one developer's month, pulled from the Aliz AI usage dashboard. Earlier this month I wrote You Don't Deserve Fable (Yet), which argued that routing discipline comes before reaching for the biggest model on the menu. This post is a month of receipts for that argument, including the parts where I plainly haven't applied my own advice. It's also a mixture of before and after: the routing discipline was still forming while the month ran, so this isn't a steady state. Some of the chart describes habits I'd already changed, and some of it describes habits I was in the middle of changing.

You Don't Deserve Fable (Yet)

· 5 min read
Gergely Sipos
Frontend Architect

Claude Fable is the most capable model Anthropic has shipped — above Opus class, frontier reasoning, genuinely impressive on hard problems. It's also the fastest way to burn through your token budget if you haven't built the discipline to use it correctly. Most developers will reach for it because it's the best, use it for tasks Haiku could handle, and wonder why their costs exploded. This post is about earning the right to use Fable by mastering the cheaper models first.

Copilot Token Efficiency: What the Platform Is Doing Behind the Scenes

· 4 min read
Gergely Sipos
Frontend Architect

Since June 1, every Copilot token has a dollar sign attached. You're thinking about model selection, prompt size, and whether that 200-line file really needs to be in context. Good — but while you optimize your side, the VS Code team has been shipping infrastructure-level improvements that cut token consumption and latency without any user action. Ryan Caldwell and Bhavya U published a deep dive on these changes, and the numbers are worth knowing. The platform is doing heavy lifting so you can focus on the strategies you control.