AI
Models, August 2026: price wars, daily drivers, and what I actually pick
As of 2026-08-01. I’ll refresh this kind of note about once a month, give or take a week when the labs decide to throw a parade.
Every few weeks someone asks me which model is “best,” as if there is a crown in a glass case and we all agree who owns it.
There isn’t. There are jobs. Build. Research. Poke holes. Keep sync from lying. Do a hundred cheap passes in the background without setting money on fire. July did not hand me a new monarch. It made agents cheaper, more pickable, and a little more honest about which product surface they actually live on. For someone who works with a roster instead of one magic window, that is the interesting news.
I’m going to stamp this with a date and tell you what I’m actually reaching for as of today. If a vendor’s chart disagrees with my take Tuesday, the chart can wait.
July, without the parade music
Mid-month, Grok 4.5 showed up as xAI’s coding-and-agents bet, trained alongside Cursor as a launch partner. That wording matters. I’m not going to invent a story about whose session logs went into the soup. What I care about is that it is in the picker, it is priced like something you can leave running, and it fits the harder lanes on my bench when Composer is not enough. (Full transparency: I consider Cursor my primary driver from a tools perspective right now, though I use several others as appropriate.)
Anthropic’s Claude Opus 5 landed as the new daily Opus story: close enough to Fable-class work for a lot of real days, without making Mythos the default personality of the week. Cursor’s Router arrived for Teams and Enterprise Auto mode. I’m on a solo plan, so I still choose by hand. Org routing is a different essay for a different paycheck.
OpenAI’s GPT-5.6 Sol / Terra / Luna finally left preview and then got repriced at the end of the month. Luna in particular got aggressively cheap on the API. Useful. Also easy to oversell. Sol is the ChatGPT paid-flagship roll-out. Terra and Luna are mostly Work / Codex / API creatures. They are not “whatever ChatGPT happens to be today.” And no, Cursor’s Auto-review did not secretly become Luna while I wasn’t looking. Different harness, different button.
Google refreshed the workhorse tier with Gemini 3.6 Flash and Flash-Lite. Meta put Muse Spark on a Model API in public preview. I’m interested; I’m also not going to tattoo geo and list prices on a launch post that doesn’t print them. DeepSeek’s late-July V4 Flash 0731 keeps the open-weight / cheap-API lane alive. Treat the new API surface like the public beta it is until they say otherwise.
And hanging over the whole month: Fable 5 got pulled under export controls in June and came back globally in early July, while Mythos stayed gated. That is less “which model do I click Tuesday” and more “why did the leaderboards feel drunk for three weeks.” Governance Notes can own the longer version. Here it is just context.
Benchmarks are directional. Diff review is still the sport.
What I’m actually clicking
On a normal week I do not chase the loudest name on X. I match a model to a lane, the same way I match an agent to a job.
For everyday building in Cursor, I still start with Composer. It is the default implementation lane for a reason. When the work gets gnarly (tests with teeth, sync edges, long tool loops), I reach for Grok 4.5. Headline API rates start around $2 / $6 per million tokens; long prompts bill higher, so check xAI’s table before you assume the brochure number. When I want Opus-class judgment without turning the day into a premium cosplay, Claude Opus 5 earns a seat.
Cheap background work lives where the cheap models actually live: Luna and Gemini Flash-Lite on their own stacks. Outside Cursor, Gemini 3.6 Flash (or Terra when I want that family) is my usual synthesis workhorse. Open-weight curiosity still goes to DeepSeek V4 Flash. It is not the brain of Synesis, and honestly doesn’t touch any of my “real” code, just some prototyping. It is the “am I being silly about cloud dependence” lab.
I also keep GitHub Copilot around for inline, in-editor help when that is the lighter hammer or I’m in VS Code. It is not my agent roster and it is not where multi-step bench work lives. Different tool, different verb. It’s quick. It’s painless when I don’t need anything deep. It’s middle of the pack when it comes to how I rate it on my bench.
Premium stays a deliberate ask, and I do my best to keep it that way. If I cannot say why I truly needed to step up, I should not have stepped up. Are the big guns in my arsenal? Sure. Do I generally need them? Not usually, honestly.
Watching what’s next
Gemini 3.5 Pro is announced and not GA. Gemini 4 is pre-training talk, which is not a ship date. DeepSeek’s fuller Pro story is still a follow-up. Cursor Router on individual plans is wish territory. Mythos access is policy, not marketing.
If it is not generally available, it does not get to pretend it is on my daily picker. I have been burned by “coming soon” enthusiasm before. I prefer a lack of excitement or overly hyped anticipation. At least, I prefer that until something actually gives me a true reason to be enthusiastic, which isn’t likely to be before it’s released in most cases.
Why bother writing this down
Models rotate. My jobs do not rotate as fast. A dated note keeps me honest: what moved, what I still click, what I’m refusing to overclaim. Sources over soundtrack. Yes, I re-check dates and prices before this post lands. I contain multitudes: curiosity, caffeine, and an unreasonable need for the number to still be true on publish day.
Next month I will either refresh the list or if nothing significant has changed, wander into the cousin questions: are local models worth the iron, and when do open weights beat closed ones for the work I actually do.
Until then: instruments, not vibes. Pick the tool for the cut. Ignore the crown.
~ Trish