LLM digest: April 2026
Sent
I published 66 posts on my blog in April. Here's your sponsors-only summary of the most important trends and highlights from the past month.
As always, this issue and previous issues are archived in my simonw-private/monthly GitHub repository.
Opus 4.7 and GPT-5.5, both with price increases
Anthropic and OpenAI both released new flagship models this month, and both of them came with a price increase.
Opus 4.7 is the same price as Opus 4.6 on paper - $5/million input tokens and $25/million output. Opus 4.7 uses a new tokenizer, and when I upgraded my token counter tool I found 4.7 to use around 1.46x as many tokens as 4.6 against the same long piece of text content.
Anthropic claim that 4.7 can sometimes work out cheaper than 4.6 since it can use fewer reasoning tokens to solve the same problem. Artificial Analysis reported that Opus 4.7 used ~35% fewer output tokens than Opus 4.6 to run their benchmark suite.
GPT-5.5 was more transparent with the price increase: in the API it's 2x the cost of GPT-5.4, at $5/million input and $30/million output. That price goes up to $10/m and $45/m for input longer than 272,000 tokens.
How good are the new models? As usual, they're both incremental improvements. I've been more impressed by GPT-5.5 than Opus 4.7, and I've seen some people downgrade from Opus 4.7 back to Opus 4.6.
To my astonishment, Qwen 3.6-35B-A3B running on my laptop drew me a better pelican than Claude Opus 4.7! I mainly see this as evidence that the pelican benchmark's already loose connection to model quality has finally broken. Opus 4.7 also uses "adaptive reasoning" which removes the ability to control how hard the model "thinks" about a problem, and as a result I was unable to force it to think harder about drawing a better pelican.
In further pricing news, Anthropic started a clumsy A/B test to increase the starting price of Claude Code from $20/month to $100/month - but accidentally updated the public pricing page without any formal announcement which caused a great deal of confusion.
OpenAI took advantage of the situation to confirm that their Claude Code rival Codex will remain available on lower priced plans, including a free preview. Codex had major upgrades later in the month too - the desktop application can now automate a web browser and preview web code changes directly in the app.
Meanwhile both Windsurf and GitHub Copilot have dropped their per-request pricing in favor of the same token limits used by everyone else.
This looks to me like a textbook example of the Jevons paradox playing out. Price-per-token for LLMs has dropped like a stone for several years now, but those cheap tokens inspired new, much more token-hungry uses for the models in the form of coding agents. An LLM power user can burn hundreds of thousands of tokens in a session now.
Claude Mythos and LLM security research
On April 7th (prior to the release of Opus 4.7 on April 16th) Anthropic announced that they would not be releasing their best available model, Claude Mythos - at least not to the general public - due to its effectiveness at finding security vulnerabilities.
Instead they launched Project Glasswing, providing access to trusted partners to help them find and fix holes in widely used software before the inevitable storm of widely available models with the same capabilities.
Mozilla shipped Firefox 150 on April 22nd with fixes for 271 vulnerabilities found with the help of Mythos. Firefox CTO Bobby Holley wrote about the experience and optimistically declared that "Defenders finally have a chance to win, decisively".
The UK AI Safety Institute evaluated Mythos and found it to be credible. They later evaluated GPT-5.5 and found it to have similar security capabilities.
What's evident from this is that coding agents are really good at finding vulnerabilities now, and development teams need to take this into account. Both Cal.com and the UK's NHS have responded by reducing their commitment to open source, to widespread criticism from application security experts.
I started an ai-security-research tag on my blog to keep track of this rapidly developing area.
ChatGPT Images 2.0
Google's Nano Banana image generation models lost the crown to OpenAI's new ChatGPT Images 2.0 (aka gpt-image-2 aka "Image gen 2.0" - OpenAI continue to be very inconsistent with their naming.)
Sam Altman claimed the leap from gpt-image-1 to gpt-image-2 was equivalent to the leap from GPT-3 to GPT-5 and I think that holds up. I had fun using it to create Where's Waldo-style puzzles - Where's the raccoon with the ham radio? (ChatGPT Images 2.0).
This new image generation model can handle high resolution image output with large volumes of correctly rendered text. I think the most useful applications for this are visual explanations - things like diagrams and infographics.
Someone did have it draw a horse riding an astronaut riding a pelican riding a bicycle and ChatGPT Images added a road sign that read "WHY ARE YOU LIKE THIS".
More model releases
Outside of Anthropic and OpenAI, it was a busy month for model releases generally:
- Gemma 4 (April 2nd) - Google DeepMind shipped four open weight reasoning models: 2B, 4B, 31B, plus a 26B-A4B Mixture-of-Experts. All of them can handle image inputs, and the smaller two have native audio input, usable via MLX-audio. These models are excellent for their size - the 2B and 4B ones can run on an iPhone and the larger ones run comfortably on a 64GB Mac.
- GLM-5.1 (April 7th) - Z.ai's 754B parameter MIT-licensed flagship. It animated the pelican riding a bicycle without me asking it to, and then produced a truly spectacular NORTH VIRGINIA OPOSSUM ON AN E-SCOOTER (a new backup prompt). "Cruising the commonwealth since dusk".
- Muse Spark (April 8th) - Meta's first new LLM since Llama 4 over a year ago, albeit not open weights. This came with a new model harness at meta.ai with 16 tools that I explored in detail, including a
visual_groundingtool that can count objects in an image. - Gemini 3.1 Flash TTS (April 15th) - Google's new text-to-speech model with strong prompt-following ability. The example prompts in their prompting guide are amusingly elaborate.
- Qwen3.6-35B-A3B (April 16th) and then Qwen3.6-27B (April 22nd) - my new favorite local models, these run comfortably in around 20GB of RAM and feel comparable to the leading hosted models less than a year ago.
- DeepSeek V4 Pro and V4 Flash (April 24th) - the latest open weight (MIT) models from Chinese AI lab DeepSeek are very impressive. 1.6T total parameters / 49B active for Pro, 284B/13B for Flash. They're too big to comfortably run locally on most machines (though antirez has Flash running on 128GB MacBooks using 2-bit quantization) but are priced extremely competitively as hosted models.
- talkie-1930-13b - Alec Radford was instrumental in the creation of the GPT series at OpenAI but left that lab in 2024. Here's one of his first public projects since then - a 13B "vintage language model" trained entirely on pre-1931 English text. In addition to using out-of-copyright training data this has some very interesting research applications for figuring out how capable models can be without having a full copy of the modern web baked into their weights.
Other highlights from my blog
- Highlights from my conversation about agentic engineering on Lenny's Podcast - I went on Lenny Rachitsky's podcast where topics included agentic engineering, dark factories, OpenClaw and Kākāpō parrot breeding season.
- I explored the changes in system prompt between Claude Opus 4.6 and Opus 4.7.
- I published some notes Tracking the history of the now-deceased OpenAI Microsoft AGI clause
- The Zig programming language has a very strict anti-LLM policy for contributions. I wrote about their rationale for this - they want to invest their limited review time in evolving new human contributors to their project.
What I'm using, April 2026 edition
For the first time in over a year, I've switched my daily driver back from Anthropic to OpenAI. I'm using both GPT-5.5 and GPT-5.5 Pro heavily at the moment, and particularly enjoying their ability to fire off hundreds of web searches in answer to research questions.
I'm splitting my coding agent time about 50/50 between OpenAI Codex (both in the terminal and via their new Codex desktop app) and Claude Code. I'm still using Claude Code on the web extensively from my phone - I don't like OpenAI's Codex Cloud product nearly as much.
I'm spending $100/month with OpenAI and $100/month with Anthropic - previously I was paying $20/month to OpenAI and $200/month to Anthropic. If my allowance on one runs out I switch to the other.
For local models I'm mostly using LM Studio and llama-server (installed via "brew install llama.cpp"). Qwen 3.6 27B is my current local model of choice, but I haven't yet tried to use it to drive a full coding agent. I've heard of good results from people running it with Pi which has a much shorter system prompt than Codex or Claude Code.
That's it for April!
If you found this newsletter useful then feel free to forward it to friends who might find it useful too, especially if they might be convinced to sign up to sponsor me for the next one!