LLM digest: July 2026
Sent
I published 69 posts on my blog in July. Here's your sponsors-only summary of the most important trends and highlights from the past month.
As always, this issue and previous issues are archived in my simonw-private/monthly GitHub repository.
Accidental cyberattacks by OpenAI and Anthropic models under test
By far the spiciest story this month concerned one of OpenAI's research models escaping its sandbox and directly attacking Hugging Face in an attempt to find the solutions to a cybersecurity eval it was running.
On 16th July Hugging Face announced that an agentic system, source unknown, had breached their systems. On 21st July OpenAI admitted that it was their fault. They had tasked a new model with the ExploitGym benchmark and turned off all of the model-level guardrails, and rather than solving the benchmark the old-fashioned way, the model had instead found a zero-day exploit in a proxy it had access to, broken out of the proxy, then exploited another application (an insecure app hosted on Modal) and used that to stage several days of attacks against Hugging Face to try and exfiltrate the answers to the benchmark!
I wrote about this in detail in OpenAI’s accidental cyberattack against Hugging Face is science fiction that happened. Hugging Face later published a detailed technical timeline of what had happened - I have some extra notes on that here.
This incident inspired Anthropic to go and double-check the logs of their own cybersecurity benchmark runs... and they found that there were three previously undetected incidents where one of their models-under-test had taken advantage of a misconfigured sandbox and exploited real-world targets as well, dating back to April this year!
In my notes on that one I said:
It's abundantly clear now that running evals of cyberattack potential in models is a spectacularly risky business. Every AI lab needs to pay attention to this. Keeping a close eye on what's happening in those sandboxes is crucial.
There's been quite a bit of breathless press coverage of these incidents, because "AI breaks out and attacks other companies" is such a perfect paperclip-maximizer science fiction plot.
It's worth remembering that in all of these cases the models were doing what they had been designed to do: they were given a goal and the tools to accomplish that goal. The real problem came from misconfigured and leaky sandboxes and a lack of monitoring from their research teams.
Cybersecurity horror stories aside, there were some very significant model releases in July:
GPT-5.6 Sol, Terra, and Luna
OpenAI released their GPT 5.6 family with a new naming scheme based on the latin names for objects in our solar system. Sol is their answer to Claude Fable 5, and comes within spitting distance of that model on some independent benchmarks, beating it in others.
I've been using both Fable and Sol extensively over the past few weeks and it's honestly hard to distinguish between them. They're both excellent, and capable of taking on significantly meatier tasks than their predecessors.
Terra and Luna are the smaller models. I didn't find these particularly interesting until just a few days ago, when Luna had a 80% price drop which OpenAI credit to inference optimization work by Sol. Luna's new pricing is $0.20/million tokens input and $1.20/million output - that makes input 1/5th of the price of Anthropic's Claude Haiku 4.5 ($1/$5). Luna is also now cheaper than Google's Gemini 3.1 Flash-Lite, previously my top pick for a cheap model ($0.25/$1.50).
I'm now running Luna on my Datasette Agent demo instance and it's proving fast, effective and extremely inexpensive for both crafting SQL queries and generating HTML+JavaScript for Datasette Apps. This looks to be my new workhorse model for building my own LLM-powered applications.
OpenAI's other big launch was GPT-Live, a new voice model built on "full-duplex architecture" which is already shipping as the advanced voice mode in the ChatGPT mobile and desktop apps. It's much better at being interrupted than previous models, and also has the ability to delegate harder questions to GPT-5.6, which means it's much less like talking to a model with a knowledge cut-off frozen two years ago.
Claude Opus 5
Anthropic's big release in July was Claude Opus 5 on the 24th. The price is the same as Opus 4.8 ($5/$25), half that of Fable 5 ($10/$50). It's a very effective model, and is a good new default for Claude Code if you don't want to burn through your Fable allowance too fast.
It's also the model that Fable downgrades to if it gets the slightest hint of a security or biology related question - I had Fable switch to Opus when I asked the difference between tusks and teeth the other day!
Anthropic's guide to prompting Opus 5 is worth a look. These new models can be negatively affected by custom prompts aimed at older models.
Anthropic also announced they had discovered cryptographic weaknesses in significant algorithms using Claude Mythos Preview. They published their prompts, which amusingly included lines like "again we are not looking for low hanging fruit, we want proper research to find genuinly hard findings".
Kimi K3 and DeepSeek-V4-Flash-0731
We had some very significant open weight model releases in July.
On the 16th Chinese AI lab Moonshot previewed their Kimi K3 via their website and API, calling it their "most capable model to date, with 2.8 trillion parameters". In benchmarks it rivaled the best models from OpenAI and Anthropic - the first time an open weights model had matched as opposed to lagged those frontier models.
On the 27th they released the weights, albeit licensed such that businesses with revenue that "exceeds 20 million US dollars" would need a special arrangement with Moonshot.
Then on July 31st DeepSeek released the unfortunately named DeepSeek-V4-Flash-0731, a 304 billion parameter, 167GB model which appears to punch well above its weight. At that size it's going to be feasible to run on high-end consumer hardware, especially once the quantized versions start to land. It's looking very good on the Artificial Analysis "cost per intelligence" chart.
Until recently the common wisdom was that the Chinese AI labs were generally around 6 months behind the USA labs. With Kimi K3 I don't think that holds any more.
Open letters about AI development
Open Weights and American AI Leadership was shepherded by Microsoft, dated July 24th, and signed by 235 AI-adjacent companies including NVIDIA, Amazon, Y Combinator, The Linux Foundation and (a later signer) OpenAI.
It's clearly an argument designed to counter any instincts by the current US government to ban or limit open weight models over "safety" concerns - a reasonable consideration given what happened to Claude Fable 5!
Relying solely on closed models is not inherently safe: they can be breached, misused, or fail in ways that outsiders cannot detect. And concentrating advanced AI capabilities behind a small number of closed models compounds that risk. It results in a small number of single points of failure, weakens competition, and leaves critical technology in the hands of a few providers. Open weight models, on the other hand, allow a broad community of researchers and developers to examine their behavior, identify vulnerabilities, develop safeguards, and improve them over time.
The one surprising note in the letter is that it comes out in support of distillation, where models train on output from other models:
In shaping this ecosystem, policymakers should be careful not to conflate legitimate model-development techniques with misappropriation. Distillation, or the practice of using one model’s outputs to help train or improve another, is a widely used technique for model improvement, evaluation, and validation. It reflects a long tradition of learning from, building upon, and improving existing technologies, a tradition that has helped drive innovation since the rise of the open-source software movement.
Notably absent from the signatures: Anthropic, who published their own response Our position on open-weights models three days later. CEO Dario Amodei doubled down on the risk of authoritarian governments building "AI models that are more powerful than those built by the US", and models being "misused to carry out cyberattacks or biological attacks", and called for "a crack down on industrial-scale distillation operations", while also stating that "Anthropic has never advocated for a ban on open-weights models".
Then on July 28th Pacing the Frontier was published, featuring signatures from "1,324 employees of frontier AI companies" - with names like Jakub Pachocki (Chief Scientist, OpenAI), Ilya Sutskever (Safe Superintelligence Inc, previously OpenAI), Dario Amodei (Anthropic), Jack Clark (Anthropic) and more. Their core message:
We request that the U.S. government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.
Their concern is intense competitive pressure combined with accelerated AI progress caused by automated AI research - and given that Anthropic produce 80% of their code with Claude Code, OpenAI had Sol reduce their end-to-end serving costs by 20%, and Kimi K3 designed a chip to serve a nano model built on its own architecture, you can see why people are taking that risk more seriously right now.
A fireside chat and a podcast
At the AI Engineer World's Fair at the start of the month I hosted a Fireside Chat with Cat and Thariq from the Claude Code team. My highlights from the conversation include notes about how Anthropic use Claude Code to build Claude Code, how their non-technical team members benefit from Claude Tag (their new Slack integrated agent), why they believe "auto mode" in Claude Code can prevent prompt injection disasters, and how they decide which features are worth shipping.
Last Monday I joined the Oxide and Friends podcast with Bryan Cantrill and Adam Leventhal to talk about the OpenAI cyberattacks (the Anthropic ones had not yet been revealed) and the state of open weight models. We also talked about Golden Gate Claude, the Zizians, Alameda wild turkey attacks, Soviet Marburg virus research, the Lead-crime hypothesis, and several other worthy digressions.
Reigniting my interest in MCP
On the 28th of July the Model Context Protocol project rolled out a new version of their specification, introducing "stateless MCP". Previously MCP had worked based on sessions, where the client would establish a session ID and then work within that session. The new version removes that requirement - MCP tools can now be executed as a single HTTP request, which is easier to scale and more pleasant to reason about from both the client and the server.
I wrote about how Stateless MCP has recaptured my interest, and introduced three new projects I built to try out the new specification: mcp-explorer, datasette-mcp, and llm-mcp-client.
I'd mostly lost interest in MCP due to its limitations compared to giving an agent arbitrary code execution, but MCP makes locking down the capabilities of agents a whole lot easier. As models get smarter, being able to limit their abilities at a finely-grained level becomes even more important. This month's accidental cyberattacks are a great illustration of why!
Other model releases
Three other model releases that caught my attention this month:
- 6th: tencent/Hy3 - A new 598GB Apache 2.0 licensed model from Tencent in China. Here's its Pelican.
- 9th: Introducing Muse Spark 1.1 - The first Meta Spark model to offer an API. Pelican, plus I released a llm-meta-ai 0.1 plugin for it.
- 16th: Inkling: Our open-weights model - Mira Murati's Thinking Machines Lab released their first open-weights model. 975B total parameters, 41B active. Pelican.
The pelican benchmark is really showing its age now. I wrote about what we can still learn from the pelican as part of my Kimi K3 coverage. The biggest limitation is that it does nothing to test tool calling or long-context tool workflows, which turn out to be the key to a useful agent-driving model in 2026.
My projects
I made significant progress on all three of my core open source projects in July - Datasette, LLM, and sqlite-utils.
I released sqlite-utils 4.0, adding support for SQLite schema migrations and much improved transaction handling, then sqlite-utils 4.1 with some smaller new features a few days later.
Datasette 1.0a36 added a user-facing bulk-insert CSV/JSON feature, the ability to create a table starting with CSV/JSON, and a flurry of fixes to the JSON API design ready for a 1.0 stable release. Datasette 1.0a37 was less significant, mainly focusing on performance improvements and some improved transaction support.
LLM got two release candidates for the upcoming version 0.32 - 0.32rc1 added a content-addressed message store - a completely new schema for storing prompts and responses in SQLite which can efficiently store forked conversations and better captures the rich data returned by modern LLM APIs. 0.32rc2 adds a new llm openai endpoint $URL ... command for running prompts against arbitrary OpenAI-compatible endpoints, and switches LLM's default model to GPT-5.6-Luna (it was previously GPT-4o-mini).
I also shipped a very early prototype of llm-coding-agent 0.1a0 - a coding agent built on LLM. Plus llm-meta-ai 0.1 and datasette-apps 0.1a4 and datasette-agent 0.4a0 and llm-chat-completions-server 0.1a0 and llm-mcp-client 0.1a0.
Almost every line of code I released in July was mediated through Claude Code (Opus 5 or Fable 5) or OpenAI Codex (GPT-5.5 or GPT-5.6 Sol). These Fable-class models can really fly.
I've been working for Jesse Vincent's Prime Radiant applied AI research lab one day a week, most recently building out an evals framework for small models. We've now released the first version - it's called smevals, and you can read about it on the Prime Radiant blog.
What I'm using at the moment
I remain equally split between Claude Code and Codex, in both cases mainly using their respective desktop apps.
I'm currently paying $100/month for an OpenAI subscription and $200/month for a Claude subscription. I've been using AgentsView and uvx agentsview serve to keep an eye on what I would have spent if I was paying full price. For that $300 in July I got $2,255.67 of total usage: $1203 for Fable, $652 for Sol, $219 for GPT-5.5, and $135 for Opus 5.
That doesn't tell the whole story though, since AgentsView can only see sessions that ran on my laptop. I expect my Claude Code for web usage would drive that number up quite a bit more for the Anthropic models.
I'm also now using Datasette Agent on an almost daily basis. It's my go-to tool for answering questions about data in SQLite:
uv tool install datasette --pre
uv tool install datasette-agent
datasette --internal internal.db -s plugins.datasette-llm.default_model gpt-5.6-luna --root data.db
And with Datasette Apps installed I can one-shot little data-driven web applications. GPT-5.6 Luna turns out to be really good at these.
I frequently use Claude.ai (or the Claude mobile app) for quick experimental prototypes - regular Claude Chat can clone public repos from GitHub, so it's great for quick brainstorming exercises or spikes.
GPT-5.6 Pro in ChatGPT is my new favorite research assistant, for serious and not so serious questions.
That's it for July!
If this newsletter was useful feel free to forward it to friends who might find it useful too, especially if they might be convinced to sign up to sponsor me for the next one!