signalSimon Willison's Blog2026-09-28
2026 in LLMs (so far)
In a keynote at WeAreDevelopers World Congress, the speaker reviewed 2026's LLM trends so far, emphasizing that November's Claude Opus 4.5 and GPT-5.1 releases made coding agents like Claude Code and Codex reliable for daily use. He also mentioned his pelican SVG benchmark, the first commit to the Warelay repository, and that 40 of 277 conference sessions addressed sandboxing or agent security.
- for who
- Developers and teams adopting LLM-based coding agents
- why now
- Coding agents with Claude Opus 4.5 and GPT-5.1 are now reliable enough for daily use.
- what changes
- Developers can now trust coding agents for daily tasks, allowing them to pursue more ambitious projects.
- to do
- Use Claude Opus 4.5 or GPT-5.1 with coding agents to handle daily coding tasks reliably.
key points
- Claude Opus 4.5 and GPT-5.1 made coding agents reliable for daily use
- Pelican SVG benchmark remains a challenge even for state-of-the-art models
- 40 of 277 conference sessions addressed sandboxing or agent security
#llm trends#coding agents#model releases
score
score 7 out of 10. 0-10: how dense the facts are, multiplied by how much you can do with them after reading. 8+ means the topic's evidence bar is met: benchmarks and availability for a new model, amount and investors for a funding round, revenue figures for a solo-money story. Below 5 an item does not enter the digest. A press release scores 3 or less, a reprint loses 2, anything older than 14 days loses 1, a headline that misleads loses 3.
read the source