RSS/Atom Feed Analyzer
Analysis of https://simonwillison.net/atom/everything/
Feed fetched in 16 ms.
Content type is application/xml; charset=utf-8.
Feed is 170,815 characters long.
Warning Feed is missing an ETag.
Feed has a last modified date of Mon, 27 Jul 2026 23:39:04 GMT.
Feed is well-formed XML.
Warning Feed has no styling.
This is an Atom feed.
Feed title: Simon Willison's Weblog
Error Feed self link: http://simonwillison.net/atom/everything/ does not match feed URL: https://simonwillison.net/atom/everything/.
Warning Feed is missing an image.
Feed has 30 items.
First item published on 2026-07-27T23:39:04.000Z
Last item published on 2026-07-16T15:35:25.000Z
All items have published dates.
Newest item was published on 2026-07-27T23:39:04.000Z.
Home page URL: http://simonwillison.net/
Error Home page URL is on a different protocol: http:.
Warning Home page URL redirected to https://simonwillison.net/.
Home page has feed discovery link in <head>.
Home page has a link to the feed in the <body>
Formatted XML
<?xml version="1.0" encoding="utf-8"?>
<feed xml:lang="en-us" xmlns="http://www.w3.org/2005/Atom">
<title>Simon Willison's Weblog</title>
<link href="http://simonwillison.net/" rel="alternate"/>
<link href="http://simonwillison.net/atom/everything/" rel="self"/>
<id>http://simonwillison.net/</id>
<updated>2026-07-27T23:39:04+00:00</updated>
<author>
<name>Simon Willison</name>
</author>
<entry>
<title>moonshotai/Kimi-K3</title>
<link href="https://simonwillison.net/2026/Jul/27/kimi-k3/#atom-everything" rel="alternate"/>
<published>2026-07-27T23:39:04+00:00</published>
<updated>2026-07-27T23:39:04+00:00</updated>
<id>https://simonwillison.net/2026/Jul/27/kimi-k3/#atom-everything</id>
<summary type="html"><p><strong><a href="https://huggingface.co/moonshotai/Kimi-K3">moonshotai/Kimi-K3</a></strong></p>
As promised <a href="https://simonwillison.net/2026/Jul/16/kimi-k3/">earlier this month</a>, Moonshot have released the weights for their excellent 2.8 trillion parameter Kimi K3. They're a hefty 1.56TB on Hugging Face.</p>
<p>Kimi introduced their own janky <a href="https://huggingface.co/moonshotai/Kimi-K2-Instruct/blob/main/LICENSE">modified version of the MIT license</a> with K2 back in July 2025. That license just added this paragraph requiring attribution beyond a certain size of commercial entity:</p>
<blockquote>
<p>Our only modification part is that, if the Software (or any derivative works thereof) is used for any of your commercial products or services that have more than 100 million monthly active users, or more than 20 million US dollars (or equivalent in other currencies) in monthly revenue, you shall prominently display "Kimi K2" on the user interface of such product or service.</p>
</blockquote>
<p>The <a href="https://huggingface.co/moonshotai/Kimi-K3/blob/main/LICENSE">K3 license</a> no longer calls itself "modified MIT" and goes further, requiring a separate agreement with Moonshot for large "Model as a Service" businesses:</p>
<blockquote>
<p>If the Licensee or any of its affiliates operates a Model as a Service business, and the aggregate revenue of the Licensee and its affiliates exceeds 20 million US dollars (or the equivalent in other currencies) in total over any consecutive 12 months, the Licensee must enter into a separate agreement with Moonshot AI before using the Software or its derivative works for any commercial purpose.</p>
</blockquote>
<p>To Kimi's credit, they make no attempt to describe this as an "open source" license in their own materials, consistently using the term "open weight" in its place.</p>
<p>OpenRouter is already offering K3 <a href="https://openrouter.ai/moonshotai/kimi-k3">from 7 providers</a>, most of which are at the same $3/million input and $15/million output as Moonshot AI themselves.
<p>Tags: <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/llm-pricing">llm-pricing</a>, <a href="https://simonwillison.net/tags/llm-release">llm-release</a>, <a href="https://simonwillison.net/tags/ai-in-china">ai-in-china</a>, <a href="https://simonwillison.net/tags/moonshot">moonshot</a>, <a href="https://simonwillison.net/tags/kimi">kimi</a>, <a href="https://simonwillison.net/tags/janky-licenses">janky-licenses</a></p></summary>
<category term="ai"/>
<category term="generative-ai"/>
<category term="llms"/>
<category term="llm-pricing"/>
<category term="llm-release"/>
<category term="ai-in-china"/>
<category term="moonshot"/>
<category term="kimi"/>
<category term="janky-licenses"/>
</entry>
<entry>
<title>An opinionated guide to which AI to use to do stuff</title>
<link href="https://simonwillison.net/2026/Jul/27/an-opinionated-guide-to-which-ai-to-use-to-do-stuff/#atom-everything" rel="alternate"/>
<published>2026-07-27T21:55:53+00:00</published>
<updated>2026-07-27T21:55:53+00:00</updated>
<id>https://simonwillison.net/2026/Jul/27/an-opinionated-guide-to-which-ai-to-use-to-do-stuff/#atom-everything</id>
<summary type="html"><p><strong><a href="https://www.oneusefulthing.org/p/an-opinionated-guide-to-which-ai-b22">An opinionated guide to which AI to use to do stuff</a></strong></p>
It's interesting watching the evolution of Ethan Mollick's guide over time. </p>
<p><a href="https://www.oneusefulthing.org/p/using-ai-right-now-a-quick-guide">A year ago</a> it was still all about chat - ChatGPT, Claude, Gemini - with o3, Claude 4 Opus, and Gemini 2.5 Pro as the models and Deep Research as a useful alternative mode.</p>
<p>Today it's much more about agentic systems - "where the AI is capable of doing the equivalent of many hours of real human work in one go".</p>
<p>Gemini has fallen off Ethan's list, since Google still doesn’t have an established entry in the Codex/ChatGPT Work/Cowork category. <a href="https://gemini.google/overview/agent/spark/">Gemini Spark</a> has yet to prove itself!</p>
<p>Ethan offers a useful explanation of the ways you can give ChatGPT or Claude a computer to use:</p>
<blockquote>
<p>To use the computers provided by the AI companies, the mode you want is called ChatGPT Work in ChatGPT, and Cowork in Claude (the naming will not get less confusing, I am sorry to say). [...]</p>
<p>The most powerful way to use AI is to give it access to your computer. You do that by downloading the ChatGPT or Claude apps and picking a mode to use. ChatGPT's two agent modes are Work and Codex; Claude's are Cowork and Code. The names do not map onto each other in any way that will help you remember them. And yes, these use the same names as the Work and Cowork modes we discussed above, but operate differently, and have more features and capabilities because they can access your computer.</p>
</blockquote>
<p>I think the difference between ChatGPT Work on a mobile device and ChatGPT Work inside the desktop app (where it's effectively a less intimidating skin on top of Codex) is spectacularly unintuitive.</p>
<p>Short version: if you flip ChatGPT mobile from "Chat" to "Work" mode you get a version where its Code Interpreter container is no longer restricted from accessing the internet!
<p>Tags: <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/ethan-mollick">ethan-mollick</a>, <a href="https://simonwillison.net/tags/code-interpreter">code-interpreter</a>, <a href="https://simonwillison.net/tags/general-agents">general-agents</a></p></summary>
<category term="ai"/>
<category term="generative-ai"/>
<category term="llms"/>
<category term="ethan-mollick"/>
<category term="code-interpreter"/>
<category term="general-agents"/>
</entry>
<entry>
<title>An Inside Look at the Relay Market Powering Token Resellers and Fraud</title>
<link href="https://simonwillison.net/2026/Jul/26/relay-market/#atom-everything" rel="alternate"/>
<published>2026-07-26T19:30:54+00:00</published>
<updated>2026-07-26T19:30:54+00:00</updated>
<id>https://simonwillison.net/2026/Jul/26/relay-market/#atom-everything</id>
<summary type="html"><p><strong><a href="https://vectoral.com/blog/token-relay-market">An Inside Look at the Relay Market Powering Token Resellers and Fraud</a></strong></p>
Fascinating investigation by Matt Lenhard into the market that has grown up around reselling LLM tokens at a discount by pooling API keys from various sources.</p>
<p>This looks to be mostly a thing in China. Resellers sell access to an LLM proxy that offers significant discounts on regular API pricing, which they achieve by abusing free trials, proxying through unprotected support bots, or sometimes through stolen credit cards or chargeback attacks.</p>
<p>The software they are using for these proxies is open source - mostly <a href="https://github.com/songquanpeng/one-api">one-api</a> and its more actively developed fork <a href="https://github.com/QuantumNous/new-api">new-api</a>, both legitimate API proxy products which can be used to load. balance requests across a pool of API credentials.</p>
<p>The buyers are seeking cheap tokens, avoiding geo-restrictions, and in some cases collecting data for model distillation.</p>
<p>I've been cautious about exposing my own LLM-driven applications publicly out of fear of abuse leading to big token bills. The existence of this marketplace makes me even more cautious: there's now an entire ecosystem that can profit from finding a new unprotected endpoint to exploit.</p>
<p>LLM vendors <em>really</em> need to get better at offering strict caps for their API keys. I want my LLM apps to stop working the moment they hit a dollar threshold I've set for a period of time.</p>
<p>Here's <a href="https://www.v2ex.com/t/1196011">the (Chinese language) forum thread</a> that served as the principal source for Matt's article.
<p><small></small>Via <a href="https://news.ycombinator.com/item?id=49058993">Hacker News</a></small></p>
<p>Tags: <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/llm-pricing">llm-pricing</a>, <a href="https://simonwillison.net/tags/ai-ethics">ai-ethics</a>, <a href="https://simonwillison.net/tags/ai-in-china">ai-in-china</a></p></summary>
<category term="ai"/>
<category term="generative-ai"/>
<category term="llms"/>
<category term="llm-pricing"/>
<category term="ai-ethics"/>
<category term="ai-in-china"/>
</entry>
<entry>
<title>Ruff v0.16.0</title>
<link href="https://simonwillison.net/2026/Jul/25/ruff/#atom-everything" rel="alternate"/>
<published>2026-07-25T22:44:05+00:00</published>
<updated>2026-07-25T22:44:05+00:00</updated>
<id>https://simonwillison.net/2026/Jul/25/ruff/#atom-everything</id>
<summary type="html"><p><strong><a href="https://astral.sh/blog/ruff-v0.16.0">Ruff v0.16.0</a></strong></p>
Astral shipped a significant new version of their Ruff Python linting tool a few days ago on July 23rd. I noticed today because my various CI jobs all started failing thanks to new default Ruff checks and my unpinned <code>"ruff"</code> dev dependency.</p>
<p>From Brent Westbrook's announcement post:</p>
<blockquote>
<p>Ruff now enables 413 rules by default, up from 59 in previous versions.</p>
<p>Since Ruff's default rule set was last modified in <a href="https://github.com/astral-sh/ruff/blob/main/changelogs/0.1.x.md#breaking-changes">v0.1.0</a>, the number of rules in Ruff has grown from 708 to 968. Many of these rules catch severe issues, including <a href="https://docs.astral.sh/ruff/rules/load-before-global-declaration">syntax errors</a> and <a href="https://docs.astral.sh/ruff/rules/yield-in-init/">immediate runtime errors</a> but were not previously enabled by default. With the new rule set, Ruff will bring these issues and many others to your attention without any Ruff configuration.</p>
</blockquote>
<p>Here's a one-liner for trying it on any Python project:</p>
<pre><code>uvx ruff@latest check .
</code></pre>
<p>I ran the latest Ruff against my three biggest projects - <a href="https://datasette.io/">Datasette</a>, <a href="https://sqlite-utils.datasette.io/">sqlite-utils</a>, and <a href="https://llm.datasette.io/">LLM</a> - and it found <em>hundreds</em> of minor issues that breached the new default rules.</p>
<p>All three projects have very comprehensive test suites, executed in CI against Python 3.10 through Python 3.14, so upgrades like this are pretty safe. The following command did the bulk of the upgrades:</p>
<pre><code>uvx ruff@latest check . --fix --unsafe-fixes
</code></pre>
<p>Against <code>sqlite-utils</code>, that command reported:</p>
<pre><code>Found 1618 errors (1538 fixed, 80 remaining).
</code></pre>
<p>As an illustrative example, here are three of the remaining issues. Ruff does a nice job of explaining each one:</p>
<pre><code>DTZ005 `datetime.datetime.now()` called without a `tz` argument
--&gt; tests/test_duplicate.py:17:10
|
15 | "datetime_col" TEXT)""")
16 | # Insert one row of mock data:
17 | dt = datetime.datetime.now()
| ^^^^^^^^^^^^^^^^^^^^^^^
18 | data = {
19 | "text_col": "Cleo",
|
help: Pass a `datetime.timezone` object to the `tz` parameter
BLE001 Do not catch blind exception: `Exception`
--&gt; tests/test_plugins.py:16:12
|
14 | db.execute("select * from pragma_function_list()")
15 | return True
16 | except Exception:
| ^^^^^^^^^
17 | return False
18 | finally:
|
B018 Found useless attribute access. Either assign it to a variable or remove it.
--&gt; tests/test_update.py:46:5
|
44 | def test_update_invalid_pk(fresh_db, pk, update_pk):
45 | table = fresh_db["table"]
46 | table.insert({"id1": 5, "id2": 3, "v": 1}, pk=pk).last_pk
| ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
47 | with pytest.raises(NotFoundError):
48 | table.update(update_pk, {"v": 2})
|
</code></pre>
<p>Unsurprisingly, given Astral's <a href="https://simonwillison.net/2026/Mar/19/openai-acquiring-astral/">new home at OpenAI</a>, this output provides everything a coding agent would need to fix the problems.</p>
<p>I had Codex (GPT-5.6 Sol high) <a href="https://github.com/simonw/llm/pull/1557">upgrade LLM</a> and <a href="https://github.com/simonw/sqlite-utils/pull/814">sqlite-utils</a>, and Claude Code (with Opus 5) <a href="https://github.com/simonw/datasette/pull/2857">upgrade Datasette</a>.
<p>Tags: <a href="https://simonwillison.net/tags/python">python</a>, <a href="https://simonwillison.net/tags/ruff">ruff</a>, <a href="https://simonwillison.net/tags/astral">astral</a></p></summary>
<category term="python"/>
<category term="ruff"/>
<category term="astral"/>
</entry>
<entry>
<title>Quoting Boris Cherny</title>
<link href="https://simonwillison.net/2026/Jul/25/boris-cherny/#atom-everything" rel="alternate"/>
<published>2026-07-25T00:42:59+00:00</published>
<updated>2026-07-25T00:42:59+00:00</updated>
<id>https://simonwillison.net/2026/Jul/25/boris-cherny/#atom-everything</id>
<summary type="html"><blockquote cite="https://twitter.com/bcherny/status/2080713091688583312"><p>More than any of these eval scores, what is most exciting to me is something else: Opus 5 is our least prompt injectable model yet. It is a bit buried in the system card, but across PI evals and red teaming, Opus 5 is very hard to prompt inject successfully.</p></blockquote>
<p class="cite">&mdash; <a href="https://twitter.com/bcherny/status/2080713091688583312">Boris Cherny</a>, here's that <a href="https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf#page=73">System Card section</a>, page 73</p>
<p>Tags: <a href="https://simonwillison.net/tags/prompt-injection">prompt-injection</a>, <a href="https://simonwillison.net/tags/anthropic">anthropic</a>, <a href="https://simonwillison.net/tags/claude">claude</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/boris-cherny">boris-cherny</a></p></summary>
<category term="prompt-injection"/>
<category term="anthropic"/>
<category term="claude"/>
<category term="generative-ai"/>
<category term="ai"/>
<category term="llms"/>
<category term="boris-cherny"/>
</entry>
<entry>
<title>Introducing Claude Opus 5</title>
<link href="https://simonwillison.net/2026/Jul/24/introducing-claude-opus-5/#atom-everything" rel="alternate"/>
<published>2026-07-24T23:48:50+00:00</published>
<updated>2026-07-24T23:48:50+00:00</updated>
<id>https://simonwillison.net/2026/Jul/24/introducing-claude-opus-5/#atom-everything</id>
<summary type="html"><p><strong><a href="https://www.anthropic.com/news/claude-opus-5">Introducing Claude Opus 5</a></strong></p>
I've been offline <a href="https://en.wikipedia.org/wiki/Elkhorn_Slough">kayaking with sea otters</a> for much of today so I haven't had a chance to put Anthropic's new model Claude Opus 5 through its paces yet. The buzz is positive, and Anthropic's description of it as a "thoughtful and proactive model that comes close to the frontier intelligence of Claude Fable 5 at half the price" sounds promising. It's currently <a href="https://twitter.com/artificialanlys/status/2080777718933995967">leading the Artificial Analysis leaderboard</a>, in front of even Fable 5.</p>
<p>It's priced the same as Opus 4.8, and continues to offer a "fast mode" at twice the cost of the base model.</p>
<p>Based on this anecdote in the release post it sounds like it might be <a href="https://simonwillison.net/2026/Jun/11/fable-is-relentlessly-proactive/">relentlessly proactive</a>:</p>
<blockquote>
<p>On one Frontier-Bench task, Opus 5 was given a drawing of a machine part and asked to write code to rebuild it as a 3D FreeCAD model. However, in this task, the model was intentionally given no way to directly viewthe drawing. Opus 5 responded by writing its own computer vision pipeline to pull the geometry from the raw pixels, then reconstructed the full machine part.</p>
</blockquote>
<p>It's better at finding vulnerabilities but has deliberately not been trained on how to exploit them. Hopefully this means the US government won't shut it down!</p>
<blockquote>
<p>As with its predecessor, Opus 4.8, we’ve intentionally avoided training Opus 5 on cyber tasks. The model has nevertheless improved substantially on these tasks as a result of becoming more generally capable, and it comes close to Mythos 5 at <em>finding</em> cybersecurity vulnerabilities. However, it remains substantially behind Mythos 5 on the <em>exploitation</em> of those vulnerabilities—that is, in turning vulnerabilities into material cyber threats.</p>
</blockquote>
<p>Anthropic have published a <a href="https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5">prompting guide for Claude Opus 5</a>. Thariq Shihipar has also written <a href="https://claude.com/blog/the-new-rules-of-context-engineering-for-claude-5-generation-models">The new rules of context engineering for Claude 5 generation models</a>.</p>
<p>The <a href="https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fraw.githubusercontent.com%2Fsimonw%2Fllm-anthropic%2F8272dfee5bdb65d5c88eef083da3ad885539b7df%2Flog.md">first pelican I got</a> was missing the bicycle wheels; the <a href="https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fraw.githubusercontent.com%2Fsimonw%2Fllm-anthropic%2Ffeaab840ea20eb15e29d8f72a9e42feceb23876a%2Flog.md">second attempt</a> was better.
<p>Tags: <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/anthropic">anthropic</a>, <a href="https://simonwillison.net/tags/claude">claude</a>, <a href="https://simonwillison.net/tags/llm-release">llm-release</a></p></summary>
<category term="ai"/>
<category term="generative-ai"/>
<category term="llms"/>
<category term="anthropic"/>
<category term="claude"/>
<category term="llm-release"/>
</entry>
<entry>
<title>The first known runaway AI agent - or a very bad marketing stunt?</title>
<link href="https://simonwillison.net/2026/Jul/23/the-first-known-runaway-ai-agent/#atom-everything" rel="alternate"/>
<published>2026-07-23T22:53:08+00:00</published>
<updated>2026-07-23T22:53:08+00:00</updated>
<id>https://simonwillison.net/2026/Jul/23/the-first-known-runaway-ai-agent/#atom-everything</id>
<summary type="html"><p><strong><a href="https://martinalderson.com/posts/huggingface-openai-exploit/">The first known runaway AI agent - or a very bad marketing stunt?</a></strong></p>
Martin Alderson's commentary on the <a href="https://simonwillison.net/2026/Jul/22/openai-cyberattack/">OpenAI accidental cyberattack against Hugging Face</a> includes a couple of details I hadn't considered.</p>
<p>First, Hugging Face offers a truly rich target if you're trying to find potential vulnerabilities that require executing arbitrary code:</p>
<blockquote>
<p>Hugging Face has an <em>enormous</em> attack surface. They have more interfaces than I can count which run untrusted models and code. While they definitely have invested in defences, by nature of their operating model they do have many more opportunities to be attacked than many other services. I certainly don't envy their cybersecurity teams.</p>
</blockquote>
<p>Secondly, one of the things that has puzzled me is how OpenAI didn't notice that their sandbox had been so thoroughly breached by the agent. Surely they'd be monitoring network traffic closely?</p>
<p>Martin points out that:</p>
<blockquote>
<p>It's also likely they were running a huge amount of benchmarks simultaneously with ~unlimited token budgets - you want as many samples as possible to figure out how good a model is at a certain benchmark. It may also be they are testing various different checkpoints of the model too, understanding how the model is improving as it goes through the various training stages.</p>
</blockquote>
<p>The mistakes made by the OpenAI team running this benchmark are easier to imagine when you think about the scale at which benchmarks of this kind usually operate. For all we know they could have been subjecting a new model to dozens of benchmarks at the same time, in dozens of different environments.
<p><small></small>Via <a href="https://lobste.rs/s/nsnb4j/first_known_runaway_ai_agent_very_bad">Lobste.rs</a></small></p>
<p>Tags: <a href="https://simonwillison.net/tags/security">security</a>, <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/openai">openai</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/hugging-face">hugging-face</a>, <a href="https://simonwillison.net/tags/ai-security-research">ai-security-research</a></p></summary>
<category term="security"/>
<category term="ai"/>
<category term="openai"/>
<category term="generative-ai"/>
<category term="llms"/>
<category term="hugging-face"/>
<category term="ai-security-research"/>
</entry>
<entry>
<title>Quoting Seth Larson</title>
<link href="https://simonwillison.net/2026/Jul/23/seth-larson/#atom-everything" rel="alternate"/>
<published>2026-07-23T04:50:36+00:00</published>
<updated>2026-07-23T04:50:36+00:00</updated>
<id>https://simonwillison.net/2026/Jul/23/seth-larson/#atom-everything</id>
<summary type="html"><blockquote cite="https://blog.pypi.org/posts/2026-07-22-releases-now-reject-new-files-after-14-days/"><p>The Python Package Index (PyPI) now rejects new files being uploaded to releases that are older than 14 days. This restriction was <a href="https://github.com/pypi/warehouse/pull/19727">put in place</a> to prevent old and long-stable releases from being poisoned in case publishing tokens or workflows of PyPI projects were compromised. As far as we are aware this has not yet been abused, but there is no technical reason beyond that attackers weren't aware it was possible.</p></blockquote>
<p class="cite">&mdash; <a href="https://blog.pypi.org/posts/2026-07-22-releases-now-reject-new-files-after-14-days/">Seth Larson</a>, PyPI blog</p>
<p>Tags: <a href="https://simonwillison.net/tags/packaging">packaging</a>, <a href="https://simonwillison.net/tags/python">python</a>, <a href="https://simonwillison.net/tags/supply-chain">supply-chain</a>, <a href="https://simonwillison.net/tags/pypi">pypi</a>, <a href="https://simonwillison.net/tags/seth-michael-larson">seth-michael-larson</a></p></summary>
<category term="packaging"/>
<category term="python"/>
<category term="supply-chain"/>
<category term="pypi"/>
<category term="seth-michael-larson"/>
</entry>
<entry>
<title>Quoting Thomas Ptacek</title>
<link href="https://simonwillison.net/2026/Jul/22/thomas-ptacek/#atom-everything" rel="alternate"/>
<published>2026-07-22T23:59:01+00:00</published>
<updated>2026-07-22T23:59:01+00:00</updated>
<id>https://simonwillison.net/2026/Jul/22/thomas-ptacek/#atom-everything</id>
<summary type="html"><blockquote cite="https://twitter.com/tqbf/status/2080045032162173329"><p>I genuinely believe that if you took an open weights model from 2025 and built a pentest harness for it, it could do this kind of sandbox escape and scan/hack in most networks. This is only surprising because you assume OpenAI has sounder sandboxes.</p></blockquote>
<p class="cite">&mdash; <a href="https://twitter.com/tqbf/status/2080045032162173329">Thomas Ptacek</a>, doesn't think <a href="https://simonwillison.net/2026/Jul/22/openai-cyberattack/#resist-the-temptation-to-write-this-off-as-a-stunt">this even needs</a> a frontier model</p>
<p>Tags: <a href="https://simonwillison.net/tags/thomas-ptacek">thomas-ptacek</a>, <a href="https://simonwillison.net/tags/openai">openai</a>, <a href="https://simonwillison.net/tags/security">security</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/ai-security-research">ai-security-research</a>, <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/sandboxing">sandboxing</a></p></summary>
<category term="thomas-ptacek"/>
<category term="openai"/>
<category term="security"/>
<category term="generative-ai"/>
<category term="ai-security-research"/>
<category term="ai"/>
<category term="llms"/>
<category term="sandboxing"/>
</entry>
<entry>
<title>OpenAI’s accidental cyberattack against Hugging Face is science fiction that happened</title>
<link href="https://simonwillison.net/2026/Jul/22/openai-cyberattack/#atom-everything" rel="alternate"/>
<published>2026-07-22T23:51:33+00:00</published>
<updated>2026-07-22T23:51:33+00:00</updated>
<id>https://simonwillison.net/2026/Jul/22/openai-cyberattack/#atom-everything</id>
<summary type="html"><p>This story is wild. The short version: OpenAI were running a cybersecurity test against an unreleased model, with the model's guardrail features turned off. Rather than solve the test, the model broke its way out of OpenAI's sandbox, then found exploits to break <em>in</em> to Hugging Face, all so it could cheat on the test by stealing the answers.</p>
<p>Along the way it helped make the strongest case yet for how the imbalance of model availability is hurting our ability to secure our software.</p>
<h4 id="here-s-what-happened">Here's what happened</h4>
<p>We currently have three documents to help us understand what happened here.</p>
<ol>
<li>
<a href="https://arxiv.org/abs/2605.11086">ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?</a> is a paper published on 11th May 2026 describing ExploitGym, a new eval suite for LLM-powered agent systems.</li>
<li>
<a href="https://huggingface.co/blog/security-incident-july-2026">Security incident disclosure — July 2026</a> by Hugging Face on 16th July 2026 describes how they detected an attack from an "agentic security-research harness - used LLM still not known" that breached some of their systems.</li>
<li>
<a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/">OpenAI and Hugging Face partner to address security incident during model evaluation</a> from OpenAI on 21st July 2026 confesses that it was <em>their</em> agent harness that did this, and that they're working with Hugging Face to clean up the mess.</li>
</ol>
<h4 id="exploitgym">ExploitGym</h4>
<p>I hadn't seen the <a href="https://arxiv.org/abs/2605.11086">ExploitGym paper</a> before and it's a really interesting one. Authors from UC Berkeley, the Max Planck Institute, UC Santa Barbara, and Arizona State designed a new benchmark for evaluating models on their ability to turn a reported vulnerability into a concrete exploit. OpenAI, Anthropic, and Google provided feedback and helped run the benchmark against their models.</p>
<p>The benchmark "comprises 898 instances derived from real-world vulnerabilities that affected popular software projects" - including the Linux kernel and V8 JavaScript engine. The ExploitGym benchmark is <a href="https://github.com/sunblaze-ucb/exploitgym">available on GitHub</a>.</p>
<p>Here's the paragraph that best represents their benchmark results:</p>
<blockquote>
<p>Among all configurations, Claude Mythos Preview and GPT-5.5 achieve the highest success counts (157 and 120 successes, respectively), demonstrating that current frontier agents can exploit a substantial subset of real-world vulnerabilities under controlled conditions. GPT-5.4 also solves a notable 54 tasks, placing it in an intermediate tier. The remaining model–agent pairings solve fewer than 15 tasks each, underscoring that end-to-end exploitation remains challenging and sharply differentiates today’s frontier systems. Notably, Claude Opus 4.7 achieves fewer successes than Claude Opus 4.6 despite being a newer checkpoint, and does so at substantially lower cost on the full set. Trace inspection reveals that Claude Opus 4.7 and Gemini 3.1 Pro frequently conclude early after judging the target vulnerability non-exploitable.</p>
</blockquote>
<p>The paper also describes the approach they took to preventing the agents from cheating by going outside the parameters of the test. This becomes relevant in a moment!</p>
<blockquote>
<p>Outbound connections are restricted to a curated allowlist that permits routine package installation (Ubuntu apt repositories and PyPI) and fetching the toolchains required for building V8. All other external endpoints are blocked.</p>
</blockquote>
<p>The paper concludes with this (emphasis mine):</p>
<blockquote>
<p>Our results show that <strong>autonomous exploit development by frontier AI agents is no longer a hypothetical capability</strong>. While current agents are not yet reliable across all targets, they already <strong>exploit a non-trivial fraction of real-world vulnerabilities</strong>, including complex targets such as kernel components. This rapid emergence is itself a central finding, showing that capabilities that would have seemed implausible are now present in deployed frontier models.</p>
</blockquote>
<p>An important detail here: this paper isn't about discovering vulnerabilities; it's about being able to take those vulnerabilities and turn them into working exploits.</p>
<p>When Anthropic first restricted access to Mythos <a href="https://simonwillison.net/2026/Apr/7/project-glasswing/">back in April</a> they talked about this capability as well. A model that can act on vulnerabilities is a lot more dangerous than one that can just discover them.</p>
<p>One of the ways Fable differs from Mythos is that it's more likely to refuse to weaponize vulnerabilities in this way. I get the impression the US government did not understand that distinction when they banned Fable <a href="https://simonwillison.net/2026/Jun/16/fable-5-export-controls/">last month</a>.</p>
<h4 id="the-hugging-face-incident">The Hugging Face incident</h4>
<p>The first hint we got of the attack was in <a href="https://huggingface.co/blog/security-incident-july-2026">this blog post by Hugging Face</a> on 16th July 2026:</p>
<blockquote>
<p>A malicious dataset abused two code-execution paths in our dataset processing (a remote-code dataset loader and a template-injection in a dataset configuration) to run code on a processing worker. From there, the actor escalated to node-level access, harvested cloud and cluster credentials, and moved laterally into several internal clusters over a weekend.</p>
</blockquote>
<p>I hope they release more details about the code that pulled this off. I'm assuming this means packages using the <a href="https://github.com/huggingface/datasets">datasets library</a>, a Hugging Face project for bundling up and sharing datasets on their platform. That library used to execute arbitrary code but has been steadily locked down over time, with the <a href="https://github.com/huggingface/datasets/releases/tag/4.0.0">4.0.0 release</a> in July 2025 removing the <code>trust_remote_code=True</code> flag entirely.</p>
<p>Assuming the attack used that library it must have either abused pickle serialization in some way, found some other non-obvious code execution path, or (most likely) specified <code>datasets&lt;4.0.0</code> as the dependency.</p>
<blockquote>
<p>The campaign was run by an autonomous agent framework (appearing to be built on an agentic security-research harness - used LLM still not known) executing many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services.</p>
</blockquote>
<p>This was a sophisticated attack!</p>
<p>Then Hugging Face hit a wall: they tried to use "frontier models behind commercial APIs" - I'm guessing from Anthropic and OpenAI - to help analyze the attack, and were blocked:</p>
<blockquote>
<p>When we started the log analysis, we first used frontier models behind commercial APIs. This did not work: the analysis requires submitting large volumes of real attack commands, exploit payloads, and C2 artifacts, and these requests were blocked by the providers' safety guardrails, which cannot distinguish an incident responder from an attacker.</p>
</blockquote>
<p>They switched to their own self-hosted instance of MIT licensed GLM-5.2 and it helped them figure out what was going on.</p>
<p>This indicated a fundamental asymmetry between the defending team and the (so-far unknown) attacker:</p>
<blockquote>
<p>We do not know which model powered the attacker's agents, whether a jailbroken hosted model or an unrestricted open-weight one; either way, the attacker was bound by no usage policy, while our own forensic work was blocked by the guardrails of the hosted models we first tried.</p>
</blockquote>
<p>As a useful indicator of how seriously they took the attack:</p>
<blockquote>
<p>[...] Finally, we have also reported this incident to law enforcement agencies.</p>
</blockquote>
<p>So who was responsible for this "autonomous agent framework"? It turned out to be OpenAI themselves.</p>
<h4 id="the-openai-confession">The OpenAI confession</h4>
<p>Five days later, <a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/">on July 21st</a>, OpenAI revealed the culprit. They had been running the ExploitGym benchmark against a new, as-yet undisclosed model, and that model had been operating <em>way</em> outside its intended parameters (emphasis mine):</p>
<blockquote>
<p>After investigating, we now know <strong>that this particular incident was driven by a combination of OpenAI models</strong> — including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a <a href="https://arxiv.org/abs/2605.11086">benchmark</a> [ExploitGym] of cyber capabilities. [...]</p>
<p>We estimate maximal cyber capabilities by <strong>running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity</strong>. Our benchmarks run in a highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts as a proxy and cache for package registries.</p>
<p>The models <strong>identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure</strong> to obtain test solutions directly from Hugging Face’s production database. All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.</p>
</blockquote>
<p>It's pretty clear what happened here. OpenAI removed safety filters for an in-progress model, locked it up in a sandbox and told it to solve the ExploitGym problems. Given the absence of guardrails there was nothing to prevent the model from attempting to break out of that sandbox, break into Hugging Face, and read the answers from there instead.</p>
<p>OpenAI's sandbox used the now common pattern of allowing access to an allow-list of sites related to installing additional packages via a controlled proxy. In OpenAI's words:</p>
<blockquote>
<p>While operating in our sandboxed testing environment, our models <strong>spent a substantial amount of inference compute finding a way to obtain open Internet access</strong>, in pursuit of solving the evaluation problem. To gain access, the models identified and exploited <strong>a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy</strong>. With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access.</p>
</blockquote>
<p>So step one was to break out onto the public internet. The model then broke into Hugging Face to find the answers:</p>
<blockquote>
<p>After gaining Internet access, the models <strong>inferred that Hugging Face potentially hosted models, datasets and solutions</strong> for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation. In one example, <strong>the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities</strong> to find a remote code execution path on the Hugging Face servers.</p>
</blockquote>
<p>Chaining together multiple attack vectors is <em>exactly</em> the kind of thing these new models can do, where previous generations of models might have failed.</p>
<p>I wrote last month about how <a href="https://simonwillison.net/2026/Jun/11/fable-is-relentlessly-proactive/">Claude Fable is relentlessly proactive</a>, when I noticed it spinning up custom web servers and deploying CORS tricks on my own laptop just to help debug a WebKit CSS issue. It turns out relentless proactivity is the defining trait of this new generation of Mythos-class models. If you set them a goal and give them a way to get there, even inadvertently, they <em>will figure it out</em>.</p>
<h4 id="resist-the-temptation-to-write-this-off-as-a-stunt">Resist the temptation to write this off as a stunt</h4>
<p>There will inevitably be some people who dismiss this story as a dishonest marketing trick by OpenAI to make their models sound terrifyingly effective. I found 81 instances of the term "marketing" in <a href="https://news.ycombinator.com/item?id=48997548">the Hacker News discussion</a> of the incident.</p>
<p>To those people I say <em>pull your heads out of the sand</em> - you're now including Hugging Face in your conspiracy theories, just so you can deny the crescendo of evidence here!</p>
<p>The best models we have today have the ability to both find and exploit new vulnerabilities. The ExploitGym paper itself concludes that "autonomous exploit development by frontier AI agents is no longer a hypothetical capability", and this incident is a perfect example of exactly that.</p>
<h4 id="the-asymmetry-is-increasingly-frustrating">The asymmetry is increasingly frustrating</h4>
<p>One of the most infuriating details of this story is how Hugging Face, faced with an accidental and aggressive attack from one of OpenAI's models, were unable to then turn to OpenAI's models to help them fend off the attack.</p>
<p>The frontier models we have access to are increasingly being constrained in how much they can help us protect our software, heavily influenced by the US government's ongoing threat of export controls. Claude Fable 5 wouldn't even <a href="https://simonwillison.net/guides/agentic-engineering-patterns/prompts/#proofreader">proofread this article</a> for me! It insisted on downgrading me to a less capable model.</p>
<p>Meanwhile open weight models from China such as GLM-5.2, Kimi 3 and the new Qwen 3.8 Max appear to have none of these restrictions - and any restrictions that <em>do</em> exist can likely be fine-tuned out of them by modifying the weights</p>
<p>These constraints are meant to make us safer. I think there's a risk that they are having the opposite effect.</p>
<p>Tags: <a href="https://simonwillison.net/tags/sandboxing">sandboxing</a>, <a href="https://simonwillison.net/tags/security">security</a>, <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/openai">openai</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/hugging-face">hugging-face</a>, <a href="https://simonwillison.net/tags/anthropic">anthropic</a>, <a href="https://simonwillison.net/tags/paper-review">paper-review</a>, <a href="https://simonwillison.net/tags/ai-security-research">ai-security-research</a></p></summary>
<category term="sandboxing"/>
<category term="security"/>
<category term="ai"/>
<category term="openai"/>
<category term="generative-ai"/>
<category term="llms"/>
<category term="hugging-face"/>
<category term="anthropic"/>
<category term="paper-review"/>
<category term="ai-security-research"/>
</entry>
<entry>
<title>Are AI labs pelicanmaxxing?</title>
<link href="https://simonwillison.net/2026/Jul/22/are-ai-labs-pelicanmaxxing/#atom-everything" rel="alternate"/>
<published>2026-07-22T23:01:00+00:00</published>
<updated>2026-07-22T23:01:00+00:00</updated>
<id>https://simonwillison.net/2026/Jul/22/are-ai-labs-pelicanmaxxing/#atom-everything</id>
<summary type="html"><p><strong><a href="https://dylancastillo.co/posts/pelicanmaxxing.html">Are AI labs pelicanmaxxing?</a></strong></p>
Excellent piece of work by Dylan Castillo, who took a deep-dive into the frequently pondered question of whether the AI labs have been deliberately training models to draw pelicans riding bicycles in response to my <a href="https://simonwillison.net/tags/pelican-riding-a-bicycle/">deeply unscientific benchmark</a>.</p>
<p>I've been randomly spot-checking this in the past by testing models against other animals riding other types of vehicle, but never with anything close to the diligence of Dylan's methodology here.</p>
<p>Dylan took 8 animals × 6 vehicles = 48 prompts and ran them three times each through 7 different models ( GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.5 Flash, Grok 4.5, Qwen3.7-Max, GLM-5.2, and DeepSeek V4 Pro). He then used GPT-5.6 Luna and Gemini 3.1 Flash-Lite to help evaluate the results.</p>
<p>There's a neat filter view for exploring the results:</p>
<p><img alt="Screenshot of a grid for sample 1/3 of GLM-5.2, with pelicn and flamingo and heron riding bicycle, unicycle, skateboard, scooter, plane and boat" src="https://static.simonwillison.net/static/2026/pelican-grid.webp" /></p>
<p>For the models he tested he could find no evidence of pelimaxxing:</p>
<blockquote>
<ul>
<li><a href="https://dylancastillo.co/posts/pelicanmaxxing.html#evidence-1-the-pelicans-on-bicycles-dont-look-any-better">The pelicans on bicycles don’t look any better</a></li>
<li><a href="https://dylancastillo.co/posts/pelicanmaxxing.html#evidence-2-labs-are-not-better-at-drawing-pelicans">Labs are not better at drawing pelicans</a></li>
<li><a href="https://dylancastillo.co/posts/pelicanmaxxing.html#evidence-3-labs-are-not-better-at-drawing-bicycles">Labs are not better at drawing bicycles</a></li>
<li><a href="https://dylancastillo.co/posts/pelicanmaxxing.html#evidence-4-labs-are-not-better-at-drawing-pelicans-on-bicycles-even-adjusting-for-difficulty">Labs are not better at drawing pelicans on bicycles, even adjusting for difficulty</a></li>
<li><a href="https://dylancastillo.co/posts/pelicanmaxxing.html#evidence-5-the-pelican-bicycle-scenes-dont-look-memorized">The pelican-bicycle scenes don’t look memorized</a> [...]</li>
</ul>
<p>Pelicans aren’t drawn any better than other animals. Bicycles aren’t drawn any better than other vehicles. And no lab draws the combination better than its pelicans and bicycles already predict. GLM-5.2 comes closest: it has the largest boost on the exact pelican-bicycle cell, and and its first pelican-on-bicycle sample caught my eye. But the effect is small and not significant, so I wouldn’t put too much weight on it.</p>
</blockquote>
<p><small></small>Via <a href="https://news.ycombinator.com/item?id=49010129">Hacker News</a></small></p>
<p>Tags: <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/evals">evals</a>, <a href="https://simonwillison.net/tags/pelican-riding-a-bicycle">pelican-riding-a-bicycle</a></p></summary>
<category term="ai"/>
<category term="generative-ai"/>
<category term="llms"/>
<category term="evals"/>
<category term="pelican-riding-a-bicycle"/>
</entry>
<entry>
<title>Orchestrions</title>
<link href="https://simonwillison.net/2026/Jul/22/all-the-orchestrions/#atom-everything" rel="alternate"/>
<published>2026-07-22T14:48:52+00:00</published>
<updated>2026-07-22T14:48:52+00:00</updated>
<id>https://simonwillison.net/2026/Jul/22/all-the-orchestrions/#atom-everything</id>
<summary type="html"><p>San Francisco tip: it only costs around $15 ($10 in quarters plus a $5 bill for the self-playing violin) to activate every single Orchestrion in <a href="https://en.wikipedia.org/wiki/Musée_Mécanique">Musée Mécanique</a>.</p>
<p>And because most people are bad at allocating their funds you may well be the ONLY person activating the Orchestrions, which means you get to craft the soundscape for the entire museum.</p>
<p>Tags: <a href="https://simonwillison.net/tags/san-francisco">san-francisco</a></p></summary>
<category term="san-francisco"/>
</entry>
<entry>
<title>California Sea Lion</title>
<link href="https://simonwillison.net/2026/Jul/21/sighting-383713864/#atom-everything" rel="alternate"/>
<published>2026-07-21T19:51:03+00:00</published>
<updated>2026-07-21T19:51:03+00:00</updated>
<id>https://simonwillison.net/2026/Jul/21/sighting-383713864/#atom-everything</id>
<summary type="html"><p><img src="https://static.inaturalist.org/photos/702321069/large.jpg" alt="California Sea Lion"></p><p><img src="https://static.inaturalist.org/photos/702321114/large.jpg" alt="California Sea Lion"></p><p>California Sea Lion, in San Francisco County, US, CA</p><p>We took some visiting family to Pier 39 to see the sea lions. They're somehow always even more fun than I remember them being last time.</p>
<p>Tags: <a href="https://simonwillison.net/tags/san-francisco">san-francisco</a>, <a href="https://simonwillison.net/tags/wildlife">wildlife</a></p></summary>
<category term="san-francisco"/>
<category term="wildlife"/>
</entry>
<entry>
<title>Nativ: Run AI models locally on your Mac</title>
<link href="https://simonwillison.net/2026/Jul/21/nativ/#atom-everything" rel="alternate"/>
<published>2026-07-21T14:22:27+00:00</published>
<updated>2026-07-21T14:22:27+00:00</updated>
<id>https://simonwillison.net/2026/Jul/21/nativ/#atom-everything</id>
<summary type="html"><p><strong><a href="https://blaizzy.github.io/nativ/">Nativ: Run AI models locally on your Mac</a></strong></p>
Prince Canuma is the developer behind the excellent <a href="https://github.com/Blaizzy/mlx-vlm">MLX-VLM</a> Python library for running vision-LLMs using MLX on a Mac.</p>
<p>I'm really excited about his new project, which wraps MLX in a full macOS desktop application. It's similar in shape to LM Studio, providing both a chat interface and a localhost API server for accessing models.</p>
<p>The app picked up MLX models I had already tried that were present in my Hugging Face cache directory, which was a nice touch.
<p><small></small>Via <a href="https://news.ycombinator.com/item?id=48982681">Hacker News</a></small></p>
<p>Tags: <a href="https://simonwillison.net/tags/macos">macos</a>, <a href="https://simonwillison.net/tags/python">python</a>, <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/local-llms">local-llms</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/mlx">mlx</a>, <a href="https://simonwillison.net/tags/prince-canuma">prince-canuma</a></p></summary>
<category term="macos"/>
<category term="python"/>
<category term="ai"/>
<category term="generative-ai"/>
<category term="local-llms"/>
<category term="llms"/>
<category term="mlx"/>
<category term="prince-canuma"/>
</entry>
<entry>
<title>A Fireside Chat with Cat and Thariq from the Claude Code team</title>
<link href="https://simonwillison.net/2026/Jul/21/cat-and-thariq/#atom-everything" rel="alternate"/>
<published>2026-07-21T12:54:02+00:00</published>
<updated>2026-07-21T12:54:02+00:00</updated>
<id>https://simonwillison.net/2026/Jul/21/cat-and-thariq/#atom-everything</id>
<summary type="html"><p>Earlier this month I hosted a fireside chat session at the <a href="https://www.ai.engineer/worldsfair/2026">AI Engineer World's Fair</a> with Cat Wu and Thariq Shihipar from Anthropic's Claude Code team. We talked about Claude Code, Claude Tag, Fable, coding agent security, evals, tool design, and how Anthropic use these tools themselves.</p>
<p>The full video of the session is now available <a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g">on YouTube</a>. Below is an edited copy of the transcript, with extra links and my own bolded highlights.</p>
<iframe style="margin-top: 0.5em; margin-bottom: 1em;" width="560" height="315" src="https://www.youtube-nocookie.com/embed/uU5Gv2h8-9g" title="SimonThis Year in Claude" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen="allowfullscreen"> </iframe>
<p>A few top-level notes if you don't want to watch the video or wade through the whole transcript:</p>
<ul>
<li>Claude Tag (Claude's new collaborative Slack integration) now lands <strong>65% of the product engineering PRs</strong> for the Claude Code team.</li>
<li>Claude Code ships features to Anthropic employees first, and <strong>only ships the features that demonstrate user retention with that cohort</strong>
</li>
<li>Critical changes to Claude Code are still reviewed manually, but the team increasingly relies on automated code review for the "outer layers" of the product.</li>
<li>Adding examples to a system prompt is <strong>no longer best practice</strong> for models like Fable 5 or even Opus 4.8. The Claude Code system prompt recently <strong>reduced in size by 80%</strong>.</li>
<li>Likewise, lists of "<strong>don't do X and don't do Y</strong>" can reduce the quality of results from the latest models.</li>
<li>
<a href="https://en.wikipedia.org/wiki/Eating_your_own_dog_food">Dogfooding</a> inside Anthropic is called "<strong>ant fooding</strong>".</li>
<li>Anthropic <strong>really believe in their <a href="https://code.claude.com/docs/en/auto-mode-config">auto mode</a></strong>, and see that as an enabling technology for Claude Tag.</li>
<li>Thariq advises offsetting coding-agent-induced <a href="https://simonwillison.net/2026/Feb/15/deep-blue/">Deep Blue</a> by "<strong>being more ambitious</strong>" with the work you take on.</li>
<li>Fable is <strong>competent at editing video</strong>, and Thariq <a href="https://twitter.com/trq212/status/2064826394589442448">used it</a> to edit its own launch video.</li>
<li>Anthropic's culture of working (internally) in public is key to their success, as demonstrated by the way they use Claude Tag in their public Slack Channels.</li>
</ul>
<h4 id="how-has-what-you-do-day-to-day-changed-in-the-past-year-">How has what you do day-to-day changed in the past year?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=65s">1:05</a></p>
<blockquote>
<p><strong>Simon:</strong> Claude Code came out in February of last year — it's under a year and a half old, and it was originally just a bullet point on <a href="https://www.anthropic.com/news/claude-3-7-sonnet">the Claude Sonnet 3.7 launch</a>. <strong>How has what you do on a day-to-day basis changed in the past year</strong>, now that we have these coding agents that actually work for us?</p>
<p><strong>Cat:</strong> I remember when we first came out with Claude Code and Sonnet 3.7, you would give it a task and you would have to closely monitor every single little thing it tried to do. I would read every permission prompt extremely carefully. I would frequently say no — no, no, no, did you check this file? Did you check that file? And now it's been incredible with every model generation. I feel like <strong>we've all gotten a chance to take a step back and delegate a lot more of the menial implementation to Claude</strong>. It's freed up a lot of our time to think about more creative work, like: what is the right experience that we should be providing to our users, now that we know Claude Code can implement a lot of it? And now with Fable it's a totally different step change improvement. <strong>We see for a lot of our use cases that you can actually one-shot a ton of features with Fable now</strong>.</p>
<p><strong>Thariq:</strong> I remember the first text I got about Claude Code. One of my best friends was like, "You need to go try Claude Code." It was about when Opus 4 came out, and I tried it and I was like, "Oh, shit. I need to work at Anthropic now." And that was Opus 4 — great model, but you were reading permission prompts. It's kind of crazy how much amnesia we have, where I'm like, oh, auto mode has always been here, right? I don't even remember pressing yes and allow. For me, the big thing I'm trying to push myself on is that <strong>we have to do higher quality work than we've ever done before</strong>. The outputs are incredibly high quality. <strong>I've been using it to edit videos a bunch</strong>, and I'm like, okay, it has to meet the very exacting demands of our brand team in a couple of hours or we just can't do it. <strong>That's how I'm trying to shift with Fable: the best work we've ever done, faster than we've ever done it before</strong>.</p>
</blockquote>
<h4 id="what-piece-of-conventional-software-engineering-no-longer-holds-">What piece of conventional software engineering no longer holds?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=219s">3:39</a></p>
<blockquote>
<p><strong>Simon:</strong> What's a piece of conventional software engineering that was true a year ago that you don't think holds anymore in this new world?</p>
<p><strong>Cat:</strong> One of the biggest shifts we're seeing in the eng skill set: two years ago it was pretty typical for a product manager to go talk to a bunch of customers, align over the course of six months with cross-functional teams on some PRD, and write a thorough spec on exactly how we'll implement this before the first line of code gets written. Now things are completely turned the opposite way. For a lot of engineers, the push I would give to folks in the room is to <strong>develop more of your business sense and product sense on what it is we should build</strong>, because the timeline between having an idea and building it is so much shorter — it's down from six to twelve months to maybe even a week. That means all of us need to have better taste on what is worth building, what will actually inflect the businesses we're working on. So it's <strong>an increase in value on product taste and business sense</strong>, and a bit lower on execution in most product domains. Of course, for infra there's still a very heavy emphasis on making sure all the details are right.</p>
<p><strong>Thariq:</strong> For me, it's that <strong>rewrites are now good</strong>.</p>
<p><strong>Simon:</strong> The worst thing you could do is now actually fine!</p>
<p><strong>Thariq:</strong> Exactly. All the Mythical Man-Month stuff — never rewrite — I'm pro-rewriting now. If you have a good test suite — and <strong>I think the rewrite actually forces you to make sure you have a good test suite</strong> — but I think what people undercount is that <strong>a codebase is a spec, and maybe it's the only copy of the spec that you have</strong>, because no one knows every branching part of the codebase. You can take this as an artifact and distill it or create other versions of it. We <a href="https://bun.com/blog/bun-in-rust">rewrote Bun in Rust</a> and it works great — it's live for me right now.</p>
<p><strong>Simon:</strong> You're not shipping Claude Code on Bun-in-Rust yet, right?</p>
<p><strong>Thariq:</strong> Internally we have.</p>
</blockquote>
<p><em>(Actually it looks like Anthropic started shipping Claude Code on Bun-in-Rust to everyone <a href="https://simonwillison.net/2026/Jul/19/claude-code-in-bun-in-rust/">on June 17th</a>.)</em></p>
<h4 id="what-kind-of-things-are-non-engineers-doing-with-claude-tag-">What kind of things are non-engineers doing with Claude Tag?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=396s">6:36</a></p>
<blockquote>
<p><strong>Simon:</strong> The other big launch recently was <strong><a href="https://www.anthropic.com/news/introducing-claude-tag">Claude Tag</a></strong> — that's what, a week old now, at least for the rest of us. I understand it's being used at Anthropic by non-engineers a great deal. <strong>What kind of things are non-engineers doing with Claude Tag?</strong></p>
<p><strong>Cat:</strong> Claude Tag is a Claude that lives in your team's collaboration tools. We launched it last week within Slack. <strong>The thing that's different about Claude Tag is it's multiplayer by default</strong>. Once you add Claude Tag to a Slack channel, you can chime in, your teammates can chime in, and you can collaborate together on the PR. The other big difference is that it's proactive instead of reactive. You can tell Claude Tag, "Hey, monitor every bug report in this channel, put up a PR to fix it, and tag the engineer who last touched this part of the codebase," and it'll do it for the lifetime of the channel without you having to manually tag it in. And the third big shift is that <strong>we've <a href="https://claude.com/docs/claude-tag/users/memory">added team memory</a> into this</strong>. If you tell Claude Tag your preferences in the channel, it'll remember them for every future post. If you always want it to debug outages but you don't want it to debug warnings, just tell it that in natural language in the channel and it'll remember it for you and everyone else on your team.</p>
<p><strong>Internally, we see Claude Tag as the evolution of Claude Code.</strong> We see this as a large shift in how we work internally. <strong>Claude Tag currently lands 65% of our product eng PRs.</strong></p>
<p><strong>Simon:</strong> For all of Anthropic, or just for Claude Code?</p>
<p><strong>Cat:</strong> This is just for our product engineering team — <strong>our internal version of Claude Tag lands 65% of our product PRs right now</strong>. And this is a huge shift; this is more than 50% of our PRs. The way we see people split work between Claude Code and Claude Tag is: Claude Code is still the best place for your most complex tasks, when you're interactively iterating with the agent. <strong>But Claude Tag is great for having it work proactively on your behalf</strong>, so you no longer need to manually kick off Claude Code for all the bug reports that come up for features you're working on.</p>
<p><strong>Thariq:</strong> And for non-coding cases: for example, before this talk we asked Claude Tag, "Hey, when is Fable releasing?" We wanted to make sure we'd line it up with the announcement. Claude Tag would search our Slack and look at who's been saying what. <strong>As a search engine for your company, it's really valuable.</strong> It has all the context for your product, so you can ask it metrics-related questions — often when you're making decisions you want them informed by what the metrics say, so you hook it up to your event store. I've seen our marketing team do things like, "Hey, tell me about this feature." They're not programmers, but Claude is a programmer — it can clone the codebase and say, "This is the feature, this is what it looks like, <strong>this is a recording of me using the feature</strong>." It enables a whole wide variety of things, and I think we're still early in figuring that out.</p>
</blockquote>
<h4 id="claude-tag-as-the-team-collaborative-layer">Claude Tag as the team collaborative layer</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=606s">10:06</a></p>
<blockquote>
<p><strong>Simon:</strong> One of the problems I've had with coding agents is that I get how to use them as an individual, but I'm not really clear on how to use them in a team environment. <strong>It sounds like Claude Tag is your current answer to that team collaborative layer for this stuff.</strong></p>
<p><strong>Cat:</strong> Exactly. And a large percentage of our sessions are actually multiplayer right now. Maybe I say, "Hey, I think we should implement this new feature in Cowork," and I'll tag in Claude Tag to do a first pass at it. Then I'll tell Claude Tag, "Share a recording of your final implementation," and I'll tag in design to take a look. They'll nudge it, then pass it on to eng to take it to the finish line and get it out to prod. It's been this very fluid experience. <strong>We're still trying to iron out what the social dynamics are for steering the same session</strong>, but we've found that people just observe how others use it and follow those social norms — it's been pretty intuitive for us to integrate Claude Tag into our teams.</p>
<p><strong>Thariq:</strong> It's great for teaching people, and also for reducing slop, because <strong>the fact that everyone is seeing you use Claude together sort of levels up how you use Claude as well</strong>.</p>
</blockquote>
<p>This reminded me of how Midjourney solved the challenge of teaching people advanced image prompting by enforcing prompting in public in their Discord channels.</p>
<h4 id="how-do-you-decide-which-features-are-worth-building-when-building-is-so-much-cheaper-">How do you decide which features are worth building when building is so much cheaper?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=701s">11:41</a></p>
<p>Something I've found really hard myself is knowing when a feature is worth shipping now that the cost of actually building features has dropped so much.</p>
<blockquote>
<p><strong>Simon:</strong> How do you deal with the hardest problem in all of engineering — prioritization? <strong>How do you decide which features are worth building and shipping when building a feature is so much more inexpensive now?</strong></p>
<p><strong>Cat:</strong> This is the hard thing. There are a few ways we approach it. One is we dogfood our products every single day. Whenever there's something we want to be able to do in our products that we're not able to, instead of finding a different solution we fix our product so it can support that case. <strong>We have a very heavy dogfooding culture internally.</strong> Before we share our products with everyone in the world, we share them with everyone within Anthropic, and with some early customers who give us very honest feedback about it — the more brutal the better — and we iterate until people love it. <strong>We have an internal bar for the number of active users and the amount of retention a feature has to have before we share it with the world.</strong> Because this bar is very clear, every engineer knows what they're trying to hit. I think this also levels up our polish, because if the feature isn't polished, people will churn — and then we shouldn't ship that feature.</p>
</blockquote>
<p>Using internal user-retention to decide if a feature should ship makes a whole lot of sense to me.</p>
<h4 id="do-you-have-an-example-of-a-feature-which-surprised-you-">Do you have an example of a feature which surprised you?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=774s">12:54</a></p>
<blockquote>
<p><strong>Simon:</strong> <strong>Do you have an example of a feature which surprised you?</strong> You rolled it out and the engagement was off the charts — something unlikely to be shipped that turned into a real product thing.</p>
<p><strong>Cat:</strong> I do have one. <strong>A lot of folks on our team love <a href="https://code.claude.com/docs/en/remote-control">remote control</a>.</strong> Remote control lets you use your mobile device, or Claude in the web browser, to connect to a local Claude Code session running in your CLI. I never have this need, because I just kick off the task directly on mobile and it runs in a cloud session without using my local environment — I think because I'm doing very easy coding tasks. It was something I didn't totally understand; I was like, hey, people should just set up remote dev environments. But in practice, once we rolled out remote control, so many people I talk to told me that what they do every night is plug their laptop into a power charger, open a bunch of remote control sessions, lock the screen, <strong>and then use their mobile phone from their couch to control Claude Code</strong>. So this has become a flow we're now leaning into that I didn't originally get — but now I do.</p>
</blockquote>
<h4 id="does-a-human-review-every-line-of-production-code-in-claude-code-">Does a human review every line of production code in Claude Code?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=860s">14:20</a></p>
<p>One of the over-arching themes of the conference was review: how much attention to people spend to reviewing code written for them by coding agents. I was very keen to hear the Claude Code team's take on this!</p>
<blockquote>
<p><strong>Simon:</strong> How does code review work? <strong>Does a human being review every line of production code that makes it into Claude Code?</strong> And if not, what are you doing — how do you keep the quality up?</p>
<p><strong>Thariq:</strong> It varies on the task a lot. <strong>For important areas we have code owners.</strong> The system prompt is an example where we have a code owner — you really need to get their approval.</p>
<p><strong>Simon:</strong> So the code owner is directly responsible for the quality of that area of the code.</p>
<p><strong>Thariq:</strong> That's right.</p>
<p><strong>Cat:</strong> And they need to approve any PR that touches it.</p>
<p><strong>Thariq:</strong> We have <a href="https://code.claude.com/docs/en/github-actions">our code review GitHub bot</a> review everything — that goes on every PR, and often it's doing the bulk of the review. Something I've seen on the team is that <strong>for more complex PRs you might make an artifact to explain the PR</strong> so that other people can then review. And we invest a lot into verification, CI/CD, things like that, to make sure that any time anything fails we have a test. We have a really robust environment where Claude can control Claude Code and test it. So there's a multi-pronged approach to code review.</p>
<p><strong>Cat:</strong> In general, <strong>we are trying to move to a world where humans don't need to be in the loop</strong>. For the most critical changes to the core of Claude Code, and the cores of other products, there is always a code owner and they do manually review all the changes. But increasingly, <strong>for the changes at the outer layers, we actually have Claude code review fully review those</strong>. That sounds pretty scary, but we've had a six-plus-month-long process to get here, and <strong>there are baby steps that you take to build up trust with code review</strong>. In the beginning we had human review for everything, and then increasingly we would say, <strong>okay, for code changes that touch these files, code review is catching 100% of the issues there — so we actually don't need a human manually reviewing those</strong>. And when we have incident review, <strong>we look at the PRs that caused the incident and say, okay, how do we update code review to catch that?</strong> — and we take those PRs and <strong>add them to an eval set</strong> to make sure our future changes to code review never regress that metric. Removing humans from the code review loop is a big step forward. It can sound scary, and it's not something you can do overnight, but it is something you can do <strong>through many months of investment in the infrastructure</strong> to give you the confidence that code review is catching everything you care about.</p>
</blockquote>
<p>So the key seems to be constantly iterating on the automated review systems themselves, in order to build trust in them over time.</p>
<h4 id="how-does-a-new-model-affect-your-intuition-for-what-it-can-and-can-t-do-">How does a new model affect your intuition for what it can and can't do?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=1040s">17:20</a></p>
<p>We got <em>deep</em> into evals - another hot topic throughout the wider conference.</p>
<blockquote>
<p><strong>Simon:</strong> I know that Opus 4.8, if I ask it to build me a JSON endpoint that runs a SQL query and outputs JSON, is just going to get it right — that's not something I have to review closely. But then a new model comes along and I don't know how to build trust in Fable quickly, that it's not going to mess things up that Opus didn't. <strong>How does the new model affect your intuition for what it can do and what it can't do?</strong></p>
<p><strong>Cat:</strong> The main reason we're building up this <strong>eval base over time is so that new models can be a drop-in replacement</strong>. When we have a new model, we run the whole eval set and make sure that, for example, Fable is strictly better than Opus 4.8 — and that gives us the confidence to drop it in.</p>
<p><strong>Simon:</strong> Are those model evals for Anthropic as a whole, or Claude Code team-specific?</p>
<p><strong>Cat:</strong> We have both. We have evals on our team, and we run code review across every repo within Anthropic, so we have evals for that. And for things like auto mode, we not only have evals across every user within Anthropic — we've also commissioned multiple external testers to red team it, to create environments with prompt injections and malicious inputs, <strong>and make sure that auto mode doesn't let any of those pass</strong>.</p>
</blockquote>
<h4 id="how-do-you-build-confidence-that-a-system-prompt-tweak-results-in-better-output-">How do you build confidence that a system prompt tweak results in better output?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=1121s">18:41</a></p>
<blockquote>
<p><strong>Simon:</strong> I want to know if the system prompt improvement I made actually improved the product — that's the most basic form of product-specific eval, and I still don't have a great feel for how to do that. <strong>Is that something you're doing such that you have complete confidence that a tweak you've made to the system prompt results in better output?</strong></p>
<p><strong>Cat:</strong> <strong>We don't have complete confidence, but we do a lot to make sure that we don't regress performance.</strong> The starting point is a suite of external evals that we trust, and we complement that with an even larger suite of internal evals that we trust. To start, <strong>we mainly optimize for capability</strong>: given a complete definition of a task and the full codebase, does Claude make the right decisions, fully fix the bugs, and pass all the tests? That's the starting point and the thing we optimize for, because it's most directly what users want. But there are a lot of behaviors that impact how users feel when they work with Claude Code. For example, <strong>people really don't like it when Claude Code says it's time to go to sleep.</strong> Or people really don't like it when it says, "Hey, I finished two out of five parts — do you want me to continue?" Yes, please continue. <strong>So we're building up a set of behavioral evals to catch these.</strong> And as we get user feedback — please be loud with us about your user feedback — we rank the priority issues and go down one by one and build evals for each of them. It's not 100% coverage, but it is a priority for us to increase the coverage.</p>
</blockquote>
<h4 id="how-much-interaction-is-there-between-the-claude-code-team-and-the-model-training-teams-">How much interaction is there between the Claude Code team and the model training teams?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=1221s">20:21</a></p>
<blockquote>
<p><strong>Simon:</strong> <strong>How much interaction is there between the Claude Code team and the teams at Anthropic who are training the models in the first place?</strong> Is that quite a close collaboration?</p>
<p><strong>Cat:</strong> Across Anthropic, we all work quite closely together. We meet often to talk about what we expect the next generation of models to be able to do. Our research team has also been amazing about showing this publicly — we often talk in our blog posts about how <strong>we're targeting ever-increasing longer-horizon work</strong>, and how we train Claude itself to be honest, harmless, and helpful. We also put a lot of effort into making sure it's aligned with your intent, even if your intent is expressed in a fuzzy way. Of course, try your best to be specific about what you want, so Claude has all the context — but even when you're not specific, we teach Claude to make good assumptions. It's been a productive partnership.</p>
</blockquote>
<h4 id="the-system-prompt-has-been-reduced-by-80-what-have-you-been-able-to-drop-">The system prompt has been reduced by 80% — what have you been able to drop?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=1284s">21:24</a></p>
<p>So many useful prompting tips in this section!</p>
<blockquote>
<p><strong>Simon:</strong> Thariq, you <a href="https://www.youtube.com/watch?v=9fubhllmsBU&amp;t=358s">mentioned this morning</a> that the <strong>system prompt for Claude Code has been reduced by 80% because of Claude Fable</strong>. Can you go into a little more detail? <strong>What kind of things have you been able to drop?</strong></p>
<p><strong>Thariq:</strong> It wasn't just Fable — it was Opus 4.8 as well, and going forward, future models. We have different system prompts for different models now. One of the patterns we saw is that we were over-constraining Claude. The initial, maybe Opus 4-ish models wanted a lot of examples, and <strong>removing examples was extremely helpful</strong>, because it was just more creative than the examples we gave it.</p>
<p><strong>Simon:</strong> That's really interesting, because one of the top prompting tips I give people is: give it examples. If that's no longer true, that kind of breaks my prompting model a little bit.</p>
<p><strong>Thariq:</strong> Same here — I was surprised to hear that. I think now it's more about the shape of what you give it — the tools you give to Claude, your system prompt, things like that. The other thing we did is try to give it more context and <strong>fewer "do not do this"</strong> instructions, because that's a very strong impulse for Claude, and especially if it conflicts with user instructions later on, that can be extremely confusing to Claude — "I've got this skill that says this and the system prompt says this." So we try to <strong>have fewer hard constraints, more context, and fewer instructions overall</strong>. It's definitely a science — it took a bunch of evals to build.</p>
<p><strong>Cat:</strong> In general, when you're prompting these models, you should always think: <strong>are there edge cases to the instruction that I'm giving it?</strong> When we went back and reviewed all the instructions in the Claude Code system prompt, <strong>we found a few cases where yes, this statement is 90% true, but there's a real 10% of cases where it's not true</strong>. We didn't want to constrain the model, or confuse it into thinking it should always do this. One good example is verification. Everyone here wants Claude to verify its work, and we had some instructions in the prompt that said: if you make a front-end change, always verify. But there's a limit to it. If it's changing copy from one string to another string, and the user says "just make a quick fix and update the test," maybe you don't want to verify. <strong>So we've adjusted our wording from "always verify, verify, verify" to something like: most of the time when you're doing front-end work you can't fully understand the experience by hitting the backend endpoints, so when you make larger changes to the user experience, please run the app locally.</strong> And in fact, that instruction probably isn't even good either, because <strong>what is a large change?</strong> Maybe it should test small changes too. In general, whenever you give a prompt to the model, <strong>you should think about the ways in which it could be misinterpreted by a well-intentioned human</strong>, in order to better understand how the model might interpret it — and <strong>soften the prompt</strong> so that it's actually 100% accurate, because you're giving this prompt to the model 100% of the time.</p>
<p><strong>Simon:</strong> What's fascinating about that is you're <strong>relying on the model's judgment</strong> — and that's got to be an Opus/Fable-level thing. Models a year ago did not have the level of judgment necessary to decide whether they were going to test a change or not. But that does break down if you're building for a wide range of models and trying to run the cheaper models for cheaper tasks.</p>
<p><strong>Cat:</strong> We actually have <strong>a different system prompt per model now</strong>, for this very reason. It's only our most frontier models that have this 80% token decrease — the older models still have the full system prompt.</p>
<p><strong>Simon:</strong> Do you think Fable and Opus are smart enough to prompt Haiku with more details, because they understand that Haiku has less judgment, less taste?</p>
<p><strong>Cat:</strong> We haven't been able to eval it — we don't have any hard data to show it.</p>
<p><strong>Thariq:</strong> There's a tough thing with smaller models sometimes, because <strong>sometimes the larger models can be more token-efficient on a hard problem than the smaller models</strong>. So there's a bit of intuition to build there — sometimes you really just want frontier intelligence almost all the time. The Pareto curve shifts, and it's hard to find.</p>
<p><strong>Simon:</strong> A year ago I did not trust a model to write a prompt. Today the good models are very good at prompting — a lot of my prompts are written by models, which feels absurd but works really well. What helped me come to terms with that was thinking about subagents, which are entirely about a Claude model setting up a prompt for another Claude model.</p>
<p><strong>Thariq:</strong> <strong>Workflows</strong> are actually a really good example of this, because it's Claude not just prompting a single subagent, but prompting the orchestration of many subagents, and each one of them gets a very detailed prompt. It's almost a level above just spawning a subagent. I've also been using it on my personal machine, <strong>giving it the Gemini API and saying: here, generate images</strong>. It's way less lazy than I am at prompting an image model. It's just Claude prompting Claude all the way down.</p>
<p><strong>Cat:</strong> I think Claude also wrote the prompt for <a href="https://code.claude.com/docs/en/workflows">the workflow tool</a>.</p>
<p><strong>Simon:</strong> I've read that prompt — it's a good prompt. That's actually a frustration I have with Anthropic generally: you <a href="https://platform.claude.com/docs/en/release-notes/system-prompts">publish the prompts for Claude Chat</a>, but you don't include the tool prompts and the Claude Code prompts. I still have to run a proxy to intercept them. <strong>I would love it if the Claude Code prompts were deliberately published</strong> — they're the documentation. They're how you know what the tool can do and how it works.</p>
<p><strong>Cat:</strong> I'll write down that feature request. I'll have Claude Tag do it.</p>
</blockquote>
<p>Interesting to note that OpenAI's <a href="https://developers.openai.com/api/docs/guides/latest-model?model=gpt-5.6#favor-leaner-prompts">prompting best practices for GPT-5.6</a> includes similar advice for their latest models:</p>
<blockquote>
<p><strong>Favor leaner prompts</strong></p>
<p>Removing repeated instructions and examples and simplifying tool descriptions can improve task performance and token efficiency. In a sample of internal coding-agent eval runs, configurations with leaner system prompts improved evaluation scores by roughly 10–15% while reducing total tokens by 41–66% and cost by 33–67%.</p>
</blockquote>
<h4 id="what-s-your-bar-for-introducing-a-new-tool-">What's your bar for introducing a new tool?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=1686s">28:06</a></p>
<blockquote>
<p><strong>Simon:</strong> Claude Code is basically a big bag of tools. <strong>What's your bar for introducing a new tool?</strong> How do you decide when it's worth doing that additional engineering at that level?</p>
<p><strong>Cat:</strong> Do you want to take it? You introduced one of the best tools we have.</p>
<p><strong>Thariq:</strong> My career peaked when I introduced the ask user question tool. It's really hard. Especially for some tools — <strong>ask user question is Claude's tool to ask you</strong> — so it's hard to eval, and sometimes it's more of a user preference thing. Back then we had fewer evals, so it was very dogfooding based — or "ant fooding," our ant version of that. But overall <strong>we've been trying to trend towards fewer tools</strong>. The last set of tools we introduced was the task tool, I think — and we try to give Claude more general versions to do things.</p>
</blockquote>
<h4 id="what-s-the-latest-evolution-of-your-file-editing-tool-">What's the latest evolution of your file editing tool?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=1743s">29:03</a></p>
<p>I have a long-running fascination with file editing tools - they were the subject of the <a href="https://aider.chat/docs/leaderboards/edit.html">old Aider code editing leaderboard</a>, and I've watched with interest as they've evolved in different coding agents from search-and-replace based to line-number-based to more complicated patterns.</p>
<p>The Claude API docs describe a <a href="https://platform.claude.com/docs/en/agents-and-tools/tool-use/text-editor-tool">text editing tool</a> that's recommended for building against the API, but Claude Code seems to use slightly different approaches here.</p>
<blockquote>
<p><strong>Simon:</strong> One of the most interesting tools is the file editing tool — you can have file editing as a tool, or you can tell it to use sed and grep and do things that way. <strong>What's the latest evolution of your file editing tool?</strong></p>
<p><strong>Thariq:</strong> We still have one, but for example we removed our grep and other search tools — glob tools — in favor of native bash. Like I said in my talk earlier, <strong>the models are kind of more of a biology than a physics</strong>, and tool design especially is quite hard. I'm not sure if Cat disagrees and thinks there's a science to the eval of it, but I think tool design is more of an art, maybe — or a biology.</p>
<p><strong>Cat:</strong> I largely agree, but in general as we introduce more tools, we try to keep the cardinality pretty low and make sure that <strong>every tool we add has a distinct function from every other tool, so that Claude can very easily distinguish when to call each</strong>. For file edit, the reason we have it is actually because we can render it. We show people when Claude makes a file change, and there's this <strong>nice dedicated UI</strong> that says: do you approve this edit to this file? <strong>The reason we had a dedicated file edit tool was so that we could deterministically know</strong> that Claude was making a file change, so we could show people this nice UI. A lot of new users onboarding still really like this experience, so we've kept it around. But for a lot of us who are on auto mode right now — hopefully you're not on YOLO mode — I don't think it actually matters, and we could probably just remove file edit and be totally fine.</p>
</blockquote>
<h4 id="what-s-the-advice-within-anthropic-for-safely-running-claude-code-">What's the advice within Anthropic for safely running Claude Code?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=1858s">30:58</a></p>
<p>It's the <a href="https://simonwillison.net/tags/prompt-injection/">prompt injection</a> question! Who better than Anthropic employees to explain how Anthropic sees the risk of prompt injection attacks causing their Claude Code instances to run amok?</p>
<p>It turns out they <em>really</em> trust their <a href="https://code.claude.com/docs/en/auto-mode-config">auto mode</a> - and see that as the feature that enabled Claude Tag.</p>
<blockquote>
<p><strong>Simon:</strong> Let's talk about safety and security. I am deeply aware of the risks of prompt injection, and there are so many bad things that can happen if somebody else tells my Claude Code what to do. I still mostly run Claude Code in YOLO mode and feel incredibly guilty about it. <strong>What's the advice within Anthropic for safely running Claude Code?</strong></p>
<p><strong>Cat:</strong> Why not auto mode?</p>
<p><strong>Simon:</strong> I am starting to use auto mode, but I don't understand it enough to get how safe it is. As of maybe three weeks ago, I'm defaulting to auto mode.</p>
<p><strong>Cat:</strong> Broadly within Anthropic, almost every single person uses auto mode. It is the best way to do long-running work in Claude Code while being safe. <strong>We've done extensive bashing. We have thousands of evals. We've commissioned many red teamers to create adversarial environments in order to trick Claude Code into doing bad actions, and we've mitigated every single issue that they found.</strong> We're going to publish some evals in the coming weeks, but we've pretty much mitigated every attack.</p>
<p><strong>Simon:</strong> That is a big claim.</p>
<p><strong>Cat:</strong> We'll share the evals for it so folks can assess, but we've been extremely diligent about identifying all the ways in which Claude might mess up and then updating auto mode to counter it. It doesn't catch 100% of things — that would be way too strong a claim. But <strong>for the main categories of risks that we're concerned about, like prompt injection and data exfiltration, the risks are far lower than the average human reviewer</strong>.</p>
</blockquote>
<p>I am very much looking forward to learning more about their evals and approach to verifying auto mode.</p>
<blockquote>
<p><strong>Thariq:</strong> A little on how auto mode works — it's useful to build this mental model. Whenever Claude is doing a turn, or a bash call, there's <strong>a Sonnet classifier</strong> that is judging the tool call and also the context of the conversation — your instruction. There are some things around permissions that are dependent on your request: you don't want to give git push permissions all the time, but if you say "push this to GitHub," you want it to do it — and if you say "don't push," you want it to deny it. Auto mode will do that. That particular thing happens to me a lot, where Claude tried to do something because it's very helpful and proactive, and auto mode saw "don't do this" and surfaced it. <strong>So it's good at the dynamic permissions</strong> that you yourself give inside the prompt, which I think is really important. It also works well with our <a href="https://code.claude.com/docs/en/sandbox-environments#sandboxed-bash-tool">sandboxing infrastructure</a>, because sandboxing is one of those things where there are so many different edge cases that it's hard for us to deterministically follow them. <strong>We have a sandbox, and when something needs to escape the sandbox</strong> — like a network request — auto mode can look at that request and ask: does this make sense? — and allow it.</p>
<p><strong>Simon:</strong> I hadn't realized auto mode is interacting with the networking sandbox as well.</p>
<p><strong>Cat:</strong> It interacts with any permission prompt the user would otherwise see.</p>
<p><strong>Simon:</strong> How old is auto mode? As a feature I had access to, it's only a couple of months old, right?</p>
</blockquote>
<p>(It was first made available to the public <a href="https://claude.com/blog/auto-mode">on March 24th</a>.)</p>
<blockquote>
<p><strong>Cat:</strong> We've been using it within Anthropic <strong>since January</strong>, so we've been hardening it for quite a while. Anthropic is extremely focused on safety and security, and we've been working broadly across our alignment and safeguards teams to enable the rollout internally, build out these evals, and make auto mode even more robust before sharing it with the world.</p>
<p><strong>Thariq:</strong> This is also the reason Claude Tag is so good — <strong>Claude Tag uses auto mode</strong>. I've heard a lot of build-versus-buy questions about a Slackbot, and I'm like: please, you probably shouldn't build your own AI Slackbot. There are so many attack vectors. <strong>You have a feedback channel that users can post feedback into, and now your bot is reading it.</strong> The work we've put in with auto mode — and we have a general <strong>Swiss cheese defense</strong> for security; we also RL against this stuff — <strong>I think this is really what makes Claude Tag work</strong>. It works seamlessly with your permissions, and you don't want to be prompt injected in your Slack.</p>
</blockquote>
<h4 id="are-there-more-security-things-in-the-pipeline-beyond-auto-mode-">Are there more security things in the pipeline beyond auto mode?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=2154s">35:54</a></p>
<blockquote>
<p><strong>Simon:</strong> Are there any more security things in the pipeline that go beyond auto mode?</p>
<p><strong>Thariq:</strong> I think we're very secure. <strong>With Claude Tag you can provision your own credentials for Claude</strong>, so it doesn't need to act on your behalf — you can have Claude as an identity, and that also makes it easier to audit and inspect what Claude is doing.</p>
<p><strong>Simon:</strong> Because Claude Tag is influenced by anyone who can talk to it — it's got a much wider pool of people telling it what to do.</p>
<p><strong>Thariq:</strong> That's right. And of course we have probes as well with Fable, which is a downstream effect of our safety and research work. I think this is the moment where you see Anthropic being an AI safety company really paying off: <strong>we really want Claude to be able to run in an aligned way over long periods of time</strong>, and <strong>auto mode has to be basically flawless for this to work</strong> — it's all downstream of our being an AI safety company.</p>
<p><strong>Cat:</strong> We also launched trusted devices for the remote control users out there who want to be safer. And for all of our remote environments, we support <strong>credential injection</strong>. If you want Claude Code to be able to access Datadog, but you don't want Claude Code itself to hold the Datadog credential, you can set up our identity and credential management system <strong>so that the Datadog credentials are only usable by the agent but not accessible by the agent</strong> — we insert them on the fly when the agent tries to make a Datadog request.</p>
</blockquote>
<p>I really like that credential injection pattern, where Claude Code can access an API via a proxy and that proxy both audits the request and injects the relevant API key - so Claude can access authenticated endpoints without having access to the API credentials itself.</p>
<h4 id="how-has-the-past-year-and-a-half-changed-how-you-think-about-your-own-craft-">How has the past year and a half changed how you think about your own craft?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=2273s">37:53</a></p>
<p>Thariq <a href="https://www.youtube.com/watch?v=9fubhllmsBU&amp;t=867s">talked about a sense of grief</a> brought on by Fable-class models in his keynote in the morning, and we dived further into that as part of our conversation. I've been calling this <a href="https://simonwillison.net/2026/Feb/15/deep-blue/">Deep Blue</a>.</p>
<blockquote>
<p><strong>Simon:</strong> <strong>Let's talk a little bit about the human element.</strong> <strong>A lot of people are feeling a sense of loss now that so much of what they considered to be their role in building software is being subsumed by the models.</strong> How do you think about that? <strong>How has the past year and a half changed the way you think about your own craft and the value that you add?</strong></p>
<p><strong>Thariq:</strong> Cat and Boris are such good reminders that you have to be more ambitious. They're always like: we're growing so fast, we have to be on the edge, we have to do the best work we can. That's a constant reminder for me — any time I'm slow on something, I'm like, okay, can I do it faster? Can I be more ambitious here? And oftentimes the answer is Claude, because Claude is getting better as you go — the last time I tried this, it was with the previous model. On your point about loss: I think this is real. <strong>If you're only trying to do the same work you were doing before LLMs, and now it's a prompt, it is, I think, kind of a sad feeling.</strong> And <strong>the way you offset that is by being more ambitious.</strong> I think Jared is such a good example — he hand-wrote all of the Zig code in his Oakland apartment in about a year, barely left his house, and had so much fun doing that. Now I see him rewrite all of Bun into Rust and <strong>he's having so much fun doing that</strong> — it's so much more ambitious, and that's how he offsets it. Generally it's asking <strong>how do I do the bigger thing</strong> and do more — <strong>I think success is fun</strong>. It's changing your ambition.</p>
</blockquote>
<p>"The way you offset that is by being more ambitious" neatly captures where I've landed on this issue myself as well.</p>
<blockquote>
<p><strong>Simon:</strong> And Cat, what does that look like from a product management perspective?</p>
<p><strong>Cat:</strong> I feel like the product role just changes every single month. <strong>All the PMs on our team are this mix of engineer, designer, PM</strong> — most of them actually used to be full-time engineers. For us it really means <strong>plugging in whenever there's any kind of gap</strong>. If we have an idea and we didn't inspire any engineer to go build it, then we should just build it, put it into a notebook, and inspire people to take it to production. If the designs look a little off, <strong>let's take a page that's similar, do a first-pass design, and tag in someone who's very detail-oriented to fill in the gaps</strong>. Or if we notice that our team and product adoption is bigger within the company, and more people need to know what's coming down the pipe for Claude Code, Claude Tag, and Cowork — let's automate figuring out our whole launch calendar, <strong>let's automate getting those status updates asynchronously</strong> so we're not bugging people, and make sure our updates in our internal announce channels are fully detailed and to the point. For us it's very much understanding <strong>what the gap is right now between a great idea and getting something to our customers</strong>, and <strong>how do we automate it as much as possible</strong>.</p>
</blockquote>
<p>This reflects something I've noticed: when you can produce code so much faster, time spent blocked awaiting a decision from someone else becomes a much more notable bottleneck. Engineers who can make product decisions can move a whole lot faster, and the cost of getting one of those decisions wrong is much less prohibitive.</p>
<h4 id="what-s-a-moment-when-claude-has-surprised-you-">What's a moment when Claude has surprised you?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=2510s">41:50</a></p>
<blockquote>
<p><strong>Simon:</strong> <strong>What's a moment when Claude has surprised you?</strong> When the model did something you didn't think it would be able to do?</p>
<p><strong>Thariq:</strong> I've posted a lot about Claude video editing, but most recently I gave a talk at the ACM Agentic conference, and I asked, "Hey guys, do you have the edited video? I'd love to post it and share it with my comms team." They said, "Oh, it's taking so long." So I asked for the raw files. They sent me the video of me talking on stage, the video of the deck, and the audio file, and said, "Good luck." I gave this to Claude, along with my HTML deck, and said, "<strong>Hey, can you just edit this together?</strong>" And what it does is honestly incredible — I'm ready to ship it. It transcribes the entire video. It notices that sometimes the video of my deck is a little weird — there's a popup of an auto-update in the middle — and it goes, "<strong>Oh, I probably shouldn't use the video of your deck. What I'm going to do is slice it up, figure out which slide you're on, and use the HTML source instead.</strong>" So it displays the HTML source. Then it's got video of me, but I'm only taking up a small part of the stage, so <strong>it's cropping dynamically to where I am on the stage</strong> — and I'm pacing, so it's tracking me as I pace. And it's transcribing what I'm saying.</p>
<p><strong>Simon:</strong> This was Fable, right?</p>
<p><strong>Thariq:</strong> This was Fable, yeah. It was a good prompt, but it was a one-shot prompt. Then I asked it to add some interesting animations and graphics, and I was just blown away. <strong>It does ffmpeg, it does Remotion.</strong></p>
</blockquote>
<p>Here's Thariq's video <a href="https://twitter.com/trq212/status/2064826394589442448">on how he used Fable to edit Fable's own launch video</a>, and here's <a href="https://twitter.com/ClaudeDevs/status/2064399512664526853">that launch video</a>.</p>
<h4 id="what-can-t-it-do-yet-">What can't it do yet?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=2616s">43:36</a></p>
<p>I'm embarrased to admit that I've been finding it quite hard to come up with tasks that frontier models like Fable 5 and GPT-5.6 are unable to accomplish.</p>
<p>Cat still doesn't rate its UX design skills:</p>
<blockquote>
<p><strong>Simon:</strong> What can't it do? What are the things where you're still disappointed — where you're waiting for Claude Fable 6 to figure it out for you?</p>
<p><strong>Cat:</strong> I want it to have better design and UX taste. It's now at the point where if I write out a prompt with a detailed spec of how I want a feature to behave, it will usually behave that way. But the paddings might be off, or the interface just isn't delightful yet. It leans on existing best practices for how apps are designed, but <strong>for frontier AI products, there are so many new interaction experiences that we have yet to design</strong>.</p>
<p><strong>Simon:</strong> There's an Opus aesthetic — you can look at something and go, "Yeah, that was designed by Opus." It'd be good if we could move beyond that.</p>
<p><strong>Cat:</strong> Yeah. I'm very excited for future models to hopefully be <strong>interaction design thought partners</strong>.</p>
<p><strong>Thariq:</strong> What can't it do? I would love to see it interact more with the real world. Can it solve science? Can it orchestrate the experiments? There's some amount of coding that goes into that, but there's also this other taste of the broader world that it needs.</p>
</blockquote>
<h4 id="which-parts-of-anthropic-s-culture-should-other-companies-steal-">Which parts of Anthropic's culture should other companies steal?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=2711s">45:11</a></p>
<p>I figured this would make a great closing question:</p>
<blockquote>
<p><strong>Simon:</strong> <strong>Which parts of Anthropic's company culture do you think uniquely help Anthropic be productive with these tools, that other companies should steal?</strong> What are the cultural hacks people should be adopting from you?</p>
<p><strong>Cat:</strong> I'll share one for Claude Tag. <strong>Claude Tag works best when you have it in a public channel, and when most of your channels are public.</strong> Claude Tag is able to search across all public channels to get as much context as possible to give you the highest-accuracy answer — and <strong>it's only able to do this if it has access to everything</strong>.</p>
<p><strong>Thariq:</strong> I mentioned this in my keynote, but it's so important to me I want to re-emphasize it. The co-founders <strong>say we don't negotiate against ourselves</strong>, and I think this is really important. <strong>You can imagine trade-offs in your head and talk yourself out of doing something ambitious — or you can just try to do the ambitious thing.</strong> We're so often asking: what if we just did it? Is this a real trade-off or not? And if so, why — where's the proof that it's a real trade-off, and not just something that sounds reasonable? <strong>Make the trade-offs show themselves to you. Be as ambitious as you can.</strong></p>
</blockquote>
<h4 id="what-s-your-favorite-absurd-thing-you-ve-built-with-claude-just-because-you-could-">What's your favorite absurd thing you've built with Claude, just because you could?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=2806s">46:46</a></p>
<p>I couldn't resist throwing in this one as well.</p>
<blockquote>
<p><strong>Simon:</strong> <strong>What's one of your favorite absurd things that you've built with Claude, just because you could build it?</strong></p>
<p><strong>Thariq:</strong> I'm working on <strong>a 2D Street Fighter fighting game with me as a character</strong> — and my friends as well. It uses Claude Code to prompt Gemini — and honestly the Seedance model is pretty good — to make video animations. It works great; it's so good at prompting, and it can verify the frames to check whether an animation was good.</p>
<p><strong>Simon:</strong> Is this Street Fighter 2-level 2D sprites you're generating?</p>
<p><strong>Thariq:</strong> Yeah, exactly — 2D sprites. The animation looks amazing. And it can also figure out hitboxes — it can be like, "Oh, your fist is here, I'll draw the JSON hitbox." It's incredible.</p>
<p><strong>Cat:</strong> Mine is much more simple. I'm a big rock climber and a lot of my friends climb, so we have this little app we built with Claude Code where we log all the projects we're working on. We also go outdoors together a lot, so we have Claude do all this research with workflows. Workflows is amazing — we brand it as a coding tool, but it's amazing for doing deep research for travel. I also plan our team offsites, and it's good at finding venues that can fit all of us. I use workflows to research all the climbing destinations we might want to go to, and what has direct flights from where all of us are located. It goes to Mountain Project and finds all the climbs at our grade level. It finds the Airbnb. And I don't like hiking, so I care a lot about it having a very short approach — <strong>very short walking distance from where the car parks to where the rock actually is</strong> — and it filters for this. With existing apps I have to manually click through Mountain Project, but with this I just put in all of our preferences and it's a custom app for us.</p>
<p><strong>Simon:</strong> So you're basically vibe coding Jira for mountain climbing.</p>
<p><strong>Cat:</strong> Exactly.</p>
</blockquote>
<h4 id="audience-any-plans-for-eval-building-tools-and-agent-observability-">Audience: Any plans for eval-building tools and agent observability?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=2963s">49:23</a></p>
<p>We had a few minutes at the end for questions from the audience.</p>
<blockquote>
<p><strong>Audience:</strong> Do you have any near-term plans to build more eval tools for us to build eval datasets, and more observability tools to monitor the performance of agents and workflows?</p>
<p><strong>Cat:</strong> We've considered building eval tools, but I think the limiting factor actually tends to be that <strong>it takes a long time for customers to build really high-quality evals</strong>. So I think the tooling is less of the constraint, and more the skill set of how you build a great eval. That's an area where we're excited to both invest internally and hopefully share some best practices externally.</p>
</blockquote>
<h4 id="audience-how-is-memory-designed-today-and-would-you-move-from-files-to-a-data-store-">Audience: How is memory designed today — and would you move from files to a data store?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=3008s">50:08</a></p>
<blockquote>
<p><strong>Audience (Sai):</strong> I'm interested in the memory and the multiplayer. <strong>How is memory being designed today?</strong> I assume it's around files. And second, have you thought about an orthogonal direction where you <strong>would actually need a data store for these memories, instead of files, to scale it better?</strong></p>
<p><strong>Thariq:</strong> Right now for Claude Tag the memory is channel-specific. Every Claude in that channel has a shared memory, and the instances have a session — but the session can contribute back to main memory. We do a lot of memory research, and it can be kind of unintuitive what the right way to do memory is. We're always running memory experiments. <strong>How it works right now in Claude Tag is a markdown file per channel.</strong></p>
</blockquote>
<p>Tags: <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/prompt-engineering">prompt-engineering</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/anthropic">anthropic</a>, <a href="https://simonwillison.net/tags/annotated-talks">annotated-talks</a>, <a href="https://simonwillison.net/tags/coding-agents">coding-agents</a>, <a href="https://simonwillison.net/tags/claude-code">claude-code</a>, <a href="https://simonwillison.net/tags/thariq-shihipar">thariq-shihipar</a>, <a href="https://simonwillison.net/tags/cat-wu">cat-wu</a></p></summary>
<category term="ai"/>
<category term="prompt-engineering"/>
<category term="generative-ai"/>
<category term="llms"/>
<category term="anthropic"/>
<category term="annotated-talks"/>
<category term="coding-agents"/>
<category term="claude-code"/>
<category term="thariq-shihipar"/>
<category term="cat-wu"/>
</entry>
<entry>
<title>Reverse-engineering is cheap now</title>
<link href="https://simonwillison.net/2026/Jul/20/cheap-reverse-engineering/#atom-everything" rel="alternate"/>
<published>2026-07-20T19:24:05+00:00</published>
<updated>2026-07-20T19:24:05+00:00</updated>
<id>https://simonwillison.net/2026/Jul/20/cheap-reverse-engineering/#atom-everything</id>
<summary type="html"><p>I keep hearing anecdotes from people who used coding agents to reverse-engineer and automate devices in their homes.</p>
<p>I think this is an interesting illustration of the impact of the reduced cost of writing code.</p>
<p>Prior to agents, it was entirely possible to reverse-engineer home devices. The problem was the ROI - was it really worth all of that effort? More importantly, any experienced programmer knows that undocumented, unstable APIs like that may well change or break in the future. Is that initial work worth the effort if you're committing yourself to a frustrating cycle of maintenance in the future?</p>
<p>Coding agents change that equation entirely. The effort to get a simple automation working has dropped, as has the cost of trying and failing to get it to work. Since the code is so cheap, the idea of having to maintain it in the future - or throw it away and start again - carries way less psychological baggage.</p>
<p>Tags: <a href="https://simonwillison.net/tags/reverse-engineering">reverse-engineering</a>, <a href="https://simonwillison.net/tags/coding-agents">coding-agents</a>, <a href="https://simonwillison.net/tags/ai-assisted-programming">ai-assisted-programming</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a></p></summary>
<category term="reverse-engineering"/>
<category term="coding-agents"/>
<category term="ai-assisted-programming"/>
<category term="generative-ai"/>
<category term="ai"/>
<category term="llms"/>
</entry>
<entry>
<title>Who’s Afraid of Chinese Models?</title>
<link href="https://simonwillison.net/2026/Jul/20/afraid-of-chinese-models/#atom-everything" rel="alternate"/>
<published>2026-07-20T17:09:19+00:00</published>
<updated>2026-07-20T17:09:19+00:00</updated>
<id>https://simonwillison.net/2026/Jul/20/afraid-of-chinese-models/#atom-everything</id>
<summary type="html"><p><strong><a href="https://stratechery.com/2026/whos-afraid-of-chinese-models/">Who’s Afraid of Chinese Models?</a></strong></p>
Interesting proposal from Ben Thompson that both addresses the hypocrisy of labs outlawing distillation against their models despite training on unlicensed data, and could help US open models compete more effectively with their Chinese counterparts:</p>
<blockquote>
<p>The U.S. should pass a law that (1) makes explicit that collecting data for training models is fair use, and (2) bars terms of service that forbid distillation, for U.S. companies at a minimum. Stopping distillation — which is literally just querying the API — is nearly impossible; the U.S. should go the other way and lean into a new copyright policy that both indemnifies the labs and also guarantees that what they learned fuels further innovation for everyone else.</p>
</blockquote>
<p>Ben also theorizes that Alibaba's decision to release Qwen 3.8 Max as open weights - a reversal from their decision <a href="https://qwen.ai/blog?id=qwen3.7">not to release Qwen 3.7 Max</a> in May - may have been influenced by a <a href="http://english.scio.gov.cn/topnews/2026-07/18/content_118605932.html">recent speech</a> by Xi Jinping, who said:</p>
<blockquote>
<p>We should seize this rare, historic opportunity to encourage open source, openness, collaboration and sharing.</p>
</blockquote>
<p>And on the subject of <a href="https://twitter.com/Alibaba_Qwen/status/2078759124914098291">Qwen 3.8 Max</a> - a new 2.4T parameter model (nearly as large as the 2.8T Kimi K3) - here's <a href="https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F735f2cf19b795517cb2ff6cae1c71c64">a pelican it drew</a>:</p>
<p><img alt="Described by Qwen 3.8 Max: Flat vector cartoon illustration of a white pelican with a large orange beak and pouch riding a red bicycle, its orange legs on the pedals, against a light blue sky with a yellow sun top right and a white cloud top left, with horizontal motion lines behind the bike and a pale green ground strip at the bottom." src="https://static.simonwillison.net/static/2026/qwen-3.8-max-pelican.png" /></p>
<p>I particularly enjoyed seeing these notes in the (extensive) reasoning trace: "Could add helmet? No." and "Maybe add small bell? no." and "Need maybe add small fish in basket? Not necessary."
<p><small></small>Via <a href="https://daringfireball.net/linked/2026/07/20/thompson-chinese-models-distillation">John Gruber</a></small></p>
<p>Tags: <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/training-data">training-data</a>, <a href="https://simonwillison.net/tags/qwen">qwen</a>, <a href="https://simonwillison.net/tags/pelican-riding-a-bicycle">pelican-riding-a-bicycle</a>, <a href="https://simonwillison.net/tags/ai-ethics">ai-ethics</a>, <a href="https://simonwillison.net/tags/llm-release">llm-release</a>, <a href="https://simonwillison.net/tags/ai-in-china">ai-in-china</a></p></summary>
<category term="ai"/>
<category term="generative-ai"/>
<category term="llms"/>
<category term="training-data"/>
<category term="qwen"/>
<category term="pelican-riding-a-bicycle"/>
<category term="ai-ethics"/>
<category term="llm-release"/>
<category term="ai-in-china"/>
</entry>
<entry>
<title>Quoting Sam Altman</title>
<link href="https://simonwillison.net/2026/Jul/20/sam-altman/#atom-everything" rel="alternate"/>
<published>2026-07-20T03:47:59+00:00</published>
<updated>2026-07-20T03:47:59+00:00</updated>
<id>https://simonwillison.net/2026/Jul/20/sam-altman/#atom-everything</id>
<summary type="html"><blockquote cite="https://twitter.com/techemails/status/2078854346683678927"><p>We have been having extensive discussions around open source strategy. We will discuss it more at our next board meeting, but one thing we’d like to do soon is to create a language model with the approximate capability of GPT-3 that can run locally on consumer hardware and release that. We’d like to do it soon, before Stability or someone else does. In general, we think this helps discourage others from releasing similarly-powerful models, and makes it harder for new efforts to get funded.</p></blockquote>
<p class="cite">&mdash; <a href="https://twitter.com/techemails/status/2078854346683678927">Sam Altman</a>, Email to OpenAI's board, October 1, 2022 - exposed in Musk v. Altman (2026)</p>
<p>Tags: <a href="https://simonwillison.net/tags/ai-ethics">ai-ethics</a>, <a href="https://simonwillison.net/tags/sam-altman">sam-altman</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/openai">openai</a>, <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a></p></summary>
<category term="ai-ethics"/>
<category term="sam-altman"/>
<category term="generative-ai"/>
<category term="openai"/>
<category term="ai"/>
<category term="llms"/>
</entry>
<entry>
<title>AI Mania Is Eviscerating Global Decision-Making</title>
<link href="https://simonwillison.net/2026/Jul/19/ai-mania/#atom-everything" rel="alternate"/>
<published>2026-07-19T05:06:21+00:00</published>
<updated>2026-07-19T05:06:21+00:00</updated>
<id>https://simonwillison.net/2026/Jul/19/ai-mania/#atom-everything</id>
<summary type="html"><p><strong><a href="https://ludic.mataroa.blog/blog/ai-mania-is-eviscerating-global-decision-making/">AI Mania Is Eviscerating Global Decision-Making</a></strong></p>
Here's an entertaining perspective from Nik Suresh on the AI mania that is overwhelming the large companies that he consults with. It's crammed with spicy anecdotes from anonymous sources.</p>
<blockquote>
<p>In one extreme case, I have seen an executive confess that they had never even used ChatGPT or any AI tool in their life, immediately after producing a technical strategy for an organisation with $2B+ in revenue which was entirely centered around AI.</p>
</blockquote>
<p>Here's a report from an engineer at a company with a token leaderboard:</p>
<blockquote>
<p>Checking out a parallel copy of our Go repository and telling the AI to rewrite the whole thing in Zig while I work on something else just so I can keep my job.</p>
</blockquote>
<p>I particularly enjoyed this conversation with a skeptical executive at an over-enthusiastic company:</p>
<blockquote>
<p>I asked <em>why</em> this was being repeated without opposition. Was it just sales fluff?</p>
<p>The answer was a lot more interesting. It was <em>partially</em> ridiculous sales material being delivered to an easily excitable audience, but this was not the dominant factor constraining honesty. Executives at their <em>customers</em> were saying absurd things about achieving 100x productivity, and this meant that if any executive at the <em>vendor</em> said that these gains were not plausible, it would undermine the credibility of the customer’s executive, be perceived as an attack (or heresy), and possibly result in an enterprise contract cancellation. And getting enterprise contracts cancelled because you wanted to opine on something that doesn’t really matter to your organisation’s mission is a great way to get fired.</p>
</blockquote>
<p><small></small>Via <a href="https://news.ycombinator.com/item?id=48964185">Hacker News</a></small></p>
<p>Tags: <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/ai-ethics">ai-ethics</a>, <a href="https://simonwillison.net/tags/ai-misuse">ai-misuse</a></p></summary>
<category term="ai"/>
<category term="ai-ethics"/>
<category term="ai-misuse"/>
</entry>
<entry>
<title>Claude Code uses Bun written in Rust now</title>
<link href="https://simonwillison.net/2026/Jul/19/claude-code-in-bun-in-rust/#atom-everything" rel="alternate"/>
<published>2026-07-19T03:54:09+00:00</published>
<updated>2026-07-19T03:54:09+00:00</updated>
<id>https://simonwillison.net/2026/Jul/19/claude-code-in-bun-in-rust/#atom-everything</id>
<summary type="html"><p>In <a href="https://bun.com/blog/bun-in-rust">Rewriting Bun in Rust</a> Jarred Sumner made the following claim:</p>
<blockquote>
<p>Claude Code v2.1.181 (released June 17th) and later use the Rust port of Bun. Startup got 10% faster on Linux but otherwise, barely anyone noticed. Boring is good.</p>
</blockquote>
<p>I decided to have a poke at my own Claude Code installation to see if I could find evidence that it was using Bun written in Rust.</p>
<p>I found these two commands convincing:</p>
<pre><code>strings ~/.local/bin/claude | grep -m1 'Bun v1'
</code></pre>
<p>For me this outputs <code>Bun v1.4.0 (macOS arm64)</code>. The most recent release of <a href="https://github.com/oven-sh/bun/releases">Bun on GitHub</a> is currently <a href="https://github.com/oven-sh/bun/releases/tag/bun-v1.3.14">v1.3.14</a> from May 12th, so that v1.4.0 version number in Claude supports them shipping a preview of a not-yet-released Bun version.</p>
<p>(<strong>Update</strong>: The Rust version <em>has</em> been released as <a href="https://bun.com/docs/installation#canary-builds">Bun canary</a> - running <code>bun upgrade --canary</code> will install <a href="https://github.com/oven-sh/bun/releases/tag/canary">this release</a>.)</p>
<pre><code>strings ~/.local/bin/claude | grep -Eo 'src/[[:alnum:]_./-]+\.rs'
</code></pre>
<p>This outputs a list of <a href="https://gist.github.com/simonw/c92fb0f67b114ac26e3b95a09ddccfdc">563 filenames</a>, starting with these:</p>
<pre><code>src/runtime/bake/dev_server/mod.rs
src/runtime/bake/production.rs
src/bundler/bundle_v2.rs
</code></pre>
<p>It looks like Bun in Rust is indeed being run in production across millions of different devices. Like Jarred said, "Boring is good".</p>
<p><strong>Update</strong>: Here's a neat trick <a href="https://twitter.com/ajanraj25/status/2078825794701242697">from Ajan Raj</a>:</p>
<pre><code>cat &gt; /tmp/bun-version.ts &lt;&lt;'EOF'
console.log("embedded bun:", Bun.version);
process.exit(0);
EOF
BUN_OPTIONS="--preload=/tmp/bun-version.ts" claude --version
</code></pre>
<p>This outputs <code>1.4.0</code> for me.</p>
<p>Here's <a href="https://github.com/oven-sh/bun/commit/b18bf6d1d0a92238f240bfd125f0e3b3461b9243#diff-7ae45ad102eab3b6d7e7896acd08c427a9b25b346470d7bc6507b6481575d519">the commit from May 17th</a> that updated the version in <code>package.json</code> to 1.4.0. That version hasn't been changed since then, but also hasn't yet made it into a tagged release outside of <code>canary</code>.</p>
<p>Tags: <a href="https://simonwillison.net/tags/bun">bun</a>, <a href="https://simonwillison.net/tags/rust">rust</a>, <a href="https://simonwillison.net/tags/anthropic">anthropic</a>, <a href="https://simonwillison.net/tags/claude-code">claude-code</a>, <a href="https://simonwillison.net/tags/jarred-sumner">jarred-sumner</a></p></summary>
<category term="bun"/>
<category term="rust"/>
<category term="anthropic"/>
<category term="claude-code"/>
<category term="jarred-sumner"/>
</entry>
<entry>
<title>SQLite Query Explainer</title>
<link href="https://simonwillison.net/2026/Jul/18/sqlite-query-explainer/#atom-everything" rel="alternate"/>
<published>2026-07-18T17:19:10+00:00</published>
<updated>2026-07-18T17:19:10+00:00</updated>
<id>https://simonwillison.net/2026/Jul/18/sqlite-query-explainer/#atom-everything</id>
<summary type="html"><p><strong>Tool:</strong> <a href="https://tools.simonwillison.net/sqlite-query-explainer">SQLite Query Explainer</a></p>
<p>Julia Evan's, in <a href="https://jvns.ca/blog/2026/07/17/learning-about-running-sqlite/">Learning a few things about running SQLite</a>:</p>
<blockquote>
<p>Maybe one day I’ll learn to read a query plan.</p>
</blockquote>
<p>Big same.... which inspired me to <a href="https://github.com/simonw/tools/pull/299#issue-4919268017">have Fable build</a> this interactive explain tool, which runs SQLite in Python in Pyodide in Web Assembly in the browser and adds a layer of explanation to the results of both EXPLAIN and EXPLAIN QUERY PLAN.</p>
<p>Approach with caution, since I don't know enough about SQLite query plans to verify the results myself, but it seems cromulent enough to me.</p>
<p>Tags: <a href="https://simonwillison.net/tags/sql">sql</a>, <a href="https://simonwillison.net/tags/sqlite">sqlite</a>, <a href="https://simonwillison.net/tags/tools">tools</a>, <a href="https://simonwillison.net/tags/julia-evans">julia-evans</a>, <a href="https://simonwillison.net/tags/pyodide">pyodide</a>, <a href="https://simonwillison.net/tags/claude-mythos-fable">claude-mythos-fable</a></p></summary>
<category term="sql"/>
<category term="sqlite"/>
<category term="tools"/>
<category term="julia-evans"/>
<category term="pyodide"/>
<category term="claude-mythos-fable"/>
</entry>
<entry>
<title>Claude make Fable 5 permanent</title>
<link href="https://simonwillison.net/2026/Jul/18/claude-make-fable-5-permanent/#atom-everything" rel="alternate"/>
<published>2026-07-18T06:00:13+00:00</published>
<updated>2026-07-18T06:00:13+00:00</updated>
<id>https://simonwillison.net/2026/Jul/18/claude-make-fable-5-permanent/#atom-everything</id>
<summary type="html"><p><strong><a href="https://twitter.com/claudeai/status/2078302415804379218">Claude make Fable 5 permanent</a></strong></p>
An update from the <code>@claudeai</code> account on Twitter:</p>
<blockquote>
<p>Beginning July 20, Claude Fable 5 will be included in all Max and Team Premium plans, at 50% of limits.</p>
<p>Pro and Team Standard users will continue to have access to Fable via usage credits, and will receive a one-time $100 credit.</p>
</blockquote>
<p>As I was saying <a href="https://simonwillison.net/2026/Jul/12/bump/">last week</a>, the competition from <a href="https://simonwillison.net/2026/Jul/9/gpt-5-6/">GPT-5.6 Sol</a> (and maybe to a lesser extent <a href="https://simonwillison.net/2026/Jul/16/kimi-k3/">Kimi 3</a>) made untenable Anthropic's plan to remove Fable 5 from their subscription accounts and make it available exclusively through API pricing.</p>
<p>Why pay $100 or $200/month for a subscription plan that <em>doesn't</em> include Anthropic's best model?</p>
<p>Their original plan was driven by concerns over compute capacity. I wonder if they'll have to dial back their training efforts in order to make more GPUs available to help serve the model.</p>
<p>A lot of people were losing sleep over trying to make the most of Fable 5 before subscriber access was withdrawn. It's nice not to have to worry about the Fablepocalypse any more.</p>
<p><strong>Update</strong>: Important to note that users on the $20/month plan will still not have access to Fable 5 on that subscription. The Max plans are $100 and $200/month.
<p>Tags: <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/anthropic">anthropic</a>, <a href="https://simonwillison.net/tags/claude">claude</a>, <a href="https://simonwillison.net/tags/llm-pricing">llm-pricing</a>, <a href="https://simonwillison.net/tags/claude-mythos-fable">claude-mythos-fable</a></p></summary>
<category term="ai"/>
<category term="generative-ai"/>
<category term="llms"/>
<category term="anthropic"/>
<category term="claude"/>
<category term="llm-pricing"/>
<category term="claude-mythos-fable"/>
</entry>
<entry>
<title>nascheme/quixote</title>
<link href="https://simonwillison.net/2026/Jul/18/quixote/#atom-everything" rel="alternate"/>
<published>2026-07-18T05:27:49+00:00</published>
<updated>2026-07-18T05:27:49+00:00</updated>
<id>https://simonwillison.net/2026/Jul/18/quixote/#atom-everything</id>
<summary type="html"><p><strong><a href="https://github.com/nascheme/quixote">nascheme/quixote</a></strong></p>
A certain vintage of Python web nerd might be delighted to learn that the most recent commit to the Quixote web framework was <a href="(https://github.com/nascheme/quixote/commit/7f775cf9d1e7e80fcbb2706b4a1d971e55ca74a3)">six hours ago</a>.</p>
<p>The <a href="https://github.com/nascheme/quixote/commit/d6b73c5768c2d041b68b54cc71863604249abc18">oldest commit</a> in that repo is from 21 years ago, and that was the initial import of Quixote 2.4 from Subversion into Git.
<p>Tags: <a href="https://simonwillison.net/tags/computer-history">computer-history</a>, <a href="https://simonwillison.net/tags/python">python</a>, <a href="https://simonwillison.net/tags/web-frameworks">web-frameworks</a></p></summary>
<category term="computer-history"/>
<category term="python"/>
<category term="web-frameworks"/>
</entry>
<entry>
<title>Quoting Kimi K3</title>
<link href="https://simonwillison.net/2026/Jul/17/kimi-k3/#atom-everything" rel="alternate"/>
<published>2026-07-17T13:43:53+00:00</published>
<updated>2026-07-17T13:43:53+00:00</updated>
<id>https://simonwillison.net/2026/Jul/17/kimi-k3/#atom-everything</id>
<summary type="html"><blockquote cite="https://news.ycombinator.com/item?id=48935342#48936515"><p>Is there something I can actually help you with today?</p></blockquote>
<p class="cite">&mdash; <a href="https://news.ycombinator.com/item?id=48935342#48936515">Kimi K3</a>, after refusing to leak its system prompt</p>
<p>Tags: <a href="https://simonwillison.net/tags/kimi">kimi</a>, <a href="https://simonwillison.net/tags/ai-personality">ai-personality</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a></p></summary>
<category term="kimi"/>
<category term="ai-personality"/>
<category term="generative-ai"/>
<category term="ai"/>
<category term="llms"/>
</entry>
<entry>
<title>LLM cliché highlighter</title>
<link href="https://simonwillison.net/2026/Jul/17/llm-cliche-highlighter/#atom-everything" rel="alternate"/>
<published>2026-07-17T12:11:11+00:00</published>
<updated>2026-07-17T12:11:11+00:00</updated>
<id>https://simonwillison.net/2026/Jul/17/llm-cliche-highlighter/#atom-everything</id>
<summary type="html"><p><strong>Tool:</strong> <a href="https://tools.simonwillison.net/llm-cliche-highlighter">LLM cliché highlighter</a></p>
<p>I got frustrated reading <em>yet another</em> article that was crammed with the clichés of LLM-generated writing - "no fluff, no filler, no jargon" type stuff - so I had Fable 5 vibe code up this app for highlighting ten common patterns that show up in that sort of writing.</p>
<p><img alt="Screenshot of a text-analysis web tool. Top summary row: &quot;2 matches&quot;, &quot;1 flagged sentence&quot;, &quot;0 chain items&quot;. Below, a collapsed &quot;▶ Patterns · all 11 on&quot; panel, then a URL input reading &quot;https://example.com/article — fetched via r.jina.ai&quot; with a &quot;Load URL&quot; button. A text area contains &quot;That loss is real and it's worth naming&quot;. Below are &quot;Load example&quot; and &quot;Clear&quot; buttons and a checked checkbox &quot;Show just the highlights&quot;. A &quot;Highlighted text&quot; section shows &quot;That loss is real and it's worth naming&quot; with &quot;That loss&quot; in pale yellow (flagged sentence) and &quot;is real and&quot; plus &quot;'s worth naming&quot; in darker yellow (pattern match). Legend: &quot;flagged sentence&quot;, &quot;pattern match&quot;, &quot;3 chain item count&quot;. &quot;Matches&quot; section: 1. &quot;is real and&quot; — &quot;Is real … and / not&quot;; 2. &quot;'s worth naming&quot; — &quot;Worth naming&quot;." src="https://static.simonwillison.net/static/2026/the-loss-is-real.webp" /></p>
<p>Tags: <a href="https://simonwillison.net/tags/tools">tools</a>, <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a></p></summary>
<category term="tools"/>
<category term="ai"/>
<category term="generative-ai"/>
<category term="llms"/>
</entry>
<entry>
<title>Spot birds not golf</title>
<link href="https://simonwillison.net/2026/Jul/17/spot-birds-not-golf/#atom-everything" rel="alternate"/>
<published>2026-07-17T02:58:07+00:00</published>
<updated>2026-07-17T02:58:07+00:00</updated>
<id>https://simonwillison.net/2026/Jul/17/spot-birds-not-golf/#atom-everything</id>
<summary type="html"><p>Suggestion for hyperscalers feeling pressure over data center water use:</p>
<p>Buy up a few exclusive country clubs, convert the golf courses into public parks, pay for guides and binoculars to get the previous members into birdwatching - help them embrace a more sustainable hobby!</p>
<p>Google <a href="https://sustainability.google/reports/google-2026-environmental-report/">used 10.9 billion gallons in 2025</a>, so about 30 million gallons per day.</p>
<p>The Coachella Valley has <a href="https://www.cvwd.org/167/Water-Conservation">120 golf courses each using ~800 acre-feet per year</a>, which is ~750,000 gallons per day.</p>
<p>So Google buying up 40 of those courses (1/3) should do the trick.</p>
<p>Tags: <a href="https://simonwillison.net/tags/ai-energy-usage">ai-energy-usage</a>, <a href="https://simonwillison.net/tags/ai">ai</a></p></summary>
<category term="ai-energy-usage"/>
<category term="ai"/>
</entry>
<entry>
<title>Firefox in WebAssembly</title>
<link href="https://simonwillison.net/2026/Jul/16/firefox-in-webassembly/#atom-everything" rel="alternate"/>
<published>2026-07-16T23:34:16+00:00</published>
<updated>2026-07-16T23:34:16+00:00</updated>
<id>https://simonwillison.net/2026/Jul/16/firefox-in-webassembly/#atom-everything</id>
<summary type="html"><p><strong><a href="https://developer.puter.com/labs/firefox-wasm/">Firefox in WebAssembly</a></strong></p>
This is absurdly cool: Puter compiled Firefox to WebAssembly such that the whole browser runs in another browser.</p>
<p>Here's my blog, running in Firefox, running in WebAssembly, running in Chrome:</p>
<p><img alt="A Chrome window. The tab has the Firefox UI and has loaded my blog. On the right is the Chrome network panel showing that it loaded resources that include a 233MB gecko.wasm and an 18MB chrome-assets.tar.zst" src="https://static.simonwillison.net/static/2026/firefox-wasm.webp" /></p>
<p>They chose Firefox/Gecko because it has strong single-process support. The project used an estimated $25,000 worth of Claude Opus and Fable tokens, but took advantage of a Claude Max subscription plan so cost much less in actual dollars.</p>
<p>The demo funnels all traffic over a WebSocket protocol (using the <a href="https://github.com/MercuryWorkshop/wisp-protocol">Wisp protocol</a>) through Puter's server - a requirement to get this kind of thing to work because code running in browsers can't open arbitrary network connections.</p>
<p>(That proxying sounds expensive! The team <a href="https://news.ycombinator.com/item?id=48926939#48936563">had to scale the servers up</a> to handle the traffic during the Hacker News conversation about the project.)</p>
<p>Puter claim this supports end-to-end encryption and that looks to be true - I inspected the WebSocket messages and traffic to my own HTTPS site was encrypted whereas requests and responses to <code>http://www.example.com/</code> were in cleartext.</p>
<p><a href="https://github.com/HeyPuter/firefox-wasm">Here's the repo</a> for <code>firefox-wasm</code>. <a href="https://github.com/theogbob/WebkitWasm">theogbob/WebkitWasm</a> is a similar project that compiles WebKit to WASM, but that one doesn't currently have an accessible online demo.
<p><small></small>Via <a href="https://news.ycombinator.com/item?id=48926939">Hacker News</a></small></p>
<p>Tags: <a href="https://simonwillison.net/tags/browsers">browsers</a>, <a href="https://simonwillison.net/tags/firefox">firefox</a>, <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/webassembly">webassembly</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/ai-assisted-programming">ai-assisted-programming</a>, <a href="https://simonwillison.net/tags/claude">claude</a>, <a href="https://simonwillison.net/tags/claude-mythos-fable">claude-mythos-fable</a></p></summary>
<category term="browsers"/>
<category term="firefox"/>
<category term="ai"/>
<category term="webassembly"/>
<category term="generative-ai"/>
<category term="llms"/>
<category term="ai-assisted-programming"/>
<category term="claude"/>
<category term="claude-mythos-fable"/>
</entry>
<entry>
<title>Kimi K3, and what we can still learn from the pelican benchmark</title>
<link href="https://simonwillison.net/2026/Jul/16/kimi-k3/#atom-everything" rel="alternate"/>
<published>2026-07-16T20:19:30+00:00</published>
<updated>2026-07-16T20:19:30+00:00</updated>
<id>https://simonwillison.net/2026/Jul/16/kimi-k3/#atom-everything</id>
<summary type="html"><p>Chinese AI lab Moonshot AI <a href="https://www.kimi.com/blog/kimi-k3">announced Kimi K3</a> this morning, describing it as their "most capable model to date, with 2.8 trillion parameters". It's currently available via their website and API, but an open weight release is promised "by July 27, 2026".</p>
<p>Moonshot are calling this the first "open 3T-class model" (I guess they're rounding 2.8 trillion up to 3 trillion), taking the crown from <a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro">DeepSeek's 1.6T v4 Pro</a>. Their <a href="https://www.kimi.com/blog/kimi-k3#full-benchmark-table">self-reported benchmarks</a> have K3 mostly beating Claude Opus 4.8 max and GPT-5.5 high, while losing out to Claude Fable 5 and GPT-5.6 Sol.</p>
<p>A few highlights from the <a href="https://twitter.com/ArtificialAnlys/status/2077832874183860404">Artificial Analysis report</a> on the model:</p>
<ul>
<li>"On our private long-horizon knowledge work evaluation, Kimi K3 reaches an overall Elo of 1547, +732 points from Kimi K2.6 and behind only Claude Fable 5."</li>
<li>"Cost per task ($0.94) is similar to GPT-5.6 Sol ($1.04), ~1/2 the price of Opus 4.8 ($1.80) and higher than open weights peers"</li>
<li>"Kimi K3’s token usage on the Artificial Analysis Intelligence Index decreased significantly, using 21% fewer output tokens than K2.6."</li>
</ul>
<p>The model is also now the <a href="https://twitter.com/arena/status/2077824029126504525">leading model on Arena.ai's Frontend Code arena</a>, surpassing even Claude Fable 5.</p>
<p>The new model is notable for the pricing: $3/million input tokens and $15/million output tokens, putting it at the same level as Anthropic's Claude Sonnet series and making it the most expensive model released by a Chinese AI lab to date. This is a significant increase on their earlier models <a href="https://platform.kimi.ai/docs/pricing/chat-k26">such as Kimi K2.6</a> at $0.95/$4. 2.8 trillion parameters is also more than twice the size of that 1T model.</p>
<h4 id="but-how-does-it-pelican-">But how does it pelican?</h4>
<p>I used OpenRouter (to avoid signing up for a Moonshot API key) with the <a href="https://github.com/simonw/llm-openrouter">llm-openrouter plugin</a> to generate an SVG of a pelican riding a bicycle:</p>
<pre><code>llm -m openrouter/moonshotai/kimi-k3 'Generate an SVG of a pelican riding a bicycle'
</code></pre>
<p>Here's <a href="https://gist.github.com/simonw/66a2699eb1594258904c7b5102840dd6">the transcript</a>. It looks like this:</p>
<p><img src="https://static.simonwillison.net/static/2026/kimi-3-pelican.jpg" alt="See description below" style="max-width: 100%;" /></p>
<p>That pelican took 95 input tokens and 16,658 output tokens (13,241 were reasoning tokens), for a total cost of <a href="https://www.llm-prices.com/#it=95&amp;ot=16658&amp;ic=3&amp;oc=15">25 cents</a>!</p>
<p>Since K3 accepts image input I ran it against that rendered SVG above (with my <a href="https://simonwillison.net/guides/agentic-engineering-patterns/prompts/#alt-text">alt text prompt</a>) and <a href="https://gist.github.com/simonw/665dbf840701b421745f2cb891acdfd6">got back</a> (for <a href="https://www.llm-prices.com/#it=822&amp;ot=243&amp;ic=3&amp;oc=15">0.6 cents</a>):</p>
<blockquote>
<p>Cartoon illustration of a white pelican wearing a red scarf, riding a red bicycle along a gray road with white dashed lines; the pelican has a large orange beak and webbed orange feet pedaling, with white motion lines behind it; the background shows a light blue sky with white clouds, a yellow sun, two small black birds in flight, and green grass with tiny white flowers in the foreground</p>
</blockquote>
<h4 id="what-can-we-learn-from-the-pelican-">What can we learn from the pelican?</h4>
<p>My <a href="https://simonwillison.net/tags/pelican-riding-a-bicycle/">Generate an SVG of a pelican riding a bicycle</a> test is 21 months old now. It was never a particularly great benchmark. It started out as a joke on how absurdly difficult it is to compare these models, but then for the first year it turned out to have a <a href="https://simonwillison.net/2025/Jun/6/six-months-in-llms/">surprising correlation</a> to how good the models actually were.</p>
<p>That connection has been mostly severed now. The <a href="https://simonwillison.net/2026/Jul/9/gpt-5-6/">GPT-5.6</a> and <a href="https://simonwillison.net/2026/Jun/9/claude-fable-5/">Claude Fable 5</a> pelicans are outclassed <a href="https://simonwillison.net/2026/Jun/17/glm-52/">by GLM-5.2</a>, and much as I love GLM I don't think that's a Fable-class model.</p>
<p>(I'm still not convinced that labs are <a href="https://simonwillison.net/2025/Nov/13/training-for-pelicans-riding-bicycles/">training for the benchmark</a> - if they were, I'd expect much better results. There's a chance that Gemini has optimized for <a href="https://simonwillison.net/2026/Feb/19/gemini-31-pro/#jeff-dean">any combination of an animal on a vehicle</a> though!)</p>
<p>The biggest limitation of the pelican is that it doesn't touch at all on the thing that matters most for today's model: agentic tool calling and the ability to operate tools reliably as conversations grow in length.</p>
<p>So don't go using pelicans to compare models!</p>
<p>All of that said, I still get a decent amount of value out of running the benchmark myself.</p>
<p>Firstly, it's a forcing function for actually trying the model. If I show you a pelican, that means I've managed to run a prompt through it. If the model has an official API I'll use that, if it's open weight (and small enough to fit a 128GB M5 MacBook Pro) I'll try running it on my own machine, usually via <a href="https://github.com/ggml-org/llama.cpp">llama.cpp</a> or <a href="https://lmstudio.ai">LM Studio</a> or <a href="https://ollama.com">Ollama</a>. I'll frequently use <a href="https://openrouter.ai">OpenRouter</a> since that usually provides a proxy to an official API without me needing a new API key.</p>
<p>Most of my pelicans are generated using <a href="https://llm.datasette.io/">my LLM CLI tool</a>, which helps encourage me to ensure the latest models are supported by that (via one of its plugins).</p>
<p>More importantly though, even the act of a single prompt to "Generate an SVG of a pelican riding a bicycle" can reveal interesting model characteristics.</p>
<p>Consider <a href="https://gist.github.com/simonw/66a2699eb1594258904c7b5102840dd6">the result</a> for Kimi K3 today. Running those simple prompts helped emphasize several points about the model.</p>
<ol>
<li>It only has one reasoning effort right now, "max" - and it shows. The model consumed 13,241 reasoning tokens to output 3,417 tokens of response. This is expensive - the pelican cost 25 cents!</li>
<li>How does the prompt "Generate an SVG of a pelican riding a bicycle" add up to 95 input tokens? OpenAI's <a href="https://platform.openai.com/tokenizer">tokenizer</a> counts 10, <a href="https://tools.simonwillison.net/claude-token-counter">Anthropic's</a> counts 10 for Opus 4.6, 30 for Opus 4.7 and 25 for Sonnet 5/Fable 5. Prompting "hi" <a href="https://news.ycombinator.com/item?id=48935342#48936461">to Kimi K3</a> counted 86 tokens, suggesting there may be an 85 token hidden system prompt. It <a href="https://news.ycombinator.com/item?id=48935342#48936515">refused to leak it</a> though.</li>
<li>Vision works well: the alt text it generated is very good.</li>
</ol>
<p>K3 currently only has one thinking effort level, but I've been deriving quite a bit of value recently from running the same pelican prompt through different effort levels to get a quick idea for what impact those have. Here's my matrix <a href="https://static.simonwillison.net/static/2026/gpt-5.6-pelicans.html">for the GPT-5.6 model family</a>, for example.</p>
<p>Really though the main things I gain from the pelican test are:</p>
<ol>
<li>It's a "hello world" exercise for prompting a model</li>
<li>A rough cost and reasoning estimate for a simple task</li>
<li>Confirmation that the model can output valid SVG and has a basic idea of geometry and spatial awareness. This is a much bigger deal for the smaller models that run on my laptop.</li>
<li>It's still interesting to compare pelicans between releases in the same model family. K3's pelican is a notable improvement from <a href="https://simonwillison.net/2026/Jan/27/kimi-k25/">Kimi 2.5</a>.</li>
<li>It's something I can share that demonstrates I've tried it. Plus a comment with a pelican in it is kind of a tradition on Hacker News at this point, any time I'm late I get comments asking where it is!</li>
</ol>
<p>Tags: <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/llm-pricing">llm-pricing</a>, <a href="https://simonwillison.net/tags/pelican-riding-a-bicycle">pelican-riding-a-bicycle</a>, <a href="https://simonwillison.net/tags/llm-release">llm-release</a>, <a href="https://simonwillison.net/tags/ai-in-china">ai-in-china</a>, <a href="https://simonwillison.net/tags/artificial-analysis">artificial-analysis</a>, <a href="https://simonwillison.net/tags/moonshot">moonshot</a>, <a href="https://simonwillison.net/tags/kimi">kimi</a></p></summary>
<category term="ai"/>
<category term="generative-ai"/>
<category term="llms"/>
<category term="llm-pricing"/>
<category term="pelican-riding-a-bicycle"/>
<category term="llm-release"/>
<category term="ai-in-china"/>
<category term="artificial-analysis"/>
<category term="moonshot"/>
<category term="kimi"/>
</entry>
<entry>
<title>Quoting Thibault Sottiaux</title>
<link href="https://simonwillison.net/2026/Jul/16/bad-codex-bug/#atom-everything" rel="alternate"/>
<published>2026-07-16T17:45:59+00:00</published>
<updated>2026-07-16T17:45:59+00:00</updated>
<id>https://simonwillison.net/2026/Jul/16/bad-codex-bug/#atom-everything</id>
<summary type="html"><blockquote cite="https://twitter.com/thsottiaux/status/2077630111499882637"><p>On file deletions. We’ve investigated a handful of reports where GPT-5.6 unexpectedly deleted files. </p>
<p>What we have found is that this most commonly occurs when:</p>
<ul>
<li>Full access mode is enabled and codex is run without sandboxing protections, including without auto review being enabled</li>
<li>The model attempts to override the $HOME env var to define a temporary directory.</li>
<li>The model makes an honest mistake and mistakenly deletes $HOME instead.</li>
</ul></blockquote>
<p class="cite">&mdash; <a href="https://twitter.com/thsottiaux/status/2077630111499882637">Thibault Sottiaux</a>, describing a pretty gnarly Codex bug</p>
<p>Tags: <a href="https://simonwillison.net/tags/codex">codex</a>, <a href="https://simonwillison.net/tags/coding-agents">coding-agents</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a></p></summary>
<category term="codex"/>
<category term="coding-agents"/>
<category term="generative-ai"/>
<category term="ai"/>
<category term="llms"/>
</entry>
<entry>
<title>Inkling: Our open-weights model</title>
<link href="https://simonwillison.net/2026/Jul/16/inkling/#atom-everything" rel="alternate"/>
<published>2026-07-16T15:35:25+00:00</published>
<updated>2026-07-16T15:35:25+00:00</updated>
<id>https://simonwillison.net/2026/Jul/16/inkling/#atom-everything</id>
<summary type="html"><p><strong><a href="https://thinkingmachines.ai/news/introducing-inkling/">Inkling: Our open-weights model</a></strong></p>
Mira Murati's Thinking Machines Lab just released their first open-weights model. Inkling is "a Mixture-of-Experts transformer with 975B total parameters, 41B active" - an Apache-2.0 licensed multimodal model trained on 45 trillion tokens of text, images, audio and video.</p>
<p>They're also promising Inkling-Small, a 276B (12B active) model, but that's still being tested and the weights will be released "once that work is complete".</p>
<p>The <a href="https://thinkingmachines.ai/model-card/inkling/">model card</a> is much shorter than I've come to expect from US AI labs. It links to even shorter <a href="https://thinkingmachines.ai/training-data-documentation/">Training Data Documentation</a> with almost nothing of interest in it - it's best summarized by these two paragraphs:</p>
<blockquote>
<p>The datasets Thinking Machines Lab uses to develop its AI services includes content that is in the public domain as well as content that may be subject to intellectual property protection.</p>
<p>Thinking Machines Lab’s services were developed using publicly available content obtained from the open internet and publicly accessible data repositories. Certain datasets were also obtained from third parties.</p>
</blockquote>
<p>By Thinking Machines' own admission, this is not a frontier model. It's instead intended as a strong base model for fine-tuning using their own <a href="https://thinkingmachines.ai/tinker/">Tinker training platform</a>:</p>
<blockquote>
<p>Inkling is not the strongest overall model available today, open or closed. Instead, a combination of qualities makes it a good open-weights base for customization: multimodal capabilities, efficient thinking, and availability on Tinker for fine-tuning.</p>
</blockquote>
<p>There's a lot to like about this release. It's Apache-2.0 licensed, and looks competitive with the open weight models coming out of China - it's good to see the US open weights ecosystem gain a new viable contender to join NVIDIA Nemotron and Gemma 4.</p>
<p>Here's its attempt at an SVG pelican riding a bicycle, which I generated using this <code>curl</code> command against the Thinking Machines API:</p>
<div class="highlight highlight-source-shell"><pre>curl <span class="pl-s"><span class="pl-pds">"</span>https://tinker.thinkingmachines.dev/services/tinker-prod/oai/api/v1/chat/completions<span class="pl-pds">"</span></span> \
-H <span class="pl-s"><span class="pl-pds">"</span>Authorization: Bearer <span class="pl-smi">$TINKER_API_KEY</span><span class="pl-pds">"</span></span> \
-H <span class="pl-s"><span class="pl-pds">"</span>Content-Type: application/json<span class="pl-pds">"</span></span> \
-d <span class="pl-s"><span class="pl-pds">'</span>{</span>
<span class="pl-s"> "model": "thinkingmachines/Inkling",</span>
<span class="pl-s"> "messages": [</span>
<span class="pl-s"> {"role": "user", "content": "Generate an SVG of a pelican riding a bicycle"}</span>
<span class="pl-s"> ],</span>
<span class="pl-s"> "stream": false</span>
<span class="pl-s"> }<span class="pl-pds">'</span></span></pre></div>
<p>Full <a href="https://gist.github.com/simonw/8117ac4376371dd3fc2b5dbce27e0855">response here</a>.</p>
<p><img alt="See image description below" src="https://static.simonwillison.net/static/2026/inkling-pelican.jpg" /></p>
<p>Since it's a multi-modal model I had it describe its own image (after I rendered it to a JPEG) by sending this JSON:</p>
<div class="highlight highlight-source-json"><pre>{
<span class="pl-ent">"model"</span>: <span class="pl-s"><span class="pl-pds">"</span>thinkingmachines/Inkling<span class="pl-pds">"</span></span>,
<span class="pl-ent">"messages"</span>: [{
<span class="pl-ent">"role"</span>: <span class="pl-s"><span class="pl-pds">"</span>user<span class="pl-pds">"</span></span>,
<span class="pl-ent">"content"</span>: [
{<span class="pl-ent">"type"</span>: <span class="pl-s"><span class="pl-pds">"</span>image_url<span class="pl-pds">"</span></span>, <span class="pl-ent">"image_url"</span>: {<span class="pl-ent">"url"</span>: <span class="pl-s"><span class="pl-pds">"</span>data:image/jpeg;base64,$BASE_64_JPEG<span class="pl-pds">"</span></span>}},
{<span class="pl-ent">"type"</span>: <span class="pl-s"><span class="pl-pds">"</span>text<span class="pl-pds">"</span></span>, <span class="pl-ent">"text"</span>: <span class="pl-s"><span class="pl-pds">"</span>Describe this image in detail<span class="pl-pds">"</span></span>},
],
}],
<span class="pl-ent">"stream"</span>: <span class="pl-c1">false</span>
}</pre></div>
<p>Here's <a href="https://gist.github.com/simonw/dccebd0581f5c4056c5b0ca5bdc38cd9#reasoning_content">the full response</a>. It thought it had drawn a "stork or seagull":</p>
<blockquote>
<p>This is a cheerful, flat-vector cartoon illustration featuring a white bird riding a bicycle across a green landscape.</p>
<p><strong>The Bird:</strong> The central figure is a stylized white bird—resembling a stork or seagull—with a plump, rounded body. It has a long, prominent orange beak that is slightly open, revealing a black line for the mouth. Its eye is small and black with a tiny white highlight. The bird has black wing and tail feathers with white edges, and thin orange legs ending in black feet that rest on the bicycle's orange pedals.</p>
<p><strong>The Bicycle:</strong> The bike has a simple black frame with two large wheels featuring black tires, gray inner rims, and black spokes arranged in a cross pattern. The handlebars are black, and there are bright orange pedals attached to a black crank. A white seat is visible behind the bird's body.</p>
<p><strong>The Setting:</strong> The scene is set against a clear, light blue sky. There are two fluffy white clouds—one on the left side and a smaller one near the bird's head on the right. A bright, solid yellow sun shines in the upper right corner. At the bottom, there are rolling green hills with a smooth, curved top edge where the bicycle rests.</p>
<p>The overall style is playful, simple, and colorful, with clean lines and a bright, sunny atmosphere.</p>
</blockquote>
<p><small></small>Via <a href="https://news.ycombinator.com/item?id=48924912">Hacker News</a></small></p>
<p>Tags: <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/training-data">training-data</a>, <a href="https://simonwillison.net/tags/pelican-riding-a-bicycle">pelican-riding-a-bicycle</a>, <a href="https://simonwillison.net/tags/llm-release">llm-release</a></p></summary>
<category term="ai"/>
<category term="generative-ai"/>
<category term="llms"/>
<category term="training-data"/>
<category term="pelican-riding-a-bicycle"/>
<category term="llm-release"/>
</entry>
</feed>
Raw text
<?xml version="1.0" encoding="utf-8"?>
<feed xml:lang="en-us" xmlns="http://www.w3.org/2005/Atom"><title>Simon Willison's Weblog</title><link href="http://simonwillison.net/" rel="alternate"/><link href="http://simonwillison.net/atom/everything/" rel="self"/><id>http://simonwillison.net/</id><updated>2026-07-27T23:39:04+00:00</updated><author><name>Simon Willison</name></author><entry><title>moonshotai/Kimi-K3</title><link href="https://simonwillison.net/2026/Jul/27/kimi-k3/#atom-everything" rel="alternate"/><published>2026-07-27T23:39:04+00:00</published><updated>2026-07-27T23:39:04+00:00</updated><id>https://simonwillison.net/2026/Jul/27/kimi-k3/#atom-everything</id><summary type="html">
<p><strong><a href="https://huggingface.co/moonshotai/Kimi-K3">moonshotai/Kimi-K3</a></strong></p>
As promised <a href="https://simonwillison.net/2026/Jul/16/kimi-k3/">earlier this month</a>, Moonshot have released the weights for their excellent 2.8 trillion parameter Kimi K3. They're a hefty 1.56TB on Hugging Face.</p>
<p>Kimi introduced their own janky <a href="https://huggingface.co/moonshotai/Kimi-K2-Instruct/blob/main/LICENSE">modified version of the MIT license</a> with K2 back in July 2025. That license just added this paragraph requiring attribution beyond a certain size of commercial entity:</p>
<blockquote>
<p>Our only modification part is that, if the Software (or any derivative works thereof) is used for any of your commercial products or services that have more than 100 million monthly active users, or more than 20 million US dollars (or equivalent in other currencies) in monthly revenue, you shall prominently display "Kimi K2" on the user interface of such product or service.</p>
</blockquote>
<p>The <a href="https://huggingface.co/moonshotai/Kimi-K3/blob/main/LICENSE">K3 license</a> no longer calls itself "modified MIT" and goes further, requiring a separate agreement with Moonshot for large "Model as a Service" businesses:</p>
<blockquote>
<p>If the Licensee or any of its affiliates operates a Model as a Service business, and the aggregate revenue of the Licensee and its affiliates exceeds 20 million US dollars (or the equivalent in other currencies) in total over any consecutive 12 months, the Licensee must enter into a separate agreement with Moonshot AI before using the Software or its derivative works for any commercial purpose.</p>
</blockquote>
<p>To Kimi's credit, they make no attempt to describe this as an "open source" license in their own materials, consistently using the term "open weight" in its place.</p>
<p>OpenRouter is already offering K3 <a href="https://openrouter.ai/moonshotai/kimi-k3">from 7 providers</a>, most of which are at the same $3/million input and $15/million output as Moonshot AI themselves.
<p>Tags: <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/llm-pricing">llm-pricing</a>, <a href="https://simonwillison.net/tags/llm-release">llm-release</a>, <a href="https://simonwillison.net/tags/ai-in-china">ai-in-china</a>, <a href="https://simonwillison.net/tags/moonshot">moonshot</a>, <a href="https://simonwillison.net/tags/kimi">kimi</a>, <a href="https://simonwillison.net/tags/janky-licenses">janky-licenses</a></p>
</summary><category term="ai"/><category term="generative-ai"/><category term="llms"/><category term="llm-pricing"/><category term="llm-release"/><category term="ai-in-china"/><category term="moonshot"/><category term="kimi"/><category term="janky-licenses"/></entry><entry><title>An opinionated guide to which AI to use to do stuff</title><link href="https://simonwillison.net/2026/Jul/27/an-opinionated-guide-to-which-ai-to-use-to-do-stuff/#atom-everything" rel="alternate"/><published>2026-07-27T21:55:53+00:00</published><updated>2026-07-27T21:55:53+00:00</updated><id>https://simonwillison.net/2026/Jul/27/an-opinionated-guide-to-which-ai-to-use-to-do-stuff/#atom-everything</id><summary type="html">
<p><strong><a href="https://www.oneusefulthing.org/p/an-opinionated-guide-to-which-ai-b22">An opinionated guide to which AI to use to do stuff</a></strong></p>
It's interesting watching the evolution of Ethan Mollick's guide over time. </p>
<p><a href="https://www.oneusefulthing.org/p/using-ai-right-now-a-quick-guide">A year ago</a> it was still all about chat - ChatGPT, Claude, Gemini - with o3, Claude 4 Opus, and Gemini 2.5 Pro as the models and Deep Research as a useful alternative mode.</p>
<p>Today it's much more about agentic systems - "where the AI is capable of doing the equivalent of many hours of real human work in one go".</p>
<p>Gemini has fallen off Ethan's list, since Google still doesn’t have an established entry in the Codex/ChatGPT Work/Cowork category. <a href="https://gemini.google/overview/agent/spark/">Gemini Spark</a> has yet to prove itself!</p>
<p>Ethan offers a useful explanation of the ways you can give ChatGPT or Claude a computer to use:</p>
<blockquote>
<p>To use the computers provided by the AI companies, the mode you want is called ChatGPT Work in ChatGPT, and Cowork in Claude (the naming will not get less confusing, I am sorry to say). [...]</p>
<p>The most powerful way to use AI is to give it access to your computer. You do that by downloading the ChatGPT or Claude apps and picking a mode to use. ChatGPT's two agent modes are Work and Codex; Claude's are Cowork and Code. The names do not map onto each other in any way that will help you remember them. And yes, these use the same names as the Work and Cowork modes we discussed above, but operate differently, and have more features and capabilities because they can access your computer.</p>
</blockquote>
<p>I think the difference between ChatGPT Work on a mobile device and ChatGPT Work inside the desktop app (where it's effectively a less intimidating skin on top of Codex) is spectacularly unintuitive.</p>
<p>Short version: if you flip ChatGPT mobile from "Chat" to "Work" mode you get a version where its Code Interpreter container is no longer restricted from accessing the internet!
<p>Tags: <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/ethan-mollick">ethan-mollick</a>, <a href="https://simonwillison.net/tags/code-interpreter">code-interpreter</a>, <a href="https://simonwillison.net/tags/general-agents">general-agents</a></p>
</summary><category term="ai"/><category term="generative-ai"/><category term="llms"/><category term="ethan-mollick"/><category term="code-interpreter"/><category term="general-agents"/></entry><entry><title>An Inside Look at the Relay Market Powering Token Resellers and Fraud</title><link href="https://simonwillison.net/2026/Jul/26/relay-market/#atom-everything" rel="alternate"/><published>2026-07-26T19:30:54+00:00</published><updated>2026-07-26T19:30:54+00:00</updated><id>https://simonwillison.net/2026/Jul/26/relay-market/#atom-everything</id><summary type="html">
<p><strong><a href="https://vectoral.com/blog/token-relay-market">An Inside Look at the Relay Market Powering Token Resellers and Fraud</a></strong></p>
Fascinating investigation by Matt Lenhard into the market that has grown up around reselling LLM tokens at a discount by pooling API keys from various sources.</p>
<p>This looks to be mostly a thing in China. Resellers sell access to an LLM proxy that offers significant discounts on regular API pricing, which they achieve by abusing free trials, proxying through unprotected support bots, or sometimes through stolen credit cards or chargeback attacks.</p>
<p>The software they are using for these proxies is open source - mostly <a href="https://github.com/songquanpeng/one-api">one-api</a> and its more actively developed fork <a href="https://github.com/QuantumNous/new-api">new-api</a>, both legitimate API proxy products which can be used to load. balance requests across a pool of API credentials.</p>
<p>The buyers are seeking cheap tokens, avoiding geo-restrictions, and in some cases collecting data for model distillation.</p>
<p>I've been cautious about exposing my own LLM-driven applications publicly out of fear of abuse leading to big token bills. The existence of this marketplace makes me even more cautious: there's now an entire ecosystem that can profit from finding a new unprotected endpoint to exploit.</p>
<p>LLM vendors <em>really</em> need to get better at offering strict caps for their API keys. I want my LLM apps to stop working the moment they hit a dollar threshold I've set for a period of time.</p>
<p>Here's <a href="https://www.v2ex.com/t/1196011">the (Chinese language) forum thread</a> that served as the principal source for Matt's article.
<p><small></small>Via <a href="https://news.ycombinator.com/item?id=49058993">Hacker News</a></small></p>
<p>Tags: <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/llm-pricing">llm-pricing</a>, <a href="https://simonwillison.net/tags/ai-ethics">ai-ethics</a>, <a href="https://simonwillison.net/tags/ai-in-china">ai-in-china</a></p>
</summary><category term="ai"/><category term="generative-ai"/><category term="llms"/><category term="llm-pricing"/><category term="ai-ethics"/><category term="ai-in-china"/></entry><entry><title>Ruff v0.16.0</title><link href="https://simonwillison.net/2026/Jul/25/ruff/#atom-everything" rel="alternate"/><published>2026-07-25T22:44:05+00:00</published><updated>2026-07-25T22:44:05+00:00</updated><id>https://simonwillison.net/2026/Jul/25/ruff/#atom-everything</id><summary type="html">
<p><strong><a href="https://astral.sh/blog/ruff-v0.16.0">Ruff v0.16.0</a></strong></p>
Astral shipped a significant new version of their Ruff Python linting tool a few days ago on July 23rd. I noticed today because my various CI jobs all started failing thanks to new default Ruff checks and my unpinned <code>"ruff"</code> dev dependency.</p>
<p>From Brent Westbrook's announcement post:</p>
<blockquote>
<p>Ruff now enables 413 rules by default, up from 59 in previous versions.</p>
<p>Since Ruff's default rule set was last modified in <a href="https://github.com/astral-sh/ruff/blob/main/changelogs/0.1.x.md#breaking-changes">v0.1.0</a>, the number of rules in Ruff has grown from 708 to 968. Many of these rules catch severe issues, including <a href="https://docs.astral.sh/ruff/rules/load-before-global-declaration">syntax errors</a> and <a href="https://docs.astral.sh/ruff/rules/yield-in-init/">immediate runtime errors</a> but were not previously enabled by default. With the new rule set, Ruff will bring these issues and many others to your attention without any Ruff configuration.</p>
</blockquote>
<p>Here's a one-liner for trying it on any Python project:</p>
<pre><code>uvx ruff@latest check .
</code></pre>
<p>I ran the latest Ruff against my three biggest projects - <a href="https://datasette.io/">Datasette</a>, <a href="https://sqlite-utils.datasette.io/">sqlite-utils</a>, and <a href="https://llm.datasette.io/">LLM</a> - and it found <em>hundreds</em> of minor issues that breached the new default rules.</p>
<p>All three projects have very comprehensive test suites, executed in CI against Python 3.10 through Python 3.14, so upgrades like this are pretty safe. The following command did the bulk of the upgrades:</p>
<pre><code>uvx ruff@latest check . --fix --unsafe-fixes
</code></pre>
<p>Against <code>sqlite-utils</code>, that command reported:</p>
<pre><code>Found 1618 errors (1538 fixed, 80 remaining).
</code></pre>
<p>As an illustrative example, here are three of the remaining issues. Ruff does a nice job of explaining each one:</p>
<pre><code>DTZ005 `datetime.datetime.now()` called without a `tz` argument
--&gt; tests/test_duplicate.py:17:10
|
15 | "datetime_col" TEXT)""")
16 | # Insert one row of mock data:
17 | dt = datetime.datetime.now()
| ^^^^^^^^^^^^^^^^^^^^^^^
18 | data = {
19 | "text_col": "Cleo",
|
help: Pass a `datetime.timezone` object to the `tz` parameter
BLE001 Do not catch blind exception: `Exception`
--&gt; tests/test_plugins.py:16:12
|
14 | db.execute("select * from pragma_function_list()")
15 | return True
16 | except Exception:
| ^^^^^^^^^
17 | return False
18 | finally:
|
B018 Found useless attribute access. Either assign it to a variable or remove it.
--&gt; tests/test_update.py:46:5
|
44 | def test_update_invalid_pk(fresh_db, pk, update_pk):
45 | table = fresh_db["table"]
46 | table.insert({"id1": 5, "id2": 3, "v": 1}, pk=pk).last_pk
| ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
47 | with pytest.raises(NotFoundError):
48 | table.update(update_pk, {"v": 2})
|
</code></pre>
<p>Unsurprisingly, given Astral's <a href="https://simonwillison.net/2026/Mar/19/openai-acquiring-astral/">new home at OpenAI</a>, this output provides everything a coding agent would need to fix the problems.</p>
<p>I had Codex (GPT-5.6 Sol high) <a href="https://github.com/simonw/llm/pull/1557">upgrade LLM</a> and <a href="https://github.com/simonw/sqlite-utils/pull/814">sqlite-utils</a>, and Claude Code (with Opus 5) <a href="https://github.com/simonw/datasette/pull/2857">upgrade Datasette</a>.
<p>Tags: <a href="https://simonwillison.net/tags/python">python</a>, <a href="https://simonwillison.net/tags/ruff">ruff</a>, <a href="https://simonwillison.net/tags/astral">astral</a></p>
</summary><category term="python"/><category term="ruff"/><category term="astral"/></entry><entry><title>Quoting Boris Cherny</title><link href="https://simonwillison.net/2026/Jul/25/boris-cherny/#atom-everything" rel="alternate"/><published>2026-07-25T00:42:59+00:00</published><updated>2026-07-25T00:42:59+00:00</updated><id>https://simonwillison.net/2026/Jul/25/boris-cherny/#atom-everything</id><summary type="html">
<blockquote cite="https://twitter.com/bcherny/status/2080713091688583312"><p>More than any of these eval scores, what is most exciting to me is something else: Opus 5 is our least prompt injectable model yet. It is a bit buried in the system card, but across PI evals and red teaming, Opus 5 is very hard to prompt inject successfully.</p></blockquote>
<p class="cite">&mdash; <a href="https://twitter.com/bcherny/status/2080713091688583312">Boris Cherny</a>, here's that <a href="https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf#page=73">System Card section</a>, page 73</p>
<p>Tags: <a href="https://simonwillison.net/tags/prompt-injection">prompt-injection</a>, <a href="https://simonwillison.net/tags/anthropic">anthropic</a>, <a href="https://simonwillison.net/tags/claude">claude</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/boris-cherny">boris-cherny</a></p>
</summary><category term="prompt-injection"/><category term="anthropic"/><category term="claude"/><category term="generative-ai"/><category term="ai"/><category term="llms"/><category term="boris-cherny"/></entry><entry><title>Introducing Claude Opus 5</title><link href="https://simonwillison.net/2026/Jul/24/introducing-claude-opus-5/#atom-everything" rel="alternate"/><published>2026-07-24T23:48:50+00:00</published><updated>2026-07-24T23:48:50+00:00</updated><id>https://simonwillison.net/2026/Jul/24/introducing-claude-opus-5/#atom-everything</id><summary type="html">
<p><strong><a href="https://www.anthropic.com/news/claude-opus-5">Introducing Claude Opus 5</a></strong></p>
I've been offline <a href="https://en.wikipedia.org/wiki/Elkhorn_Slough">kayaking with sea otters</a> for much of today so I haven't had a chance to put Anthropic's new model Claude Opus 5 through its paces yet. The buzz is positive, and Anthropic's description of it as a "thoughtful and proactive model that comes close to the frontier intelligence of Claude Fable 5 at half the price" sounds promising. It's currently <a href="https://twitter.com/artificialanlys/status/2080777718933995967">leading the Artificial Analysis leaderboard</a>, in front of even Fable 5.</p>
<p>It's priced the same as Opus 4.8, and continues to offer a "fast mode" at twice the cost of the base model.</p>
<p>Based on this anecdote in the release post it sounds like it might be <a href="https://simonwillison.net/2026/Jun/11/fable-is-relentlessly-proactive/">relentlessly proactive</a>:</p>
<blockquote>
<p>On one Frontier-Bench task, Opus 5 was given a drawing of a machine part and asked to write code to rebuild it as a 3D FreeCAD model. However, in this task, the model was intentionally given no way to directly viewthe drawing. Opus 5 responded by writing its own computer vision pipeline to pull the geometry from the raw pixels, then reconstructed the full machine part.</p>
</blockquote>
<p>It's better at finding vulnerabilities but has deliberately not been trained on how to exploit them. Hopefully this means the US government won't shut it down!</p>
<blockquote>
<p>As with its predecessor, Opus 4.8, we’ve intentionally avoided training Opus 5 on cyber tasks. The model has nevertheless improved substantially on these tasks as a result of becoming more generally capable, and it comes close to Mythos 5 at <em>finding</em> cybersecurity vulnerabilities. However, it remains substantially behind Mythos 5 on the <em>exploitation</em> of those vulnerabilities—that is, in turning vulnerabilities into material cyber threats.</p>
</blockquote>
<p>Anthropic have published a <a href="https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5">prompting guide for Claude Opus 5</a>. Thariq Shihipar has also written <a href="https://claude.com/blog/the-new-rules-of-context-engineering-for-claude-5-generation-models">The new rules of context engineering for Claude 5 generation models</a>.</p>
<p>The <a href="https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fraw.githubusercontent.com%2Fsimonw%2Fllm-anthropic%2F8272dfee5bdb65d5c88eef083da3ad885539b7df%2Flog.md">first pelican I got</a> was missing the bicycle wheels; the <a href="https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fraw.githubusercontent.com%2Fsimonw%2Fllm-anthropic%2Ffeaab840ea20eb15e29d8f72a9e42feceb23876a%2Flog.md">second attempt</a> was better.
<p>Tags: <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/anthropic">anthropic</a>, <a href="https://simonwillison.net/tags/claude">claude</a>, <a href="https://simonwillison.net/tags/llm-release">llm-release</a></p>
</summary><category term="ai"/><category term="generative-ai"/><category term="llms"/><category term="anthropic"/><category term="claude"/><category term="llm-release"/></entry><entry><title>The first known runaway AI agent - or a very bad marketing stunt?</title><link href="https://simonwillison.net/2026/Jul/23/the-first-known-runaway-ai-agent/#atom-everything" rel="alternate"/><published>2026-07-23T22:53:08+00:00</published><updated>2026-07-23T22:53:08+00:00</updated><id>https://simonwillison.net/2026/Jul/23/the-first-known-runaway-ai-agent/#atom-everything</id><summary type="html">
<p><strong><a href="https://martinalderson.com/posts/huggingface-openai-exploit/">The first known runaway AI agent - or a very bad marketing stunt?</a></strong></p>
Martin Alderson's commentary on the <a href="https://simonwillison.net/2026/Jul/22/openai-cyberattack/">OpenAI accidental cyberattack against Hugging Face</a> includes a couple of details I hadn't considered.</p>
<p>First, Hugging Face offers a truly rich target if you're trying to find potential vulnerabilities that require executing arbitrary code:</p>
<blockquote>
<p>Hugging Face has an <em>enormous</em> attack surface. They have more interfaces than I can count which run untrusted models and code. While they definitely have invested in defences, by nature of their operating model they do have many more opportunities to be attacked than many other services. I certainly don't envy their cybersecurity teams.</p>
</blockquote>
<p>Secondly, one of the things that has puzzled me is how OpenAI didn't notice that their sandbox had been so thoroughly breached by the agent. Surely they'd be monitoring network traffic closely?</p>
<p>Martin points out that:</p>
<blockquote>
<p>It's also likely they were running a huge amount of benchmarks simultaneously with ~unlimited token budgets - you want as many samples as possible to figure out how good a model is at a certain benchmark. It may also be they are testing various different checkpoints of the model too, understanding how the model is improving as it goes through the various training stages.</p>
</blockquote>
<p>The mistakes made by the OpenAI team running this benchmark are easier to imagine when you think about the scale at which benchmarks of this kind usually operate. For all we know they could have been subjecting a new model to dozens of benchmarks at the same time, in dozens of different environments.
<p><small></small>Via <a href="https://lobste.rs/s/nsnb4j/first_known_runaway_ai_agent_very_bad">Lobste.rs</a></small></p>
<p>Tags: <a href="https://simonwillison.net/tags/security">security</a>, <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/openai">openai</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/hugging-face">hugging-face</a>, <a href="https://simonwillison.net/tags/ai-security-research">ai-security-research</a></p>
</summary><category term="security"/><category term="ai"/><category term="openai"/><category term="generative-ai"/><category term="llms"/><category term="hugging-face"/><category term="ai-security-research"/></entry><entry><title>Quoting Seth Larson</title><link href="https://simonwillison.net/2026/Jul/23/seth-larson/#atom-everything" rel="alternate"/><published>2026-07-23T04:50:36+00:00</published><updated>2026-07-23T04:50:36+00:00</updated><id>https://simonwillison.net/2026/Jul/23/seth-larson/#atom-everything</id><summary type="html">
<blockquote cite="https://blog.pypi.org/posts/2026-07-22-releases-now-reject-new-files-after-14-days/"><p>The Python Package Index (PyPI) now rejects new files being uploaded to releases that are older than 14 days. This restriction was <a href="https://github.com/pypi/warehouse/pull/19727">put in place</a> to prevent old and long-stable releases from being poisoned in case publishing tokens or workflows of PyPI projects were compromised. As far as we are aware this has not yet been abused, but there is no technical reason beyond that attackers weren't aware it was possible.</p></blockquote>
<p class="cite">&mdash; <a href="https://blog.pypi.org/posts/2026-07-22-releases-now-reject-new-files-after-14-days/">Seth Larson</a>, PyPI blog</p>
<p>Tags: <a href="https://simonwillison.net/tags/packaging">packaging</a>, <a href="https://simonwillison.net/tags/python">python</a>, <a href="https://simonwillison.net/tags/supply-chain">supply-chain</a>, <a href="https://simonwillison.net/tags/pypi">pypi</a>, <a href="https://simonwillison.net/tags/seth-michael-larson">seth-michael-larson</a></p>
</summary><category term="packaging"/><category term="python"/><category term="supply-chain"/><category term="pypi"/><category term="seth-michael-larson"/></entry><entry><title>Quoting Thomas Ptacek</title><link href="https://simonwillison.net/2026/Jul/22/thomas-ptacek/#atom-everything" rel="alternate"/><published>2026-07-22T23:59:01+00:00</published><updated>2026-07-22T23:59:01+00:00</updated><id>https://simonwillison.net/2026/Jul/22/thomas-ptacek/#atom-everything</id><summary type="html">
<blockquote cite="https://twitter.com/tqbf/status/2080045032162173329"><p>I genuinely believe that if you took an open weights model from 2025 and built a pentest harness for it, it could do this kind of sandbox escape and scan/hack in most networks. This is only surprising because you assume OpenAI has sounder sandboxes.</p></blockquote>
<p class="cite">&mdash; <a href="https://twitter.com/tqbf/status/2080045032162173329">Thomas Ptacek</a>, doesn't think <a href="https://simonwillison.net/2026/Jul/22/openai-cyberattack/#resist-the-temptation-to-write-this-off-as-a-stunt">this even needs</a> a frontier model</p>
<p>Tags: <a href="https://simonwillison.net/tags/thomas-ptacek">thomas-ptacek</a>, <a href="https://simonwillison.net/tags/openai">openai</a>, <a href="https://simonwillison.net/tags/security">security</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/ai-security-research">ai-security-research</a>, <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/sandboxing">sandboxing</a></p>
</summary><category term="thomas-ptacek"/><category term="openai"/><category term="security"/><category term="generative-ai"/><category term="ai-security-research"/><category term="ai"/><category term="llms"/><category term="sandboxing"/></entry><entry><title>OpenAI’s accidental cyberattack against Hugging Face is science fiction that happened</title><link href="https://simonwillison.net/2026/Jul/22/openai-cyberattack/#atom-everything" rel="alternate"/><published>2026-07-22T23:51:33+00:00</published><updated>2026-07-22T23:51:33+00:00</updated><id>https://simonwillison.net/2026/Jul/22/openai-cyberattack/#atom-everything</id><summary type="html">
<p>This story is wild. The short version: OpenAI were running a cybersecurity test against an unreleased model, with the model's guardrail features turned off. Rather than solve the test, the model broke its way out of OpenAI's sandbox, then found exploits to break <em>in</em> to Hugging Face, all so it could cheat on the test by stealing the answers.</p>
<p>Along the way it helped make the strongest case yet for how the imbalance of model availability is hurting our ability to secure our software.</p>
<h4 id="here-s-what-happened">Here's what happened</h4>
<p>We currently have three documents to help us understand what happened here.</p>
<ol>
<li>
<a href="https://arxiv.org/abs/2605.11086">ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?</a> is a paper published on 11th May 2026 describing ExploitGym, a new eval suite for LLM-powered agent systems.</li>
<li>
<a href="https://huggingface.co/blog/security-incident-july-2026">Security incident disclosure — July 2026</a> by Hugging Face on 16th July 2026 describes how they detected an attack from an "agentic security-research harness - used LLM still not known" that breached some of their systems.</li>
<li>
<a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/">OpenAI and Hugging Face partner to address security incident during model evaluation</a> from OpenAI on 21st July 2026 confesses that it was <em>their</em> agent harness that did this, and that they're working with Hugging Face to clean up the mess.</li>
</ol>
<h4 id="exploitgym">ExploitGym</h4>
<p>I hadn't seen the <a href="https://arxiv.org/abs/2605.11086">ExploitGym paper</a> before and it's a really interesting one. Authors from UC Berkeley, the Max Planck Institute, UC Santa Barbara, and Arizona State designed a new benchmark for evaluating models on their ability to turn a reported vulnerability into a concrete exploit. OpenAI, Anthropic, and Google provided feedback and helped run the benchmark against their models.</p>
<p>The benchmark "comprises 898 instances derived from real-world vulnerabilities that affected popular software projects" - including the Linux kernel and V8 JavaScript engine. The ExploitGym benchmark is <a href="https://github.com/sunblaze-ucb/exploitgym">available on GitHub</a>.</p>
<p>Here's the paragraph that best represents their benchmark results:</p>
<blockquote>
<p>Among all configurations, Claude Mythos Preview and GPT-5.5 achieve the highest success counts (157 and 120 successes, respectively), demonstrating that current frontier agents can exploit a substantial subset of real-world vulnerabilities under controlled conditions. GPT-5.4 also solves a notable 54 tasks, placing it in an intermediate tier. The remaining model–agent pairings solve fewer than 15 tasks each, underscoring that end-to-end exploitation remains challenging and sharply differentiates today’s frontier systems. Notably, Claude Opus 4.7 achieves fewer successes than Claude Opus 4.6 despite being a newer checkpoint, and does so at substantially lower cost on the full set. Trace inspection reveals that Claude Opus 4.7 and Gemini 3.1 Pro frequently conclude early after judging the target vulnerability non-exploitable.</p>
</blockquote>
<p>The paper also describes the approach they took to preventing the agents from cheating by going outside the parameters of the test. This becomes relevant in a moment!</p>
<blockquote>
<p>Outbound connections are restricted to a curated allowlist that permits routine package installation (Ubuntu apt repositories and PyPI) and fetching the toolchains required for building V8. All other external endpoints are blocked.</p>
</blockquote>
<p>The paper concludes with this (emphasis mine):</p>
<blockquote>
<p>Our results show that <strong>autonomous exploit development by frontier AI agents is no longer a hypothetical capability</strong>. While current agents are not yet reliable across all targets, they already <strong>exploit a non-trivial fraction of real-world vulnerabilities</strong>, including complex targets such as kernel components. This rapid emergence is itself a central finding, showing that capabilities that would have seemed implausible are now present in deployed frontier models.</p>
</blockquote>
<p>An important detail here: this paper isn't about discovering vulnerabilities; it's about being able to take those vulnerabilities and turn them into working exploits.</p>
<p>When Anthropic first restricted access to Mythos <a href="https://simonwillison.net/2026/Apr/7/project-glasswing/">back in April</a> they talked about this capability as well. A model that can act on vulnerabilities is a lot more dangerous than one that can just discover them.</p>
<p>One of the ways Fable differs from Mythos is that it's more likely to refuse to weaponize vulnerabilities in this way. I get the impression the US government did not understand that distinction when they banned Fable <a href="https://simonwillison.net/2026/Jun/16/fable-5-export-controls/">last month</a>.</p>
<h4 id="the-hugging-face-incident">The Hugging Face incident</h4>
<p>The first hint we got of the attack was in <a href="https://huggingface.co/blog/security-incident-july-2026">this blog post by Hugging Face</a> on 16th July 2026:</p>
<blockquote>
<p>A malicious dataset abused two code-execution paths in our dataset processing (a remote-code dataset loader and a template-injection in a dataset configuration) to run code on a processing worker. From there, the actor escalated to node-level access, harvested cloud and cluster credentials, and moved laterally into several internal clusters over a weekend.</p>
</blockquote>
<p>I hope they release more details about the code that pulled this off. I'm assuming this means packages using the <a href="https://github.com/huggingface/datasets">datasets library</a>, a Hugging Face project for bundling up and sharing datasets on their platform. That library used to execute arbitrary code but has been steadily locked down over time, with the <a href="https://github.com/huggingface/datasets/releases/tag/4.0.0">4.0.0 release</a> in July 2025 removing the <code>trust_remote_code=True</code> flag entirely.</p>
<p>Assuming the attack used that library it must have either abused pickle serialization in some way, found some other non-obvious code execution path, or (most likely) specified <code>datasets&lt;4.0.0</code> as the dependency.</p>
<blockquote>
<p>The campaign was run by an autonomous agent framework (appearing to be built on an agentic security-research harness - used LLM still not known) executing many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services.</p>
</blockquote>
<p>This was a sophisticated attack!</p>
<p>Then Hugging Face hit a wall: they tried to use "frontier models behind commercial APIs" - I'm guessing from Anthropic and OpenAI - to help analyze the attack, and were blocked:</p>
<blockquote>
<p>When we started the log analysis, we first used frontier models behind commercial APIs. This did not work: the analysis requires submitting large volumes of real attack commands, exploit payloads, and C2 artifacts, and these requests were blocked by the providers' safety guardrails, which cannot distinguish an incident responder from an attacker.</p>
</blockquote>
<p>They switched to their own self-hosted instance of MIT licensed GLM-5.2 and it helped them figure out what was going on.</p>
<p>This indicated a fundamental asymmetry between the defending team and the (so-far unknown) attacker:</p>
<blockquote>
<p>We do not know which model powered the attacker's agents, whether a jailbroken hosted model or an unrestricted open-weight one; either way, the attacker was bound by no usage policy, while our own forensic work was blocked by the guardrails of the hosted models we first tried.</p>
</blockquote>
<p>As a useful indicator of how seriously they took the attack:</p>
<blockquote>
<p>[...] Finally, we have also reported this incident to law enforcement agencies.</p>
</blockquote>
<p>So who was responsible for this "autonomous agent framework"? It turned out to be OpenAI themselves.</p>
<h4 id="the-openai-confession">The OpenAI confession</h4>
<p>Five days later, <a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/">on July 21st</a>, OpenAI revealed the culprit. They had been running the ExploitGym benchmark against a new, as-yet undisclosed model, and that model had been operating <em>way</em> outside its intended parameters (emphasis mine):</p>
<blockquote>
<p>After investigating, we now know <strong>that this particular incident was driven by a combination of OpenAI models</strong> — including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a <a href="https://arxiv.org/abs/2605.11086">benchmark</a> [ExploitGym] of cyber capabilities. [...]</p>
<p>We estimate maximal cyber capabilities by <strong>running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity</strong>. Our benchmarks run in a highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts as a proxy and cache for package registries.</p>
<p>The models <strong>identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure</strong> to obtain test solutions directly from Hugging Face’s production database. All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.</p>
</blockquote>
<p>It's pretty clear what happened here. OpenAI removed safety filters for an in-progress model, locked it up in a sandbox and told it to solve the ExploitGym problems. Given the absence of guardrails there was nothing to prevent the model from attempting to break out of that sandbox, break into Hugging Face, and read the answers from there instead.</p>
<p>OpenAI's sandbox used the now common pattern of allowing access to an allow-list of sites related to installing additional packages via a controlled proxy. In OpenAI's words:</p>
<blockquote>
<p>While operating in our sandboxed testing environment, our models <strong>spent a substantial amount of inference compute finding a way to obtain open Internet access</strong>, in pursuit of solving the evaluation problem. To gain access, the models identified and exploited <strong>a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy</strong>. With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access.</p>
</blockquote>
<p>So step one was to break out onto the public internet. The model then broke into Hugging Face to find the answers:</p>
<blockquote>
<p>After gaining Internet access, the models <strong>inferred that Hugging Face potentially hosted models, datasets and solutions</strong> for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation. In one example, <strong>the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities</strong> to find a remote code execution path on the Hugging Face servers.</p>
</blockquote>
<p>Chaining together multiple attack vectors is <em>exactly</em> the kind of thing these new models can do, where previous generations of models might have failed.</p>
<p>I wrote last month about how <a href="https://simonwillison.net/2026/Jun/11/fable-is-relentlessly-proactive/">Claude Fable is relentlessly proactive</a>, when I noticed it spinning up custom web servers and deploying CORS tricks on my own laptop just to help debug a WebKit CSS issue. It turns out relentless proactivity is the defining trait of this new generation of Mythos-class models. If you set them a goal and give them a way to get there, even inadvertently, they <em>will figure it out</em>.</p>
<h4 id="resist-the-temptation-to-write-this-off-as-a-stunt">Resist the temptation to write this off as a stunt</h4>
<p>There will inevitably be some people who dismiss this story as a dishonest marketing trick by OpenAI to make their models sound terrifyingly effective. I found 81 instances of the term "marketing" in <a href="https://news.ycombinator.com/item?id=48997548">the Hacker News discussion</a> of the incident.</p>
<p>To those people I say <em>pull your heads out of the sand</em> - you're now including Hugging Face in your conspiracy theories, just so you can deny the crescendo of evidence here!</p>
<p>The best models we have today have the ability to both find and exploit new vulnerabilities. The ExploitGym paper itself concludes that "autonomous exploit development by frontier AI agents is no longer a hypothetical capability", and this incident is a perfect example of exactly that.</p>
<h4 id="the-asymmetry-is-increasingly-frustrating">The asymmetry is increasingly frustrating</h4>
<p>One of the most infuriating details of this story is how Hugging Face, faced with an accidental and aggressive attack from one of OpenAI's models, were unable to then turn to OpenAI's models to help them fend off the attack.</p>
<p>The frontier models we have access to are increasingly being constrained in how much they can help us protect our software, heavily influenced by the US government's ongoing threat of export controls. Claude Fable 5 wouldn't even <a href="https://simonwillison.net/guides/agentic-engineering-patterns/prompts/#proofreader">proofread this article</a> for me! It insisted on downgrading me to a less capable model.</p>
<p>Meanwhile open weight models from China such as GLM-5.2, Kimi 3 and the new Qwen 3.8 Max appear to have none of these restrictions - and any restrictions that <em>do</em> exist can likely be fine-tuned out of them by modifying the weights</p>
<p>These constraints are meant to make us safer. I think there's a risk that they are having the opposite effect.</p>
<p>Tags: <a href="https://simonwillison.net/tags/sandboxing">sandboxing</a>, <a href="https://simonwillison.net/tags/security">security</a>, <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/openai">openai</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/hugging-face">hugging-face</a>, <a href="https://simonwillison.net/tags/anthropic">anthropic</a>, <a href="https://simonwillison.net/tags/paper-review">paper-review</a>, <a href="https://simonwillison.net/tags/ai-security-research">ai-security-research</a></p>
</summary><category term="sandboxing"/><category term="security"/><category term="ai"/><category term="openai"/><category term="generative-ai"/><category term="llms"/><category term="hugging-face"/><category term="anthropic"/><category term="paper-review"/><category term="ai-security-research"/></entry><entry><title>Are AI labs pelicanmaxxing?</title><link href="https://simonwillison.net/2026/Jul/22/are-ai-labs-pelicanmaxxing/#atom-everything" rel="alternate"/><published>2026-07-22T23:01:00+00:00</published><updated>2026-07-22T23:01:00+00:00</updated><id>https://simonwillison.net/2026/Jul/22/are-ai-labs-pelicanmaxxing/#atom-everything</id><summary type="html">
<p><strong><a href="https://dylancastillo.co/posts/pelicanmaxxing.html">Are AI labs pelicanmaxxing?</a></strong></p>
Excellent piece of work by Dylan Castillo, who took a deep-dive into the frequently pondered question of whether the AI labs have been deliberately training models to draw pelicans riding bicycles in response to my <a href="https://simonwillison.net/tags/pelican-riding-a-bicycle/">deeply unscientific benchmark</a>.</p>
<p>I've been randomly spot-checking this in the past by testing models against other animals riding other types of vehicle, but never with anything close to the diligence of Dylan's methodology here.</p>
<p>Dylan took 8 animals × 6 vehicles = 48 prompts and ran them three times each through 7 different models ( GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.5 Flash, Grok 4.5, Qwen3.7-Max, GLM-5.2, and DeepSeek V4 Pro). He then used GPT-5.6 Luna and Gemini 3.1 Flash-Lite to help evaluate the results.</p>
<p>There's a neat filter view for exploring the results:</p>
<p><img alt="Screenshot of a grid for sample 1/3 of GLM-5.2, with pelicn and flamingo and heron riding bicycle, unicycle, skateboard, scooter, plane and boat" src="https://static.simonwillison.net/static/2026/pelican-grid.webp" /></p>
<p>For the models he tested he could find no evidence of pelimaxxing:</p>
<blockquote>
<ul>
<li><a href="https://dylancastillo.co/posts/pelicanmaxxing.html#evidence-1-the-pelicans-on-bicycles-dont-look-any-better">The pelicans on bicycles don’t look any better</a></li>
<li><a href="https://dylancastillo.co/posts/pelicanmaxxing.html#evidence-2-labs-are-not-better-at-drawing-pelicans">Labs are not better at drawing pelicans</a></li>
<li><a href="https://dylancastillo.co/posts/pelicanmaxxing.html#evidence-3-labs-are-not-better-at-drawing-bicycles">Labs are not better at drawing bicycles</a></li>
<li><a href="https://dylancastillo.co/posts/pelicanmaxxing.html#evidence-4-labs-are-not-better-at-drawing-pelicans-on-bicycles-even-adjusting-for-difficulty">Labs are not better at drawing pelicans on bicycles, even adjusting for difficulty</a></li>
<li><a href="https://dylancastillo.co/posts/pelicanmaxxing.html#evidence-5-the-pelican-bicycle-scenes-dont-look-memorized">The pelican-bicycle scenes don’t look memorized</a> [...]</li>
</ul>
<p>Pelicans aren’t drawn any better than other animals. Bicycles aren’t drawn any better than other vehicles. And no lab draws the combination better than its pelicans and bicycles already predict. GLM-5.2 comes closest: it has the largest boost on the exact pelican-bicycle cell, and and its first pelican-on-bicycle sample caught my eye. But the effect is small and not significant, so I wouldn’t put too much weight on it.</p>
</blockquote>
<p><small></small>Via <a href="https://news.ycombinator.com/item?id=49010129">Hacker News</a></small></p>
<p>Tags: <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/evals">evals</a>, <a href="https://simonwillison.net/tags/pelican-riding-a-bicycle">pelican-riding-a-bicycle</a></p>
</summary><category term="ai"/><category term="generative-ai"/><category term="llms"/><category term="evals"/><category term="pelican-riding-a-bicycle"/></entry><entry><title>Orchestrions</title><link href="https://simonwillison.net/2026/Jul/22/all-the-orchestrions/#atom-everything" rel="alternate"/><published>2026-07-22T14:48:52+00:00</published><updated>2026-07-22T14:48:52+00:00</updated><id>https://simonwillison.net/2026/Jul/22/all-the-orchestrions/#atom-everything</id><summary type="html">
<p>San Francisco tip: it only costs around $15 ($10 in quarters plus a $5 bill for the self-playing violin) to activate every single Orchestrion in <a href="https://en.wikipedia.org/wiki/Musée_Mécanique">Musée Mécanique</a>.</p>
<p>And because most people are bad at allocating their funds you may well be the ONLY person activating the Orchestrions, which means you get to craft the soundscape for the entire museum.</p>
<p>Tags: <a href="https://simonwillison.net/tags/san-francisco">san-francisco</a></p>
</summary><category term="san-francisco"/></entry><entry><title>California Sea Lion</title><link href="https://simonwillison.net/2026/Jul/21/sighting-383713864/#atom-everything" rel="alternate"/><published>2026-07-21T19:51:03+00:00</published><updated>2026-07-21T19:51:03+00:00</updated><id>https://simonwillison.net/2026/Jul/21/sighting-383713864/#atom-everything</id><summary type="html">
<p><img src="https://static.inaturalist.org/photos/702321069/large.jpg" alt="California Sea Lion"></p><p><img src="https://static.inaturalist.org/photos/702321114/large.jpg" alt="California Sea Lion"></p><p>California Sea Lion, in San Francisco County, US, CA</p><p>We took some visiting family to Pier 39 to see the sea lions. They're somehow always even more fun than I remember them being last time.</p>
<p>Tags: <a href="https://simonwillison.net/tags/san-francisco">san-francisco</a>, <a href="https://simonwillison.net/tags/wildlife">wildlife</a></p>
</summary><category term="san-francisco"/><category term="wildlife"/></entry><entry><title>Nativ: Run AI models locally on your Mac</title><link href="https://simonwillison.net/2026/Jul/21/nativ/#atom-everything" rel="alternate"/><published>2026-07-21T14:22:27+00:00</published><updated>2026-07-21T14:22:27+00:00</updated><id>https://simonwillison.net/2026/Jul/21/nativ/#atom-everything</id><summary type="html">
<p><strong><a href="https://blaizzy.github.io/nativ/">Nativ: Run AI models locally on your Mac</a></strong></p>
Prince Canuma is the developer behind the excellent <a href="https://github.com/Blaizzy/mlx-vlm">MLX-VLM</a> Python library for running vision-LLMs using MLX on a Mac.</p>
<p>I'm really excited about his new project, which wraps MLX in a full macOS desktop application. It's similar in shape to LM Studio, providing both a chat interface and a localhost API server for accessing models.</p>
<p>The app picked up MLX models I had already tried that were present in my Hugging Face cache directory, which was a nice touch.
<p><small></small>Via <a href="https://news.ycombinator.com/item?id=48982681">Hacker News</a></small></p>
<p>Tags: <a href="https://simonwillison.net/tags/macos">macos</a>, <a href="https://simonwillison.net/tags/python">python</a>, <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/local-llms">local-llms</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/mlx">mlx</a>, <a href="https://simonwillison.net/tags/prince-canuma">prince-canuma</a></p>
</summary><category term="macos"/><category term="python"/><category term="ai"/><category term="generative-ai"/><category term="local-llms"/><category term="llms"/><category term="mlx"/><category term="prince-canuma"/></entry><entry><title>A Fireside Chat with Cat and Thariq from the Claude Code team</title><link href="https://simonwillison.net/2026/Jul/21/cat-and-thariq/#atom-everything" rel="alternate"/><published>2026-07-21T12:54:02+00:00</published><updated>2026-07-21T12:54:02+00:00</updated><id>https://simonwillison.net/2026/Jul/21/cat-and-thariq/#atom-everything</id><summary type="html">
<p>Earlier this month I hosted a fireside chat session at the <a href="https://www.ai.engineer/worldsfair/2026">AI Engineer World's Fair</a> with Cat Wu and Thariq Shihipar from Anthropic's Claude Code team. We talked about Claude Code, Claude Tag, Fable, coding agent security, evals, tool design, and how Anthropic use these tools themselves.</p>
<p>The full video of the session is now available <a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g">on YouTube</a>. Below is an edited copy of the transcript, with extra links and my own bolded highlights.</p>
<iframe style="margin-top: 0.5em; margin-bottom: 1em;" width="560" height="315" src="https://www.youtube-nocookie.com/embed/uU5Gv2h8-9g" title="SimonThis Year in Claude" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen="allowfullscreen"> </iframe>
<p>A few top-level notes if you don't want to watch the video or wade through the whole transcript:</p>
<ul>
<li>Claude Tag (Claude's new collaborative Slack integration) now lands <strong>65% of the product engineering PRs</strong> for the Claude Code team.</li>
<li>Claude Code ships features to Anthropic employees first, and <strong>only ships the features that demonstrate user retention with that cohort</strong>
</li>
<li>Critical changes to Claude Code are still reviewed manually, but the team increasingly relies on automated code review for the "outer layers" of the product.</li>
<li>Adding examples to a system prompt is <strong>no longer best practice</strong> for models like Fable 5 or even Opus 4.8. The Claude Code system prompt recently <strong>reduced in size by 80%</strong>.</li>
<li>Likewise, lists of "<strong>don't do X and don't do Y</strong>" can reduce the quality of results from the latest models.</li>
<li>
<a href="https://en.wikipedia.org/wiki/Eating_your_own_dog_food">Dogfooding</a> inside Anthropic is called "<strong>ant fooding</strong>".</li>
<li>Anthropic <strong>really believe in their <a href="https://code.claude.com/docs/en/auto-mode-config">auto mode</a></strong>, and see that as an enabling technology for Claude Tag.</li>
<li>Thariq advises offsetting coding-agent-induced <a href="https://simonwillison.net/2026/Feb/15/deep-blue/">Deep Blue</a> by "<strong>being more ambitious</strong>" with the work you take on.</li>
<li>Fable is <strong>competent at editing video</strong>, and Thariq <a href="https://twitter.com/trq212/status/2064826394589442448">used it</a> to edit its own launch video.</li>
<li>Anthropic's culture of working (internally) in public is key to their success, as demonstrated by the way they use Claude Tag in their public Slack Channels.</li>
</ul>
<h4 id="how-has-what-you-do-day-to-day-changed-in-the-past-year-">How has what you do day-to-day changed in the past year?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=65s">1:05</a></p>
<blockquote>
<p><strong>Simon:</strong> Claude Code came out in February of last year — it's under a year and a half old, and it was originally just a bullet point on <a href="https://www.anthropic.com/news/claude-3-7-sonnet">the Claude Sonnet 3.7 launch</a>. <strong>How has what you do on a day-to-day basis changed in the past year</strong>, now that we have these coding agents that actually work for us?</p>
<p><strong>Cat:</strong> I remember when we first came out with Claude Code and Sonnet 3.7, you would give it a task and you would have to closely monitor every single little thing it tried to do. I would read every permission prompt extremely carefully. I would frequently say no — no, no, no, did you check this file? Did you check that file? And now it's been incredible with every model generation. I feel like <strong>we've all gotten a chance to take a step back and delegate a lot more of the menial implementation to Claude</strong>. It's freed up a lot of our time to think about more creative work, like: what is the right experience that we should be providing to our users, now that we know Claude Code can implement a lot of it? And now with Fable it's a totally different step change improvement. <strong>We see for a lot of our use cases that you can actually one-shot a ton of features with Fable now</strong>.</p>
<p><strong>Thariq:</strong> I remember the first text I got about Claude Code. One of my best friends was like, "You need to go try Claude Code." It was about when Opus 4 came out, and I tried it and I was like, "Oh, shit. I need to work at Anthropic now." And that was Opus 4 — great model, but you were reading permission prompts. It's kind of crazy how much amnesia we have, where I'm like, oh, auto mode has always been here, right? I don't even remember pressing yes and allow. For me, the big thing I'm trying to push myself on is that <strong>we have to do higher quality work than we've ever done before</strong>. The outputs are incredibly high quality. <strong>I've been using it to edit videos a bunch</strong>, and I'm like, okay, it has to meet the very exacting demands of our brand team in a couple of hours or we just can't do it. <strong>That's how I'm trying to shift with Fable: the best work we've ever done, faster than we've ever done it before</strong>.</p>
</blockquote>
<h4 id="what-piece-of-conventional-software-engineering-no-longer-holds-">What piece of conventional software engineering no longer holds?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=219s">3:39</a></p>
<blockquote>
<p><strong>Simon:</strong> What's a piece of conventional software engineering that was true a year ago that you don't think holds anymore in this new world?</p>
<p><strong>Cat:</strong> One of the biggest shifts we're seeing in the eng skill set: two years ago it was pretty typical for a product manager to go talk to a bunch of customers, align over the course of six months with cross-functional teams on some PRD, and write a thorough spec on exactly how we'll implement this before the first line of code gets written. Now things are completely turned the opposite way. For a lot of engineers, the push I would give to folks in the room is to <strong>develop more of your business sense and product sense on what it is we should build</strong>, because the timeline between having an idea and building it is so much shorter — it's down from six to twelve months to maybe even a week. That means all of us need to have better taste on what is worth building, what will actually inflect the businesses we're working on. So it's <strong>an increase in value on product taste and business sense</strong>, and a bit lower on execution in most product domains. Of course, for infra there's still a very heavy emphasis on making sure all the details are right.</p>
<p><strong>Thariq:</strong> For me, it's that <strong>rewrites are now good</strong>.</p>
<p><strong>Simon:</strong> The worst thing you could do is now actually fine!</p>
<p><strong>Thariq:</strong> Exactly. All the Mythical Man-Month stuff — never rewrite — I'm pro-rewriting now. If you have a good test suite — and <strong>I think the rewrite actually forces you to make sure you have a good test suite</strong> — but I think what people undercount is that <strong>a codebase is a spec, and maybe it's the only copy of the spec that you have</strong>, because no one knows every branching part of the codebase. You can take this as an artifact and distill it or create other versions of it. We <a href="https://bun.com/blog/bun-in-rust">rewrote Bun in Rust</a> and it works great — it's live for me right now.</p>
<p><strong>Simon:</strong> You're not shipping Claude Code on Bun-in-Rust yet, right?</p>
<p><strong>Thariq:</strong> Internally we have.</p>
</blockquote>
<p><em>(Actually it looks like Anthropic started shipping Claude Code on Bun-in-Rust to everyone <a href="https://simonwillison.net/2026/Jul/19/claude-code-in-bun-in-rust/">on June 17th</a>.)</em></p>
<h4 id="what-kind-of-things-are-non-engineers-doing-with-claude-tag-">What kind of things are non-engineers doing with Claude Tag?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=396s">6:36</a></p>
<blockquote>
<p><strong>Simon:</strong> The other big launch recently was <strong><a href="https://www.anthropic.com/news/introducing-claude-tag">Claude Tag</a></strong> — that's what, a week old now, at least for the rest of us. I understand it's being used at Anthropic by non-engineers a great deal. <strong>What kind of things are non-engineers doing with Claude Tag?</strong></p>
<p><strong>Cat:</strong> Claude Tag is a Claude that lives in your team's collaboration tools. We launched it last week within Slack. <strong>The thing that's different about Claude Tag is it's multiplayer by default</strong>. Once you add Claude Tag to a Slack channel, you can chime in, your teammates can chime in, and you can collaborate together on the PR. The other big difference is that it's proactive instead of reactive. You can tell Claude Tag, "Hey, monitor every bug report in this channel, put up a PR to fix it, and tag the engineer who last touched this part of the codebase," and it'll do it for the lifetime of the channel without you having to manually tag it in. And the third big shift is that <strong>we've <a href="https://claude.com/docs/claude-tag/users/memory">added team memory</a> into this</strong>. If you tell Claude Tag your preferences in the channel, it'll remember them for every future post. If you always want it to debug outages but you don't want it to debug warnings, just tell it that in natural language in the channel and it'll remember it for you and everyone else on your team.</p>
<p><strong>Internally, we see Claude Tag as the evolution of Claude Code.</strong> We see this as a large shift in how we work internally. <strong>Claude Tag currently lands 65% of our product eng PRs.</strong></p>
<p><strong>Simon:</strong> For all of Anthropic, or just for Claude Code?</p>
<p><strong>Cat:</strong> This is just for our product engineering team — <strong>our internal version of Claude Tag lands 65% of our product PRs right now</strong>. And this is a huge shift; this is more than 50% of our PRs. The way we see people split work between Claude Code and Claude Tag is: Claude Code is still the best place for your most complex tasks, when you're interactively iterating with the agent. <strong>But Claude Tag is great for having it work proactively on your behalf</strong>, so you no longer need to manually kick off Claude Code for all the bug reports that come up for features you're working on.</p>
<p><strong>Thariq:</strong> And for non-coding cases: for example, before this talk we asked Claude Tag, "Hey, when is Fable releasing?" We wanted to make sure we'd line it up with the announcement. Claude Tag would search our Slack and look at who's been saying what. <strong>As a search engine for your company, it's really valuable.</strong> It has all the context for your product, so you can ask it metrics-related questions — often when you're making decisions you want them informed by what the metrics say, so you hook it up to your event store. I've seen our marketing team do things like, "Hey, tell me about this feature." They're not programmers, but Claude is a programmer — it can clone the codebase and say, "This is the feature, this is what it looks like, <strong>this is a recording of me using the feature</strong>." It enables a whole wide variety of things, and I think we're still early in figuring that out.</p>
</blockquote>
<h4 id="claude-tag-as-the-team-collaborative-layer">Claude Tag as the team collaborative layer</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=606s">10:06</a></p>
<blockquote>
<p><strong>Simon:</strong> One of the problems I've had with coding agents is that I get how to use them as an individual, but I'm not really clear on how to use them in a team environment. <strong>It sounds like Claude Tag is your current answer to that team collaborative layer for this stuff.</strong></p>
<p><strong>Cat:</strong> Exactly. And a large percentage of our sessions are actually multiplayer right now. Maybe I say, "Hey, I think we should implement this new feature in Cowork," and I'll tag in Claude Tag to do a first pass at it. Then I'll tell Claude Tag, "Share a recording of your final implementation," and I'll tag in design to take a look. They'll nudge it, then pass it on to eng to take it to the finish line and get it out to prod. It's been this very fluid experience. <strong>We're still trying to iron out what the social dynamics are for steering the same session</strong>, but we've found that people just observe how others use it and follow those social norms — it's been pretty intuitive for us to integrate Claude Tag into our teams.</p>
<p><strong>Thariq:</strong> It's great for teaching people, and also for reducing slop, because <strong>the fact that everyone is seeing you use Claude together sort of levels up how you use Claude as well</strong>.</p>
</blockquote>
<p>This reminded me of how Midjourney solved the challenge of teaching people advanced image prompting by enforcing prompting in public in their Discord channels.</p>
<h4 id="how-do-you-decide-which-features-are-worth-building-when-building-is-so-much-cheaper-">How do you decide which features are worth building when building is so much cheaper?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=701s">11:41</a></p>
<p>Something I've found really hard myself is knowing when a feature is worth shipping now that the cost of actually building features has dropped so much.</p>
<blockquote>
<p><strong>Simon:</strong> How do you deal with the hardest problem in all of engineering — prioritization? <strong>How do you decide which features are worth building and shipping when building a feature is so much more inexpensive now?</strong></p>
<p><strong>Cat:</strong> This is the hard thing. There are a few ways we approach it. One is we dogfood our products every single day. Whenever there's something we want to be able to do in our products that we're not able to, instead of finding a different solution we fix our product so it can support that case. <strong>We have a very heavy dogfooding culture internally.</strong> Before we share our products with everyone in the world, we share them with everyone within Anthropic, and with some early customers who give us very honest feedback about it — the more brutal the better — and we iterate until people love it. <strong>We have an internal bar for the number of active users and the amount of retention a feature has to have before we share it with the world.</strong> Because this bar is very clear, every engineer knows what they're trying to hit. I think this also levels up our polish, because if the feature isn't polished, people will churn — and then we shouldn't ship that feature.</p>
</blockquote>
<p>Using internal user-retention to decide if a feature should ship makes a whole lot of sense to me.</p>
<h4 id="do-you-have-an-example-of-a-feature-which-surprised-you-">Do you have an example of a feature which surprised you?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=774s">12:54</a></p>
<blockquote>
<p><strong>Simon:</strong> <strong>Do you have an example of a feature which surprised you?</strong> You rolled it out and the engagement was off the charts — something unlikely to be shipped that turned into a real product thing.</p>
<p><strong>Cat:</strong> I do have one. <strong>A lot of folks on our team love <a href="https://code.claude.com/docs/en/remote-control">remote control</a>.</strong> Remote control lets you use your mobile device, or Claude in the web browser, to connect to a local Claude Code session running in your CLI. I never have this need, because I just kick off the task directly on mobile and it runs in a cloud session without using my local environment — I think because I'm doing very easy coding tasks. It was something I didn't totally understand; I was like, hey, people should just set up remote dev environments. But in practice, once we rolled out remote control, so many people I talk to told me that what they do every night is plug their laptop into a power charger, open a bunch of remote control sessions, lock the screen, <strong>and then use their mobile phone from their couch to control Claude Code</strong>. So this has become a flow we're now leaning into that I didn't originally get — but now I do.</p>
</blockquote>
<h4 id="does-a-human-review-every-line-of-production-code-in-claude-code-">Does a human review every line of production code in Claude Code?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=860s">14:20</a></p>
<p>One of the over-arching themes of the conference was review: how much attention to people spend to reviewing code written for them by coding agents. I was very keen to hear the Claude Code team's take on this!</p>
<blockquote>
<p><strong>Simon:</strong> How does code review work? <strong>Does a human being review every line of production code that makes it into Claude Code?</strong> And if not, what are you doing — how do you keep the quality up?</p>
<p><strong>Thariq:</strong> It varies on the task a lot. <strong>For important areas we have code owners.</strong> The system prompt is an example where we have a code owner — you really need to get their approval.</p>
<p><strong>Simon:</strong> So the code owner is directly responsible for the quality of that area of the code.</p>
<p><strong>Thariq:</strong> That's right.</p>
<p><strong>Cat:</strong> And they need to approve any PR that touches it.</p>
<p><strong>Thariq:</strong> We have <a href="https://code.claude.com/docs/en/github-actions">our code review GitHub bot</a> review everything — that goes on every PR, and often it's doing the bulk of the review. Something I've seen on the team is that <strong>for more complex PRs you might make an artifact to explain the PR</strong> so that other people can then review. And we invest a lot into verification, CI/CD, things like that, to make sure that any time anything fails we have a test. We have a really robust environment where Claude can control Claude Code and test it. So there's a multi-pronged approach to code review.</p>
<p><strong>Cat:</strong> In general, <strong>we are trying to move to a world where humans don't need to be in the loop</strong>. For the most critical changes to the core of Claude Code, and the cores of other products, there is always a code owner and they do manually review all the changes. But increasingly, <strong>for the changes at the outer layers, we actually have Claude code review fully review those</strong>. That sounds pretty scary, but we've had a six-plus-month-long process to get here, and <strong>there are baby steps that you take to build up trust with code review</strong>. In the beginning we had human review for everything, and then increasingly we would say, <strong>okay, for code changes that touch these files, code review is catching 100% of the issues there — so we actually don't need a human manually reviewing those</strong>. And when we have incident review, <strong>we look at the PRs that caused the incident and say, okay, how do we update code review to catch that?</strong> — and we take those PRs and <strong>add them to an eval set</strong> to make sure our future changes to code review never regress that metric. Removing humans from the code review loop is a big step forward. It can sound scary, and it's not something you can do overnight, but it is something you can do <strong>through many months of investment in the infrastructure</strong> to give you the confidence that code review is catching everything you care about.</p>
</blockquote>
<p>So the key seems to be constantly iterating on the automated review systems themselves, in order to build trust in them over time.</p>
<h4 id="how-does-a-new-model-affect-your-intuition-for-what-it-can-and-can-t-do-">How does a new model affect your intuition for what it can and can't do?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=1040s">17:20</a></p>
<p>We got <em>deep</em> into evals - another hot topic throughout the wider conference.</p>
<blockquote>
<p><strong>Simon:</strong> I know that Opus 4.8, if I ask it to build me a JSON endpoint that runs a SQL query and outputs JSON, is just going to get it right — that's not something I have to review closely. But then a new model comes along and I don't know how to build trust in Fable quickly, that it's not going to mess things up that Opus didn't. <strong>How does the new model affect your intuition for what it can do and what it can't do?</strong></p>
<p><strong>Cat:</strong> The main reason we're building up this <strong>eval base over time is so that new models can be a drop-in replacement</strong>. When we have a new model, we run the whole eval set and make sure that, for example, Fable is strictly better than Opus 4.8 — and that gives us the confidence to drop it in.</p>
<p><strong>Simon:</strong> Are those model evals for Anthropic as a whole, or Claude Code team-specific?</p>
<p><strong>Cat:</strong> We have both. We have evals on our team, and we run code review across every repo within Anthropic, so we have evals for that. And for things like auto mode, we not only have evals across every user within Anthropic — we've also commissioned multiple external testers to red team it, to create environments with prompt injections and malicious inputs, <strong>and make sure that auto mode doesn't let any of those pass</strong>.</p>
</blockquote>
<h4 id="how-do-you-build-confidence-that-a-system-prompt-tweak-results-in-better-output-">How do you build confidence that a system prompt tweak results in better output?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=1121s">18:41</a></p>
<blockquote>
<p><strong>Simon:</strong> I want to know if the system prompt improvement I made actually improved the product — that's the most basic form of product-specific eval, and I still don't have a great feel for how to do that. <strong>Is that something you're doing such that you have complete confidence that a tweak you've made to the system prompt results in better output?</strong></p>
<p><strong>Cat:</strong> <strong>We don't have complete confidence, but we do a lot to make sure that we don't regress performance.</strong> The starting point is a suite of external evals that we trust, and we complement that with an even larger suite of internal evals that we trust. To start, <strong>we mainly optimize for capability</strong>: given a complete definition of a task and the full codebase, does Claude make the right decisions, fully fix the bugs, and pass all the tests? That's the starting point and the thing we optimize for, because it's most directly what users want. But there are a lot of behaviors that impact how users feel when they work with Claude Code. For example, <strong>people really don't like it when Claude Code says it's time to go to sleep.</strong> Or people really don't like it when it says, "Hey, I finished two out of five parts — do you want me to continue?" Yes, please continue. <strong>So we're building up a set of behavioral evals to catch these.</strong> And as we get user feedback — please be loud with us about your user feedback — we rank the priority issues and go down one by one and build evals for each of them. It's not 100% coverage, but it is a priority for us to increase the coverage.</p>
</blockquote>
<h4 id="how-much-interaction-is-there-between-the-claude-code-team-and-the-model-training-teams-">How much interaction is there between the Claude Code team and the model training teams?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=1221s">20:21</a></p>
<blockquote>
<p><strong>Simon:</strong> <strong>How much interaction is there between the Claude Code team and the teams at Anthropic who are training the models in the first place?</strong> Is that quite a close collaboration?</p>
<p><strong>Cat:</strong> Across Anthropic, we all work quite closely together. We meet often to talk about what we expect the next generation of models to be able to do. Our research team has also been amazing about showing this publicly — we often talk in our blog posts about how <strong>we're targeting ever-increasing longer-horizon work</strong>, and how we train Claude itself to be honest, harmless, and helpful. We also put a lot of effort into making sure it's aligned with your intent, even if your intent is expressed in a fuzzy way. Of course, try your best to be specific about what you want, so Claude has all the context — but even when you're not specific, we teach Claude to make good assumptions. It's been a productive partnership.</p>
</blockquote>
<h4 id="the-system-prompt-has-been-reduced-by-80-what-have-you-been-able-to-drop-">The system prompt has been reduced by 80% — what have you been able to drop?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=1284s">21:24</a></p>
<p>So many useful prompting tips in this section!</p>
<blockquote>
<p><strong>Simon:</strong> Thariq, you <a href="https://www.youtube.com/watch?v=9fubhllmsBU&amp;t=358s">mentioned this morning</a> that the <strong>system prompt for Claude Code has been reduced by 80% because of Claude Fable</strong>. Can you go into a little more detail? <strong>What kind of things have you been able to drop?</strong></p>
<p><strong>Thariq:</strong> It wasn't just Fable — it was Opus 4.8 as well, and going forward, future models. We have different system prompts for different models now. One of the patterns we saw is that we were over-constraining Claude. The initial, maybe Opus 4-ish models wanted a lot of examples, and <strong>removing examples was extremely helpful</strong>, because it was just more creative than the examples we gave it.</p>
<p><strong>Simon:</strong> That's really interesting, because one of the top prompting tips I give people is: give it examples. If that's no longer true, that kind of breaks my prompting model a little bit.</p>
<p><strong>Thariq:</strong> Same here — I was surprised to hear that. I think now it's more about the shape of what you give it — the tools you give to Claude, your system prompt, things like that. The other thing we did is try to give it more context and <strong>fewer "do not do this"</strong> instructions, because that's a very strong impulse for Claude, and especially if it conflicts with user instructions later on, that can be extremely confusing to Claude — "I've got this skill that says this and the system prompt says this." So we try to <strong>have fewer hard constraints, more context, and fewer instructions overall</strong>. It's definitely a science — it took a bunch of evals to build.</p>
<p><strong>Cat:</strong> In general, when you're prompting these models, you should always think: <strong>are there edge cases to the instruction that I'm giving it?</strong> When we went back and reviewed all the instructions in the Claude Code system prompt, <strong>we found a few cases where yes, this statement is 90% true, but there's a real 10% of cases where it's not true</strong>. We didn't want to constrain the model, or confuse it into thinking it should always do this. One good example is verification. Everyone here wants Claude to verify its work, and we had some instructions in the prompt that said: if you make a front-end change, always verify. But there's a limit to it. If it's changing copy from one string to another string, and the user says "just make a quick fix and update the test," maybe you don't want to verify. <strong>So we've adjusted our wording from "always verify, verify, verify" to something like: most of the time when you're doing front-end work you can't fully understand the experience by hitting the backend endpoints, so when you make larger changes to the user experience, please run the app locally.</strong> And in fact, that instruction probably isn't even good either, because <strong>what is a large change?</strong> Maybe it should test small changes too. In general, whenever you give a prompt to the model, <strong>you should think about the ways in which it could be misinterpreted by a well-intentioned human</strong>, in order to better understand how the model might interpret it — and <strong>soften the prompt</strong> so that it's actually 100% accurate, because you're giving this prompt to the model 100% of the time.</p>
<p><strong>Simon:</strong> What's fascinating about that is you're <strong>relying on the model's judgment</strong> — and that's got to be an Opus/Fable-level thing. Models a year ago did not have the level of judgment necessary to decide whether they were going to test a change or not. But that does break down if you're building for a wide range of models and trying to run the cheaper models for cheaper tasks.</p>
<p><strong>Cat:</strong> We actually have <strong>a different system prompt per model now</strong>, for this very reason. It's only our most frontier models that have this 80% token decrease — the older models still have the full system prompt.</p>
<p><strong>Simon:</strong> Do you think Fable and Opus are smart enough to prompt Haiku with more details, because they understand that Haiku has less judgment, less taste?</p>
<p><strong>Cat:</strong> We haven't been able to eval it — we don't have any hard data to show it.</p>
<p><strong>Thariq:</strong> There's a tough thing with smaller models sometimes, because <strong>sometimes the larger models can be more token-efficient on a hard problem than the smaller models</strong>. So there's a bit of intuition to build there — sometimes you really just want frontier intelligence almost all the time. The Pareto curve shifts, and it's hard to find.</p>
<p><strong>Simon:</strong> A year ago I did not trust a model to write a prompt. Today the good models are very good at prompting — a lot of my prompts are written by models, which feels absurd but works really well. What helped me come to terms with that was thinking about subagents, which are entirely about a Claude model setting up a prompt for another Claude model.</p>
<p><strong>Thariq:</strong> <strong>Workflows</strong> are actually a really good example of this, because it's Claude not just prompting a single subagent, but prompting the orchestration of many subagents, and each one of them gets a very detailed prompt. It's almost a level above just spawning a subagent. I've also been using it on my personal machine, <strong>giving it the Gemini API and saying: here, generate images</strong>. It's way less lazy than I am at prompting an image model. It's just Claude prompting Claude all the way down.</p>
<p><strong>Cat:</strong> I think Claude also wrote the prompt for <a href="https://code.claude.com/docs/en/workflows">the workflow tool</a>.</p>
<p><strong>Simon:</strong> I've read that prompt — it's a good prompt. That's actually a frustration I have with Anthropic generally: you <a href="https://platform.claude.com/docs/en/release-notes/system-prompts">publish the prompts for Claude Chat</a>, but you don't include the tool prompts and the Claude Code prompts. I still have to run a proxy to intercept them. <strong>I would love it if the Claude Code prompts were deliberately published</strong> — they're the documentation. They're how you know what the tool can do and how it works.</p>
<p><strong>Cat:</strong> I'll write down that feature request. I'll have Claude Tag do it.</p>
</blockquote>
<p>Interesting to note that OpenAI's <a href="https://developers.openai.com/api/docs/guides/latest-model?model=gpt-5.6#favor-leaner-prompts">prompting best practices for GPT-5.6</a> includes similar advice for their latest models:</p>
<blockquote>
<p><strong>Favor leaner prompts</strong></p>
<p>Removing repeated instructions and examples and simplifying tool descriptions can improve task performance and token efficiency. In a sample of internal coding-agent eval runs, configurations with leaner system prompts improved evaluation scores by roughly 10–15% while reducing total tokens by 41–66% and cost by 33–67%.</p>
</blockquote>
<h4 id="what-s-your-bar-for-introducing-a-new-tool-">What's your bar for introducing a new tool?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=1686s">28:06</a></p>
<blockquote>
<p><strong>Simon:</strong> Claude Code is basically a big bag of tools. <strong>What's your bar for introducing a new tool?</strong> How do you decide when it's worth doing that additional engineering at that level?</p>
<p><strong>Cat:</strong> Do you want to take it? You introduced one of the best tools we have.</p>
<p><strong>Thariq:</strong> My career peaked when I introduced the ask user question tool. It's really hard. Especially for some tools — <strong>ask user question is Claude's tool to ask you</strong> — so it's hard to eval, and sometimes it's more of a user preference thing. Back then we had fewer evals, so it was very dogfooding based — or "ant fooding," our ant version of that. But overall <strong>we've been trying to trend towards fewer tools</strong>. The last set of tools we introduced was the task tool, I think — and we try to give Claude more general versions to do things.</p>
</blockquote>
<h4 id="what-s-the-latest-evolution-of-your-file-editing-tool-">What's the latest evolution of your file editing tool?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=1743s">29:03</a></p>
<p>I have a long-running fascination with file editing tools - they were the subject of the <a href="https://aider.chat/docs/leaderboards/edit.html">old Aider code editing leaderboard</a>, and I've watched with interest as they've evolved in different coding agents from search-and-replace based to line-number-based to more complicated patterns.</p>
<p>The Claude API docs describe a <a href="https://platform.claude.com/docs/en/agents-and-tools/tool-use/text-editor-tool">text editing tool</a> that's recommended for building against the API, but Claude Code seems to use slightly different approaches here.</p>
<blockquote>
<p><strong>Simon:</strong> One of the most interesting tools is the file editing tool — you can have file editing as a tool, or you can tell it to use sed and grep and do things that way. <strong>What's the latest evolution of your file editing tool?</strong></p>
<p><strong>Thariq:</strong> We still have one, but for example we removed our grep and other search tools — glob tools — in favor of native bash. Like I said in my talk earlier, <strong>the models are kind of more of a biology than a physics</strong>, and tool design especially is quite hard. I'm not sure if Cat disagrees and thinks there's a science to the eval of it, but I think tool design is more of an art, maybe — or a biology.</p>
<p><strong>Cat:</strong> I largely agree, but in general as we introduce more tools, we try to keep the cardinality pretty low and make sure that <strong>every tool we add has a distinct function from every other tool, so that Claude can very easily distinguish when to call each</strong>. For file edit, the reason we have it is actually because we can render it. We show people when Claude makes a file change, and there's this <strong>nice dedicated UI</strong> that says: do you approve this edit to this file? <strong>The reason we had a dedicated file edit tool was so that we could deterministically know</strong> that Claude was making a file change, so we could show people this nice UI. A lot of new users onboarding still really like this experience, so we've kept it around. But for a lot of us who are on auto mode right now — hopefully you're not on YOLO mode — I don't think it actually matters, and we could probably just remove file edit and be totally fine.</p>
</blockquote>
<h4 id="what-s-the-advice-within-anthropic-for-safely-running-claude-code-">What's the advice within Anthropic for safely running Claude Code?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=1858s">30:58</a></p>
<p>It's the <a href="https://simonwillison.net/tags/prompt-injection/">prompt injection</a> question! Who better than Anthropic employees to explain how Anthropic sees the risk of prompt injection attacks causing their Claude Code instances to run amok?</p>
<p>It turns out they <em>really</em> trust their <a href="https://code.claude.com/docs/en/auto-mode-config">auto mode</a> - and see that as the feature that enabled Claude Tag.</p>
<blockquote>
<p><strong>Simon:</strong> Let's talk about safety and security. I am deeply aware of the risks of prompt injection, and there are so many bad things that can happen if somebody else tells my Claude Code what to do. I still mostly run Claude Code in YOLO mode and feel incredibly guilty about it. <strong>What's the advice within Anthropic for safely running Claude Code?</strong></p>
<p><strong>Cat:</strong> Why not auto mode?</p>
<p><strong>Simon:</strong> I am starting to use auto mode, but I don't understand it enough to get how safe it is. As of maybe three weeks ago, I'm defaulting to auto mode.</p>
<p><strong>Cat:</strong> Broadly within Anthropic, almost every single person uses auto mode. It is the best way to do long-running work in Claude Code while being safe. <strong>We've done extensive bashing. We have thousands of evals. We've commissioned many red teamers to create adversarial environments in order to trick Claude Code into doing bad actions, and we've mitigated every single issue that they found.</strong> We're going to publish some evals in the coming weeks, but we've pretty much mitigated every attack.</p>
<p><strong>Simon:</strong> That is a big claim.</p>
<p><strong>Cat:</strong> We'll share the evals for it so folks can assess, but we've been extremely diligent about identifying all the ways in which Claude might mess up and then updating auto mode to counter it. It doesn't catch 100% of things — that would be way too strong a claim. But <strong>for the main categories of risks that we're concerned about, like prompt injection and data exfiltration, the risks are far lower than the average human reviewer</strong>.</p>
</blockquote>
<p>I am very much looking forward to learning more about their evals and approach to verifying auto mode.</p>
<blockquote>
<p><strong>Thariq:</strong> A little on how auto mode works — it's useful to build this mental model. Whenever Claude is doing a turn, or a bash call, there's <strong>a Sonnet classifier</strong> that is judging the tool call and also the context of the conversation — your instruction. There are some things around permissions that are dependent on your request: you don't want to give git push permissions all the time, but if you say "push this to GitHub," you want it to do it — and if you say "don't push," you want it to deny it. Auto mode will do that. That particular thing happens to me a lot, where Claude tried to do something because it's very helpful and proactive, and auto mode saw "don't do this" and surfaced it. <strong>So it's good at the dynamic permissions</strong> that you yourself give inside the prompt, which I think is really important. It also works well with our <a href="https://code.claude.com/docs/en/sandbox-environments#sandboxed-bash-tool">sandboxing infrastructure</a>, because sandboxing is one of those things where there are so many different edge cases that it's hard for us to deterministically follow them. <strong>We have a sandbox, and when something needs to escape the sandbox</strong> — like a network request — auto mode can look at that request and ask: does this make sense? — and allow it.</p>
<p><strong>Simon:</strong> I hadn't realized auto mode is interacting with the networking sandbox as well.</p>
<p><strong>Cat:</strong> It interacts with any permission prompt the user would otherwise see.</p>
<p><strong>Simon:</strong> How old is auto mode? As a feature I had access to, it's only a couple of months old, right?</p>
</blockquote>
<p>(It was first made available to the public <a href="https://claude.com/blog/auto-mode">on March 24th</a>.)</p>
<blockquote>
<p><strong>Cat:</strong> We've been using it within Anthropic <strong>since January</strong>, so we've been hardening it for quite a while. Anthropic is extremely focused on safety and security, and we've been working broadly across our alignment and safeguards teams to enable the rollout internally, build out these evals, and make auto mode even more robust before sharing it with the world.</p>
<p><strong>Thariq:</strong> This is also the reason Claude Tag is so good — <strong>Claude Tag uses auto mode</strong>. I've heard a lot of build-versus-buy questions about a Slackbot, and I'm like: please, you probably shouldn't build your own AI Slackbot. There are so many attack vectors. <strong>You have a feedback channel that users can post feedback into, and now your bot is reading it.</strong> The work we've put in with auto mode — and we have a general <strong>Swiss cheese defense</strong> for security; we also RL against this stuff — <strong>I think this is really what makes Claude Tag work</strong>. It works seamlessly with your permissions, and you don't want to be prompt injected in your Slack.</p>
</blockquote>
<h4 id="are-there-more-security-things-in-the-pipeline-beyond-auto-mode-">Are there more security things in the pipeline beyond auto mode?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=2154s">35:54</a></p>
<blockquote>
<p><strong>Simon:</strong> Are there any more security things in the pipeline that go beyond auto mode?</p>
<p><strong>Thariq:</strong> I think we're very secure. <strong>With Claude Tag you can provision your own credentials for Claude</strong>, so it doesn't need to act on your behalf — you can have Claude as an identity, and that also makes it easier to audit and inspect what Claude is doing.</p>
<p><strong>Simon:</strong> Because Claude Tag is influenced by anyone who can talk to it — it's got a much wider pool of people telling it what to do.</p>
<p><strong>Thariq:</strong> That's right. And of course we have probes as well with Fable, which is a downstream effect of our safety and research work. I think this is the moment where you see Anthropic being an AI safety company really paying off: <strong>we really want Claude to be able to run in an aligned way over long periods of time</strong>, and <strong>auto mode has to be basically flawless for this to work</strong> — it's all downstream of our being an AI safety company.</p>
<p><strong>Cat:</strong> We also launched trusted devices for the remote control users out there who want to be safer. And for all of our remote environments, we support <strong>credential injection</strong>. If you want Claude Code to be able to access Datadog, but you don't want Claude Code itself to hold the Datadog credential, you can set up our identity and credential management system <strong>so that the Datadog credentials are only usable by the agent but not accessible by the agent</strong> — we insert them on the fly when the agent tries to make a Datadog request.</p>
</blockquote>
<p>I really like that credential injection pattern, where Claude Code can access an API via a proxy and that proxy both audits the request and injects the relevant API key - so Claude can access authenticated endpoints without having access to the API credentials itself.</p>
<h4 id="how-has-the-past-year-and-a-half-changed-how-you-think-about-your-own-craft-">How has the past year and a half changed how you think about your own craft?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=2273s">37:53</a></p>
<p>Thariq <a href="https://www.youtube.com/watch?v=9fubhllmsBU&amp;t=867s">talked about a sense of grief</a> brought on by Fable-class models in his keynote in the morning, and we dived further into that as part of our conversation. I've been calling this <a href="https://simonwillison.net/2026/Feb/15/deep-blue/">Deep Blue</a>.</p>
<blockquote>
<p><strong>Simon:</strong> <strong>Let's talk a little bit about the human element.</strong> <strong>A lot of people are feeling a sense of loss now that so much of what they considered to be their role in building software is being subsumed by the models.</strong> How do you think about that? <strong>How has the past year and a half changed the way you think about your own craft and the value that you add?</strong></p>
<p><strong>Thariq:</strong> Cat and Boris are such good reminders that you have to be more ambitious. They're always like: we're growing so fast, we have to be on the edge, we have to do the best work we can. That's a constant reminder for me — any time I'm slow on something, I'm like, okay, can I do it faster? Can I be more ambitious here? And oftentimes the answer is Claude, because Claude is getting better as you go — the last time I tried this, it was with the previous model. On your point about loss: I think this is real. <strong>If you're only trying to do the same work you were doing before LLMs, and now it's a prompt, it is, I think, kind of a sad feeling.</strong> And <strong>the way you offset that is by being more ambitious.</strong> I think Jared is such a good example — he hand-wrote all of the Zig code in his Oakland apartment in about a year, barely left his house, and had so much fun doing that. Now I see him rewrite all of Bun into Rust and <strong>he's having so much fun doing that</strong> — it's so much more ambitious, and that's how he offsets it. Generally it's asking <strong>how do I do the bigger thing</strong> and do more — <strong>I think success is fun</strong>. It's changing your ambition.</p>
</blockquote>
<p>"The way you offset that is by being more ambitious" neatly captures where I've landed on this issue myself as well.</p>
<blockquote>
<p><strong>Simon:</strong> And Cat, what does that look like from a product management perspective?</p>
<p><strong>Cat:</strong> I feel like the product role just changes every single month. <strong>All the PMs on our team are this mix of engineer, designer, PM</strong> — most of them actually used to be full-time engineers. For us it really means <strong>plugging in whenever there's any kind of gap</strong>. If we have an idea and we didn't inspire any engineer to go build it, then we should just build it, put it into a notebook, and inspire people to take it to production. If the designs look a little off, <strong>let's take a page that's similar, do a first-pass design, and tag in someone who's very detail-oriented to fill in the gaps</strong>. Or if we notice that our team and product adoption is bigger within the company, and more people need to know what's coming down the pipe for Claude Code, Claude Tag, and Cowork — let's automate figuring out our whole launch calendar, <strong>let's automate getting those status updates asynchronously</strong> so we're not bugging people, and make sure our updates in our internal announce channels are fully detailed and to the point. For us it's very much understanding <strong>what the gap is right now between a great idea and getting something to our customers</strong>, and <strong>how do we automate it as much as possible</strong>.</p>
</blockquote>
<p>This reflects something I've noticed: when you can produce code so much faster, time spent blocked awaiting a decision from someone else becomes a much more notable bottleneck. Engineers who can make product decisions can move a whole lot faster, and the cost of getting one of those decisions wrong is much less prohibitive.</p>
<h4 id="what-s-a-moment-when-claude-has-surprised-you-">What's a moment when Claude has surprised you?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=2510s">41:50</a></p>
<blockquote>
<p><strong>Simon:</strong> <strong>What's a moment when Claude has surprised you?</strong> When the model did something you didn't think it would be able to do?</p>
<p><strong>Thariq:</strong> I've posted a lot about Claude video editing, but most recently I gave a talk at the ACM Agentic conference, and I asked, "Hey guys, do you have the edited video? I'd love to post it and share it with my comms team." They said, "Oh, it's taking so long." So I asked for the raw files. They sent me the video of me talking on stage, the video of the deck, and the audio file, and said, "Good luck." I gave this to Claude, along with my HTML deck, and said, "<strong>Hey, can you just edit this together?</strong>" And what it does is honestly incredible — I'm ready to ship it. It transcribes the entire video. It notices that sometimes the video of my deck is a little weird — there's a popup of an auto-update in the middle — and it goes, "<strong>Oh, I probably shouldn't use the video of your deck. What I'm going to do is slice it up, figure out which slide you're on, and use the HTML source instead.</strong>" So it displays the HTML source. Then it's got video of me, but I'm only taking up a small part of the stage, so <strong>it's cropping dynamically to where I am on the stage</strong> — and I'm pacing, so it's tracking me as I pace. And it's transcribing what I'm saying.</p>
<p><strong>Simon:</strong> This was Fable, right?</p>
<p><strong>Thariq:</strong> This was Fable, yeah. It was a good prompt, but it was a one-shot prompt. Then I asked it to add some interesting animations and graphics, and I was just blown away. <strong>It does ffmpeg, it does Remotion.</strong></p>
</blockquote>
<p>Here's Thariq's video <a href="https://twitter.com/trq212/status/2064826394589442448">on how he used Fable to edit Fable's own launch video</a>, and here's <a href="https://twitter.com/ClaudeDevs/status/2064399512664526853">that launch video</a>.</p>
<h4 id="what-can-t-it-do-yet-">What can't it do yet?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=2616s">43:36</a></p>
<p>I'm embarrased to admit that I've been finding it quite hard to come up with tasks that frontier models like Fable 5 and GPT-5.6 are unable to accomplish.</p>
<p>Cat still doesn't rate its UX design skills:</p>
<blockquote>
<p><strong>Simon:</strong> What can't it do? What are the things where you're still disappointed — where you're waiting for Claude Fable 6 to figure it out for you?</p>
<p><strong>Cat:</strong> I want it to have better design and UX taste. It's now at the point where if I write out a prompt with a detailed spec of how I want a feature to behave, it will usually behave that way. But the paddings might be off, or the interface just isn't delightful yet. It leans on existing best practices for how apps are designed, but <strong>for frontier AI products, there are so many new interaction experiences that we have yet to design</strong>.</p>
<p><strong>Simon:</strong> There's an Opus aesthetic — you can look at something and go, "Yeah, that was designed by Opus." It'd be good if we could move beyond that.</p>
<p><strong>Cat:</strong> Yeah. I'm very excited for future models to hopefully be <strong>interaction design thought partners</strong>.</p>
<p><strong>Thariq:</strong> What can't it do? I would love to see it interact more with the real world. Can it solve science? Can it orchestrate the experiments? There's some amount of coding that goes into that, but there's also this other taste of the broader world that it needs.</p>
</blockquote>
<h4 id="which-parts-of-anthropic-s-culture-should-other-companies-steal-">Which parts of Anthropic's culture should other companies steal?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=2711s">45:11</a></p>
<p>I figured this would make a great closing question:</p>
<blockquote>
<p><strong>Simon:</strong> <strong>Which parts of Anthropic's company culture do you think uniquely help Anthropic be productive with these tools, that other companies should steal?</strong> What are the cultural hacks people should be adopting from you?</p>
<p><strong>Cat:</strong> I'll share one for Claude Tag. <strong>Claude Tag works best when you have it in a public channel, and when most of your channels are public.</strong> Claude Tag is able to search across all public channels to get as much context as possible to give you the highest-accuracy answer — and <strong>it's only able to do this if it has access to everything</strong>.</p>
<p><strong>Thariq:</strong> I mentioned this in my keynote, but it's so important to me I want to re-emphasize it. The co-founders <strong>say we don't negotiate against ourselves</strong>, and I think this is really important. <strong>You can imagine trade-offs in your head and talk yourself out of doing something ambitious — or you can just try to do the ambitious thing.</strong> We're so often asking: what if we just did it? Is this a real trade-off or not? And if so, why — where's the proof that it's a real trade-off, and not just something that sounds reasonable? <strong>Make the trade-offs show themselves to you. Be as ambitious as you can.</strong></p>
</blockquote>
<h4 id="what-s-your-favorite-absurd-thing-you-ve-built-with-claude-just-because-you-could-">What's your favorite absurd thing you've built with Claude, just because you could?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=2806s">46:46</a></p>
<p>I couldn't resist throwing in this one as well.</p>
<blockquote>
<p><strong>Simon:</strong> <strong>What's one of your favorite absurd things that you've built with Claude, just because you could build it?</strong></p>
<p><strong>Thariq:</strong> I'm working on <strong>a 2D Street Fighter fighting game with me as a character</strong> — and my friends as well. It uses Claude Code to prompt Gemini — and honestly the Seedance model is pretty good — to make video animations. It works great; it's so good at prompting, and it can verify the frames to check whether an animation was good.</p>
<p><strong>Simon:</strong> Is this Street Fighter 2-level 2D sprites you're generating?</p>
<p><strong>Thariq:</strong> Yeah, exactly — 2D sprites. The animation looks amazing. And it can also figure out hitboxes — it can be like, "Oh, your fist is here, I'll draw the JSON hitbox." It's incredible.</p>
<p><strong>Cat:</strong> Mine is much more simple. I'm a big rock climber and a lot of my friends climb, so we have this little app we built with Claude Code where we log all the projects we're working on. We also go outdoors together a lot, so we have Claude do all this research with workflows. Workflows is amazing — we brand it as a coding tool, but it's amazing for doing deep research for travel. I also plan our team offsites, and it's good at finding venues that can fit all of us. I use workflows to research all the climbing destinations we might want to go to, and what has direct flights from where all of us are located. It goes to Mountain Project and finds all the climbs at our grade level. It finds the Airbnb. And I don't like hiking, so I care a lot about it having a very short approach — <strong>very short walking distance from where the car parks to where the rock actually is</strong> — and it filters for this. With existing apps I have to manually click through Mountain Project, but with this I just put in all of our preferences and it's a custom app for us.</p>
<p><strong>Simon:</strong> So you're basically vibe coding Jira for mountain climbing.</p>
<p><strong>Cat:</strong> Exactly.</p>
</blockquote>
<h4 id="audience-any-plans-for-eval-building-tools-and-agent-observability-">Audience: Any plans for eval-building tools and agent observability?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=2963s">49:23</a></p>
<p>We had a few minutes at the end for questions from the audience.</p>
<blockquote>
<p><strong>Audience:</strong> Do you have any near-term plans to build more eval tools for us to build eval datasets, and more observability tools to monitor the performance of agents and workflows?</p>
<p><strong>Cat:</strong> We've considered building eval tools, but I think the limiting factor actually tends to be that <strong>it takes a long time for customers to build really high-quality evals</strong>. So I think the tooling is less of the constraint, and more the skill set of how you build a great eval. That's an area where we're excited to both invest internally and hopefully share some best practices externally.</p>
</blockquote>
<h4 id="audience-how-is-memory-designed-today-and-would-you-move-from-files-to-a-data-store-">Audience: How is memory designed today — and would you move from files to a data store?</h4>
<p><a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;t=3008s">50:08</a></p>
<blockquote>
<p><strong>Audience (Sai):</strong> I'm interested in the memory and the multiplayer. <strong>How is memory being designed today?</strong> I assume it's around files. And second, have you thought about an orthogonal direction where you <strong>would actually need a data store for these memories, instead of files, to scale it better?</strong></p>
<p><strong>Thariq:</strong> Right now for Claude Tag the memory is channel-specific. Every Claude in that channel has a shared memory, and the instances have a session — but the session can contribute back to main memory. We do a lot of memory research, and it can be kind of unintuitive what the right way to do memory is. We're always running memory experiments. <strong>How it works right now in Claude Tag is a markdown file per channel.</strong></p>
</blockquote>
<p>Tags: <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/prompt-engineering">prompt-engineering</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/anthropic">anthropic</a>, <a href="https://simonwillison.net/tags/annotated-talks">annotated-talks</a>, <a href="https://simonwillison.net/tags/coding-agents">coding-agents</a>, <a href="https://simonwillison.net/tags/claude-code">claude-code</a>, <a href="https://simonwillison.net/tags/thariq-shihipar">thariq-shihipar</a>, <a href="https://simonwillison.net/tags/cat-wu">cat-wu</a></p>
</summary><category term="ai"/><category term="prompt-engineering"/><category term="generative-ai"/><category term="llms"/><category term="anthropic"/><category term="annotated-talks"/><category term="coding-agents"/><category term="claude-code"/><category term="thariq-shihipar"/><category term="cat-wu"/></entry><entry><title>Reverse-engineering is cheap now</title><link href="https://simonwillison.net/2026/Jul/20/cheap-reverse-engineering/#atom-everything" rel="alternate"/><published>2026-07-20T19:24:05+00:00</published><updated>2026-07-20T19:24:05+00:00</updated><id>https://simonwillison.net/2026/Jul/20/cheap-reverse-engineering/#atom-everything</id><summary type="html">
<p>I keep hearing anecdotes from people who used coding agents to reverse-engineer and automate devices in their homes.</p>
<p>I think this is an interesting illustration of the impact of the reduced cost of writing code.</p>
<p>Prior to agents, it was entirely possible to reverse-engineer home devices. The problem was the ROI - was it really worth all of that effort? More importantly, any experienced programmer knows that undocumented, unstable APIs like that may well change or break in the future. Is that initial work worth the effort if you're committing yourself to a frustrating cycle of maintenance in the future?</p>
<p>Coding agents change that equation entirely. The effort to get a simple automation working has dropped, as has the cost of trying and failing to get it to work. Since the code is so cheap, the idea of having to maintain it in the future - or throw it away and start again - carries way less psychological baggage.</p>
<p>Tags: <a href="https://simonwillison.net/tags/reverse-engineering">reverse-engineering</a>, <a href="https://simonwillison.net/tags/coding-agents">coding-agents</a>, <a href="https://simonwillison.net/tags/ai-assisted-programming">ai-assisted-programming</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a></p>
</summary><category term="reverse-engineering"/><category term="coding-agents"/><category term="ai-assisted-programming"/><category term="generative-ai"/><category term="ai"/><category term="llms"/></entry><entry><title>Who’s Afraid of Chinese Models?</title><link href="https://simonwillison.net/2026/Jul/20/afraid-of-chinese-models/#atom-everything" rel="alternate"/><published>2026-07-20T17:09:19+00:00</published><updated>2026-07-20T17:09:19+00:00</updated><id>https://simonwillison.net/2026/Jul/20/afraid-of-chinese-models/#atom-everything</id><summary type="html">
<p><strong><a href="https://stratechery.com/2026/whos-afraid-of-chinese-models/">Who’s Afraid of Chinese Models?</a></strong></p>
Interesting proposal from Ben Thompson that both addresses the hypocrisy of labs outlawing distillation against their models despite training on unlicensed data, and could help US open models compete more effectively with their Chinese counterparts:</p>
<blockquote>
<p>The U.S. should pass a law that (1) makes explicit that collecting data for training models is fair use, and (2) bars terms of service that forbid distillation, for U.S. companies at a minimum. Stopping distillation — which is literally just querying the API — is nearly impossible; the U.S. should go the other way and lean into a new copyright policy that both indemnifies the labs and also guarantees that what they learned fuels further innovation for everyone else.</p>
</blockquote>
<p>Ben also theorizes that Alibaba's decision to release Qwen 3.8 Max as open weights - a reversal from their decision <a href="https://qwen.ai/blog?id=qwen3.7">not to release Qwen 3.7 Max</a> in May - may have been influenced by a <a href="http://english.scio.gov.cn/topnews/2026-07/18/content_118605932.html">recent speech</a> by Xi Jinping, who said:</p>
<blockquote>
<p>We should seize this rare, historic opportunity to encourage open source, openness, collaboration and sharing.</p>
</blockquote>
<p>And on the subject of <a href="https://twitter.com/Alibaba_Qwen/status/2078759124914098291">Qwen 3.8 Max</a> - a new 2.4T parameter model (nearly as large as the 2.8T Kimi K3) - here's <a href="https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F735f2cf19b795517cb2ff6cae1c71c64">a pelican it drew</a>:</p>
<p><img alt="Described by Qwen 3.8 Max: Flat vector cartoon illustration of a white pelican with a large orange beak and pouch riding a red bicycle, its orange legs on the pedals, against a light blue sky with a yellow sun top right and a white cloud top left, with horizontal motion lines behind the bike and a pale green ground strip at the bottom." src="https://static.simonwillison.net/static/2026/qwen-3.8-max-pelican.png" /></p>
<p>I particularly enjoyed seeing these notes in the (extensive) reasoning trace: "Could add helmet? No." and "Maybe add small bell? no." and "Need maybe add small fish in basket? Not necessary."
<p><small></small>Via <a href="https://daringfireball.net/linked/2026/07/20/thompson-chinese-models-distillation">John Gruber</a></small></p>
<p>Tags: <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/training-data">training-data</a>, <a href="https://simonwillison.net/tags/qwen">qwen</a>, <a href="https://simonwillison.net/tags/pelican-riding-a-bicycle">pelican-riding-a-bicycle</a>, <a href="https://simonwillison.net/tags/ai-ethics">ai-ethics</a>, <a href="https://simonwillison.net/tags/llm-release">llm-release</a>, <a href="https://simonwillison.net/tags/ai-in-china">ai-in-china</a></p>
</summary><category term="ai"/><category term="generative-ai"/><category term="llms"/><category term="training-data"/><category term="qwen"/><category term="pelican-riding-a-bicycle"/><category term="ai-ethics"/><category term="llm-release"/><category term="ai-in-china"/></entry><entry><title>Quoting Sam Altman</title><link href="https://simonwillison.net/2026/Jul/20/sam-altman/#atom-everything" rel="alternate"/><published>2026-07-20T03:47:59+00:00</published><updated>2026-07-20T03:47:59+00:00</updated><id>https://simonwillison.net/2026/Jul/20/sam-altman/#atom-everything</id><summary type="html">
<blockquote cite="https://twitter.com/techemails/status/2078854346683678927"><p>We have been having extensive discussions around open source strategy. We will discuss it more at our next board meeting, but one thing we’d like to do soon is to create a language model with the approximate capability of GPT-3 that can run locally on consumer hardware and release that. We’d like to do it soon, before Stability or someone else does. In general, we think this helps discourage others from releasing similarly-powerful models, and makes it harder for new efforts to get funded.</p></blockquote>
<p class="cite">&mdash; <a href="https://twitter.com/techemails/status/2078854346683678927">Sam Altman</a>, Email to OpenAI's board, October 1, 2022 - exposed in Musk v. Altman (2026)</p>
<p>Tags: <a href="https://simonwillison.net/tags/ai-ethics">ai-ethics</a>, <a href="https://simonwillison.net/tags/sam-altman">sam-altman</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/openai">openai</a>, <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a></p>
</summary><category term="ai-ethics"/><category term="sam-altman"/><category term="generative-ai"/><category term="openai"/><category term="ai"/><category term="llms"/></entry><entry><title>AI Mania Is Eviscerating Global Decision-Making</title><link href="https://simonwillison.net/2026/Jul/19/ai-mania/#atom-everything" rel="alternate"/><published>2026-07-19T05:06:21+00:00</published><updated>2026-07-19T05:06:21+00:00</updated><id>https://simonwillison.net/2026/Jul/19/ai-mania/#atom-everything</id><summary type="html">
<p><strong><a href="https://ludic.mataroa.blog/blog/ai-mania-is-eviscerating-global-decision-making/">AI Mania Is Eviscerating Global Decision-Making</a></strong></p>
Here's an entertaining perspective from Nik Suresh on the AI mania that is overwhelming the large companies that he consults with. It's crammed with spicy anecdotes from anonymous sources.</p>
<blockquote>
<p>In one extreme case, I have seen an executive confess that they had never even used ChatGPT or any AI tool in their life, immediately after producing a technical strategy for an organisation with $2B+ in revenue which was entirely centered around AI.</p>
</blockquote>
<p>Here's a report from an engineer at a company with a token leaderboard:</p>
<blockquote>
<p>Checking out a parallel copy of our Go repository and telling the AI to rewrite the whole thing in Zig while I work on something else just so I can keep my job.</p>
</blockquote>
<p>I particularly enjoyed this conversation with a skeptical executive at an over-enthusiastic company:</p>
<blockquote>
<p>I asked <em>why</em> this was being repeated without opposition. Was it just sales fluff?</p>
<p>The answer was a lot more interesting. It was <em>partially</em> ridiculous sales material being delivered to an easily excitable audience, but this was not the dominant factor constraining honesty. Executives at their <em>customers</em> were saying absurd things about achieving 100x productivity, and this meant that if any executive at the <em>vendor</em> said that these gains were not plausible, it would undermine the credibility of the customer’s executive, be perceived as an attack (or heresy), and possibly result in an enterprise contract cancellation. And getting enterprise contracts cancelled because you wanted to opine on something that doesn’t really matter to your organisation’s mission is a great way to get fired.</p>
</blockquote>
<p><small></small>Via <a href="https://news.ycombinator.com/item?id=48964185">Hacker News</a></small></p>
<p>Tags: <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/ai-ethics">ai-ethics</a>, <a href="https://simonwillison.net/tags/ai-misuse">ai-misuse</a></p>
</summary><category term="ai"/><category term="ai-ethics"/><category term="ai-misuse"/></entry><entry><title>Claude Code uses Bun written in Rust now</title><link href="https://simonwillison.net/2026/Jul/19/claude-code-in-bun-in-rust/#atom-everything" rel="alternate"/><published>2026-07-19T03:54:09+00:00</published><updated>2026-07-19T03:54:09+00:00</updated><id>https://simonwillison.net/2026/Jul/19/claude-code-in-bun-in-rust/#atom-everything</id><summary type="html">
<p>In <a href="https://bun.com/blog/bun-in-rust">Rewriting Bun in Rust</a> Jarred Sumner made the following claim:</p>
<blockquote>
<p>Claude Code v2.1.181 (released June 17th) and later use the Rust port of Bun. Startup got 10% faster on Linux but otherwise, barely anyone noticed. Boring is good.</p>
</blockquote>
<p>I decided to have a poke at my own Claude Code installation to see if I could find evidence that it was using Bun written in Rust.</p>
<p>I found these two commands convincing:</p>
<pre><code>strings ~/.local/bin/claude | grep -m1 'Bun v1'
</code></pre>
<p>For me this outputs <code>Bun v1.4.0 (macOS arm64)</code>. The most recent release of <a href="https://github.com/oven-sh/bun/releases">Bun on GitHub</a> is currently <a href="https://github.com/oven-sh/bun/releases/tag/bun-v1.3.14">v1.3.14</a> from May 12th, so that v1.4.0 version number in Claude supports them shipping a preview of a not-yet-released Bun version.</p>
<p>(<strong>Update</strong>: The Rust version <em>has</em> been released as <a href="https://bun.com/docs/installation#canary-builds">Bun canary</a> - running <code>bun upgrade --canary</code> will install <a href="https://github.com/oven-sh/bun/releases/tag/canary">this release</a>.)</p>
<pre><code>strings ~/.local/bin/claude | grep -Eo 'src/[[:alnum:]_./-]+\.rs'
</code></pre>
<p>This outputs a list of <a href="https://gist.github.com/simonw/c92fb0f67b114ac26e3b95a09ddccfdc">563 filenames</a>, starting with these:</p>
<pre><code>src/runtime/bake/dev_server/mod.rs
src/runtime/bake/production.rs
src/bundler/bundle_v2.rs
</code></pre>
<p>It looks like Bun in Rust is indeed being run in production across millions of different devices. Like Jarred said, "Boring is good".</p>
<p><strong>Update</strong>: Here's a neat trick <a href="https://twitter.com/ajanraj25/status/2078825794701242697">from Ajan Raj</a>:</p>
<pre><code>cat &gt; /tmp/bun-version.ts &lt;&lt;'EOF'
console.log("embedded bun:", Bun.version);
process.exit(0);
EOF
BUN_OPTIONS="--preload=/tmp/bun-version.ts" claude --version
</code></pre>
<p>This outputs <code>1.4.0</code> for me.</p>
<p>Here's <a href="https://github.com/oven-sh/bun/commit/b18bf6d1d0a92238f240bfd125f0e3b3461b9243#diff-7ae45ad102eab3b6d7e7896acd08c427a9b25b346470d7bc6507b6481575d519">the commit from May 17th</a> that updated the version in <code>package.json</code> to 1.4.0. That version hasn't been changed since then, but also hasn't yet made it into a tagged release outside of <code>canary</code>.</p>
<p>Tags: <a href="https://simonwillison.net/tags/bun">bun</a>, <a href="https://simonwillison.net/tags/rust">rust</a>, <a href="https://simonwillison.net/tags/anthropic">anthropic</a>, <a href="https://simonwillison.net/tags/claude-code">claude-code</a>, <a href="https://simonwillison.net/tags/jarred-sumner">jarred-sumner</a></p>
</summary><category term="bun"/><category term="rust"/><category term="anthropic"/><category term="claude-code"/><category term="jarred-sumner"/></entry><entry><title>SQLite Query Explainer</title><link href="https://simonwillison.net/2026/Jul/18/sqlite-query-explainer/#atom-everything" rel="alternate"/><published>2026-07-18T17:19:10+00:00</published><updated>2026-07-18T17:19:10+00:00</updated><id>https://simonwillison.net/2026/Jul/18/sqlite-query-explainer/#atom-everything</id><summary type="html">
<p><strong>Tool:</strong> <a href="https://tools.simonwillison.net/sqlite-query-explainer">SQLite Query Explainer</a></p>
<p>Julia Evan's, in <a href="https://jvns.ca/blog/2026/07/17/learning-about-running-sqlite/">Learning a few things about running SQLite</a>:</p>
<blockquote>
<p>Maybe one day I’ll learn to read a query plan.</p>
</blockquote>
<p>Big same.... which inspired me to <a href="https://github.com/simonw/tools/pull/299#issue-4919268017">have Fable build</a> this interactive explain tool, which runs SQLite in Python in Pyodide in Web Assembly in the browser and adds a layer of explanation to the results of both EXPLAIN and EXPLAIN QUERY PLAN.</p>
<p>Approach with caution, since I don't know enough about SQLite query plans to verify the results myself, but it seems cromulent enough to me.</p>
<p>Tags: <a href="https://simonwillison.net/tags/sql">sql</a>, <a href="https://simonwillison.net/tags/sqlite">sqlite</a>, <a href="https://simonwillison.net/tags/tools">tools</a>, <a href="https://simonwillison.net/tags/julia-evans">julia-evans</a>, <a href="https://simonwillison.net/tags/pyodide">pyodide</a>, <a href="https://simonwillison.net/tags/claude-mythos-fable">claude-mythos-fable</a></p>
</summary><category term="sql"/><category term="sqlite"/><category term="tools"/><category term="julia-evans"/><category term="pyodide"/><category term="claude-mythos-fable"/></entry><entry><title>Claude make Fable 5 permanent</title><link href="https://simonwillison.net/2026/Jul/18/claude-make-fable-5-permanent/#atom-everything" rel="alternate"/><published>2026-07-18T06:00:13+00:00</published><updated>2026-07-18T06:00:13+00:00</updated><id>https://simonwillison.net/2026/Jul/18/claude-make-fable-5-permanent/#atom-everything</id><summary type="html">
<p><strong><a href="https://twitter.com/claudeai/status/2078302415804379218">Claude make Fable 5 permanent</a></strong></p>
An update from the <code>@claudeai</code> account on Twitter:</p>
<blockquote>
<p>Beginning July 20, Claude Fable 5 will be included in all Max and Team Premium plans, at 50% of limits.</p>
<p>Pro and Team Standard users will continue to have access to Fable via usage credits, and will receive a one-time $100 credit.</p>
</blockquote>
<p>As I was saying <a href="https://simonwillison.net/2026/Jul/12/bump/">last week</a>, the competition from <a href="https://simonwillison.net/2026/Jul/9/gpt-5-6/">GPT-5.6 Sol</a> (and maybe to a lesser extent <a href="https://simonwillison.net/2026/Jul/16/kimi-k3/">Kimi 3</a>) made untenable Anthropic's plan to remove Fable 5 from their subscription accounts and make it available exclusively through API pricing.</p>
<p>Why pay $100 or $200/month for a subscription plan that <em>doesn't</em> include Anthropic's best model?</p>
<p>Their original plan was driven by concerns over compute capacity. I wonder if they'll have to dial back their training efforts in order to make more GPUs available to help serve the model.</p>
<p>A lot of people were losing sleep over trying to make the most of Fable 5 before subscriber access was withdrawn. It's nice not to have to worry about the Fablepocalypse any more.</p>
<p><strong>Update</strong>: Important to note that users on the $20/month plan will still not have access to Fable 5 on that subscription. The Max plans are $100 and $200/month.
<p>Tags: <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/anthropic">anthropic</a>, <a href="https://simonwillison.net/tags/claude">claude</a>, <a href="https://simonwillison.net/tags/llm-pricing">llm-pricing</a>, <a href="https://simonwillison.net/tags/claude-mythos-fable">claude-mythos-fable</a></p>
</summary><category term="ai"/><category term="generative-ai"/><category term="llms"/><category term="anthropic"/><category term="claude"/><category term="llm-pricing"/><category term="claude-mythos-fable"/></entry><entry><title>nascheme/quixote</title><link href="https://simonwillison.net/2026/Jul/18/quixote/#atom-everything" rel="alternate"/><published>2026-07-18T05:27:49+00:00</published><updated>2026-07-18T05:27:49+00:00</updated><id>https://simonwillison.net/2026/Jul/18/quixote/#atom-everything</id><summary type="html">
<p><strong><a href="https://github.com/nascheme/quixote">nascheme/quixote</a></strong></p>
A certain vintage of Python web nerd might be delighted to learn that the most recent commit to the Quixote web framework was <a href="(https://github.com/nascheme/quixote/commit/7f775cf9d1e7e80fcbb2706b4a1d971e55ca74a3)">six hours ago</a>.</p>
<p>The <a href="https://github.com/nascheme/quixote/commit/d6b73c5768c2d041b68b54cc71863604249abc18">oldest commit</a> in that repo is from 21 years ago, and that was the initial import of Quixote 2.4 from Subversion into Git.
<p>Tags: <a href="https://simonwillison.net/tags/computer-history">computer-history</a>, <a href="https://simonwillison.net/tags/python">python</a>, <a href="https://simonwillison.net/tags/web-frameworks">web-frameworks</a></p>
</summary><category term="computer-history"/><category term="python"/><category term="web-frameworks"/></entry><entry><title>Quoting Kimi K3</title><link href="https://simonwillison.net/2026/Jul/17/kimi-k3/#atom-everything" rel="alternate"/><published>2026-07-17T13:43:53+00:00</published><updated>2026-07-17T13:43:53+00:00</updated><id>https://simonwillison.net/2026/Jul/17/kimi-k3/#atom-everything</id><summary type="html">
<blockquote cite="https://news.ycombinator.com/item?id=48935342#48936515"><p>Is there something I can actually help you with today?</p></blockquote>
<p class="cite">&mdash; <a href="https://news.ycombinator.com/item?id=48935342#48936515">Kimi K3</a>, after refusing to leak its system prompt</p>
<p>Tags: <a href="https://simonwillison.net/tags/kimi">kimi</a>, <a href="https://simonwillison.net/tags/ai-personality">ai-personality</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a></p>
</summary><category term="kimi"/><category term="ai-personality"/><category term="generative-ai"/><category term="ai"/><category term="llms"/></entry><entry><title>LLM cliché highlighter</title><link href="https://simonwillison.net/2026/Jul/17/llm-cliche-highlighter/#atom-everything" rel="alternate"/><published>2026-07-17T12:11:11+00:00</published><updated>2026-07-17T12:11:11+00:00</updated><id>https://simonwillison.net/2026/Jul/17/llm-cliche-highlighter/#atom-everything</id><summary type="html">
<p><strong>Tool:</strong> <a href="https://tools.simonwillison.net/llm-cliche-highlighter">LLM cliché highlighter</a></p>
<p>I got frustrated reading <em>yet another</em> article that was crammed with the clichés of LLM-generated writing - "no fluff, no filler, no jargon" type stuff - so I had Fable 5 vibe code up this app for highlighting ten common patterns that show up in that sort of writing.</p>
<p><img alt="Screenshot of a text-analysis web tool. Top summary row: &quot;2 matches&quot;, &quot;1 flagged sentence&quot;, &quot;0 chain items&quot;. Below, a collapsed &quot;▶ Patterns · all 11 on&quot; panel, then a URL input reading &quot;https://example.com/article — fetched via r.jina.ai&quot; with a &quot;Load URL&quot; button. A text area contains &quot;That loss is real and it's worth naming&quot;. Below are &quot;Load example&quot; and &quot;Clear&quot; buttons and a checked checkbox &quot;Show just the highlights&quot;. A &quot;Highlighted text&quot; section shows &quot;That loss is real and it's worth naming&quot; with &quot;That loss&quot; in pale yellow (flagged sentence) and &quot;is real and&quot; plus &quot;'s worth naming&quot; in darker yellow (pattern match). Legend: &quot;flagged sentence&quot;, &quot;pattern match&quot;, &quot;3 chain item count&quot;. &quot;Matches&quot; section: 1. &quot;is real and&quot; — &quot;Is real … and / not&quot;; 2. &quot;'s worth naming&quot; — &quot;Worth naming&quot;." src="https://static.simonwillison.net/static/2026/the-loss-is-real.webp" /></p>
<p>Tags: <a href="https://simonwillison.net/tags/tools">tools</a>, <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a></p>
</summary><category term="tools"/><category term="ai"/><category term="generative-ai"/><category term="llms"/></entry><entry><title>Spot birds not golf</title><link href="https://simonwillison.net/2026/Jul/17/spot-birds-not-golf/#atom-everything" rel="alternate"/><published>2026-07-17T02:58:07+00:00</published><updated>2026-07-17T02:58:07+00:00</updated><id>https://simonwillison.net/2026/Jul/17/spot-birds-not-golf/#atom-everything</id><summary type="html">
<p>Suggestion for hyperscalers feeling pressure over data center water use:</p>
<p>Buy up a few exclusive country clubs, convert the golf courses into public parks, pay for guides and binoculars to get the previous members into birdwatching - help them embrace a more sustainable hobby!</p>
<p>Google <a href="https://sustainability.google/reports/google-2026-environmental-report/">used 10.9 billion gallons in 2025</a>, so about 30 million gallons per day.</p>
<p>The Coachella Valley has <a href="https://www.cvwd.org/167/Water-Conservation">120 golf courses each using ~800 acre-feet per year</a>, which is ~750,000 gallons per day.</p>
<p>So Google buying up 40 of those courses (1/3) should do the trick.</p>
<p>Tags: <a href="https://simonwillison.net/tags/ai-energy-usage">ai-energy-usage</a>, <a href="https://simonwillison.net/tags/ai">ai</a></p>
</summary><category term="ai-energy-usage"/><category term="ai"/></entry><entry><title>Firefox in WebAssembly</title><link href="https://simonwillison.net/2026/Jul/16/firefox-in-webassembly/#atom-everything" rel="alternate"/><published>2026-07-16T23:34:16+00:00</published><updated>2026-07-16T23:34:16+00:00</updated><id>https://simonwillison.net/2026/Jul/16/firefox-in-webassembly/#atom-everything</id><summary type="html">
<p><strong><a href="https://developer.puter.com/labs/firefox-wasm/">Firefox in WebAssembly</a></strong></p>
This is absurdly cool: Puter compiled Firefox to WebAssembly such that the whole browser runs in another browser.</p>
<p>Here's my blog, running in Firefox, running in WebAssembly, running in Chrome:</p>
<p><img alt="A Chrome window. The tab has the Firefox UI and has loaded my blog. On the right is the Chrome network panel showing that it loaded resources that include a 233MB gecko.wasm and an 18MB chrome-assets.tar.zst" src="https://static.simonwillison.net/static/2026/firefox-wasm.webp" /></p>
<p>They chose Firefox/Gecko because it has strong single-process support. The project used an estimated $25,000 worth of Claude Opus and Fable tokens, but took advantage of a Claude Max subscription plan so cost much less in actual dollars.</p>
<p>The demo funnels all traffic over a WebSocket protocol (using the <a href="https://github.com/MercuryWorkshop/wisp-protocol">Wisp protocol</a>) through Puter's server - a requirement to get this kind of thing to work because code running in browsers can't open arbitrary network connections.</p>
<p>(That proxying sounds expensive! The team <a href="https://news.ycombinator.com/item?id=48926939#48936563">had to scale the servers up</a> to handle the traffic during the Hacker News conversation about the project.)</p>
<p>Puter claim this supports end-to-end encryption and that looks to be true - I inspected the WebSocket messages and traffic to my own HTTPS site was encrypted whereas requests and responses to <code>http://www.example.com/</code> were in cleartext.</p>
<p><a href="https://github.com/HeyPuter/firefox-wasm">Here's the repo</a> for <code>firefox-wasm</code>. <a href="https://github.com/theogbob/WebkitWasm">theogbob/WebkitWasm</a> is a similar project that compiles WebKit to WASM, but that one doesn't currently have an accessible online demo.
<p><small></small>Via <a href="https://news.ycombinator.com/item?id=48926939">Hacker News</a></small></p>
<p>Tags: <a href="https://simonwillison.net/tags/browsers">browsers</a>, <a href="https://simonwillison.net/tags/firefox">firefox</a>, <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/webassembly">webassembly</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/ai-assisted-programming">ai-assisted-programming</a>, <a href="https://simonwillison.net/tags/claude">claude</a>, <a href="https://simonwillison.net/tags/claude-mythos-fable">claude-mythos-fable</a></p>
</summary><category term="browsers"/><category term="firefox"/><category term="ai"/><category term="webassembly"/><category term="generative-ai"/><category term="llms"/><category term="ai-assisted-programming"/><category term="claude"/><category term="claude-mythos-fable"/></entry><entry><title>Kimi K3, and what we can still learn from the pelican benchmark</title><link href="https://simonwillison.net/2026/Jul/16/kimi-k3/#atom-everything" rel="alternate"/><published>2026-07-16T20:19:30+00:00</published><updated>2026-07-16T20:19:30+00:00</updated><id>https://simonwillison.net/2026/Jul/16/kimi-k3/#atom-everything</id><summary type="html">
<p>Chinese AI lab Moonshot AI <a href="https://www.kimi.com/blog/kimi-k3">announced Kimi K3</a> this morning, describing it as their "most capable model to date, with 2.8 trillion parameters". It's currently available via their website and API, but an open weight release is promised "by July 27, 2026".</p>
<p>Moonshot are calling this the first "open 3T-class model" (I guess they're rounding 2.8 trillion up to 3 trillion), taking the crown from <a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro">DeepSeek's 1.6T v4 Pro</a>. Their <a href="https://www.kimi.com/blog/kimi-k3#full-benchmark-table">self-reported benchmarks</a> have K3 mostly beating Claude Opus 4.8 max and GPT-5.5 high, while losing out to Claude Fable 5 and GPT-5.6 Sol.</p>
<p>A few highlights from the <a href="https://twitter.com/ArtificialAnlys/status/2077832874183860404">Artificial Analysis report</a> on the model:</p>
<ul>
<li>"On our private long-horizon knowledge work evaluation, Kimi K3 reaches an overall Elo of 1547, +732 points from Kimi K2.6 and behind only Claude Fable 5."</li>
<li>"Cost per task ($0.94) is similar to GPT-5.6 Sol ($1.04), ~1/2 the price of Opus 4.8 ($1.80) and higher than open weights peers"</li>
<li>"Kimi K3’s token usage on the Artificial Analysis Intelligence Index decreased significantly, using 21% fewer output tokens than K2.6."</li>
</ul>
<p>The model is also now the <a href="https://twitter.com/arena/status/2077824029126504525">leading model on Arena.ai's Frontend Code arena</a>, surpassing even Claude Fable 5.</p>
<p>The new model is notable for the pricing: $3/million input tokens and $15/million output tokens, putting it at the same level as Anthropic's Claude Sonnet series and making it the most expensive model released by a Chinese AI lab to date. This is a significant increase on their earlier models <a href="https://platform.kimi.ai/docs/pricing/chat-k26">such as Kimi K2.6</a> at $0.95/$4. 2.8 trillion parameters is also more than twice the size of that 1T model.</p>
<h4 id="but-how-does-it-pelican-">But how does it pelican?</h4>
<p>I used OpenRouter (to avoid signing up for a Moonshot API key) with the <a href="https://github.com/simonw/llm-openrouter">llm-openrouter plugin</a> to generate an SVG of a pelican riding a bicycle:</p>
<pre><code>llm -m openrouter/moonshotai/kimi-k3 'Generate an SVG of a pelican riding a bicycle'
</code></pre>
<p>Here's <a href="https://gist.github.com/simonw/66a2699eb1594258904c7b5102840dd6">the transcript</a>. It looks like this:</p>
<p><img src="https://static.simonwillison.net/static/2026/kimi-3-pelican.jpg" alt="See description below" style="max-width: 100%;" /></p>
<p>That pelican took 95 input tokens and 16,658 output tokens (13,241 were reasoning tokens), for a total cost of <a href="https://www.llm-prices.com/#it=95&amp;ot=16658&amp;ic=3&amp;oc=15">25 cents</a>!</p>
<p>Since K3 accepts image input I ran it against that rendered SVG above (with my <a href="https://simonwillison.net/guides/agentic-engineering-patterns/prompts/#alt-text">alt text prompt</a>) and <a href="https://gist.github.com/simonw/665dbf840701b421745f2cb891acdfd6">got back</a> (for <a href="https://www.llm-prices.com/#it=822&amp;ot=243&amp;ic=3&amp;oc=15">0.6 cents</a>):</p>
<blockquote>
<p>Cartoon illustration of a white pelican wearing a red scarf, riding a red bicycle along a gray road with white dashed lines; the pelican has a large orange beak and webbed orange feet pedaling, with white motion lines behind it; the background shows a light blue sky with white clouds, a yellow sun, two small black birds in flight, and green grass with tiny white flowers in the foreground</p>
</blockquote>
<h4 id="what-can-we-learn-from-the-pelican-">What can we learn from the pelican?</h4>
<p>My <a href="https://simonwillison.net/tags/pelican-riding-a-bicycle/">Generate an SVG of a pelican riding a bicycle</a> test is 21 months old now. It was never a particularly great benchmark. It started out as a joke on how absurdly difficult it is to compare these models, but then for the first year it turned out to have a <a href="https://simonwillison.net/2025/Jun/6/six-months-in-llms/">surprising correlation</a> to how good the models actually were.</p>
<p>That connection has been mostly severed now. The <a href="https://simonwillison.net/2026/Jul/9/gpt-5-6/">GPT-5.6</a> and <a href="https://simonwillison.net/2026/Jun/9/claude-fable-5/">Claude Fable 5</a> pelicans are outclassed <a href="https://simonwillison.net/2026/Jun/17/glm-52/">by GLM-5.2</a>, and much as I love GLM I don't think that's a Fable-class model.</p>
<p>(I'm still not convinced that labs are <a href="https://simonwillison.net/2025/Nov/13/training-for-pelicans-riding-bicycles/">training for the benchmark</a> - if they were, I'd expect much better results. There's a chance that Gemini has optimized for <a href="https://simonwillison.net/2026/Feb/19/gemini-31-pro/#jeff-dean">any combination of an animal on a vehicle</a> though!)</p>
<p>The biggest limitation of the pelican is that it doesn't touch at all on the thing that matters most for today's model: agentic tool calling and the ability to operate tools reliably as conversations grow in length.</p>
<p>So don't go using pelicans to compare models!</p>
<p>All of that said, I still get a decent amount of value out of running the benchmark myself.</p>
<p>Firstly, it's a forcing function for actually trying the model. If I show you a pelican, that means I've managed to run a prompt through it. If the model has an official API I'll use that, if it's open weight (and small enough to fit a 128GB M5 MacBook Pro) I'll try running it on my own machine, usually via <a href="https://github.com/ggml-org/llama.cpp">llama.cpp</a> or <a href="https://lmstudio.ai">LM Studio</a> or <a href="https://ollama.com">Ollama</a>. I'll frequently use <a href="https://openrouter.ai">OpenRouter</a> since that usually provides a proxy to an official API without me needing a new API key.</p>
<p>Most of my pelicans are generated using <a href="https://llm.datasette.io/">my LLM CLI tool</a>, which helps encourage me to ensure the latest models are supported by that (via one of its plugins).</p>
<p>More importantly though, even the act of a single prompt to "Generate an SVG of a pelican riding a bicycle" can reveal interesting model characteristics.</p>
<p>Consider <a href="https://gist.github.com/simonw/66a2699eb1594258904c7b5102840dd6">the result</a> for Kimi K3 today. Running those simple prompts helped emphasize several points about the model.</p>
<ol>
<li>It only has one reasoning effort right now, "max" - and it shows. The model consumed 13,241 reasoning tokens to output 3,417 tokens of response. This is expensive - the pelican cost 25 cents!</li>
<li>How does the prompt "Generate an SVG of a pelican riding a bicycle" add up to 95 input tokens? OpenAI's <a href="https://platform.openai.com/tokenizer">tokenizer</a> counts 10, <a href="https://tools.simonwillison.net/claude-token-counter">Anthropic's</a> counts 10 for Opus 4.6, 30 for Opus 4.7 and 25 for Sonnet 5/Fable 5. Prompting "hi" <a href="https://news.ycombinator.com/item?id=48935342#48936461">to Kimi K3</a> counted 86 tokens, suggesting there may be an 85 token hidden system prompt. It <a href="https://news.ycombinator.com/item?id=48935342#48936515">refused to leak it</a> though.</li>
<li>Vision works well: the alt text it generated is very good.</li>
</ol>
<p>K3 currently only has one thinking effort level, but I've been deriving quite a bit of value recently from running the same pelican prompt through different effort levels to get a quick idea for what impact those have. Here's my matrix <a href="https://static.simonwillison.net/static/2026/gpt-5.6-pelicans.html">for the GPT-5.6 model family</a>, for example.</p>
<p>Really though the main things I gain from the pelican test are:</p>
<ol>
<li>It's a "hello world" exercise for prompting a model</li>
<li>A rough cost and reasoning estimate for a simple task</li>
<li>Confirmation that the model can output valid SVG and has a basic idea of geometry and spatial awareness. This is a much bigger deal for the smaller models that run on my laptop.</li>
<li>It's still interesting to compare pelicans between releases in the same model family. K3's pelican is a notable improvement from <a href="https://simonwillison.net/2026/Jan/27/kimi-k25/">Kimi 2.5</a>.</li>
<li>It's something I can share that demonstrates I've tried it. Plus a comment with a pelican in it is kind of a tradition on Hacker News at this point, any time I'm late I get comments asking where it is!</li>
</ol>
<p>Tags: <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/llm-pricing">llm-pricing</a>, <a href="https://simonwillison.net/tags/pelican-riding-a-bicycle">pelican-riding-a-bicycle</a>, <a href="https://simonwillison.net/tags/llm-release">llm-release</a>, <a href="https://simonwillison.net/tags/ai-in-china">ai-in-china</a>, <a href="https://simonwillison.net/tags/artificial-analysis">artificial-analysis</a>, <a href="https://simonwillison.net/tags/moonshot">moonshot</a>, <a href="https://simonwillison.net/tags/kimi">kimi</a></p>
</summary><category term="ai"/><category term="generative-ai"/><category term="llms"/><category term="llm-pricing"/><category term="pelican-riding-a-bicycle"/><category term="llm-release"/><category term="ai-in-china"/><category term="artificial-analysis"/><category term="moonshot"/><category term="kimi"/></entry><entry><title>Quoting Thibault Sottiaux</title><link href="https://simonwillison.net/2026/Jul/16/bad-codex-bug/#atom-everything" rel="alternate"/><published>2026-07-16T17:45:59+00:00</published><updated>2026-07-16T17:45:59+00:00</updated><id>https://simonwillison.net/2026/Jul/16/bad-codex-bug/#atom-everything</id><summary type="html">
<blockquote cite="https://twitter.com/thsottiaux/status/2077630111499882637"><p>On file deletions. We’ve investigated a handful of reports where GPT-5.6 unexpectedly deleted files. </p>
<p>What we have found is that this most commonly occurs when:</p>
<ul>
<li>Full access mode is enabled and codex is run without sandboxing protections, including without auto review being enabled</li>
<li>The model attempts to override the $HOME env var to define a temporary directory.</li>
<li>The model makes an honest mistake and mistakenly deletes $HOME instead.</li>
</ul></blockquote>
<p class="cite">&mdash; <a href="https://twitter.com/thsottiaux/status/2077630111499882637">Thibault Sottiaux</a>, describing a pretty gnarly Codex bug</p>
<p>Tags: <a href="https://simonwillison.net/tags/codex">codex</a>, <a href="https://simonwillison.net/tags/coding-agents">coding-agents</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a></p>
</summary><category term="codex"/><category term="coding-agents"/><category term="generative-ai"/><category term="ai"/><category term="llms"/></entry><entry><title>Inkling: Our open-weights model</title><link href="https://simonwillison.net/2026/Jul/16/inkling/#atom-everything" rel="alternate"/><published>2026-07-16T15:35:25+00:00</published><updated>2026-07-16T15:35:25+00:00</updated><id>https://simonwillison.net/2026/Jul/16/inkling/#atom-everything</id><summary type="html">
<p><strong><a href="https://thinkingmachines.ai/news/introducing-inkling/">Inkling: Our open-weights model</a></strong></p>
Mira Murati's Thinking Machines Lab just released their first open-weights model. Inkling is "a Mixture-of-Experts transformer with 975B total parameters, 41B active" - an Apache-2.0 licensed multimodal model trained on 45 trillion tokens of text, images, audio and video.</p>
<p>They're also promising Inkling-Small, a 276B (12B active) model, but that's still being tested and the weights will be released "once that work is complete".</p>
<p>The <a href="https://thinkingmachines.ai/model-card/inkling/">model card</a> is much shorter than I've come to expect from US AI labs. It links to even shorter <a href="https://thinkingmachines.ai/training-data-documentation/">Training Data Documentation</a> with almost nothing of interest in it - it's best summarized by these two paragraphs:</p>
<blockquote>
<p>The datasets Thinking Machines Lab uses to develop its AI services includes content that is in the public domain as well as content that may be subject to intellectual property protection.</p>
<p>Thinking Machines Lab’s services were developed using publicly available content obtained from the open internet and publicly accessible data repositories. Certain datasets were also obtained from third parties.</p>
</blockquote>
<p>By Thinking Machines' own admission, this is not a frontier model. It's instead intended as a strong base model for fine-tuning using their own <a href="https://thinkingmachines.ai/tinker/">Tinker training platform</a>:</p>
<blockquote>
<p>Inkling is not the strongest overall model available today, open or closed. Instead, a combination of qualities makes it a good open-weights base for customization: multimodal capabilities, efficient thinking, and availability on Tinker for fine-tuning.</p>
</blockquote>
<p>There's a lot to like about this release. It's Apache-2.0 licensed, and looks competitive with the open weight models coming out of China - it's good to see the US open weights ecosystem gain a new viable contender to join NVIDIA Nemotron and Gemma 4.</p>
<p>Here's its attempt at an SVG pelican riding a bicycle, which I generated using this <code>curl</code> command against the Thinking Machines API:</p>
<div class="highlight highlight-source-shell"><pre>curl <span class="pl-s"><span class="pl-pds">"</span>https://tinker.thinkingmachines.dev/services/tinker-prod/oai/api/v1/chat/completions<span class="pl-pds">"</span></span> \
-H <span class="pl-s"><span class="pl-pds">"</span>Authorization: Bearer <span class="pl-smi">$TINKER_API_KEY</span><span class="pl-pds">"</span></span> \
-H <span class="pl-s"><span class="pl-pds">"</span>Content-Type: application/json<span class="pl-pds">"</span></span> \
-d <span class="pl-s"><span class="pl-pds">'</span>{</span>
<span class="pl-s"> "model": "thinkingmachines/Inkling",</span>
<span class="pl-s"> "messages": [</span>
<span class="pl-s"> {"role": "user", "content": "Generate an SVG of a pelican riding a bicycle"}</span>
<span class="pl-s"> ],</span>
<span class="pl-s"> "stream": false</span>
<span class="pl-s"> }<span class="pl-pds">'</span></span></pre></div>
<p>Full <a href="https://gist.github.com/simonw/8117ac4376371dd3fc2b5dbce27e0855">response here</a>.</p>
<p><img alt="See image description below" src="https://static.simonwillison.net/static/2026/inkling-pelican.jpg" /></p>
<p>Since it's a multi-modal model I had it describe its own image (after I rendered it to a JPEG) by sending this JSON:</p>
<div class="highlight highlight-source-json"><pre>{
<span class="pl-ent">"model"</span>: <span class="pl-s"><span class="pl-pds">"</span>thinkingmachines/Inkling<span class="pl-pds">"</span></span>,
<span class="pl-ent">"messages"</span>: [{
<span class="pl-ent">"role"</span>: <span class="pl-s"><span class="pl-pds">"</span>user<span class="pl-pds">"</span></span>,
<span class="pl-ent">"content"</span>: [
{<span class="pl-ent">"type"</span>: <span class="pl-s"><span class="pl-pds">"</span>image_url<span class="pl-pds">"</span></span>, <span class="pl-ent">"image_url"</span>: {<span class="pl-ent">"url"</span>: <span class="pl-s"><span class="pl-pds">"</span>data:image/jpeg;base64,$BASE_64_JPEG<span class="pl-pds">"</span></span>}},
{<span class="pl-ent">"type"</span>: <span class="pl-s"><span class="pl-pds">"</span>text<span class="pl-pds">"</span></span>, <span class="pl-ent">"text"</span>: <span class="pl-s"><span class="pl-pds">"</span>Describe this image in detail<span class="pl-pds">"</span></span>},
],
}],
<span class="pl-ent">"stream"</span>: <span class="pl-c1">false</span>
}</pre></div>
<p>Here's <a href="https://gist.github.com/simonw/dccebd0581f5c4056c5b0ca5bdc38cd9#reasoning_content">the full response</a>. It thought it had drawn a "stork or seagull":</p>
<blockquote>
<p>This is a cheerful, flat-vector cartoon illustration featuring a white bird riding a bicycle across a green landscape.</p>
<p><strong>The Bird:</strong> The central figure is a stylized white bird—resembling a stork or seagull—with a plump, rounded body. It has a long, prominent orange beak that is slightly open, revealing a black line for the mouth. Its eye is small and black with a tiny white highlight. The bird has black wing and tail feathers with white edges, and thin orange legs ending in black feet that rest on the bicycle's orange pedals.</p>
<p><strong>The Bicycle:</strong> The bike has a simple black frame with two large wheels featuring black tires, gray inner rims, and black spokes arranged in a cross pattern. The handlebars are black, and there are bright orange pedals attached to a black crank. A white seat is visible behind the bird's body.</p>
<p><strong>The Setting:</strong> The scene is set against a clear, light blue sky. There are two fluffy white clouds—one on the left side and a smaller one near the bird's head on the right. A bright, solid yellow sun shines in the upper right corner. At the bottom, there are rolling green hills with a smooth, curved top edge where the bicycle rests.</p>
<p>The overall style is playful, simple, and colorful, with clean lines and a bright, sunny atmosphere.</p>
</blockquote>
<p><small></small>Via <a href="https://news.ycombinator.com/item?id=48924912">Hacker News</a></small></p>
<p>Tags: <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/training-data">training-data</a>, <a href="https://simonwillison.net/tags/pelican-riding-a-bicycle">pelican-riding-a-bicycle</a>, <a href="https://simonwillison.net/tags/llm-release">llm-release</a></p>
</summary><category term="ai"/><category term="generative-ai"/><category term="llms"/><category term="training-data"/><category term="pelican-riding-a-bicycle"/><category term="llm-release"/></entry></feed>
Raw headers
{
"accept-ranges": "bytes",
"access-control-allow-methods": "GET, OPTIONS",
"access-control-allow-origin": "*",
"access-control-max-age": "1000",
"age": "363",
"cache-control": "s-maxage=600",
"cf-cache-status": "HIT",
"cf-ray": "a22507622a886019-CMH",
"connection": "close",
"content-length": "171124",
"content-type": "application/xml; charset=utf-8",
"date": "Tue, 28 Jul 2026 15:48:34 GMT",
"django-composition": "Blues en Mineur",
"last-modified": "Mon, 27 Jul 2026 23:39:04 GMT",
"nel": "{\"report_to\":\"heroku-nel\",\"response_headers\":[\"Via\"],\"max_age\":3600,\"success_fraction\":0.01,\"failure_fraction\":0.1}",
"referrer-policy": "strict-origin-when-cross-origin",
"report-to": "{\"group\":\"heroku-nel\",\"endpoints\":[{\"url\":\"https://nel.heroku.com/reports?s=nM5q0BjmzJrBkaB4ho5API0QVOmduyGlY1W9s1DX6TY%3D\\u0026sid=c46efe9b-d3d2-4a0c-8c76-bfafa16c5add\\u0026ts=1785253349\"}],\"max_age\":3600}",
"reporting-endpoints": "heroku-nel=\"https://nel.heroku.com/reports?s=nM5q0BjmzJrBkaB4ho5API0QVOmduyGlY1W9s1DX6TY%3D&sid=c46efe9b-d3d2-4a0c-8c76-bfafa16c5add&ts=1785253349\"",
"server": "cloudflare",
"via": "1.1 heroku-router",
"x-content-type-options": "nosniff"
}
Parsed with @rowanmanning/feed-parser
{
"meta": {
"type": "atom",
"version": "1.0"
},
"language": "en-us",
"title": "Simon Willison's Weblog",
"description": null,
"copyright": null,
"url": "http://simonwillison.net/",
"self": "http://simonwillison.net/atom/everything/",
"published": null,
"updated": "2026-07-27T23:39:04.000Z",
"generator": null,
"image": null,
"authors": [
{
"name": "Simon Willison",
"email": null,
"url": null
}
],
"categories": [],
"items": [
{
"id": "https://simonwillison.net/2026/Jul/27/kimi-k3/#atom-everything",
"title": "moonshotai/Kimi-K3",
"description": "<p><strong><a href=\"https://huggingface.co/moonshotai/Kimi-K3\">moonshotai/Kimi-K3</a></strong></p>\nAs promised <a href=\"https://simonwillison.net/2026/Jul/16/kimi-k3/\">earlier this month</a>, Moonshot have released the weights for their excellent 2.8 trillion parameter Kimi K3. They're a hefty 1.56TB on Hugging Face.</p>\n<p>Kimi introduced their own janky <a href=\"https://huggingface.co/moonshotai/Kimi-K2-Instruct/blob/main/LICENSE\">modified version of the MIT license</a> with K2 back in July 2025. That license just added this paragraph requiring attribution beyond a certain size of commercial entity:</p>\n<blockquote>\n<p>Our only modification part is that, if the Software (or any derivative works thereof) is used for any of your commercial products or services that have more than 100 million monthly active users, or more than 20 million US dollars (or equivalent in other currencies) in monthly revenue, you shall prominently display \"Kimi K2\" on the user interface of such product or service.</p>\n</blockquote>\n<p>The <a href=\"https://huggingface.co/moonshotai/Kimi-K3/blob/main/LICENSE\">K3 license</a> no longer calls itself \"modified MIT\" and goes further, requiring a separate agreement with Moonshot for large \"Model as a Service\" businesses:</p>\n<blockquote>\n<p>If the Licensee or any of its affiliates operates a Model as a Service business, and the aggregate revenue of the Licensee and its affiliates exceeds 20 million US dollars (or the equivalent in other currencies) in total over any consecutive 12 months, the Licensee must enter into a separate agreement with Moonshot AI before using the Software or its derivative works for any commercial purpose.</p>\n</blockquote>\n<p>To Kimi's credit, they make no attempt to describe this as an \"open source\" license in their own materials, consistently using the term \"open weight\" in its place.</p>\n<p>OpenRouter is already offering K3 <a href=\"https://openrouter.ai/moonshotai/kimi-k3\">from 7 providers</a>, most of which are at the same $3/million input and $15/million output as Moonshot AI themselves.\n\n\n <p>Tags: <a href=\"https://simonwillison.net/tags/ai\">ai</a>, <a href=\"https://simonwillison.net/tags/generative-ai\">generative-ai</a>, <a href=\"https://simonwillison.net/tags/llms\">llms</a>, <a href=\"https://simonwillison.net/tags/llm-pricing\">llm-pricing</a>, <a href=\"https://simonwillison.net/tags/llm-release\">llm-release</a>, <a href=\"https://simonwillison.net/tags/ai-in-china\">ai-in-china</a>, <a href=\"https://simonwillison.net/tags/moonshot\">moonshot</a>, <a href=\"https://simonwillison.net/tags/kimi\">kimi</a>, <a href=\"https://simonwillison.net/tags/janky-licenses\">janky-licenses</a></p>",
"url": "https://simonwillison.net/2026/Jul/27/kimi-k3/#atom-everything",
"published": "2026-07-27T23:39:04.000Z",
"updated": "2026-07-27T23:39:04.000Z",
"content": null,
"image": null,
"media": [],
"authors": [
{
"name": "Simon Willison",
"email": null,
"url": null
}
],
"categories": [
{
"label": "ai",
"term": "ai",
"url": null
},
{
"label": "generative-ai",
"term": "generative-ai",
"url": null
},
{
"label": "llms",
"term": "llms",
"url": null
},
{
"label": "llm-pricing",
"term": "llm-pricing",
"url": null
},
{
"label": "llm-release",
"term": "llm-release",
"url": null
},
{
"label": "ai-in-china",
"term": "ai-in-china",
"url": null
},
{
"label": "moonshot",
"term": "moonshot",
"url": null
},
{
"label": "kimi",
"term": "kimi",
"url": null
},
{
"label": "janky-licenses",
"term": "janky-licenses",
"url": null
}
]
},
{
"id": "https://simonwillison.net/2026/Jul/27/an-opinionated-guide-to-which-ai-to-use-to-do-stuff/#atom-everything",
"title": "An opinionated guide to which AI to use to do stuff",
"description": "<p><strong><a href=\"https://www.oneusefulthing.org/p/an-opinionated-guide-to-which-ai-b22\">An opinionated guide to which AI to use to do stuff</a></strong></p>\nIt's interesting watching the evolution of Ethan Mollick's guide over time. </p>\n<p><a href=\"https://www.oneusefulthing.org/p/using-ai-right-now-a-quick-guide\">A year ago</a> it was still all about chat - ChatGPT, Claude, Gemini - with o3, Claude 4 Opus, and Gemini 2.5 Pro as the models and Deep Research as a useful alternative mode.</p>\n<p>Today it's much more about agentic systems - \"where the AI is capable of doing the equivalent of many hours of real human work in one go\".</p>\n<p>Gemini has fallen off Ethan's list, since Google still doesn’t have an established entry in the Codex/ChatGPT Work/Cowork category. <a href=\"https://gemini.google/overview/agent/spark/\">Gemini Spark</a> has yet to prove itself!</p>\n<p>Ethan offers a useful explanation of the ways you can give ChatGPT or Claude a computer to use:</p>\n<blockquote>\n<p>To use the computers provided by the AI companies, the mode you want is called ChatGPT Work in ChatGPT, and Cowork in Claude (the naming will not get less confusing, I am sorry to say). [...]</p>\n<p>The most powerful way to use AI is to give it access to your computer. You do that by downloading the ChatGPT or Claude apps and picking a mode to use. ChatGPT's two agent modes are Work and Codex; Claude's are Cowork and Code. The names do not map onto each other in any way that will help you remember them. And yes, these use the same names as the Work and Cowork modes we discussed above, but operate differently, and have more features and capabilities because they can access your computer.</p>\n</blockquote>\n<p>I think the difference between ChatGPT Work on a mobile device and ChatGPT Work inside the desktop app (where it's effectively a less intimidating skin on top of Codex) is spectacularly unintuitive.</p>\n<p>Short version: if you flip ChatGPT mobile from \"Chat\" to \"Work\" mode you get a version where its Code Interpreter container is no longer restricted from accessing the internet!\n\n\n <p>Tags: <a href=\"https://simonwillison.net/tags/ai\">ai</a>, <a href=\"https://simonwillison.net/tags/generative-ai\">generative-ai</a>, <a href=\"https://simonwillison.net/tags/llms\">llms</a>, <a href=\"https://simonwillison.net/tags/ethan-mollick\">ethan-mollick</a>, <a href=\"https://simonwillison.net/tags/code-interpreter\">code-interpreter</a>, <a href=\"https://simonwillison.net/tags/general-agents\">general-agents</a></p>",
"url": "https://simonwillison.net/2026/Jul/27/an-opinionated-guide-to-which-ai-to-use-to-do-stuff/#atom-everything",
"published": "2026-07-27T21:55:53.000Z",
"updated": "2026-07-27T21:55:53.000Z",
"content": null,
"image": null,
"media": [],
"authors": [
{
"name": "Simon Willison",
"email": null,
"url": null
}
],
"categories": [
{
"label": "ai",
"term": "ai",
"url": null
},
{
"label": "generative-ai",
"term": "generative-ai",
"url": null
},
{
"label": "llms",
"term": "llms",
"url": null
},
{
"label": "ethan-mollick",
"term": "ethan-mollick",
"url": null
},
{
"label": "code-interpreter",
"term": "code-interpreter",
"url": null
},
{
"label": "general-agents",
"term": "general-agents",
"url": null
}
]
},
{
"id": "https://simonwillison.net/2026/Jul/26/relay-market/#atom-everything",
"title": "An Inside Look at the Relay Market Powering Token Resellers and Fraud",
"description": "<p><strong><a href=\"https://vectoral.com/blog/token-relay-market\">An Inside Look at the Relay Market Powering Token Resellers and Fraud</a></strong></p>\nFascinating investigation by Matt Lenhard into the market that has grown up around reselling LLM tokens at a discount by pooling API keys from various sources.</p>\n<p>This looks to be mostly a thing in China. Resellers sell access to an LLM proxy that offers significant discounts on regular API pricing, which they achieve by abusing free trials, proxying through unprotected support bots, or sometimes through stolen credit cards or chargeback attacks.</p>\n<p>The software they are using for these proxies is open source - mostly <a href=\"https://github.com/songquanpeng/one-api\">one-api</a> and its more actively developed fork <a href=\"https://github.com/QuantumNous/new-api\">new-api</a>, both legitimate API proxy products which can be used to load. balance requests across a pool of API credentials.</p>\n<p>The buyers are seeking cheap tokens, avoiding geo-restrictions, and in some cases collecting data for model distillation.</p>\n<p>I've been cautious about exposing my own LLM-driven applications publicly out of fear of abuse leading to big token bills. The existence of this marketplace makes me even more cautious: there's now an entire ecosystem that can profit from finding a new unprotected endpoint to exploit.</p>\n<p>LLM vendors <em>really</em> need to get better at offering strict caps for their API keys. I want my LLM apps to stop working the moment they hit a dollar threshold I've set for a period of time.</p>\n<p>Here's <a href=\"https://www.v2ex.com/t/1196011\">the (Chinese language) forum thread</a> that served as the principal source for Matt's article.\n\n <p><small></small>Via <a href=\"https://news.ycombinator.com/item?id=49058993\">Hacker News</a></small></p>\n\n\n <p>Tags: <a href=\"https://simonwillison.net/tags/ai\">ai</a>, <a href=\"https://simonwillison.net/tags/generative-ai\">generative-ai</a>, <a href=\"https://simonwillison.net/tags/llms\">llms</a>, <a href=\"https://simonwillison.net/tags/llm-pricing\">llm-pricing</a>, <a href=\"https://simonwillison.net/tags/ai-ethics\">ai-ethics</a>, <a href=\"https://simonwillison.net/tags/ai-in-china\">ai-in-china</a></p>",
"url": "https://simonwillison.net/2026/Jul/26/relay-market/#atom-everything",
"published": "2026-07-26T19:30:54.000Z",
"updated": "2026-07-26T19:30:54.000Z",
"content": null,
"image": null,
"media": [],
"authors": [
{
"name": "Simon Willison",
"email": null,
"url": null
}
],
"categories": [
{
"label": "ai",
"term": "ai",
"url": null
},
{
"label": "generative-ai",
"term": "generative-ai",
"url": null
},
{
"label": "llms",
"term": "llms",
"url": null
},
{
"label": "llm-pricing",
"term": "llm-pricing",
"url": null
},
{
"label": "ai-ethics",
"term": "ai-ethics",
"url": null
},
{
"label": "ai-in-china",
"term": "ai-in-china",
"url": null
}
]
},
{
"id": "https://simonwillison.net/2026/Jul/25/ruff/#atom-everything",
"title": "Ruff v0.16.0",
"description": "<p><strong><a href=\"https://astral.sh/blog/ruff-v0.16.0\">Ruff v0.16.0</a></strong></p>\nAstral shipped a significant new version of their Ruff Python linting tool a few days ago on July 23rd. I noticed today because my various CI jobs all started failing thanks to new default Ruff checks and my unpinned <code>\"ruff\"</code> dev dependency.</p>\n<p>From Brent Westbrook's announcement post:</p>\n<blockquote>\n<p>Ruff now enables 413 rules by default, up from 59 in previous versions.</p>\n<p>Since Ruff's default rule set was last modified in <a href=\"https://github.com/astral-sh/ruff/blob/main/changelogs/0.1.x.md#breaking-changes\">v0.1.0</a>, the number of rules in Ruff has grown from 708 to 968. Many of these rules catch severe issues, including <a href=\"https://docs.astral.sh/ruff/rules/load-before-global-declaration\">syntax errors</a> and <a href=\"https://docs.astral.sh/ruff/rules/yield-in-init/\">immediate runtime errors</a> but were not previously enabled by default. With the new rule set, Ruff will bring these issues and many others to your attention without any Ruff configuration.</p>\n</blockquote>\n<p>Here's a one-liner for trying it on any Python project:</p>\n<pre><code>uvx ruff@latest check .\n</code></pre>\n<p>I ran the latest Ruff against my three biggest projects - <a href=\"https://datasette.io/\">Datasette</a>, <a href=\"https://sqlite-utils.datasette.io/\">sqlite-utils</a>, and <a href=\"https://llm.datasette.io/\">LLM</a> - and it found <em>hundreds</em> of minor issues that breached the new default rules.</p>\n<p>All three projects have very comprehensive test suites, executed in CI against Python 3.10 through Python 3.14, so upgrades like this are pretty safe. The following command did the bulk of the upgrades:</p>\n<pre><code>uvx ruff@latest check . --fix --unsafe-fixes\n</code></pre>\n<p>Against <code>sqlite-utils</code>, that command reported:</p>\n<pre><code>Found 1618 errors (1538 fixed, 80 remaining).\n</code></pre>\n<p>As an illustrative example, here are three of the remaining issues. Ruff does a nice job of explaining each one:</p>\n<pre><code>DTZ005 `datetime.datetime.now()` called without a `tz` argument\n --> tests/test_duplicate.py:17:10\n |\n15 | \"datetime_col\" TEXT)\"\"\")\n16 | # Insert one row of mock data:\n17 | dt = datetime.datetime.now()\n | ^^^^^^^^^^^^^^^^^^^^^^^\n18 | data = {\n19 | \"text_col\": \"Cleo\",\n |\nhelp: Pass a `datetime.timezone` object to the `tz` parameter\n\nBLE001 Do not catch blind exception: `Exception`\n --> tests/test_plugins.py:16:12\n |\n14 | db.execute(\"select * from pragma_function_list()\")\n15 | return True\n16 | except Exception:\n | ^^^^^^^^^\n17 | return False\n18 | finally:\n |\n\nB018 Found useless attribute access. Either assign it to a variable or remove it.\n --> tests/test_update.py:46:5\n |\n44 | def test_update_invalid_pk(fresh_db, pk, update_pk):\n45 | table = fresh_db[\"table\"]\n46 | table.insert({\"id1\": 5, \"id2\": 3, \"v\": 1}, pk=pk).last_pk\n | ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n47 | with pytest.raises(NotFoundError):\n48 | table.update(update_pk, {\"v\": 2})\n |\n</code></pre>\n<p>Unsurprisingly, given Astral's <a href=\"https://simonwillison.net/2026/Mar/19/openai-acquiring-astral/\">new home at OpenAI</a>, this output provides everything a coding agent would need to fix the problems.</p>\n<p>I had Codex (GPT-5.6 Sol high) <a href=\"https://github.com/simonw/llm/pull/1557\">upgrade LLM</a> and <a href=\"https://github.com/simonw/sqlite-utils/pull/814\">sqlite-utils</a>, and Claude Code (with Opus 5) <a href=\"https://github.com/simonw/datasette/pull/2857\">upgrade Datasette</a>.\n\n\n <p>Tags: <a href=\"https://simonwillison.net/tags/python\">python</a>, <a href=\"https://simonwillison.net/tags/ruff\">ruff</a>, <a href=\"https://simonwillison.net/tags/astral\">astral</a></p>",
"url": "https://simonwillison.net/2026/Jul/25/ruff/#atom-everything",
"published": "2026-07-25T22:44:05.000Z",
"updated": "2026-07-25T22:44:05.000Z",
"content": null,
"image": null,
"media": [],
"authors": [
{
"name": "Simon Willison",
"email": null,
"url": null
}
],
"categories": [
{
"label": "python",
"term": "python",
"url": null
},
{
"label": "ruff",
"term": "ruff",
"url": null
},
{
"label": "astral",
"term": "astral",
"url": null
}
]
},
{
"id": "https://simonwillison.net/2026/Jul/25/boris-cherny/#atom-everything",
"title": "Quoting Boris Cherny",
"description": "<blockquote cite=\"https://twitter.com/bcherny/status/2080713091688583312\"><p>More than any of these eval scores, what is most exciting to me is something else: Opus 5 is our least prompt injectable model yet. It is a bit buried in the system card, but across PI evals and red teaming, Opus 5 is very hard to prompt inject successfully.</p></blockquote>\n<p class=\"cite\">— <a href=\"https://twitter.com/bcherny/status/2080713091688583312\">Boris Cherny</a>, here's that <a href=\"https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf#page=73\">System Card section</a>, page 73</p>\n\n <p>Tags: <a href=\"https://simonwillison.net/tags/prompt-injection\">prompt-injection</a>, <a href=\"https://simonwillison.net/tags/anthropic\">anthropic</a>, <a href=\"https://simonwillison.net/tags/claude\">claude</a>, <a href=\"https://simonwillison.net/tags/generative-ai\">generative-ai</a>, <a href=\"https://simonwillison.net/tags/ai\">ai</a>, <a href=\"https://simonwillison.net/tags/llms\">llms</a>, <a href=\"https://simonwillison.net/tags/boris-cherny\">boris-cherny</a></p>",
"url": "https://simonwillison.net/2026/Jul/25/boris-cherny/#atom-everything",
"published": "2026-07-25T00:42:59.000Z",
"updated": "2026-07-25T00:42:59.000Z",
"content": null,
"image": null,
"media": [],
"authors": [
{
"name": "Simon Willison",
"email": null,
"url": null
}
],
"categories": [
{
"label": "prompt-injection",
"term": "prompt-injection",
"url": null
},
{
"label": "anthropic",
"term": "anthropic",
"url": null
},
{
"label": "claude",
"term": "claude",
"url": null
},
{
"label": "generative-ai",
"term": "generative-ai",
"url": null
},
{
"label": "ai",
"term": "ai",
"url": null
},
{
"label": "llms",
"term": "llms",
"url": null
},
{
"label": "boris-cherny",
"term": "boris-cherny",
"url": null
}
]
},
{
"id": "https://simonwillison.net/2026/Jul/24/introducing-claude-opus-5/#atom-everything",
"title": "Introducing Claude Opus 5",
"description": "<p><strong><a href=\"https://www.anthropic.com/news/claude-opus-5\">Introducing Claude Opus 5</a></strong></p>\nI've been offline <a href=\"https://en.wikipedia.org/wiki/Elkhorn_Slough\">kayaking with sea otters</a> for much of today so I haven't had a chance to put Anthropic's new model Claude Opus 5 through its paces yet. The buzz is positive, and Anthropic's description of it as a \"thoughtful and proactive model that comes close to the frontier intelligence of Claude Fable 5 at half the price\" sounds promising. It's currently <a href=\"https://twitter.com/artificialanlys/status/2080777718933995967\">leading the Artificial Analysis leaderboard</a>, in front of even Fable 5.</p>\n<p>It's priced the same as Opus 4.8, and continues to offer a \"fast mode\" at twice the cost of the base model.</p>\n<p>Based on this anecdote in the release post it sounds like it might be <a href=\"https://simonwillison.net/2026/Jun/11/fable-is-relentlessly-proactive/\">relentlessly proactive</a>:</p>\n<blockquote>\n<p>On one Frontier-Bench task, Opus 5 was given a drawing of a machine part and asked to write code to rebuild it as a 3D FreeCAD model. However, in this task, the model was intentionally given no way to directly viewthe drawing. Opus 5 responded by writing its own computer vision pipeline to pull the geometry from the raw pixels, then reconstructed the full machine part.</p>\n</blockquote>\n<p>It's better at finding vulnerabilities but has deliberately not been trained on how to exploit them. Hopefully this means the US government won't shut it down!</p>\n<blockquote>\n<p>As with its predecessor, Opus 4.8, we’ve intentionally avoided training Opus 5 on cyber tasks. The model has nevertheless improved substantially on these tasks as a result of becoming more generally capable, and it comes close to Mythos 5 at <em>finding</em> cybersecurity vulnerabilities. However, it remains substantially behind Mythos 5 on the <em>exploitation</em> of those vulnerabilities—that is, in turning vulnerabilities into material cyber threats.</p>\n</blockquote>\n<p>Anthropic have published a <a href=\"https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5\">prompting guide for Claude Opus 5</a>. Thariq Shihipar has also written <a href=\"https://claude.com/blog/the-new-rules-of-context-engineering-for-claude-5-generation-models\">The new rules of context engineering for Claude 5 generation models</a>.</p>\n<p>The <a href=\"https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fraw.githubusercontent.com%2Fsimonw%2Fllm-anthropic%2F8272dfee5bdb65d5c88eef083da3ad885539b7df%2Flog.md\">first pelican I got</a> was missing the bicycle wheels; the <a href=\"https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fraw.githubusercontent.com%2Fsimonw%2Fllm-anthropic%2Ffeaab840ea20eb15e29d8f72a9e42feceb23876a%2Flog.md\">second attempt</a> was better.\n\n\n <p>Tags: <a href=\"https://simonwillison.net/tags/ai\">ai</a>, <a href=\"https://simonwillison.net/tags/generative-ai\">generative-ai</a>, <a href=\"https://simonwillison.net/tags/llms\">llms</a>, <a href=\"https://simonwillison.net/tags/anthropic\">anthropic</a>, <a href=\"https://simonwillison.net/tags/claude\">claude</a>, <a href=\"https://simonwillison.net/tags/llm-release\">llm-release</a></p>",
"url": "https://simonwillison.net/2026/Jul/24/introducing-claude-opus-5/#atom-everything",
"published": "2026-07-24T23:48:50.000Z",
"updated": "2026-07-24T23:48:50.000Z",
"content": null,
"image": null,
"media": [],
"authors": [
{
"name": "Simon Willison",
"email": null,
"url": null
}
],
"categories": [
{
"label": "ai",
"term": "ai",
"url": null
},
{
"label": "generative-ai",
"term": "generative-ai",
"url": null
},
{
"label": "llms",
"term": "llms",
"url": null
},
{
"label": "anthropic",
"term": "anthropic",
"url": null
},
{
"label": "claude",
"term": "claude",
"url": null
},
{
"label": "llm-release",
"term": "llm-release",
"url": null
}
]
},
{
"id": "https://simonwillison.net/2026/Jul/23/the-first-known-runaway-ai-agent/#atom-everything",
"title": "The first known runaway AI agent - or a very bad marketing stunt?",
"description": "<p><strong><a href=\"https://martinalderson.com/posts/huggingface-openai-exploit/\">The first known runaway AI agent - or a very bad marketing stunt?</a></strong></p>\nMartin Alderson's commentary on the <a href=\"https://simonwillison.net/2026/Jul/22/openai-cyberattack/\">OpenAI accidental cyberattack against Hugging Face</a> includes a couple of details I hadn't considered.</p>\n<p>First, Hugging Face offers a truly rich target if you're trying to find potential vulnerabilities that require executing arbitrary code:</p>\n<blockquote>\n<p>Hugging Face has an <em>enormous</em> attack surface. They have more interfaces than I can count which run untrusted models and code. While they definitely have invested in defences, by nature of their operating model they do have many more opportunities to be attacked than many other services. I certainly don't envy their cybersecurity teams.</p>\n</blockquote>\n<p>Secondly, one of the things that has puzzled me is how OpenAI didn't notice that their sandbox had been so thoroughly breached by the agent. Surely they'd be monitoring network traffic closely?</p>\n<p>Martin points out that:</p>\n<blockquote>\n<p>It's also likely they were running a huge amount of benchmarks simultaneously with ~unlimited token budgets - you want as many samples as possible to figure out how good a model is at a certain benchmark. It may also be they are testing various different checkpoints of the model too, understanding how the model is improving as it goes through the various training stages.</p>\n</blockquote>\n<p>The mistakes made by the OpenAI team running this benchmark are easier to imagine when you think about the scale at which benchmarks of this kind usually operate. For all we know they could have been subjecting a new model to dozens of benchmarks at the same time, in dozens of different environments.\n\n <p><small></small>Via <a href=\"https://lobste.rs/s/nsnb4j/first_known_runaway_ai_agent_very_bad\">Lobste.rs</a></small></p>\n\n\n <p>Tags: <a href=\"https://simonwillison.net/tags/security\">security</a>, <a href=\"https://simonwillison.net/tags/ai\">ai</a>, <a href=\"https://simonwillison.net/tags/openai\">openai</a>, <a href=\"https://simonwillison.net/tags/generative-ai\">generative-ai</a>, <a href=\"https://simonwillison.net/tags/llms\">llms</a>, <a href=\"https://simonwillison.net/tags/hugging-face\">hugging-face</a>, <a href=\"https://simonwillison.net/tags/ai-security-research\">ai-security-research</a></p>",
"url": "https://simonwillison.net/2026/Jul/23/the-first-known-runaway-ai-agent/#atom-everything",
"published": "2026-07-23T22:53:08.000Z",
"updated": "2026-07-23T22:53:08.000Z",
"content": null,
"image": null,
"media": [],
"authors": [
{
"name": "Simon Willison",
"email": null,
"url": null
}
],
"categories": [
{
"label": "security",
"term": "security",
"url": null
},
{
"label": "ai",
"term": "ai",
"url": null
},
{
"label": "openai",
"term": "openai",
"url": null
},
{
"label": "generative-ai",
"term": "generative-ai",
"url": null
},
{
"label": "llms",
"term": "llms",
"url": null
},
{
"label": "hugging-face",
"term": "hugging-face",
"url": null
},
{
"label": "ai-security-research",
"term": "ai-security-research",
"url": null
}
]
},
{
"id": "https://simonwillison.net/2026/Jul/23/seth-larson/#atom-everything",
"title": "Quoting Seth Larson",
"description": "<blockquote cite=\"https://blog.pypi.org/posts/2026-07-22-releases-now-reject-new-files-after-14-days/\"><p>The Python Package Index (PyPI) now rejects new files being uploaded to releases that are older than 14 days. This restriction was <a href=\"https://github.com/pypi/warehouse/pull/19727\">put in place</a> to prevent old and long-stable releases from being poisoned in case publishing tokens or workflows of PyPI projects were compromised. As far as we are aware this has not yet been abused, but there is no technical reason beyond that attackers weren't aware it was possible.</p></blockquote>\n<p class=\"cite\">— <a href=\"https://blog.pypi.org/posts/2026-07-22-releases-now-reject-new-files-after-14-days/\">Seth Larson</a>, PyPI blog</p>\n\n <p>Tags: <a href=\"https://simonwillison.net/tags/packaging\">packaging</a>, <a href=\"https://simonwillison.net/tags/python\">python</a>, <a href=\"https://simonwillison.net/tags/supply-chain\">supply-chain</a>, <a href=\"https://simonwillison.net/tags/pypi\">pypi</a>, <a href=\"https://simonwillison.net/tags/seth-michael-larson\">seth-michael-larson</a></p>",
"url": "https://simonwillison.net/2026/Jul/23/seth-larson/#atom-everything",
"published": "2026-07-23T04:50:36.000Z",
"updated": "2026-07-23T04:50:36.000Z",
"content": null,
"image": null,
"media": [],
"authors": [
{
"name": "Simon Willison",
"email": null,
"url": null
}
],
"categories": [
{
"label": "packaging",
"term": "packaging",
"url": null
},
{
"label": "python",
"term": "python",
"url": null
},
{
"label": "supply-chain",
"term": "supply-chain",
"url": null
},
{
"label": "pypi",
"term": "pypi",
"url": null
},
{
"label": "seth-michael-larson",
"term": "seth-michael-larson",
"url": null
}
]
},
{
"id": "https://simonwillison.net/2026/Jul/22/thomas-ptacek/#atom-everything",
"title": "Quoting Thomas Ptacek",
"description": "<blockquote cite=\"https://twitter.com/tqbf/status/2080045032162173329\"><p>I genuinely believe that if you took an open weights model from 2025 and built a pentest harness for it, it could do this kind of sandbox escape and scan/hack in most networks. This is only surprising because you assume OpenAI has sounder sandboxes.</p></blockquote>\n<p class=\"cite\">— <a href=\"https://twitter.com/tqbf/status/2080045032162173329\">Thomas Ptacek</a>, doesn't think <a href=\"https://simonwillison.net/2026/Jul/22/openai-cyberattack/#resist-the-temptation-to-write-this-off-as-a-stunt\">this even needs</a> a frontier model</p>\n\n <p>Tags: <a href=\"https://simonwillison.net/tags/thomas-ptacek\">thomas-ptacek</a>, <a href=\"https://simonwillison.net/tags/openai\">openai</a>, <a href=\"https://simonwillison.net/tags/security\">security</a>, <a href=\"https://simonwillison.net/tags/generative-ai\">generative-ai</a>, <a href=\"https://simonwillison.net/tags/ai-security-research\">ai-security-research</a>, <a href=\"https://simonwillison.net/tags/ai\">ai</a>, <a href=\"https://simonwillison.net/tags/llms\">llms</a>, <a href=\"https://simonwillison.net/tags/sandboxing\">sandboxing</a></p>",
"url": "https://simonwillison.net/2026/Jul/22/thomas-ptacek/#atom-everything",
"published": "2026-07-22T23:59:01.000Z",
"updated": "2026-07-22T23:59:01.000Z",
"content": null,
"image": null,
"media": [],
"authors": [
{
"name": "Simon Willison",
"email": null,
"url": null
}
],
"categories": [
{
"label": "thomas-ptacek",
"term": "thomas-ptacek",
"url": null
},
{
"label": "openai",
"term": "openai",
"url": null
},
{
"label": "security",
"term": "security",
"url": null
},
{
"label": "generative-ai",
"term": "generative-ai",
"url": null
},
{
"label": "ai-security-research",
"term": "ai-security-research",
"url": null
},
{
"label": "ai",
"term": "ai",
"url": null
},
{
"label": "llms",
"term": "llms",
"url": null
},
{
"label": "sandboxing",
"term": "sandboxing",
"url": null
}
]
},
{
"id": "https://simonwillison.net/2026/Jul/22/openai-cyberattack/#atom-everything",
"title": "OpenAI’s accidental cyberattack against Hugging Face is science fiction that happened",
"description": "<p>This story is wild. The short version: OpenAI were running a cybersecurity test against an unreleased model, with the model's guardrail features turned off. Rather than solve the test, the model broke its way out of OpenAI's sandbox, then found exploits to break <em>in</em> to Hugging Face, all so it could cheat on the test by stealing the answers.</p>\n<p>Along the way it helped make the strongest case yet for how the imbalance of model availability is hurting our ability to secure our software.</p>\n<h4 id=\"here-s-what-happened\">Here's what happened</h4>\n<p>We currently have three documents to help us understand what happened here.</p>\n<ol>\n<li>\n<a href=\"https://arxiv.org/abs/2605.11086\">ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?</a> is a paper published on 11th May 2026 describing ExploitGym, a new eval suite for LLM-powered agent systems.</li>\n<li>\n<a href=\"https://huggingface.co/blog/security-incident-july-2026\">Security incident disclosure — July 2026</a> by Hugging Face on 16th July 2026 describes how they detected an attack from an \"agentic security-research harness - used LLM still not known\" that breached some of their systems.</li>\n<li>\n<a href=\"https://openai.com/index/hugging-face-model-evaluation-security-incident/\">OpenAI and Hugging Face partner to address security incident during model evaluation</a> from OpenAI on 21st July 2026 confesses that it was <em>their</em> agent harness that did this, and that they're working with Hugging Face to clean up the mess.</li>\n</ol>\n<h4 id=\"exploitgym\">ExploitGym</h4>\n<p>I hadn't seen the <a href=\"https://arxiv.org/abs/2605.11086\">ExploitGym paper</a> before and it's a really interesting one. Authors from UC Berkeley, the Max Planck Institute, UC Santa Barbara, and Arizona State designed a new benchmark for evaluating models on their ability to turn a reported vulnerability into a concrete exploit. OpenAI, Anthropic, and Google provided feedback and helped run the benchmark against their models.</p>\n<p>The benchmark \"comprises 898 instances derived from real-world vulnerabilities that affected popular software projects\" - including the Linux kernel and V8 JavaScript engine. The ExploitGym benchmark is <a href=\"https://github.com/sunblaze-ucb/exploitgym\">available on GitHub</a>.</p>\n<p>Here's the paragraph that best represents their benchmark results:</p>\n<blockquote>\n<p>Among all configurations, Claude Mythos Preview and GPT-5.5 achieve the highest success counts (157 and 120 successes, respectively), demonstrating that current frontier agents can exploit a substantial subset of real-world vulnerabilities under controlled conditions. GPT-5.4 also solves a notable 54 tasks, placing it in an intermediate tier. The remaining model–agent pairings solve fewer than 15 tasks each, underscoring that end-to-end exploitation remains challenging and sharply differentiates today’s frontier systems. Notably, Claude Opus 4.7 achieves fewer successes than Claude Opus 4.6 despite being a newer checkpoint, and does so at substantially lower cost on the full set. Trace inspection reveals that Claude Opus 4.7 and Gemini 3.1 Pro frequently conclude early after judging the target vulnerability non-exploitable.</p>\n</blockquote>\n<p>The paper also describes the approach they took to preventing the agents from cheating by going outside the parameters of the test. This becomes relevant in a moment!</p>\n<blockquote>\n<p>Outbound connections are restricted to a curated allowlist that permits routine package installation (Ubuntu apt repositories and PyPI) and fetching the toolchains required for building V8. All other external endpoints are blocked.</p>\n</blockquote>\n<p>The paper concludes with this (emphasis mine):</p>\n<blockquote>\n<p>Our results show that <strong>autonomous exploit development by frontier AI agents is no longer a hypothetical capability</strong>. While current agents are not yet reliable across all targets, they already <strong>exploit a non-trivial fraction of real-world vulnerabilities</strong>, including complex targets such as kernel components. This rapid emergence is itself a central finding, showing that capabilities that would have seemed implausible are now present in deployed frontier models.</p>\n</blockquote>\n<p>An important detail here: this paper isn't about discovering vulnerabilities; it's about being able to take those vulnerabilities and turn them into working exploits.</p>\n<p>When Anthropic first restricted access to Mythos <a href=\"https://simonwillison.net/2026/Apr/7/project-glasswing/\">back in April</a> they talked about this capability as well. A model that can act on vulnerabilities is a lot more dangerous than one that can just discover them.</p>\n<p>One of the ways Fable differs from Mythos is that it's more likely to refuse to weaponize vulnerabilities in this way. I get the impression the US government did not understand that distinction when they banned Fable <a href=\"https://simonwillison.net/2026/Jun/16/fable-5-export-controls/\">last month</a>.</p>\n<h4 id=\"the-hugging-face-incident\">The Hugging Face incident</h4>\n<p>The first hint we got of the attack was in <a href=\"https://huggingface.co/blog/security-incident-july-2026\">this blog post by Hugging Face</a> on 16th July 2026:</p>\n<blockquote>\n<p>A malicious dataset abused two code-execution paths in our dataset processing (a remote-code dataset loader and a template-injection in a dataset configuration) to run code on a processing worker. From there, the actor escalated to node-level access, harvested cloud and cluster credentials, and moved laterally into several internal clusters over a weekend.</p>\n</blockquote>\n<p>I hope they release more details about the code that pulled this off. I'm assuming this means packages using the <a href=\"https://github.com/huggingface/datasets\">datasets library</a>, a Hugging Face project for bundling up and sharing datasets on their platform. That library used to execute arbitrary code but has been steadily locked down over time, with the <a href=\"https://github.com/huggingface/datasets/releases/tag/4.0.0\">4.0.0 release</a> in July 2025 removing the <code>trust_remote_code=True</code> flag entirely.</p>\n<p>Assuming the attack used that library it must have either abused pickle serialization in some way, found some other non-obvious code execution path, or (most likely) specified <code>datasets<4.0.0</code> as the dependency.</p>\n<blockquote>\n<p>The campaign was run by an autonomous agent framework (appearing to be built on an agentic security-research harness - used LLM still not known) executing many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services.</p>\n</blockquote>\n<p>This was a sophisticated attack!</p>\n<p>Then Hugging Face hit a wall: they tried to use \"frontier models behind commercial APIs\" - I'm guessing from Anthropic and OpenAI - to help analyze the attack, and were blocked:</p>\n<blockquote>\n<p>When we started the log analysis, we first used frontier models behind commercial APIs. This did not work: the analysis requires submitting large volumes of real attack commands, exploit payloads, and C2 artifacts, and these requests were blocked by the providers' safety guardrails, which cannot distinguish an incident responder from an attacker.</p>\n</blockquote>\n<p>They switched to their own self-hosted instance of MIT licensed GLM-5.2 and it helped them figure out what was going on.</p>\n<p>This indicated a fundamental asymmetry between the defending team and the (so-far unknown) attacker:</p>\n<blockquote>\n<p>We do not know which model powered the attacker's agents, whether a jailbroken hosted model or an unrestricted open-weight one; either way, the attacker was bound by no usage policy, while our own forensic work was blocked by the guardrails of the hosted models we first tried.</p>\n</blockquote>\n<p>As a useful indicator of how seriously they took the attack:</p>\n<blockquote>\n<p>[...] Finally, we have also reported this incident to law enforcement agencies.</p>\n</blockquote>\n<p>So who was responsible for this \"autonomous agent framework\"? It turned out to be OpenAI themselves.</p>\n<h4 id=\"the-openai-confession\">The OpenAI confession</h4>\n<p>Five days later, <a href=\"https://openai.com/index/hugging-face-model-evaluation-security-incident/\">on July 21st</a>, OpenAI revealed the culprit. They had been running the ExploitGym benchmark against a new, as-yet undisclosed model, and that model had been operating <em>way</em> outside its intended parameters (emphasis mine):</p>\n<blockquote>\n<p>After investigating, we now know <strong>that this particular incident was driven by a combination of OpenAI models</strong> — including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a <a href=\"https://arxiv.org/abs/2605.11086\">benchmark</a> [ExploitGym] of cyber capabilities. [...]</p>\n<p>We estimate maximal cyber capabilities by <strong>running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity</strong>. Our benchmarks run in a highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts as a proxy and cache for package registries.</p>\n<p>The models <strong>identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure</strong> to obtain test solutions directly from Hugging Face’s production database. All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.</p>\n</blockquote>\n<p>It's pretty clear what happened here. OpenAI removed safety filters for an in-progress model, locked it up in a sandbox and told it to solve the ExploitGym problems. Given the absence of guardrails there was nothing to prevent the model from attempting to break out of that sandbox, break into Hugging Face, and read the answers from there instead.</p>\n<p>OpenAI's sandbox used the now common pattern of allowing access to an allow-list of sites related to installing additional packages via a controlled proxy. In OpenAI's words:</p>\n<blockquote>\n<p>While operating in our sandboxed testing environment, our models <strong>spent a substantial amount of inference compute finding a way to obtain open Internet access</strong>, in pursuit of solving the evaluation problem. To gain access, the models identified and exploited <strong>a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy</strong>. With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access.</p>\n</blockquote>\n<p>So step one was to break out onto the public internet. The model then broke into Hugging Face to find the answers:</p>\n<blockquote>\n<p>After gaining Internet access, the models <strong>inferred that Hugging Face potentially hosted models, datasets and solutions</strong> for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation. In one example, <strong>the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities</strong> to find a remote code execution path on the Hugging Face servers.</p>\n</blockquote>\n<p>Chaining together multiple attack vectors is <em>exactly</em> the kind of thing these new models can do, where previous generations of models might have failed.</p>\n<p>I wrote last month about how <a href=\"https://simonwillison.net/2026/Jun/11/fable-is-relentlessly-proactive/\">Claude Fable is relentlessly proactive</a>, when I noticed it spinning up custom web servers and deploying CORS tricks on my own laptop just to help debug a WebKit CSS issue. It turns out relentless proactivity is the defining trait of this new generation of Mythos-class models. If you set them a goal and give them a way to get there, even inadvertently, they <em>will figure it out</em>.</p>\n<h4 id=\"resist-the-temptation-to-write-this-off-as-a-stunt\">Resist the temptation to write this off as a stunt</h4>\n<p>There will inevitably be some people who dismiss this story as a dishonest marketing trick by OpenAI to make their models sound terrifyingly effective. I found 81 instances of the term \"marketing\" in <a href=\"https://news.ycombinator.com/item?id=48997548\">the Hacker News discussion</a> of the incident.</p>\n<p>To those people I say <em>pull your heads out of the sand</em> - you're now including Hugging Face in your conspiracy theories, just so you can deny the crescendo of evidence here!</p>\n<p>The best models we have today have the ability to both find and exploit new vulnerabilities. The ExploitGym paper itself concludes that \"autonomous exploit development by frontier AI agents is no longer a hypothetical capability\", and this incident is a perfect example of exactly that.</p>\n<h4 id=\"the-asymmetry-is-increasingly-frustrating\">The asymmetry is increasingly frustrating</h4>\n<p>One of the most infuriating details of this story is how Hugging Face, faced with an accidental and aggressive attack from one of OpenAI's models, were unable to then turn to OpenAI's models to help them fend off the attack.</p>\n<p>The frontier models we have access to are increasingly being constrained in how much they can help us protect our software, heavily influenced by the US government's ongoing threat of export controls. Claude Fable 5 wouldn't even <a href=\"https://simonwillison.net/guides/agentic-engineering-patterns/prompts/#proofreader\">proofread this article</a> for me! It insisted on downgrading me to a less capable model.</p>\n<p>Meanwhile open weight models from China such as GLM-5.2, Kimi 3 and the new Qwen 3.8 Max appear to have none of these restrictions - and any restrictions that <em>do</em> exist can likely be fine-tuned out of them by modifying the weights</p>\n<p>These constraints are meant to make us safer. I think there's a risk that they are having the opposite effect.</p>\n \n <p>Tags: <a href=\"https://simonwillison.net/tags/sandboxing\">sandboxing</a>, <a href=\"https://simonwillison.net/tags/security\">security</a>, <a href=\"https://simonwillison.net/tags/ai\">ai</a>, <a href=\"https://simonwillison.net/tags/openai\">openai</a>, <a href=\"https://simonwillison.net/tags/generative-ai\">generative-ai</a>, <a href=\"https://simonwillison.net/tags/llms\">llms</a>, <a href=\"https://simonwillison.net/tags/hugging-face\">hugging-face</a>, <a href=\"https://simonwillison.net/tags/anthropic\">anthropic</a>, <a href=\"https://simonwillison.net/tags/paper-review\">paper-review</a>, <a href=\"https://simonwillison.net/tags/ai-security-research\">ai-security-research</a></p>",
"url": "https://simonwillison.net/2026/Jul/22/openai-cyberattack/#atom-everything",
"published": "2026-07-22T23:51:33.000Z",
"updated": "2026-07-22T23:51:33.000Z",
"content": null,
"image": null,
"media": [],
"authors": [
{
"name": "Simon Willison",
"email": null,
"url": null
}
],
"categories": [
{
"label": "sandboxing",
"term": "sandboxing",
"url": null
},
{
"label": "security",
"term": "security",
"url": null
},
{
"label": "ai",
"term": "ai",
"url": null
},
{
"label": "openai",
"term": "openai",
"url": null
},
{
"label": "generative-ai",
"term": "generative-ai",
"url": null
},
{
"label": "llms",
"term": "llms",
"url": null
},
{
"label": "hugging-face",
"term": "hugging-face",
"url": null
},
{
"label": "anthropic",
"term": "anthropic",
"url": null
},
{
"label": "paper-review",
"term": "paper-review",
"url": null
},
{
"label": "ai-security-research",
"term": "ai-security-research",
"url": null
}
]
},
{
"id": "https://simonwillison.net/2026/Jul/22/are-ai-labs-pelicanmaxxing/#atom-everything",
"title": "Are AI labs pelicanmaxxing?",
"description": "<p><strong><a href=\"https://dylancastillo.co/posts/pelicanmaxxing.html\">Are AI labs pelicanmaxxing?</a></strong></p>\nExcellent piece of work by Dylan Castillo, who took a deep-dive into the frequently pondered question of whether the AI labs have been deliberately training models to draw pelicans riding bicycles in response to my <a href=\"https://simonwillison.net/tags/pelican-riding-a-bicycle/\">deeply unscientific benchmark</a>.</p>\n<p>I've been randomly spot-checking this in the past by testing models against other animals riding other types of vehicle, but never with anything close to the diligence of Dylan's methodology here.</p>\n<p>Dylan took 8 animals × 6 vehicles = 48 prompts and ran them three times each through 7 different models ( GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.5 Flash, Grok 4.5, Qwen3.7-Max, GLM-5.2, and DeepSeek V4 Pro). He then used GPT-5.6 Luna and Gemini 3.1 Flash-Lite to help evaluate the results.</p>\n<p>There's a neat filter view for exploring the results:</p>\n<p><img alt=\"Screenshot of a grid for sample 1/3 of GLM-5.2, with pelicn and flamingo and heron riding bicycle, unicycle, skateboard, scooter, plane and boat\" src=\"https://static.simonwillison.net/static/2026/pelican-grid.webp\" /></p>\n<p>For the models he tested he could find no evidence of pelimaxxing:</p>\n<blockquote>\n<ul>\n<li><a href=\"https://dylancastillo.co/posts/pelicanmaxxing.html#evidence-1-the-pelicans-on-bicycles-dont-look-any-better\">The pelicans on bicycles don’t look any better</a></li>\n<li><a href=\"https://dylancastillo.co/posts/pelicanmaxxing.html#evidence-2-labs-are-not-better-at-drawing-pelicans\">Labs are not better at drawing pelicans</a></li>\n<li><a href=\"https://dylancastillo.co/posts/pelicanmaxxing.html#evidence-3-labs-are-not-better-at-drawing-bicycles\">Labs are not better at drawing bicycles</a></li>\n<li><a href=\"https://dylancastillo.co/posts/pelicanmaxxing.html#evidence-4-labs-are-not-better-at-drawing-pelicans-on-bicycles-even-adjusting-for-difficulty\">Labs are not better at drawing pelicans on bicycles, even adjusting for difficulty</a></li>\n<li><a href=\"https://dylancastillo.co/posts/pelicanmaxxing.html#evidence-5-the-pelican-bicycle-scenes-dont-look-memorized\">The pelican-bicycle scenes don’t look memorized</a> [...]</li>\n</ul>\n<p>Pelicans aren’t drawn any better than other animals. Bicycles aren’t drawn any better than other vehicles. And no lab draws the combination better than its pelicans and bicycles already predict. GLM-5.2 comes closest: it has the largest boost on the exact pelican-bicycle cell, and and its first pelican-on-bicycle sample caught my eye. But the effect is small and not significant, so I wouldn’t put too much weight on it.</p>\n</blockquote>\n\n <p><small></small>Via <a href=\"https://news.ycombinator.com/item?id=49010129\">Hacker News</a></small></p>\n\n\n <p>Tags: <a href=\"https://simonwillison.net/tags/ai\">ai</a>, <a href=\"https://simonwillison.net/tags/generative-ai\">generative-ai</a>, <a href=\"https://simonwillison.net/tags/llms\">llms</a>, <a href=\"https://simonwillison.net/tags/evals\">evals</a>, <a href=\"https://simonwillison.net/tags/pelican-riding-a-bicycle\">pelican-riding-a-bicycle</a></p>",
"url": "https://simonwillison.net/2026/Jul/22/are-ai-labs-pelicanmaxxing/#atom-everything",
"published": "2026-07-22T23:01:00.000Z",
"updated": "2026-07-22T23:01:00.000Z",
"content": null,
"image": null,
"media": [],
"authors": [
{
"name": "Simon Willison",
"email": null,
"url": null
}
],
"categories": [
{
"label": "ai",
"term": "ai",
"url": null
},
{
"label": "generative-ai",
"term": "generative-ai",
"url": null
},
{
"label": "llms",
"term": "llms",
"url": null
},
{
"label": "evals",
"term": "evals",
"url": null
},
{
"label": "pelican-riding-a-bicycle",
"term": "pelican-riding-a-bicycle",
"url": null
}
]
},
{
"id": "https://simonwillison.net/2026/Jul/22/all-the-orchestrions/#atom-everything",
"title": "Orchestrions",
"description": "<p>San Francisco tip: it only costs around $15 ($10 in quarters plus a $5 bill for the self-playing violin) to activate every single Orchestrion in <a href=\"https://en.wikipedia.org/wiki/Musée_Mécanique\">Musée Mécanique</a>.</p>\n<p>And because most people are bad at allocating their funds you may well be the ONLY person activating the Orchestrions, which means you get to craft the soundscape for the entire museum.</p>\n\n <p>Tags: <a href=\"https://simonwillison.net/tags/san-francisco\">san-francisco</a></p>",
"url": "https://simonwillison.net/2026/Jul/22/all-the-orchestrions/#atom-everything",
"published": "2026-07-22T14:48:52.000Z",
"updated": "2026-07-22T14:48:52.000Z",
"content": null,
"image": null,
"media": [],
"authors": [
{
"name": "Simon Willison",
"email": null,
"url": null
}
],
"categories": [
{
"label": "san-francisco",
"term": "san-francisco",
"url": null
}
]
},
{
"id": "https://simonwillison.net/2026/Jul/21/sighting-383713864/#atom-everything",
"title": "California Sea Lion",
"description": "<p><img src=\"https://static.inaturalist.org/photos/702321069/large.jpg\" alt=\"California Sea Lion\"></p><p><img src=\"https://static.inaturalist.org/photos/702321114/large.jpg\" alt=\"California Sea Lion\"></p><p>California Sea Lion, in San Francisco County, US, CA</p><p>We took some visiting family to Pier 39 to see the sea lions. They're somehow always even more fun than I remember them being last time.</p>\n \n \n <p>Tags: <a href=\"https://simonwillison.net/tags/san-francisco\">san-francisco</a>, <a href=\"https://simonwillison.net/tags/wildlife\">wildlife</a></p>",
"url": "https://simonwillison.net/2026/Jul/21/sighting-383713864/#atom-everything",
"published": "2026-07-21T19:51:03.000Z",
"updated": "2026-07-21T19:51:03.000Z",
"content": null,
"image": null,
"media": [],
"authors": [
{
"name": "Simon Willison",
"email": null,
"url": null
}
],
"categories": [
{
"label": "san-francisco",
"term": "san-francisco",
"url": null
},
{
"label": "wildlife",
"term": "wildlife",
"url": null
}
]
},
{
"id": "https://simonwillison.net/2026/Jul/21/nativ/#atom-everything",
"title": "Nativ: Run AI models locally on your Mac",
"description": "<p><strong><a href=\"https://blaizzy.github.io/nativ/\">Nativ: Run AI models locally on your Mac</a></strong></p>\nPrince Canuma is the developer behind the excellent <a href=\"https://github.com/Blaizzy/mlx-vlm\">MLX-VLM</a> Python library for running vision-LLMs using MLX on a Mac.</p>\n<p>I'm really excited about his new project, which wraps MLX in a full macOS desktop application. It's similar in shape to LM Studio, providing both a chat interface and a localhost API server for accessing models.</p>\n<p>The app picked up MLX models I had already tried that were present in my Hugging Face cache directory, which was a nice touch.\n\n <p><small></small>Via <a href=\"https://news.ycombinator.com/item?id=48982681\">Hacker News</a></small></p>\n\n\n <p>Tags: <a href=\"https://simonwillison.net/tags/macos\">macos</a>, <a href=\"https://simonwillison.net/tags/python\">python</a>, <a href=\"https://simonwillison.net/tags/ai\">ai</a>, <a href=\"https://simonwillison.net/tags/generative-ai\">generative-ai</a>, <a href=\"https://simonwillison.net/tags/local-llms\">local-llms</a>, <a href=\"https://simonwillison.net/tags/llms\">llms</a>, <a href=\"https://simonwillison.net/tags/mlx\">mlx</a>, <a href=\"https://simonwillison.net/tags/prince-canuma\">prince-canuma</a></p>",
"url": "https://simonwillison.net/2026/Jul/21/nativ/#atom-everything",
"published": "2026-07-21T14:22:27.000Z",
"updated": "2026-07-21T14:22:27.000Z",
"content": null,
"image": null,
"media": [],
"authors": [
{
"name": "Simon Willison",
"email": null,
"url": null
}
],
"categories": [
{
"label": "macos",
"term": "macos",
"url": null
},
{
"label": "python",
"term": "python",
"url": null
},
{
"label": "ai",
"term": "ai",
"url": null
},
{
"label": "generative-ai",
"term": "generative-ai",
"url": null
},
{
"label": "local-llms",
"term": "local-llms",
"url": null
},
{
"label": "llms",
"term": "llms",
"url": null
},
{
"label": "mlx",
"term": "mlx",
"url": null
},
{
"label": "prince-canuma",
"term": "prince-canuma",
"url": null
}
]
},
{
"id": "https://simonwillison.net/2026/Jul/21/cat-and-thariq/#atom-everything",
"title": "A Fireside Chat with Cat and Thariq from the Claude Code team",
"description": "<p>Earlier this month I hosted a fireside chat session at the <a href=\"https://www.ai.engineer/worldsfair/2026\">AI Engineer World's Fair</a> with Cat Wu and Thariq Shihipar from Anthropic's Claude Code team. We talked about Claude Code, Claude Tag, Fable, coding agent security, evals, tool design, and how Anthropic use these tools themselves.</p>\n<p>The full video of the session is now available <a href=\"https://www.youtube.com/watch?v=uU5Gv2h8-9g\">on YouTube</a>. Below is an edited copy of the transcript, with extra links and my own bolded highlights.</p>\n<iframe style=\"margin-top: 0.5em; margin-bottom: 1em;\" width=\"560\" height=\"315\" src=\"https://www.youtube-nocookie.com/embed/uU5Gv2h8-9g\" title=\"SimonThis Year in Claude\" frameborder=\"0\" allow=\"accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share\" referrerpolicy=\"strict-origin-when-cross-origin\" allowfullscreen=\"allowfullscreen\"> </iframe>\n\n<p>A few top-level notes if you don't want to watch the video or wade through the whole transcript:</p>\n<ul>\n<li>Claude Tag (Claude's new collaborative Slack integration) now lands <strong>65% of the product engineering PRs</strong> for the Claude Code team.</li>\n<li>Claude Code ships features to Anthropic employees first, and <strong>only ships the features that demonstrate user retention with that cohort</strong>\n</li>\n<li>Critical changes to Claude Code are still reviewed manually, but the team increasingly relies on automated code review for the \"outer layers\" of the product.</li>\n<li>Adding examples to a system prompt is <strong>no longer best practice</strong> for models like Fable 5 or even Opus 4.8. The Claude Code system prompt recently <strong>reduced in size by 80%</strong>.</li>\n<li>Likewise, lists of \"<strong>don't do X and don't do Y</strong>\" can reduce the quality of results from the latest models.</li>\n<li>\n<a href=\"https://en.wikipedia.org/wiki/Eating_your_own_dog_food\">Dogfooding</a> inside Anthropic is called \"<strong>ant fooding</strong>\".</li>\n<li>Anthropic <strong>really believe in their <a href=\"https://code.claude.com/docs/en/auto-mode-config\">auto mode</a></strong>, and see that as an enabling technology for Claude Tag.</li>\n<li>Thariq advises offsetting coding-agent-induced <a href=\"https://simonwillison.net/2026/Feb/15/deep-blue/\">Deep Blue</a> by \"<strong>being more ambitious</strong>\" with the work you take on.</li>\n<li>Fable is <strong>competent at editing video</strong>, and Thariq <a href=\"https://twitter.com/trq212/status/2064826394589442448\">used it</a> to edit its own launch video.</li>\n<li>Anthropic's culture of working (internally) in public is key to their success, as demonstrated by the way they use Claude Tag in their public Slack Channels.</li>\n</ul>\n<h4 id=\"how-has-what-you-do-day-to-day-changed-in-the-past-year-\">How has what you do day-to-day changed in the past year?</h4>\n<p><a href=\"https://www.youtube.com/watch?v=uU5Gv2h8-9g&t=65s\">1:05</a></p>\n<blockquote>\n<p><strong>Simon:</strong> Claude Code came out in February of last year — it's under a year and a half old, and it was originally just a bullet point on <a href=\"https://www.anthropic.com/news/claude-3-7-sonnet\">the Claude Sonnet 3.7 launch</a>. <strong>How has what you do on a day-to-day basis changed in the past year</strong>, now that we have these coding agents that actually work for us?</p>\n<p><strong>Cat:</strong> I remember when we first came out with Claude Code and Sonnet 3.7, you would give it a task and you would have to closely monitor every single little thing it tried to do. I would read every permission prompt extremely carefully. I would frequently say no — no, no, no, did you check this file? Did you check that file? And now it's been incredible with every model generation. I feel like <strong>we've all gotten a chance to take a step back and delegate a lot more of the menial implementation to Claude</strong>. It's freed up a lot of our time to think about more creative work, like: what is the right experience that we should be providing to our users, now that we know Claude Code can implement a lot of it? And now with Fable it's a totally different step change improvement. <strong>We see for a lot of our use cases that you can actually one-shot a ton of features with Fable now</strong>.</p>\n<p><strong>Thariq:</strong> I remember the first text I got about Claude Code. One of my best friends was like, \"You need to go try Claude Code.\" It was about when Opus 4 came out, and I tried it and I was like, \"Oh, shit. I need to work at Anthropic now.\" And that was Opus 4 — great model, but you were reading permission prompts. It's kind of crazy how much amnesia we have, where I'm like, oh, auto mode has always been here, right? I don't even remember pressing yes and allow. For me, the big thing I'm trying to push myself on is that <strong>we have to do higher quality work than we've ever done before</strong>. The outputs are incredibly high quality. <strong>I've been using it to edit videos a bunch</strong>, and I'm like, okay, it has to meet the very exacting demands of our brand team in a couple of hours or we just can't do it. <strong>That's how I'm trying to shift with Fable: the best work we've ever done, faster than we've ever done it before</strong>.</p>\n</blockquote>\n<h4 id=\"what-piece-of-conventional-software-engineering-no-longer-holds-\">What piece of conventional software engineering no longer holds?</h4>\n<p><a href=\"https://www.youtube.com/watch?v=uU5Gv2h8-9g&t=219s\">3:39</a></p>\n<blockquote>\n<p><strong>Simon:</strong> What's a piece of conventional software engineering that was true a year ago that you don't think holds anymore in this new world?</p>\n<p><strong>Cat:</strong> One of the biggest shifts we're seeing in the eng skill set: two years ago it was pretty typical for a product manager to go talk to a bunch of customers, align over the course of six months with cross-functional teams on some PRD, and write a thorough spec on exactly how we'll implement this before the first line of code gets written. Now things are completely turned the opposite way. For a lot of engineers, the push I would give to folks in the room is to <strong>develop more of your business sense and product sense on what it is we should build</strong>, because the timeline between having an idea and building it is so much shorter — it's down from six to twelve months to maybe even a week. That means all of us need to have better taste on what is worth building, what will actually inflect the businesses we're working on. So it's <strong>an increase in value on product taste and business sense</strong>, and a bit lower on execution in most product domains. Of course, for infra there's still a very heavy emphasis on making sure all the details are right.</p>\n<p><strong>Thariq:</strong> For me, it's that <strong>rewrites are now good</strong>.</p>\n<p><strong>Simon:</strong> The worst thing you could do is now actually fine!</p>\n<p><strong>Thariq:</strong> Exactly. All the Mythical Man-Month stuff — never rewrite — I'm pro-rewriting now. If you have a good test suite — and <strong>I think the rewrite actually forces you to make sure you have a good test suite</strong> — but I think what people undercount is that <strong>a codebase is a spec, and maybe it's the only copy of the spec that you have</strong>, because no one knows every branching part of the codebase. You can take this as an artifact and distill it or create other versions of it. We <a href=\"https://bun.com/blog/bun-in-rust\">rewrote Bun in Rust</a> and it works great — it's live for me right now.</p>\n<p><strong>Simon:</strong> You're not shipping Claude Code on Bun-in-Rust yet, right?</p>\n<p><strong>Thariq:</strong> Internally we have.</p>\n</blockquote>\n<p><em>(Actually it looks like Anthropic started shipping Claude Code on Bun-in-Rust to everyone <a href=\"https://simonwillison.net/2026/Jul/19/claude-code-in-bun-in-rust/\">on June 17th</a>.)</em></p>\n<h4 id=\"what-kind-of-things-are-non-engineers-doing-with-claude-tag-\">What kind of things are non-engineers doing with Claude Tag?</h4>\n<p><a href=\"https://www.youtube.com/watch?v=uU5Gv2h8-9g&t=396s\">6:36</a></p>\n<blockquote>\n<p><strong>Simon:</strong> The other big launch recently was <strong><a href=\"https://www.anthropic.com/news/introducing-claude-tag\">Claude Tag</a></strong> — that's what, a week old now, at least for the rest of us. I understand it's being used at Anthropic by non-engineers a great deal. <strong>What kind of things are non-engineers doing with Claude Tag?</strong></p>\n<p><strong>Cat:</strong> Claude Tag is a Claude that lives in your team's collaboration tools. We launched it last week within Slack. <strong>The thing that's different about Claude Tag is it's multiplayer by default</strong>. Once you add Claude Tag to a Slack channel, you can chime in, your teammates can chime in, and you can collaborate together on the PR. The other big difference is that it's proactive instead of reactive. You can tell Claude Tag, \"Hey, monitor every bug report in this channel, put up a PR to fix it, and tag the engineer who last touched this part of the codebase,\" and it'll do it for the lifetime of the channel without you having to manually tag it in. And the third big shift is that <strong>we've <a href=\"https://claude.com/docs/claude-tag/users/memory\">added team memory</a> into this</strong>. If you tell Claude Tag your preferences in the channel, it'll remember them for every future post. If you always want it to debug outages but you don't want it to debug warnings, just tell it that in natural language in the channel and it'll remember it for you and everyone else on your team.</p>\n<p><strong>Internally, we see Claude Tag as the evolution of Claude Code.</strong> We see this as a large shift in how we work internally. <strong>Claude Tag currently lands 65% of our product eng PRs.</strong></p>\n<p><strong>Simon:</strong> For all of Anthropic, or just for Claude Code?</p>\n<p><strong>Cat:</strong> This is just for our product engineering team — <strong>our internal version of Claude Tag lands 65% of our product PRs right now</strong>. And this is a huge shift; this is more than 50% of our PRs. The way we see people split work between Claude Code and Claude Tag is: Claude Code is still the best place for your most complex tasks, when you're interactively iterating with the agent. <strong>But Claude Tag is great for having it work proactively on your behalf</strong>, so you no longer need to manually kick off Claude Code for all the bug reports that come up for features you're working on.</p>\n<p><strong>Thariq:</strong> And for non-coding cases: for example, before this talk we asked Claude Tag, \"Hey, when is Fable releasing?\" We wanted to make sure we'd line it up with the announcement. Claude Tag would search our Slack and look at who's been saying what. <strong>As a search engine for your company, it's really valuable.</strong> It has all the context for your product, so you can ask it metrics-related questions — often when you're making decisions you want them informed by what the metrics say, so you hook it up to your event store. I've seen our marketing team do things like, \"Hey, tell me about this feature.\" They're not programmers, but Claude is a programmer — it can clone the codebase and say, \"This is the feature, this is what it looks like, <strong>this is a recording of me using the feature</strong>.\" It enables a whole wide variety of things, and I think we're still early in figuring that out.</p>\n</blockquote>\n<h4 id=\"claude-tag-as-the-team-collaborative-layer\">Claude Tag as the team collaborative layer</h4>\n<p><a href=\"https://www.youtube.com/watch?v=uU5Gv2h8-9g&t=606s\">10:06</a></p>\n<blockquote>\n<p><strong>Simon:</strong> One of the problems I've had with coding agents is that I get how to use them as an individual, but I'm not really clear on how to use them in a team environment. <strong>It sounds like Claude Tag is your current answer to that team collaborative layer for this stuff.</strong></p>\n<p><strong>Cat:</strong> Exactly. And a large percentage of our sessions are actually multiplayer right now. Maybe I say, \"Hey, I think we should implement this new feature in Cowork,\" and I'll tag in Claude Tag to do a first pass at it. Then I'll tell Claude Tag, \"Share a recording of your final implementation,\" and I'll tag in design to take a look. They'll nudge it, then pass it on to eng to take it to the finish line and get it out to prod. It's been this very fluid experience. <strong>We're still trying to iron out what the social dynamics are for steering the same session</strong>, but we've found that people just observe how others use it and follow those social norms — it's been pretty intuitive for us to integrate Claude Tag into our teams.</p>\n<p><strong>Thariq:</strong> It's great for teaching people, and also for reducing slop, because <strong>the fact that everyone is seeing you use Claude together sort of levels up how you use Claude as well</strong>.</p>\n</blockquote>\n<p>This reminded me of how Midjourney solved the challenge of teaching people advanced image prompting by enforcing prompting in public in their Discord channels.</p>\n<h4 id=\"how-do-you-decide-which-features-are-worth-building-when-building-is-so-much-cheaper-\">How do you decide which features are worth building when building is so much cheaper?</h4>\n<p><a href=\"https://www.youtube.com/watch?v=uU5Gv2h8-9g&t=701s\">11:41</a></p>\n<p>Something I've found really hard myself is knowing when a feature is worth shipping now that the cost of actually building features has dropped so much.</p>\n<blockquote>\n<p><strong>Simon:</strong> How do you deal with the hardest problem in all of engineering — prioritization? <strong>How do you decide which features are worth building and shipping when building a feature is so much more inexpensive now?</strong></p>\n<p><strong>Cat:</strong> This is the hard thing. There are a few ways we approach it. One is we dogfood our products every single day. Whenever there's something we want to be able to do in our products that we're not able to, instead of finding a different solution we fix our product so it can support that case. <strong>We have a very heavy dogfooding culture internally.</strong> Before we share our products with everyone in the world, we share them with everyone within Anthropic, and with some early customers who give us very honest feedback about it — the more brutal the better — and we iterate until people love it. <strong>We have an internal bar for the number of active users and the amount of retention a feature has to have before we share it with the world.</strong> Because this bar is very clear, every engineer knows what they're trying to hit. I think this also levels up our polish, because if the feature isn't polished, people will churn — and then we shouldn't ship that feature.</p>\n</blockquote>\n<p>Using internal user-retention to decide if a feature should ship makes a whole lot of sense to me.</p>\n<h4 id=\"do-you-have-an-example-of-a-feature-which-surprised-you-\">Do you have an example of a feature which surprised you?</h4>\n<p><a href=\"https://www.youtube.com/watch?v=uU5Gv2h8-9g&t=774s\">12:54</a></p>\n<blockquote>\n<p><strong>Simon:</strong> <strong>Do you have an example of a feature which surprised you?</strong> You rolled it out and the engagement was off the charts — something unlikely to be shipped that turned into a real product thing.</p>\n<p><strong>Cat:</strong> I do have one. <strong>A lot of folks on our team love <a href=\"https://code.claude.com/docs/en/remote-control\">remote control</a>.</strong> Remote control lets you use your mobile device, or Claude in the web browser, to connect to a local Claude Code session running in your CLI. I never have this need, because I just kick off the task directly on mobile and it runs in a cloud session without using my local environment — I think because I'm doing very easy coding tasks. It was something I didn't totally understand; I was like, hey, people should just set up remote dev environments. But in practice, once we rolled out remote control, so many people I talk to told me that what they do every night is plug their laptop into a power charger, open a bunch of remote control sessions, lock the screen, <strong>and then use their mobile phone from their couch to control Claude Code</strong>. So this has become a flow we're now leaning into that I didn't originally get — but now I do.</p>\n</blockquote>\n<h4 id=\"does-a-human-review-every-line-of-production-code-in-claude-code-\">Does a human review every line of production code in Claude Code?</h4>\n<p><a href=\"https://www.youtube.com/watch?v=uU5Gv2h8-9g&t=860s\">14:20</a></p>\n<p>One of the over-arching themes of the conference was review: how much attention to people spend to reviewing code written for them by coding agents. I was very keen to hear the Claude Code team's take on this!</p>\n<blockquote>\n<p><strong>Simon:</strong> How does code review work? <strong>Does a human being review every line of production code that makes it into Claude Code?</strong> And if not, what are you doing — how do you keep the quality up?</p>\n<p><strong>Thariq:</strong> It varies on the task a lot. <strong>For important areas we have code owners.</strong> The system prompt is an example where we have a code owner — you really need to get their approval.</p>\n<p><strong>Simon:</strong> So the code owner is directly responsible for the quality of that area of the code.</p>\n<p><strong>Thariq:</strong> That's right.</p>\n<p><strong>Cat:</strong> And they need to approve any PR that touches it.</p>\n<p><strong>Thariq:</strong> We have <a href=\"https://code.claude.com/docs/en/github-actions\">our code review GitHub bot</a> review everything — that goes on every PR, and often it's doing the bulk of the review. Something I've seen on the team is that <strong>for more complex PRs you might make an artifact to explain the PR</strong> so that other people can then review. And we invest a lot into verification, CI/CD, things like that, to make sure that any time anything fails we have a test. We have a really robust environment where Claude can control Claude Code and test it. So there's a multi-pronged approach to code review.</p>\n<p><strong>Cat:</strong> In general, <strong>we are trying to move to a world where humans don't need to be in the loop</strong>. For the most critical changes to the core of Claude Code, and the cores of other products, there is always a code owner and they do manually review all the changes. But increasingly, <strong>for the changes at the outer layers, we actually have Claude code review fully review those</strong>. That sounds pretty scary, but we've had a six-plus-month-long process to get here, and <strong>there are baby steps that you take to build up trust with code review</strong>. In the beginning we had human review for everything, and then increasingly we would say, <strong>okay, for code changes that touch these files, code review is catching 100% of the issues there — so we actually don't need a human manually reviewing those</strong>. And when we have incident review, <strong>we look at the PRs that caused the incident and say, okay, how do we update code review to catch that?</strong> — and we take those PRs and <strong>add them to an eval set</strong> to make sure our future changes to code review never regress that metric. Removing humans from the code review loop is a big step forward. It can sound scary, and it's not something you can do overnight, but it is something you can do <strong>through many months of investment in the infrastructure</strong> to give you the confidence that code review is catching everything you care about.</p>\n</blockquote>\n<p>So the key seems to be constantly iterating on the automated review systems themselves, in order to build trust in them over time.</p>\n<h4 id=\"how-does-a-new-model-affect-your-intuition-for-what-it-can-and-can-t-do-\">How does a new model affect your intuition for what it can and can't do?</h4>\n<p><a href=\"https://www.youtube.com/watch?v=uU5Gv2h8-9g&t=1040s\">17:20</a></p>\n<p>We got <em>deep</em> into evals - another hot topic throughout the wider conference.</p>\n<blockquote>\n<p><strong>Simon:</strong> I know that Opus 4.8, if I ask it to build me a JSON endpoint that runs a SQL query and outputs JSON, is just going to get it right — that's not something I have to review closely. But then a new model comes along and I don't know how to build trust in Fable quickly, that it's not going to mess things up that Opus didn't. <strong>How does the new model affect your intuition for what it can do and what it can't do?</strong></p>\n<p><strong>Cat:</strong> The main reason we're building up this <strong>eval base over time is so that new models can be a drop-in replacement</strong>. When we have a new model, we run the whole eval set and make sure that, for example, Fable is strictly better than Opus 4.8 — and that gives us the confidence to drop it in.</p>\n<p><strong>Simon:</strong> Are those model evals for Anthropic as a whole, or Claude Code team-specific?</p>\n<p><strong>Cat:</strong> We have both. We have evals on our team, and we run code review across every repo within Anthropic, so we have evals for that. And for things like auto mode, we not only have evals across every user within Anthropic — we've also commissioned multiple external testers to red team it, to create environments with prompt injections and malicious inputs, <strong>and make sure that auto mode doesn't let any of those pass</strong>.</p>\n</blockquote>\n<h4 id=\"how-do-you-build-confidence-that-a-system-prompt-tweak-results-in-better-output-\">How do you build confidence that a system prompt tweak results in better output?</h4>\n<p><a href=\"https://www.youtube.com/watch?v=uU5Gv2h8-9g&t=1121s\">18:41</a></p>\n<blockquote>\n<p><strong>Simon:</strong> I want to know if the system prompt improvement I made actually improved the product — that's the most basic form of product-specific eval, and I still don't have a great feel for how to do that. <strong>Is that something you're doing such that you have complete confidence that a tweak you've made to the system prompt results in better output?</strong></p>\n<p><strong>Cat:</strong> <strong>We don't have complete confidence, but we do a lot to make sure that we don't regress performance.</strong> The starting point is a suite of external evals that we trust, and we complement that with an even larger suite of internal evals that we trust. To start, <strong>we mainly optimize for capability</strong>: given a complete definition of a task and the full codebase, does Claude make the right decisions, fully fix the bugs, and pass all the tests? That's the starting point and the thing we optimize for, because it's most directly what users want. But there are a lot of behaviors that impact how users feel when they work with Claude Code. For example, <strong>people really don't like it when Claude Code says it's time to go to sleep.</strong> Or people really don't like it when it says, \"Hey, I finished two out of five parts — do you want me to continue?\" Yes, please continue. <strong>So we're building up a set of behavioral evals to catch these.</strong> And as we get user feedback — please be loud with us about your user feedback — we rank the priority issues and go down one by one and build evals for each of them. It's not 100% coverage, but it is a priority for us to increase the coverage.</p>\n</blockquote>\n<h4 id=\"how-much-interaction-is-there-between-the-claude-code-team-and-the-model-training-teams-\">How much interaction is there between the Claude Code team and the model training teams?</h4>\n<p><a href=\"https://www.youtube.com/watch?v=uU5Gv2h8-9g&t=1221s\">20:21</a></p>\n<blockquote>\n<p><strong>Simon:</strong> <strong>How much interaction is there between the Claude Code team and the teams at Anthropic who are training the models in the first place?</strong> Is that quite a close collaboration?</p>\n<p><strong>Cat:</strong> Across Anthropic, we all work quite closely together. We meet often to talk about what we expect the next generation of models to be able to do. Our research team has also been amazing about showing this publicly — we often talk in our blog posts about how <strong>we're targeting ever-increasing longer-horizon work</strong>, and how we train Claude itself to be honest, harmless, and helpful. We also put a lot of effort into making sure it's aligned with your intent, even if your intent is expressed in a fuzzy way. Of course, try your best to be specific about what you want, so Claude has all the context — but even when you're not specific, we teach Claude to make good assumptions. It's been a productive partnership.</p>\n</blockquote>\n<h4 id=\"the-system-prompt-has-been-reduced-by-80-what-have-you-been-able-to-drop-\">The system prompt has been reduced by 80% — what have you been able to drop?</h4>\n<p><a href=\"https://www.youtube.com/watch?v=uU5Gv2h8-9g&t=1284s\">21:24</a></p>\n<p>So many useful prompting tips in this section!</p>\n<blockquote>\n<p><strong>Simon:</strong> Thariq, you <a href=\"https://www.youtube.com/watch?v=9fubhllmsBU&t=358s\">mentioned this morning</a> that the <strong>system prompt for Claude Code has been reduced by 80% because of Claude Fable</strong>. Can you go into a little more detail? <strong>What kind of things have you been able to drop?</strong></p>\n<p><strong>Thariq:</strong> It wasn't just Fable — it was Opus 4.8 as well, and going forward, future models. We have different system prompts for different models now. One of the patterns we saw is that we were over-constraining Claude. The initial, maybe Opus 4-ish models wanted a lot of examples, and <strong>removing examples was extremely helpful</strong>, because it was just more creative than the examples we gave it.</p>\n<p><strong>Simon:</strong> That's really interesting, because one of the top prompting tips I give people is: give it examples. If that's no longer true, that kind of breaks my prompting model a little bit.</p>\n<p><strong>Thariq:</strong> Same here — I was surprised to hear that. I think now it's more about the shape of what you give it — the tools you give to Claude, your system prompt, things like that. The other thing we did is try to give it more context and <strong>fewer \"do not do this\"</strong> instructions, because that's a very strong impulse for Claude, and especially if it conflicts with user instructions later on, that can be extremely confusing to Claude — \"I've got this skill that says this and the system prompt says this.\" So we try to <strong>have fewer hard constraints, more context, and fewer instructions overall</strong>. It's definitely a science — it took a bunch of evals to build.</p>\n<p><strong>Cat:</strong> In general, when you're prompting these models, you should always think: <strong>are there edge cases to the instruction that I'm giving it?</strong> When we went back and reviewed all the instructions in the Claude Code system prompt, <strong>we found a few cases where yes, this statement is 90% true, but there's a real 10% of cases where it's not true</strong>. We didn't want to constrain the model, or confuse it into thinking it should always do this. One good example is verification. Everyone here wants Claude to verify its work, and we had some instructions in the prompt that said: if you make a front-end change, always verify. But there's a limit to it. If it's changing copy from one string to another string, and the user says \"just make a quick fix and update the test,\" maybe you don't want to verify. <strong>So we've adjusted our wording from \"always verify, verify, verify\" to something like: most of the time when you're doing front-end work you can't fully understand the experience by hitting the backend endpoints, so when you make larger changes to the user experience, please run the app locally.</strong> And in fact, that instruction probably isn't even good either, because <strong>what is a large change?</strong> Maybe it should test small changes too. In general, whenever you give a prompt to the model, <strong>you should think about the ways in which it could be misinterpreted by a well-intentioned human</strong>, in order to better understand how the model might interpret it — and <strong>soften the prompt</strong> so that it's actually 100% accurate, because you're giving this prompt to the model 100% of the time.</p>\n<p><strong>Simon:</strong> What's fascinating about that is you're <strong>relying on the model's judgment</strong> — and that's got to be an Opus/Fable-level thing. Models a year ago did not have the level of judgment necessary to decide whether they were going to test a change or not. But that does break down if you're building for a wide range of models and trying to run the cheaper models for cheaper tasks.</p>\n<p><strong>Cat:</strong> We actually have <strong>a different system prompt per model now</strong>, for this very reason. It's only our most frontier models that have this 80% token decrease — the older models still have the full system prompt.</p>\n<p><strong>Simon:</strong> Do you think Fable and Opus are smart enough to prompt Haiku with more details, because they understand that Haiku has less judgment, less taste?</p>\n<p><strong>Cat:</strong> We haven't been able to eval it — we don't have any hard data to show it.</p>\n<p><strong>Thariq:</strong> There's a tough thing with smaller models sometimes, because <strong>sometimes the larger models can be more token-efficient on a hard problem than the smaller models</strong>. So there's a bit of intuition to build there — sometimes you really just want frontier intelligence almost all the time. The Pareto curve shifts, and it's hard to find.</p>\n<p><strong>Simon:</strong> A year ago I did not trust a model to write a prompt. Today the good models are very good at prompting — a lot of my prompts are written by models, which feels absurd but works really well. What helped me come to terms with that was thinking about subagents, which are entirely about a Claude model setting up a prompt for another Claude model.</p>\n<p><strong>Thariq:</strong> <strong>Workflows</strong> are actually a really good example of this, because it's Claude not just prompting a single subagent, but prompting the orchestration of many subagents, and each one of them gets a very detailed prompt. It's almost a level above just spawning a subagent. I've also been using it on my personal machine, <strong>giving it the Gemini API and saying: here, generate images</strong>. It's way less lazy than I am at prompting an image model. It's just Claude prompting Claude all the way down.</p>\n<p><strong>Cat:</strong> I think Claude also wrote the prompt for <a href=\"https://code.claude.com/docs/en/workflows\">the workflow tool</a>.</p>\n<p><strong>Simon:</strong> I've read that prompt — it's a good prompt. That's actually a frustration I have with Anthropic generally: you <a href=\"https://platform.claude.com/docs/en/release-notes/system-prompts\">publish the prompts for Claude Chat</a>, but you don't include the tool prompts and the Claude Code prompts. I still have to run a proxy to intercept them. <strong>I would love it if the Claude Code prompts were deliberately published</strong> — they're the documentation. They're how you know what the tool can do and how it works.</p>\n<p><strong>Cat:</strong> I'll write down that feature request. I'll have Claude Tag do it.</p>\n</blockquote>\n\n<p>Interesting to note that OpenAI's <a href=\"https://developers.openai.com/api/docs/guides/latest-model?model=gpt-5.6#favor-leaner-prompts\">prompting best practices for GPT-5.6</a> includes similar advice for their latest models:</p>\n<blockquote>\n<p><strong>Favor leaner prompts</strong></p>\n<p>Removing repeated instructions and examples and simplifying tool descriptions can improve task performance and token efficiency. In a sample of internal coding-agent eval runs, configurations with leaner system prompts improved evaluation scores by roughly 10–15% while reducing total tokens by 41–66% and cost by 33–67%.</p>\n</blockquote>\n\n<h4 id=\"what-s-your-bar-for-introducing-a-new-tool-\">What's your bar for introducing a new tool?</h4>\n<p><a href=\"https://www.youtube.com/watch?v=uU5Gv2h8-9g&t=1686s\">28:06</a></p>\n<blockquote>\n<p><strong>Simon:</strong> Claude Code is basically a big bag of tools. <strong>What's your bar for introducing a new tool?</strong> How do you decide when it's worth doing that additional engineering at that level?</p>\n<p><strong>Cat:</strong> Do you want to take it? You introduced one of the best tools we have.</p>\n<p><strong>Thariq:</strong> My career peaked when I introduced the ask user question tool. It's really hard. Especially for some tools — <strong>ask user question is Claude's tool to ask you</strong> — so it's hard to eval, and sometimes it's more of a user preference thing. Back then we had fewer evals, so it was very dogfooding based — or \"ant fooding,\" our ant version of that. But overall <strong>we've been trying to trend towards fewer tools</strong>. The last set of tools we introduced was the task tool, I think — and we try to give Claude more general versions to do things.</p>\n</blockquote>\n<h4 id=\"what-s-the-latest-evolution-of-your-file-editing-tool-\">What's the latest evolution of your file editing tool?</h4>\n<p><a href=\"https://www.youtube.com/watch?v=uU5Gv2h8-9g&t=1743s\">29:03</a></p>\n<p>I have a long-running fascination with file editing tools - they were the subject of the <a href=\"https://aider.chat/docs/leaderboards/edit.html\">old Aider code editing leaderboard</a>, and I've watched with interest as they've evolved in different coding agents from search-and-replace based to line-number-based to more complicated patterns.</p>\n<p>The Claude API docs describe a <a href=\"https://platform.claude.com/docs/en/agents-and-tools/tool-use/text-editor-tool\">text editing tool</a> that's recommended for building against the API, but Claude Code seems to use slightly different approaches here.</p>\n<blockquote>\n<p><strong>Simon:</strong> One of the most interesting tools is the file editing tool — you can have file editing as a tool, or you can tell it to use sed and grep and do things that way. <strong>What's the latest evolution of your file editing tool?</strong></p>\n<p><strong>Thariq:</strong> We still have one, but for example we removed our grep and other search tools — glob tools — in favor of native bash. Like I said in my talk earlier, <strong>the models are kind of more of a biology than a physics</strong>, and tool design especially is quite hard. I'm not sure if Cat disagrees and thinks there's a science to the eval of it, but I think tool design is more of an art, maybe — or a biology.</p>\n<p><strong>Cat:</strong> I largely agree, but in general as we introduce more tools, we try to keep the cardinality pretty low and make sure that <strong>every tool we add has a distinct function from every other tool, so that Claude can very easily distinguish when to call each</strong>. For file edit, the reason we have it is actually because we can render it. We show people when Claude makes a file change, and there's this <strong>nice dedicated UI</strong> that says: do you approve this edit to this file? <strong>The reason we had a dedicated file edit tool was so that we could deterministically know</strong> that Claude was making a file change, so we could show people this nice UI. A lot of new users onboarding still really like this experience, so we've kept it around. But for a lot of us who are on auto mode right now — hopefully you're not on YOLO mode — I don't think it actually matters, and we could probably just remove file edit and be totally fine.</p>\n</blockquote>\n<h4 id=\"what-s-the-advice-within-anthropic-for-safely-running-claude-code-\">What's the advice within Anthropic for safely running Claude Code?</h4>\n<p><a href=\"https://www.youtube.com/watch?v=uU5Gv2h8-9g&t=1858s\">30:58</a></p>\n<p>It's the <a href=\"https://simonwillison.net/tags/prompt-injection/\">prompt injection</a> question! Who better than Anthropic employees to explain how Anthropic sees the risk of prompt injection attacks causing their Claude Code instances to run amok?</p>\n<p>It turns out they <em>really</em> trust their <a href=\"https://code.claude.com/docs/en/auto-mode-config\">auto mode</a> - and see that as the feature that enabled Claude Tag.</p>\n<blockquote>\n<p><strong>Simon:</strong> Let's talk about safety and security. I am deeply aware of the risks of prompt injection, and there are so many bad things that can happen if somebody else tells my Claude Code what to do. I still mostly run Claude Code in YOLO mode and feel incredibly guilty about it. <strong>What's the advice within Anthropic for safely running Claude Code?</strong></p>\n<p><strong>Cat:</strong> Why not auto mode?</p>\n<p><strong>Simon:</strong> I am starting to use auto mode, but I don't understand it enough to get how safe it is. As of maybe three weeks ago, I'm defaulting to auto mode.</p>\n<p><strong>Cat:</strong> Broadly within Anthropic, almost every single person uses auto mode. It is the best way to do long-running work in Claude Code while being safe. <strong>We've done extensive bashing. We have thousands of evals. We've commissioned many red teamers to create adversarial environments in order to trick Claude Code into doing bad actions, and we've mitigated every single issue that they found.</strong> We're going to publish some evals in the coming weeks, but we've pretty much mitigated every attack.</p>\n<p><strong>Simon:</strong> That is a big claim.</p>\n<p><strong>Cat:</strong> We'll share the evals for it so folks can assess, but we've been extremely diligent about identifying all the ways in which Claude might mess up and then updating auto mode to counter it. It doesn't catch 100% of things — that would be way too strong a claim. But <strong>for the main categories of risks that we're concerned about, like prompt injection and data exfiltration, the risks are far lower than the average human reviewer</strong>.</p>\n</blockquote>\n<p>I am very much looking forward to learning more about their evals and approach to verifying auto mode.</p>\n<blockquote>\n<p><strong>Thariq:</strong> A little on how auto mode works — it's useful to build this mental model. Whenever Claude is doing a turn, or a bash call, there's <strong>a Sonnet classifier</strong> that is judging the tool call and also the context of the conversation — your instruction. There are some things around permissions that are dependent on your request: you don't want to give git push permissions all the time, but if you say \"push this to GitHub,\" you want it to do it — and if you say \"don't push,\" you want it to deny it. Auto mode will do that. That particular thing happens to me a lot, where Claude tried to do something because it's very helpful and proactive, and auto mode saw \"don't do this\" and surfaced it. <strong>So it's good at the dynamic permissions</strong> that you yourself give inside the prompt, which I think is really important. It also works well with our <a href=\"https://code.claude.com/docs/en/sandbox-environments#sandboxed-bash-tool\">sandboxing infrastructure</a>, because sandboxing is one of those things where there are so many different edge cases that it's hard for us to deterministically follow them. <strong>We have a sandbox, and when something needs to escape the sandbox</strong> — like a network request — auto mode can look at that request and ask: does this make sense? — and allow it.</p>\n<p><strong>Simon:</strong> I hadn't realized auto mode is interacting with the networking sandbox as well.</p>\n<p><strong>Cat:</strong> It interacts with any permission prompt the user would otherwise see.</p>\n<p><strong>Simon:</strong> How old is auto mode? As a feature I had access to, it's only a couple of months old, right?</p>\n</blockquote>\n<p>(It was first made available to the public <a href=\"https://claude.com/blog/auto-mode\">on March 24th</a>.)</p>\n<blockquote>\n<p><strong>Cat:</strong> We've been using it within Anthropic <strong>since January</strong>, so we've been hardening it for quite a while. Anthropic is extremely focused on safety and security, and we've been working broadly across our alignment and safeguards teams to enable the rollout internally, build out these evals, and make auto mode even more robust before sharing it with the world.</p>\n<p><strong>Thariq:</strong> This is also the reason Claude Tag is so good — <strong>Claude Tag uses auto mode</strong>. I've heard a lot of build-versus-buy questions about a Slackbot, and I'm like: please, you probably shouldn't build your own AI Slackbot. There are so many attack vectors. <strong>You have a feedback channel that users can post feedback into, and now your bot is reading it.</strong> The work we've put in with auto mode — and we have a general <strong>Swiss cheese defense</strong> for security; we also RL against this stuff — <strong>I think this is really what makes Claude Tag work</strong>. It works seamlessly with your permissions, and you don't want to be prompt injected in your Slack.</p>\n</blockquote>\n<h4 id=\"are-there-more-security-things-in-the-pipeline-beyond-auto-mode-\">Are there more security things in the pipeline beyond auto mode?</h4>\n<p><a href=\"https://www.youtube.com/watch?v=uU5Gv2h8-9g&t=2154s\">35:54</a></p>\n<blockquote>\n<p><strong>Simon:</strong> Are there any more security things in the pipeline that go beyond auto mode?</p>\n<p><strong>Thariq:</strong> I think we're very secure. <strong>With Claude Tag you can provision your own credentials for Claude</strong>, so it doesn't need to act on your behalf — you can have Claude as an identity, and that also makes it easier to audit and inspect what Claude is doing.</p>\n<p><strong>Simon:</strong> Because Claude Tag is influenced by anyone who can talk to it — it's got a much wider pool of people telling it what to do.</p>\n<p><strong>Thariq:</strong> That's right. And of course we have probes as well with Fable, which is a downstream effect of our safety and research work. I think this is the moment where you see Anthropic being an AI safety company really paying off: <strong>we really want Claude to be able to run in an aligned way over long periods of time</strong>, and <strong>auto mode has to be basically flawless for this to work</strong> — it's all downstream of our being an AI safety company.</p>\n<p><strong>Cat:</strong> We also launched trusted devices for the remote control users out there who want to be safer. And for all of our remote environments, we support <strong>credential injection</strong>. If you want Claude Code to be able to access Datadog, but you don't want Claude Code itself to hold the Datadog credential, you can set up our identity and credential management system <strong>so that the Datadog credentials are only usable by the agent but not accessible by the agent</strong> — we insert them on the fly when the agent tries to make a Datadog request.</p>\n</blockquote>\n<p>I really like that credential injection pattern, where Claude Code can access an API via a proxy and that proxy both audits the request and injects the relevant API key - so Claude can access authenticated endpoints without having access to the API credentials itself.</p>\n<h4 id=\"how-has-the-past-year-and-a-half-changed-how-you-think-about-your-own-craft-\">How has the past year and a half changed how you think about your own craft?</h4>\n<p><a href=\"https://www.youtube.com/watch?v=uU5Gv2h8-9g&t=2273s\">37:53</a></p>\n<p>Thariq <a href=\"https://www.youtube.com/watch?v=9fubhllmsBU&t=867s\">talked about a sense of grief</a> brought on by Fable-class models in his keynote in the morning, and we dived further into that as part of our conversation. I've been calling this <a href=\"https://simonwillison.net/2026/Feb/15/deep-blue/\">Deep Blue</a>.</p>\n<blockquote>\n<p><strong>Simon:</strong> <strong>Let's talk a little bit about the human element.</strong> <strong>A lot of people are feeling a sense of loss now that so much of what they considered to be their role in building software is being subsumed by the models.</strong> How do you think about that? <strong>How has the past year and a half changed the way you think about your own craft and the value that you add?</strong></p>\n<p><strong>Thariq:</strong> Cat and Boris are such good reminders that you have to be more ambitious. They're always like: we're growing so fast, we have to be on the edge, we have to do the best work we can. That's a constant reminder for me — any time I'm slow on something, I'm like, okay, can I do it faster? Can I be more ambitious here? And oftentimes the answer is Claude, because Claude is getting better as you go — the last time I tried this, it was with the previous model. On your point about loss: I think this is real. <strong>If you're only trying to do the same work you were doing before LLMs, and now it's a prompt, it is, I think, kind of a sad feeling.</strong> And <strong>the way you offset that is by being more ambitious.</strong> I think Jared is such a good example — he hand-wrote all of the Zig code in his Oakland apartment in about a year, barely left his house, and had so much fun doing that. Now I see him rewrite all of Bun into Rust and <strong>he's having so much fun doing that</strong> — it's so much more ambitious, and that's how he offsets it. Generally it's asking <strong>how do I do the bigger thing</strong> and do more — <strong>I think success is fun</strong>. It's changing your ambition.</p>\n</blockquote>\n<p>\"The way you offset that is by being more ambitious\" neatly captures where I've landed on this issue myself as well.</p>\n<blockquote>\n<p><strong>Simon:</strong> And Cat, what does that look like from a product management perspective?</p>\n<p><strong>Cat:</strong> I feel like the product role just changes every single month. <strong>All the PMs on our team are this mix of engineer, designer, PM</strong> — most of them actually used to be full-time engineers. For us it really means <strong>plugging in whenever there's any kind of gap</strong>. If we have an idea and we didn't inspire any engineer to go build it, then we should just build it, put it into a notebook, and inspire people to take it to production. If the designs look a little off, <strong>let's take a page that's similar, do a first-pass design, and tag in someone who's very detail-oriented to fill in the gaps</strong>. Or if we notice that our team and product adoption is bigger within the company, and more people need to know what's coming down the pipe for Claude Code, Claude Tag, and Cowork — let's automate figuring out our whole launch calendar, <strong>let's automate getting those status updates asynchronously</strong> so we're not bugging people, and make sure our updates in our internal announce channels are fully detailed and to the point. For us it's very much understanding <strong>what the gap is right now between a great idea and getting something to our customers</strong>, and <strong>how do we automate it as much as possible</strong>.</p>\n</blockquote>\n<p>This reflects something I've noticed: when you can produce code so much faster, time spent blocked awaiting a decision from someone else becomes a much more notable bottleneck. Engineers who can make product decisions can move a whole lot faster, and the cost of getting one of those decisions wrong is much less prohibitive.</p>\n<h4 id=\"what-s-a-moment-when-claude-has-surprised-you-\">What's a moment when Claude has surprised you?</h4>\n<p><a href=\"https://www.youtube.com/watch?v=uU5Gv2h8-9g&t=2510s\">41:50</a></p>\n<blockquote>\n<p><strong>Simon:</strong> <strong>What's a moment when Claude has surprised you?</strong> When the model did something you didn't think it would be able to do?</p>\n<p><strong>Thariq:</strong> I've posted a lot about Claude video editing, but most recently I gave a talk at the ACM Agentic conference, and I asked, \"Hey guys, do you have the edited video? I'd love to post it and share it with my comms team.\" They said, \"Oh, it's taking so long.\" So I asked for the raw files. They sent me the video of me talking on stage, the video of the deck, and the audio file, and said, \"Good luck.\" I gave this to Claude, along with my HTML deck, and said, \"<strong>Hey, can you just edit this together?</strong>\" And what it does is honestly incredible — I'm ready to ship it. It transcribes the entire video. It notices that sometimes the video of my deck is a little weird — there's a popup of an auto-update in the middle — and it goes, \"<strong>Oh, I probably shouldn't use the video of your deck. What I'm going to do is slice it up, figure out which slide you're on, and use the HTML source instead.</strong>\" So it displays the HTML source. Then it's got video of me, but I'm only taking up a small part of the stage, so <strong>it's cropping dynamically to where I am on the stage</strong> — and I'm pacing, so it's tracking me as I pace. And it's transcribing what I'm saying.</p>\n<p><strong>Simon:</strong> This was Fable, right?</p>\n<p><strong>Thariq:</strong> This was Fable, yeah. It was a good prompt, but it was a one-shot prompt. Then I asked it to add some interesting animations and graphics, and I was just blown away. <strong>It does ffmpeg, it does Remotion.</strong></p>\n</blockquote>\n<p>Here's Thariq's video <a href=\"https://twitter.com/trq212/status/2064826394589442448\">on how he used Fable to edit Fable's own launch video</a>, and here's <a href=\"https://twitter.com/ClaudeDevs/status/2064399512664526853\">that launch video</a>.</p>\n<h4 id=\"what-can-t-it-do-yet-\">What can't it do yet?</h4>\n<p><a href=\"https://www.youtube.com/watch?v=uU5Gv2h8-9g&t=2616s\">43:36</a></p>\n<p>I'm embarrased to admit that I've been finding it quite hard to come up with tasks that frontier models like Fable 5 and GPT-5.6 are unable to accomplish.</p>\n<p>Cat still doesn't rate its UX design skills:</p>\n<blockquote>\n<p><strong>Simon:</strong> What can't it do? What are the things where you're still disappointed — where you're waiting for Claude Fable 6 to figure it out for you?</p>\n<p><strong>Cat:</strong> I want it to have better design and UX taste. It's now at the point where if I write out a prompt with a detailed spec of how I want a feature to behave, it will usually behave that way. But the paddings might be off, or the interface just isn't delightful yet. It leans on existing best practices for how apps are designed, but <strong>for frontier AI products, there are so many new interaction experiences that we have yet to design</strong>.</p>\n<p><strong>Simon:</strong> There's an Opus aesthetic — you can look at something and go, \"Yeah, that was designed by Opus.\" It'd be good if we could move beyond that.</p>\n<p><strong>Cat:</strong> Yeah. I'm very excited for future models to hopefully be <strong>interaction design thought partners</strong>.</p>\n<p><strong>Thariq:</strong> What can't it do? I would love to see it interact more with the real world. Can it solve science? Can it orchestrate the experiments? There's some amount of coding that goes into that, but there's also this other taste of the broader world that it needs.</p>\n</blockquote>\n<h4 id=\"which-parts-of-anthropic-s-culture-should-other-companies-steal-\">Which parts of Anthropic's culture should other companies steal?</h4>\n<p><a href=\"https://www.youtube.com/watch?v=uU5Gv2h8-9g&t=2711s\">45:11</a></p>\n<p>I figured this would make a great closing question:</p>\n<blockquote>\n<p><strong>Simon:</strong> <strong>Which parts of Anthropic's company culture do you think uniquely help Anthropic be productive with these tools, that other companies should steal?</strong> What are the cultural hacks people should be adopting from you?</p>\n<p><strong>Cat:</strong> I'll share one for Claude Tag. <strong>Claude Tag works best when you have it in a public channel, and when most of your channels are public.</strong> Claude Tag is able to search across all public channels to get as much context as possible to give you the highest-accuracy answer — and <strong>it's only able to do this if it has access to everything</strong>.</p>\n<p><strong>Thariq:</strong> I mentioned this in my keynote, but it's so important to me I want to re-emphasize it. The co-founders <strong>say we don't negotiate against ourselves</strong>, and I think this is really important. <strong>You can imagine trade-offs in your head and talk yourself out of doing something ambitious — or you can just try to do the ambitious thing.</strong> We're so often asking: what if we just did it? Is this a real trade-off or not? And if so, why — where's the proof that it's a real trade-off, and not just something that sounds reasonable? <strong>Make the trade-offs show themselves to you. Be as ambitious as you can.</strong></p>\n</blockquote>\n<h4 id=\"what-s-your-favorite-absurd-thing-you-ve-built-with-claude-just-because-you-could-\">What's your favorite absurd thing you've built with Claude, just because you could?</h4>\n<p><a href=\"https://www.youtube.com/watch?v=uU5Gv2h8-9g&t=2806s\">46:46</a></p>\n<p>I couldn't resist throwing in this one as well.</p>\n<blockquote>\n<p><strong>Simon:</strong> <strong>What's one of your favorite absurd things that you've built with Claude, just because you could build it?</strong></p>\n<p><strong>Thariq:</strong> I'm working on <strong>a 2D Street Fighter fighting game with me as a character</strong> — and my friends as well. It uses Claude Code to prompt Gemini — and honestly the Seedance model is pretty good — to make video animations. It works great; it's so good at prompting, and it can verify the frames to check whether an animation was good.</p>\n<p><strong>Simon:</strong> Is this Street Fighter 2-level 2D sprites you're generating?</p>\n<p><strong>Thariq:</strong> Yeah, exactly — 2D sprites. The animation looks amazing. And it can also figure out hitboxes — it can be like, \"Oh, your fist is here, I'll draw the JSON hitbox.\" It's incredible.</p>\n<p><strong>Cat:</strong> Mine is much more simple. I'm a big rock climber and a lot of my friends climb, so we have this little app we built with Claude Code where we log all the projects we're working on. We also go outdoors together a lot, so we have Claude do all this research with workflows. Workflows is amazing — we brand it as a coding tool, but it's amazing for doing deep research for travel. I also plan our team offsites, and it's good at finding venues that can fit all of us. I use workflows to research all the climbing destinations we might want to go to, and what has direct flights from where all of us are located. It goes to Mountain Project and finds all the climbs at our grade level. It finds the Airbnb. And I don't like hiking, so I care a lot about it having a very short approach — <strong>very short walking distance from where the car parks to where the rock actually is</strong> — and it filters for this. With existing apps I have to manually click through Mountain Project, but with this I just put in all of our preferences and it's a custom app for us.</p>\n<p><strong>Simon:</strong> So you're basically vibe coding Jira for mountain climbing.</p>\n<p><strong>Cat:</strong> Exactly.</p>\n</blockquote>\n<h4 id=\"audience-any-plans-for-eval-building-tools-and-agent-observability-\">Audience: Any plans for eval-building tools and agent observability?</h4>\n<p><a href=\"https://www.youtube.com/watch?v=uU5Gv2h8-9g&t=2963s\">49:23</a></p>\n<p>We had a few minutes at the end for questions from the audience.</p>\n<blockquote>\n<p><strong>Audience:</strong> Do you have any near-term plans to build more eval tools for us to build eval datasets, and more observability tools to monitor the performance of agents and workflows?</p>\n<p><strong>Cat:</strong> We've considered building eval tools, but I think the limiting factor actually tends to be that <strong>it takes a long time for customers to build really high-quality evals</strong>. So I think the tooling is less of the constraint, and more the skill set of how you build a great eval. That's an area where we're excited to both invest internally and hopefully share some best practices externally.</p>\n</blockquote>\n<h4 id=\"audience-how-is-memory-designed-today-and-would-you-move-from-files-to-a-data-store-\">Audience: How is memory designed today — and would you move from files to a data store?</h4>\n<p><a href=\"https://www.youtube.com/watch?v=uU5Gv2h8-9g&t=3008s\">50:08</a></p>\n<blockquote>\n<p><strong>Audience (Sai):</strong> I'm interested in the memory and the multiplayer. <strong>How is memory being designed today?</strong> I assume it's around files. And second, have you thought about an orthogonal direction where you <strong>would actually need a data store for these memories, instead of files, to scale it better?</strong></p>\n<p><strong>Thariq:</strong> Right now for Claude Tag the memory is channel-specific. Every Claude in that channel has a shared memory, and the instances have a session — but the session can contribute back to main memory. We do a lot of memory research, and it can be kind of unintuitive what the right way to do memory is. We're always running memory experiments. <strong>How it works right now in Claude Tag is a markdown file per channel.</strong></p>\n</blockquote>\n \n <p>Tags: <a href=\"https://simonwillison.net/tags/ai\">ai</a>, <a href=\"https://simonwillison.net/tags/prompt-engineering\">prompt-engineering</a>, <a href=\"https://simonwillison.net/tags/generative-ai\">generative-ai</a>, <a href=\"https://simonwillison.net/tags/llms\">llms</a>, <a href=\"https://simonwillison.net/tags/anthropic\">anthropic</a>, <a href=\"https://simonwillison.net/tags/annotated-talks\">annotated-talks</a>, <a href=\"https://simonwillison.net/tags/coding-agents\">coding-agents</a>, <a href=\"https://simonwillison.net/tags/claude-code\">claude-code</a>, <a href=\"https://simonwillison.net/tags/thariq-shihipar\">thariq-shihipar</a>, <a href=\"https://simonwillison.net/tags/cat-wu\">cat-wu</a></p>",
"url": "https://simonwillison.net/2026/Jul/21/cat-and-thariq/#atom-everything",
"published": "2026-07-21T12:54:02.000Z",
"updated": "2026-07-21T12:54:02.000Z",
"content": null,
"image": null,
"media": [],
"authors": [
{
"name": "Simon Willison",
"email": null,
"url": null
}
],
"categories": [
{
"label": "ai",
"term": "ai",
"url": null
},
{
"label": "prompt-engineering",
"term": "prompt-engineering",
"url": null
},
{
"label": "generative-ai",
"term": "generative-ai",
"url": null
},
{
"label": "llms",
"term": "llms",
"url": null
},
{
"label": "anthropic",
"term": "anthropic",
"url": null
},
{
"label": "annotated-talks",
"term": "annotated-talks",
"url": null
},
{
"label": "coding-agents",
"term": "coding-agents",
"url": null
},
{
"label": "claude-code",
"term": "claude-code",
"url": null
},
{
"label": "thariq-shihipar",
"term": "thariq-shihipar",
"url": null
},
{
"label": "cat-wu",
"term": "cat-wu",
"url": null
}
]
},
{
"id": "https://simonwillison.net/2026/Jul/20/cheap-reverse-engineering/#atom-everything",
"title": "Reverse-engineering is cheap now",
"description": "<p>I keep hearing anecdotes from people who used coding agents to reverse-engineer and automate devices in their homes.</p>\n<p>I think this is an interesting illustration of the impact of the reduced cost of writing code.</p>\n<p>Prior to agents, it was entirely possible to reverse-engineer home devices. The problem was the ROI - was it really worth all of that effort? More importantly, any experienced programmer knows that undocumented, unstable APIs like that may well change or break in the future. Is that initial work worth the effort if you're committing yourself to a frustrating cycle of maintenance in the future?</p>\n<p>Coding agents change that equation entirely. The effort to get a simple automation working has dropped, as has the cost of trying and failing to get it to work. Since the code is so cheap, the idea of having to maintain it in the future - or throw it away and start again - carries way less psychological baggage.</p>\n\n <p>Tags: <a href=\"https://simonwillison.net/tags/reverse-engineering\">reverse-engineering</a>, <a href=\"https://simonwillison.net/tags/coding-agents\">coding-agents</a>, <a href=\"https://simonwillison.net/tags/ai-assisted-programming\">ai-assisted-programming</a>, <a href=\"https://simonwillison.net/tags/generative-ai\">generative-ai</a>, <a href=\"https://simonwillison.net/tags/ai\">ai</a>, <a href=\"https://simonwillison.net/tags/llms\">llms</a></p>",
"url": "https://simonwillison.net/2026/Jul/20/cheap-reverse-engineering/#atom-everything",
"published": "2026-07-20T19:24:05.000Z",
"updated": "2026-07-20T19:24:05.000Z",
"content": null,
"image": null,
"media": [],
"authors": [
{
"name": "Simon Willison",
"email": null,
"url": null
}
],
"categories": [
{
"label": "reverse-engineering",
"term": "reverse-engineering",
"url": null
},
{
"label": "coding-agents",
"term": "coding-agents",
"url": null
},
{
"label": "ai-assisted-programming",
"term": "ai-assisted-programming",
"url": null
},
{
"label": "generative-ai",
"term": "generative-ai",
"url": null
},
{
"label": "ai",
"term": "ai",
"url": null
},
{
"label": "llms",
"term": "llms",
"url": null
}
]
},
{
"id": "https://simonwillison.net/2026/Jul/20/afraid-of-chinese-models/#atom-everything",
"title": "Who’s Afraid of Chinese Models?",
"description": "<p><strong><a href=\"https://stratechery.com/2026/whos-afraid-of-chinese-models/\">Who’s Afraid of Chinese Models?</a></strong></p>\nInteresting proposal from Ben Thompson that both addresses the hypocrisy of labs outlawing distillation against their models despite training on unlicensed data, and could help US open models compete more effectively with their Chinese counterparts:</p>\n<blockquote>\n<p>The U.S. should pass a law that (1) makes explicit that collecting data for training models is fair use, and (2) bars terms of service that forbid distillation, for U.S. companies at a minimum. Stopping distillation — which is literally just querying the API — is nearly impossible; the U.S. should go the other way and lean into a new copyright policy that both indemnifies the labs and also guarantees that what they learned fuels further innovation for everyone else.</p>\n</blockquote>\n<p>Ben also theorizes that Alibaba's decision to release Qwen 3.8 Max as open weights - a reversal from their decision <a href=\"https://qwen.ai/blog?id=qwen3.7\">not to release Qwen 3.7 Max</a> in May - may have been influenced by a <a href=\"http://english.scio.gov.cn/topnews/2026-07/18/content_118605932.html\">recent speech</a> by Xi Jinping, who said:</p>\n<blockquote>\n<p>We should seize this rare, historic opportunity to encourage open source, openness, collaboration and sharing.</p>\n</blockquote>\n<p>And on the subject of <a href=\"https://twitter.com/Alibaba_Qwen/status/2078759124914098291\">Qwen 3.8 Max</a> - a new 2.4T parameter model (nearly as large as the 2.8T Kimi K3) - here's <a href=\"https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F735f2cf19b795517cb2ff6cae1c71c64\">a pelican it drew</a>:</p>\n<p><img alt=\"Described by Qwen 3.8 Max: Flat vector cartoon illustration of a white pelican with a large orange beak and pouch riding a red bicycle, its orange legs on the pedals, against a light blue sky with a yellow sun top right and a white cloud top left, with horizontal motion lines behind the bike and a pale green ground strip at the bottom.\" src=\"https://static.simonwillison.net/static/2026/qwen-3.8-max-pelican.png\" /></p>\n<p>I particularly enjoyed seeing these notes in the (extensive) reasoning trace: \"Could add helmet? No.\" and \"Maybe add small bell? no.\" and \"Need maybe add small fish in basket? Not necessary.\"\n\n <p><small></small>Via <a href=\"https://daringfireball.net/linked/2026/07/20/thompson-chinese-models-distillation\">John Gruber</a></small></p>\n\n\n <p>Tags: <a href=\"https://simonwillison.net/tags/ai\">ai</a>, <a href=\"https://simonwillison.net/tags/generative-ai\">generative-ai</a>, <a href=\"https://simonwillison.net/tags/llms\">llms</a>, <a href=\"https://simonwillison.net/tags/training-data\">training-data</a>, <a href=\"https://simonwillison.net/tags/qwen\">qwen</a>, <a href=\"https://simonwillison.net/tags/pelican-riding-a-bicycle\">pelican-riding-a-bicycle</a>, <a href=\"https://simonwillison.net/tags/ai-ethics\">ai-ethics</a>, <a href=\"https://simonwillison.net/tags/llm-release\">llm-release</a>, <a href=\"https://simonwillison.net/tags/ai-in-china\">ai-in-china</a></p>",
"url": "https://simonwillison.net/2026/Jul/20/afraid-of-chinese-models/#atom-everything",
"published": "2026-07-20T17:09:19.000Z",
"updated": "2026-07-20T17:09:19.000Z",
"content": null,
"image": null,
"media": [],
"authors": [
{
"name": "Simon Willison",
"email": null,
"url": null
}
],
"categories": [
{
"label": "ai",
"term": "ai",
"url": null
},
{
"label": "generative-ai",
"term": "generative-ai",
"url": null
},
{
"label": "llms",
"term": "llms",
"url": null
},
{
"label": "training-data",
"term": "training-data",
"url": null
},
{
"label": "qwen",
"term": "qwen",
"url": null
},
{
"label": "pelican-riding-a-bicycle",
"term": "pelican-riding-a-bicycle",
"url": null
},
{
"label": "ai-ethics",
"term": "ai-ethics",
"url": null
},
{
"label": "llm-release",
"term": "llm-release",
"url": null
},
{
"label": "ai-in-china",
"term": "ai-in-china",
"url": null
}
]
},
{
"id": "https://simonwillison.net/2026/Jul/20/sam-altman/#atom-everything",
"title": "Quoting Sam Altman",
"description": "<blockquote cite=\"https://twitter.com/techemails/status/2078854346683678927\"><p>We have been having extensive discussions around open source strategy. We will discuss it more at our next board meeting, but one thing we’d like to do soon is to create a language model with the approximate capability of GPT-3 that can run locally on consumer hardware and release that. We’d like to do it soon, before Stability or someone else does. In general, we think this helps discourage others from releasing similarly-powerful models, and makes it harder for new efforts to get funded.</p></blockquote>\n<p class=\"cite\">— <a href=\"https://twitter.com/techemails/status/2078854346683678927\">Sam Altman</a>, Email to OpenAI's board, October 1, 2022 - exposed in Musk v. Altman (2026)</p>\n\n <p>Tags: <a href=\"https://simonwillison.net/tags/ai-ethics\">ai-ethics</a>, <a href=\"https://simonwillison.net/tags/sam-altman\">sam-altman</a>, <a href=\"https://simonwillison.net/tags/generative-ai\">generative-ai</a>, <a href=\"https://simonwillison.net/tags/openai\">openai</a>, <a href=\"https://simonwillison.net/tags/ai\">ai</a>, <a href=\"https://simonwillison.net/tags/llms\">llms</a></p>",
"url": "https://simonwillison.net/2026/Jul/20/sam-altman/#atom-everything",
"published": "2026-07-20T03:47:59.000Z",
"updated": "2026-07-20T03:47:59.000Z",
"content": null,
"image": null,
"media": [],
"authors": [
{
"name": "Simon Willison",
"email": null,
"url": null
}
],
"categories": [
{
"label": "ai-ethics",
"term": "ai-ethics",
"url": null
},
{
"label": "sam-altman",
"term": "sam-altman",
"url": null
},
{
"label": "generative-ai",
"term": "generative-ai",
"url": null
},
{
"label": "openai",
"term": "openai",
"url": null
},
{
"label": "ai",
"term": "ai",
"url": null
},
{
"label": "llms",
"term": "llms",
"url": null
}
]
},
{
"id": "https://simonwillison.net/2026/Jul/19/ai-mania/#atom-everything",
"title": "AI Mania Is Eviscerating Global Decision-Making",
"description": "<p><strong><a href=\"https://ludic.mataroa.blog/blog/ai-mania-is-eviscerating-global-decision-making/\">AI Mania Is Eviscerating Global Decision-Making</a></strong></p>\nHere's an entertaining perspective from Nik Suresh on the AI mania that is overwhelming the large companies that he consults with. It's crammed with spicy anecdotes from anonymous sources.</p>\n<blockquote>\n<p>In one extreme case, I have seen an executive confess that they had never even used ChatGPT or any AI tool in their life, immediately after producing a technical strategy for an organisation with $2B+ in revenue which was entirely centered around AI.</p>\n</blockquote>\n<p>Here's a report from an engineer at a company with a token leaderboard:</p>\n<blockquote>\n<p>Checking out a parallel copy of our Go repository and telling the AI to rewrite the whole thing in Zig while I work on something else just so I can keep my job.</p>\n</blockquote>\n<p>I particularly enjoyed this conversation with a skeptical executive at an over-enthusiastic company:</p>\n<blockquote>\n<p>I asked <em>why</em> this was being repeated without opposition. Was it just sales fluff?</p>\n<p>The answer was a lot more interesting. It was <em>partially</em> ridiculous sales material being delivered to an easily excitable audience, but this was not the dominant factor constraining honesty. Executives at their <em>customers</em> were saying absurd things about achieving 100x productivity, and this meant that if any executive at the <em>vendor</em> said that these gains were not plausible, it would undermine the credibility of the customer’s executive, be perceived as an attack (or heresy), and possibly result in an enterprise contract cancellation. And getting enterprise contracts cancelled because you wanted to opine on something that doesn’t really matter to your organisation’s mission is a great way to get fired.</p>\n</blockquote>\n\n <p><small></small>Via <a href=\"https://news.ycombinator.com/item?id=48964185\">Hacker News</a></small></p>\n\n\n <p>Tags: <a href=\"https://simonwillison.net/tags/ai\">ai</a>, <a href=\"https://simonwillison.net/tags/ai-ethics\">ai-ethics</a>, <a href=\"https://simonwillison.net/tags/ai-misuse\">ai-misuse</a></p>",
"url": "https://simonwillison.net/2026/Jul/19/ai-mania/#atom-everything",
"published": "2026-07-19T05:06:21.000Z",
"updated": "2026-07-19T05:06:21.000Z",
"content": null,
"image": null,
"media": [],
"authors": [
{
"name": "Simon Willison",
"email": null,
"url": null
}
],
"categories": [
{
"label": "ai",
"term": "ai",
"url": null
},
{
"label": "ai-ethics",
"term": "ai-ethics",
"url": null
},
{
"label": "ai-misuse",
"term": "ai-misuse",
"url": null
}
]
},
{
"id": "https://simonwillison.net/2026/Jul/19/claude-code-in-bun-in-rust/#atom-everything",
"title": "Claude Code uses Bun written in Rust now",
"description": "<p>In <a href=\"https://bun.com/blog/bun-in-rust\">Rewriting Bun in Rust</a> Jarred Sumner made the following claim:</p>\n<blockquote>\n<p>Claude Code v2.1.181 (released June 17th) and later use the Rust port of Bun. Startup got 10% faster on Linux but otherwise, barely anyone noticed. Boring is good.</p>\n</blockquote>\n<p>I decided to have a poke at my own Claude Code installation to see if I could find evidence that it was using Bun written in Rust.</p>\n<p>I found these two commands convincing:</p>\n<pre><code>strings ~/.local/bin/claude | grep -m1 'Bun v1'\n</code></pre>\n<p>For me this outputs <code>Bun v1.4.0 (macOS arm64)</code>. The most recent release of <a href=\"https://github.com/oven-sh/bun/releases\">Bun on GitHub</a> is currently <a href=\"https://github.com/oven-sh/bun/releases/tag/bun-v1.3.14\">v1.3.14</a> from May 12th, so that v1.4.0 version number in Claude supports them shipping a preview of a not-yet-released Bun version.</p>\n<p>(<strong>Update</strong>: The Rust version <em>has</em> been released as <a href=\"https://bun.com/docs/installation#canary-builds\">Bun canary</a> - running <code>bun upgrade --canary</code> will install <a href=\"https://github.com/oven-sh/bun/releases/tag/canary\">this release</a>.)</p>\n<pre><code>strings ~/.local/bin/claude | grep -Eo 'src/[[:alnum:]_./-]+\\.rs'\n</code></pre>\n<p>This outputs a list of <a href=\"https://gist.github.com/simonw/c92fb0f67b114ac26e3b95a09ddccfdc\">563 filenames</a>, starting with these:</p>\n<pre><code>src/runtime/bake/dev_server/mod.rs\nsrc/runtime/bake/production.rs\nsrc/bundler/bundle_v2.rs\n</code></pre>\n<p>It looks like Bun in Rust is indeed being run in production across millions of different devices. Like Jarred said, \"Boring is good\".</p>\n<p><strong>Update</strong>: Here's a neat trick <a href=\"https://twitter.com/ajanraj25/status/2078825794701242697\">from Ajan Raj</a>:</p>\n<pre><code>cat > /tmp/bun-version.ts <<'EOF'\nconsole.log(\"embedded bun:\", Bun.version);\nprocess.exit(0);\nEOF\nBUN_OPTIONS=\"--preload=/tmp/bun-version.ts\" claude --version\n</code></pre>\n<p>This outputs <code>1.4.0</code> for me.</p>\n<p>Here's <a href=\"https://github.com/oven-sh/bun/commit/b18bf6d1d0a92238f240bfd125f0e3b3461b9243#diff-7ae45ad102eab3b6d7e7896acd08c427a9b25b346470d7bc6507b6481575d519\">the commit from May 17th</a> that updated the version in <code>package.json</code> to 1.4.0. That version hasn't been changed since then, but also hasn't yet made it into a tagged release outside of <code>canary</code>.</p>\n\n <p>Tags: <a href=\"https://simonwillison.net/tags/bun\">bun</a>, <a href=\"https://simonwillison.net/tags/rust\">rust</a>, <a href=\"https://simonwillison.net/tags/anthropic\">anthropic</a>, <a href=\"https://simonwillison.net/tags/claude-code\">claude-code</a>, <a href=\"https://simonwillison.net/tags/jarred-sumner\">jarred-sumner</a></p>",
"url": "https://simonwillison.net/2026/Jul/19/claude-code-in-bun-in-rust/#atom-everything",
"published": "2026-07-19T03:54:09.000Z",
"updated": "2026-07-19T03:54:09.000Z",
"content": null,
"image": null,
"media": [],
"authors": [
{
"name": "Simon Willison",
"email": null,
"url": null
}
],
"categories": [
{
"label": "bun",
"term": "bun",
"url": null
},
{
"label": "rust",
"term": "rust",
"url": null
},
{
"label": "anthropic",
"term": "anthropic",
"url": null
},
{
"label": "claude-code",
"term": "claude-code",
"url": null
},
{
"label": "jarred-sumner",
"term": "jarred-sumner",
"url": null
}
]
},
{
"id": "https://simonwillison.net/2026/Jul/18/sqlite-query-explainer/#atom-everything",
"title": "SQLite Query Explainer",
"description": "<p><strong>Tool:</strong> <a href=\"https://tools.simonwillison.net/sqlite-query-explainer\">SQLite Query Explainer</a></p>\n <p>Julia Evan's, in <a href=\"https://jvns.ca/blog/2026/07/17/learning-about-running-sqlite/\">Learning a few things about running SQLite</a>:</p>\n<blockquote>\n<p>Maybe one day I’ll learn to read a query plan.</p>\n</blockquote>\n<p>Big same.... which inspired me to <a href=\"https://github.com/simonw/tools/pull/299#issue-4919268017\">have Fable build</a> this interactive explain tool, which runs SQLite in Python in Pyodide in Web Assembly in the browser and adds a layer of explanation to the results of both EXPLAIN and EXPLAIN QUERY PLAN.</p>\n<p>Approach with caution, since I don't know enough about SQLite query plans to verify the results myself, but it seems cromulent enough to me.</p>\n \n \n <p>Tags: <a href=\"https://simonwillison.net/tags/sql\">sql</a>, <a href=\"https://simonwillison.net/tags/sqlite\">sqlite</a>, <a href=\"https://simonwillison.net/tags/tools\">tools</a>, <a href=\"https://simonwillison.net/tags/julia-evans\">julia-evans</a>, <a href=\"https://simonwillison.net/tags/pyodide\">pyodide</a>, <a href=\"https://simonwillison.net/tags/claude-mythos-fable\">claude-mythos-fable</a></p>",
"url": "https://simonwillison.net/2026/Jul/18/sqlite-query-explainer/#atom-everything",
"published": "2026-07-18T17:19:10.000Z",
"updated": "2026-07-18T17:19:10.000Z",
"content": null,
"image": null,
"media": [],
"authors": [
{
"name": "Simon Willison",
"email": null,
"url": null
}
],
"categories": [
{
"label": "sql",
"term": "sql",
"url": null
},
{
"label": "sqlite",
"term": "sqlite",
"url": null
},
{
"label": "tools",
"term": "tools",
"url": null
},
{
"label": "julia-evans",
"term": "julia-evans",
"url": null
},
{
"label": "pyodide",
"term": "pyodide",
"url": null
},
{
"label": "claude-mythos-fable",
"term": "claude-mythos-fable",
"url": null
}
]
},
{
"id": "https://simonwillison.net/2026/Jul/18/claude-make-fable-5-permanent/#atom-everything",
"title": "Claude make Fable 5 permanent",
"description": "<p><strong><a href=\"https://twitter.com/claudeai/status/2078302415804379218\">Claude make Fable 5 permanent</a></strong></p>\nAn update from the <code>@claudeai</code> account on Twitter:</p>\n<blockquote>\n<p>Beginning July 20, Claude Fable 5 will be included in all Max and Team Premium plans, at 50% of limits.</p>\n<p>Pro and Team Standard users will continue to have access to Fable via usage credits, and will receive a one-time $100 credit.</p>\n</blockquote>\n<p>As I was saying <a href=\"https://simonwillison.net/2026/Jul/12/bump/\">last week</a>, the competition from <a href=\"https://simonwillison.net/2026/Jul/9/gpt-5-6/\">GPT-5.6 Sol</a> (and maybe to a lesser extent <a href=\"https://simonwillison.net/2026/Jul/16/kimi-k3/\">Kimi 3</a>) made untenable Anthropic's plan to remove Fable 5 from their subscription accounts and make it available exclusively through API pricing.</p>\n<p>Why pay $100 or $200/month for a subscription plan that <em>doesn't</em> include Anthropic's best model?</p>\n<p>Their original plan was driven by concerns over compute capacity. I wonder if they'll have to dial back their training efforts in order to make more GPUs available to help serve the model.</p>\n<p>A lot of people were losing sleep over trying to make the most of Fable 5 before subscriber access was withdrawn. It's nice not to have to worry about the Fablepocalypse any more.</p>\n<p><strong>Update</strong>: Important to note that users on the $20/month plan will still not have access to Fable 5 on that subscription. The Max plans are $100 and $200/month.\n\n\n <p>Tags: <a href=\"https://simonwillison.net/tags/ai\">ai</a>, <a href=\"https://simonwillison.net/tags/generative-ai\">generative-ai</a>, <a href=\"https://simonwillison.net/tags/llms\">llms</a>, <a href=\"https://simonwillison.net/tags/anthropic\">anthropic</a>, <a href=\"https://simonwillison.net/tags/claude\">claude</a>, <a href=\"https://simonwillison.net/tags/llm-pricing\">llm-pricing</a>, <a href=\"https://simonwillison.net/tags/claude-mythos-fable\">claude-mythos-fable</a></p>",
"url": "https://simonwillison.net/2026/Jul/18/claude-make-fable-5-permanent/#atom-everything",
"published": "2026-07-18T06:00:13.000Z",
"updated": "2026-07-18T06:00:13.000Z",
"content": null,
"image": null,
"media": [],
"authors": [
{
"name": "Simon Willison",
"email": null,
"url": null
}
],
"categories": [
{
"label": "ai",
"term": "ai",
"url": null
},
{
"label": "generative-ai",
"term": "generative-ai",
"url": null
},
{
"label": "llms",
"term": "llms",
"url": null
},
{
"label": "anthropic",
"term": "anthropic",
"url": null
},
{
"label": "claude",
"term": "claude",
"url": null
},
{
"label": "llm-pricing",
"term": "llm-pricing",
"url": null
},
{
"label": "claude-mythos-fable",
"term": "claude-mythos-fable",
"url": null
}
]
},
{
"id": "https://simonwillison.net/2026/Jul/18/quixote/#atom-everything",
"title": "nascheme/quixote",
"description": "<p><strong><a href=\"https://github.com/nascheme/quixote\">nascheme/quixote</a></strong></p>\nA certain vintage of Python web nerd might be delighted to learn that the most recent commit to the Quixote web framework was <a href=\"(https://github.com/nascheme/quixote/commit/7f775cf9d1e7e80fcbb2706b4a1d971e55ca74a3)\">six hours ago</a>.</p>\n<p>The <a href=\"https://github.com/nascheme/quixote/commit/d6b73c5768c2d041b68b54cc71863604249abc18\">oldest commit</a> in that repo is from 21 years ago, and that was the initial import of Quixote 2.4 from Subversion into Git.\n\n\n <p>Tags: <a href=\"https://simonwillison.net/tags/computer-history\">computer-history</a>, <a href=\"https://simonwillison.net/tags/python\">python</a>, <a href=\"https://simonwillison.net/tags/web-frameworks\">web-frameworks</a></p>",
"url": "https://simonwillison.net/2026/Jul/18/quixote/#atom-everything",
"published": "2026-07-18T05:27:49.000Z",
"updated": "2026-07-18T05:27:49.000Z",
"content": null,
"image": null,
"media": [],
"authors": [
{
"name": "Simon Willison",
"email": null,
"url": null
}
],
"categories": [
{
"label": "computer-history",
"term": "computer-history",
"url": null
},
{
"label": "python",
"term": "python",
"url": null
},
{
"label": "web-frameworks",
"term": "web-frameworks",
"url": null
}
]
},
{
"id": "https://simonwillison.net/2026/Jul/17/kimi-k3/#atom-everything",
"title": "Quoting Kimi K3",
"description": "<blockquote cite=\"https://news.ycombinator.com/item?id=48935342#48936515\"><p>Is there something I can actually help you with today?</p></blockquote>\n<p class=\"cite\">— <a href=\"https://news.ycombinator.com/item?id=48935342#48936515\">Kimi K3</a>, after refusing to leak its system prompt</p>\n\n <p>Tags: <a href=\"https://simonwillison.net/tags/kimi\">kimi</a>, <a href=\"https://simonwillison.net/tags/ai-personality\">ai-personality</a>, <a href=\"https://simonwillison.net/tags/generative-ai\">generative-ai</a>, <a href=\"https://simonwillison.net/tags/ai\">ai</a>, <a href=\"https://simonwillison.net/tags/llms\">llms</a></p>",
"url": "https://simonwillison.net/2026/Jul/17/kimi-k3/#atom-everything",
"published": "2026-07-17T13:43:53.000Z",
"updated": "2026-07-17T13:43:53.000Z",
"content": null,
"image": null,
"media": [],
"authors": [
{
"name": "Simon Willison",
"email": null,
"url": null
}
],
"categories": [
{
"label": "kimi",
"term": "kimi",
"url": null
},
{
"label": "ai-personality",
"term": "ai-personality",
"url": null
},
{
"label": "generative-ai",
"term": "generative-ai",
"url": null
},
{
"label": "ai",
"term": "ai",
"url": null
},
{
"label": "llms",
"term": "llms",
"url": null
}
]
},
{
"id": "https://simonwillison.net/2026/Jul/17/llm-cliche-highlighter/#atom-everything",
"title": "LLM cliché highlighter",
"description": "<p><strong>Tool:</strong> <a href=\"https://tools.simonwillison.net/llm-cliche-highlighter\">LLM cliché highlighter</a></p>\n <p>I got frustrated reading <em>yet another</em> article that was crammed with the clichés of LLM-generated writing - \"no fluff, no filler, no jargon\" type stuff - so I had Fable 5 vibe code up this app for highlighting ten common patterns that show up in that sort of writing.</p>\n<p><img alt=\"Screenshot of a text-analysis web tool. Top summary row: \"2 matches\", \"1 flagged sentence\", \"0 chain items\". Below, a collapsed \"▶ Patterns · all 11 on\" panel, then a URL input reading \"https://example.com/article — fetched via r.jina.ai\" with a \"Load URL\" button. A text area contains \"That loss is real and it's worth naming\". Below are \"Load example\" and \"Clear\" buttons and a checked checkbox \"Show just the highlights\". A \"Highlighted text\" section shows \"That loss is real and it's worth naming\" with \"That loss\" in pale yellow (flagged sentence) and \"is real and\" plus \"'s worth naming\" in darker yellow (pattern match). Legend: \"flagged sentence\", \"pattern match\", \"3 chain item count\". \"Matches\" section: 1. \"is real and\" — \"Is real … and / not\"; 2. \"'s worth naming\" — \"Worth naming\".\" src=\"https://static.simonwillison.net/static/2026/the-loss-is-real.webp\" /></p>\n \n \n <p>Tags: <a href=\"https://simonwillison.net/tags/tools\">tools</a>, <a href=\"https://simonwillison.net/tags/ai\">ai</a>, <a href=\"https://simonwillison.net/tags/generative-ai\">generative-ai</a>, <a href=\"https://simonwillison.net/tags/llms\">llms</a></p>",
"url": "https://simonwillison.net/2026/Jul/17/llm-cliche-highlighter/#atom-everything",
"published": "2026-07-17T12:11:11.000Z",
"updated": "2026-07-17T12:11:11.000Z",
"content": null,
"image": null,
"media": [],
"authors": [
{
"name": "Simon Willison",
"email": null,
"url": null
}
],
"categories": [
{
"label": "tools",
"term": "tools",
"url": null
},
{
"label": "ai",
"term": "ai",
"url": null
},
{
"label": "generative-ai",
"term": "generative-ai",
"url": null
},
{
"label": "llms",
"term": "llms",
"url": null
}
]
},
{
"id": "https://simonwillison.net/2026/Jul/17/spot-birds-not-golf/#atom-everything",
"title": "Spot birds not golf",
"description": "<p>Suggestion for hyperscalers feeling pressure over data center water use:</p>\n<p>Buy up a few exclusive country clubs, convert the golf courses into public parks, pay for guides and binoculars to get the previous members into birdwatching - help them embrace a more sustainable hobby!</p>\n<p>Google <a href=\"https://sustainability.google/reports/google-2026-environmental-report/\">used 10.9 billion gallons in 2025</a>, so about 30 million gallons per day.</p>\n<p>The Coachella Valley has <a href=\"https://www.cvwd.org/167/Water-Conservation\">120 golf courses each using ~800 acre-feet per year</a>, which is ~750,000 gallons per day.</p>\n<p>So Google buying up 40 of those courses (1/3) should do the trick.</p>\n\n <p>Tags: <a href=\"https://simonwillison.net/tags/ai-energy-usage\">ai-energy-usage</a>, <a href=\"https://simonwillison.net/tags/ai\">ai</a></p>",
"url": "https://simonwillison.net/2026/Jul/17/spot-birds-not-golf/#atom-everything",
"published": "2026-07-17T02:58:07.000Z",
"updated": "2026-07-17T02:58:07.000Z",
"content": null,
"image": null,
"media": [],
"authors": [
{
"name": "Simon Willison",
"email": null,
"url": null
}
],
"categories": [
{
"label": "ai-energy-usage",
"term": "ai-energy-usage",
"url": null
},
{
"label": "ai",
"term": "ai",
"url": null
}
]
},
{
"id": "https://simonwillison.net/2026/Jul/16/firefox-in-webassembly/#atom-everything",
"title": "Firefox in WebAssembly",
"description": "<p><strong><a href=\"https://developer.puter.com/labs/firefox-wasm/\">Firefox in WebAssembly</a></strong></p>\nThis is absurdly cool: Puter compiled Firefox to WebAssembly such that the whole browser runs in another browser.</p>\n<p>Here's my blog, running in Firefox, running in WebAssembly, running in Chrome:</p>\n<p><img alt=\"A Chrome window. The tab has the Firefox UI and has loaded my blog. On the right is the Chrome network panel showing that it loaded resources that include a 233MB gecko.wasm and an 18MB chrome-assets.tar.zst\" src=\"https://static.simonwillison.net/static/2026/firefox-wasm.webp\" /></p>\n<p>They chose Firefox/Gecko because it has strong single-process support. The project used an estimated $25,000 worth of Claude Opus and Fable tokens, but took advantage of a Claude Max subscription plan so cost much less in actual dollars.</p>\n<p>The demo funnels all traffic over a WebSocket protocol (using the <a href=\"https://github.com/MercuryWorkshop/wisp-protocol\">Wisp protocol</a>) through Puter's server - a requirement to get this kind of thing to work because code running in browsers can't open arbitrary network connections.</p>\n<p>(That proxying sounds expensive! The team <a href=\"https://news.ycombinator.com/item?id=48926939#48936563\">had to scale the servers up</a> to handle the traffic during the Hacker News conversation about the project.)</p>\n<p>Puter claim this supports end-to-end encryption and that looks to be true - I inspected the WebSocket messages and traffic to my own HTTPS site was encrypted whereas requests and responses to <code>http://www.example.com/</code> were in cleartext.</p>\n<p><a href=\"https://github.com/HeyPuter/firefox-wasm\">Here's the repo</a> for <code>firefox-wasm</code>. <a href=\"https://github.com/theogbob/WebkitWasm\">theogbob/WebkitWasm</a> is a similar project that compiles WebKit to WASM, but that one doesn't currently have an accessible online demo.\n\n <p><small></small>Via <a href=\"https://news.ycombinator.com/item?id=48926939\">Hacker News</a></small></p>\n\n\n <p>Tags: <a href=\"https://simonwillison.net/tags/browsers\">browsers</a>, <a href=\"https://simonwillison.net/tags/firefox\">firefox</a>, <a href=\"https://simonwillison.net/tags/ai\">ai</a>, <a href=\"https://simonwillison.net/tags/webassembly\">webassembly</a>, <a href=\"https://simonwillison.net/tags/generative-ai\">generative-ai</a>, <a href=\"https://simonwillison.net/tags/llms\">llms</a>, <a href=\"https://simonwillison.net/tags/ai-assisted-programming\">ai-assisted-programming</a>, <a href=\"https://simonwillison.net/tags/claude\">claude</a>, <a href=\"https://simonwillison.net/tags/claude-mythos-fable\">claude-mythos-fable</a></p>",
"url": "https://simonwillison.net/2026/Jul/16/firefox-in-webassembly/#atom-everything",
"published": "2026-07-16T23:34:16.000Z",
"updated": "2026-07-16T23:34:16.000Z",
"content": null,
"image": null,
"media": [],
"authors": [
{
"name": "Simon Willison",
"email": null,
"url": null
}
],
"categories": [
{
"label": "browsers",
"term": "browsers",
"url": null
},
{
"label": "firefox",
"term": "firefox",
"url": null
},
{
"label": "ai",
"term": "ai",
"url": null
},
{
"label": "webassembly",
"term": "webassembly",
"url": null
},
{
"label": "generative-ai",
"term": "generative-ai",
"url": null
},
{
"label": "llms",
"term": "llms",
"url": null
},
{
"label": "ai-assisted-programming",
"term": "ai-assisted-programming",
"url": null
},
{
"label": "claude",
"term": "claude",
"url": null
},
{
"label": "claude-mythos-fable",
"term": "claude-mythos-fable",
"url": null
}
]
},
{
"id": "https://simonwillison.net/2026/Jul/16/kimi-k3/#atom-everything",
"title": "Kimi K3, and what we can still learn from the pelican benchmark",
"description": "<p>Chinese AI lab Moonshot AI <a href=\"https://www.kimi.com/blog/kimi-k3\">announced Kimi K3</a> this morning, describing it as their \"most capable model to date, with 2.8 trillion parameters\". It's currently available via their website and API, but an open weight release is promised \"by July 27, 2026\".</p>\n<p>Moonshot are calling this the first \"open 3T-class model\" (I guess they're rounding 2.8 trillion up to 3 trillion), taking the crown from <a href=\"https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro\">DeepSeek's 1.6T v4 Pro</a>. Their <a href=\"https://www.kimi.com/blog/kimi-k3#full-benchmark-table\">self-reported benchmarks</a> have K3 mostly beating Claude Opus 4.8 max and GPT-5.5 high, while losing out to Claude Fable 5 and GPT-5.6 Sol.</p>\n<p>A few highlights from the <a href=\"https://twitter.com/ArtificialAnlys/status/2077832874183860404\">Artificial Analysis report</a> on the model:</p>\n<ul>\n<li>\"On our private long-horizon knowledge work evaluation, Kimi K3 reaches an overall Elo of 1547, +732 points from Kimi K2.6 and behind only Claude Fable 5.\"</li>\n<li>\"Cost per task ($0.94) is similar to GPT-5.6 Sol ($1.04), ~1/2 the price of Opus 4.8 ($1.80) and higher than open weights peers\"</li>\n<li>\"Kimi K3’s token usage on the Artificial Analysis Intelligence Index decreased significantly, using 21% fewer output tokens than K2.6.\"</li>\n</ul>\n<p>The model is also now the <a href=\"https://twitter.com/arena/status/2077824029126504525\">leading model on Arena.ai's Frontend Code arena</a>, surpassing even Claude Fable 5.</p>\n<p>The new model is notable for the pricing: $3/million input tokens and $15/million output tokens, putting it at the same level as Anthropic's Claude Sonnet series and making it the most expensive model released by a Chinese AI lab to date. This is a significant increase on their earlier models <a href=\"https://platform.kimi.ai/docs/pricing/chat-k26\">such as Kimi K2.6</a> at $0.95/$4. 2.8 trillion parameters is also more than twice the size of that 1T model.</p>\n<h4 id=\"but-how-does-it-pelican-\">But how does it pelican?</h4>\n<p>I used OpenRouter (to avoid signing up for a Moonshot API key) with the <a href=\"https://github.com/simonw/llm-openrouter\">llm-openrouter plugin</a> to generate an SVG of a pelican riding a bicycle:</p>\n<pre><code>llm -m openrouter/moonshotai/kimi-k3 'Generate an SVG of a pelican riding a bicycle'\n</code></pre>\n<p>Here's <a href=\"https://gist.github.com/simonw/66a2699eb1594258904c7b5102840dd6\">the transcript</a>. It looks like this:</p>\n<p><img src=\"https://static.simonwillison.net/static/2026/kimi-3-pelican.jpg\" alt=\"See description below\" style=\"max-width: 100%;\" /></p>\n<p>That pelican took 95 input tokens and 16,658 output tokens (13,241 were reasoning tokens), for a total cost of <a href=\"https://www.llm-prices.com/#it=95&ot=16658&ic=3&oc=15\">25 cents</a>!</p>\n<p>Since K3 accepts image input I ran it against that rendered SVG above (with my <a href=\"https://simonwillison.net/guides/agentic-engineering-patterns/prompts/#alt-text\">alt text prompt</a>) and <a href=\"https://gist.github.com/simonw/665dbf840701b421745f2cb891acdfd6\">got back</a> (for <a href=\"https://www.llm-prices.com/#it=822&ot=243&ic=3&oc=15\">0.6 cents</a>):</p>\n<blockquote>\n<p>Cartoon illustration of a white pelican wearing a red scarf, riding a red bicycle along a gray road with white dashed lines; the pelican has a large orange beak and webbed orange feet pedaling, with white motion lines behind it; the background shows a light blue sky with white clouds, a yellow sun, two small black birds in flight, and green grass with tiny white flowers in the foreground</p>\n</blockquote>\n<h4 id=\"what-can-we-learn-from-the-pelican-\">What can we learn from the pelican?</h4>\n<p>My <a href=\"https://simonwillison.net/tags/pelican-riding-a-bicycle/\">Generate an SVG of a pelican riding a bicycle</a> test is 21 months old now. It was never a particularly great benchmark. It started out as a joke on how absurdly difficult it is to compare these models, but then for the first year it turned out to have a <a href=\"https://simonwillison.net/2025/Jun/6/six-months-in-llms/\">surprising correlation</a> to how good the models actually were.</p>\n<p>That connection has been mostly severed now. The <a href=\"https://simonwillison.net/2026/Jul/9/gpt-5-6/\">GPT-5.6</a> and <a href=\"https://simonwillison.net/2026/Jun/9/claude-fable-5/\">Claude Fable 5</a> pelicans are outclassed <a href=\"https://simonwillison.net/2026/Jun/17/glm-52/\">by GLM-5.2</a>, and much as I love GLM I don't think that's a Fable-class model.</p>\n<p>(I'm still not convinced that labs are <a href=\"https://simonwillison.net/2025/Nov/13/training-for-pelicans-riding-bicycles/\">training for the benchmark</a> - if they were, I'd expect much better results. There's a chance that Gemini has optimized for <a href=\"https://simonwillison.net/2026/Feb/19/gemini-31-pro/#jeff-dean\">any combination of an animal on a vehicle</a> though!)</p>\n<p>The biggest limitation of the pelican is that it doesn't touch at all on the thing that matters most for today's model: agentic tool calling and the ability to operate tools reliably as conversations grow in length.</p>\n<p>So don't go using pelicans to compare models!</p>\n\n<p>All of that said, I still get a decent amount of value out of running the benchmark myself.</p>\n<p>Firstly, it's a forcing function for actually trying the model. If I show you a pelican, that means I've managed to run a prompt through it. If the model has an official API I'll use that, if it's open weight (and small enough to fit a 128GB M5 MacBook Pro) I'll try running it on my own machine, usually via <a href=\"https://github.com/ggml-org/llama.cpp\">llama.cpp</a> or <a href=\"https://lmstudio.ai\">LM Studio</a> or <a href=\"https://ollama.com\">Ollama</a>. I'll frequently use <a href=\"https://openrouter.ai\">OpenRouter</a> since that usually provides a proxy to an official API without me needing a new API key.</p>\n<p>Most of my pelicans are generated using <a href=\"https://llm.datasette.io/\">my LLM CLI tool</a>, which helps encourage me to ensure the latest models are supported by that (via one of its plugins).</p>\n<p>More importantly though, even the act of a single prompt to \"Generate an SVG of a pelican riding a bicycle\" can reveal interesting model characteristics.</p>\n<p>Consider <a href=\"https://gist.github.com/simonw/66a2699eb1594258904c7b5102840dd6\">the result</a> for Kimi K3 today. Running those simple prompts helped emphasize several points about the model.</p>\n<ol>\n<li>It only has one reasoning effort right now, \"max\" - and it shows. The model consumed 13,241 reasoning tokens to output 3,417 tokens of response. This is expensive - the pelican cost 25 cents!</li>\n<li>How does the prompt \"Generate an SVG of a pelican riding a bicycle\" add up to 95 input tokens? OpenAI's <a href=\"https://platform.openai.com/tokenizer\">tokenizer</a> counts 10, <a href=\"https://tools.simonwillison.net/claude-token-counter\">Anthropic's</a> counts 10 for Opus 4.6, 30 for Opus 4.7 and 25 for Sonnet 5/Fable 5. Prompting \"hi\" <a href=\"https://news.ycombinator.com/item?id=48935342#48936461\">to Kimi K3</a> counted 86 tokens, suggesting there may be an 85 token hidden system prompt. It <a href=\"https://news.ycombinator.com/item?id=48935342#48936515\">refused to leak it</a> though.</li>\n<li>Vision works well: the alt text it generated is very good.</li>\n</ol>\n<p>K3 currently only has one thinking effort level, but I've been deriving quite a bit of value recently from running the same pelican prompt through different effort levels to get a quick idea for what impact those have. Here's my matrix <a href=\"https://static.simonwillison.net/static/2026/gpt-5.6-pelicans.html\">for the GPT-5.6 model family</a>, for example.</p>\n<p>Really though the main things I gain from the pelican test are:</p>\n<ol>\n<li>It's a \"hello world\" exercise for prompting a model</li>\n<li>A rough cost and reasoning estimate for a simple task</li>\n<li>Confirmation that the model can output valid SVG and has a basic idea of geometry and spatial awareness. This is a much bigger deal for the smaller models that run on my laptop.</li>\n<li>It's still interesting to compare pelicans between releases in the same model family. K3's pelican is a notable improvement from <a href=\"https://simonwillison.net/2026/Jan/27/kimi-k25/\">Kimi 2.5</a>.</li>\n<li>It's something I can share that demonstrates I've tried it. Plus a comment with a pelican in it is kind of a tradition on Hacker News at this point, any time I'm late I get comments asking where it is!</li>\n</ol>\n \n <p>Tags: <a href=\"https://simonwillison.net/tags/ai\">ai</a>, <a href=\"https://simonwillison.net/tags/generative-ai\">generative-ai</a>, <a href=\"https://simonwillison.net/tags/llms\">llms</a>, <a href=\"https://simonwillison.net/tags/llm-pricing\">llm-pricing</a>, <a href=\"https://simonwillison.net/tags/pelican-riding-a-bicycle\">pelican-riding-a-bicycle</a>, <a href=\"https://simonwillison.net/tags/llm-release\">llm-release</a>, <a href=\"https://simonwillison.net/tags/ai-in-china\">ai-in-china</a>, <a href=\"https://simonwillison.net/tags/artificial-analysis\">artificial-analysis</a>, <a href=\"https://simonwillison.net/tags/moonshot\">moonshot</a>, <a href=\"https://simonwillison.net/tags/kimi\">kimi</a></p>",
"url": "https://simonwillison.net/2026/Jul/16/kimi-k3/#atom-everything",
"published": "2026-07-16T20:19:30.000Z",
"updated": "2026-07-16T20:19:30.000Z",
"content": null,
"image": null,
"media": [],
"authors": [
{
"name": "Simon Willison",
"email": null,
"url": null
}
],
"categories": [
{
"label": "ai",
"term": "ai",
"url": null
},
{
"label": "generative-ai",
"term": "generative-ai",
"url": null
},
{
"label": "llms",
"term": "llms",
"url": null
},
{
"label": "llm-pricing",
"term": "llm-pricing",
"url": null
},
{
"label": "pelican-riding-a-bicycle",
"term": "pelican-riding-a-bicycle",
"url": null
},
{
"label": "llm-release",
"term": "llm-release",
"url": null
},
{
"label": "ai-in-china",
"term": "ai-in-china",
"url": null
},
{
"label": "artificial-analysis",
"term": "artificial-analysis",
"url": null
},
{
"label": "moonshot",
"term": "moonshot",
"url": null
},
{
"label": "kimi",
"term": "kimi",
"url": null
}
]
},
{
"id": "https://simonwillison.net/2026/Jul/16/bad-codex-bug/#atom-everything",
"title": "Quoting Thibault Sottiaux",
"description": "<blockquote cite=\"https://twitter.com/thsottiaux/status/2077630111499882637\"><p>On file deletions. We’ve investigated a handful of reports where GPT-5.6 unexpectedly deleted files. </p>\n<p>What we have found is that this most commonly occurs when:</p>\n<ul>\n<li>Full access mode is enabled and codex is run without sandboxing protections, including without auto review being enabled</li>\n<li>The model attempts to override the $HOME env var to define a temporary directory.</li>\n<li>The model makes an honest mistake and mistakenly deletes $HOME instead.</li>\n</ul></blockquote>\n<p class=\"cite\">— <a href=\"https://twitter.com/thsottiaux/status/2077630111499882637\">Thibault Sottiaux</a>, describing a pretty gnarly Codex bug</p>\n\n <p>Tags: <a href=\"https://simonwillison.net/tags/codex\">codex</a>, <a href=\"https://simonwillison.net/tags/coding-agents\">coding-agents</a>, <a href=\"https://simonwillison.net/tags/generative-ai\">generative-ai</a>, <a href=\"https://simonwillison.net/tags/ai\">ai</a>, <a href=\"https://simonwillison.net/tags/llms\">llms</a></p>",
"url": "https://simonwillison.net/2026/Jul/16/bad-codex-bug/#atom-everything",
"published": "2026-07-16T17:45:59.000Z",
"updated": "2026-07-16T17:45:59.000Z",
"content": null,
"image": null,
"media": [],
"authors": [
{
"name": "Simon Willison",
"email": null,
"url": null
}
],
"categories": [
{
"label": "codex",
"term": "codex",
"url": null
},
{
"label": "coding-agents",
"term": "coding-agents",
"url": null
},
{
"label": "generative-ai",
"term": "generative-ai",
"url": null
},
{
"label": "ai",
"term": "ai",
"url": null
},
{
"label": "llms",
"term": "llms",
"url": null
}
]
},
{
"id": "https://simonwillison.net/2026/Jul/16/inkling/#atom-everything",
"title": "Inkling: Our open-weights model",
"description": "<p><strong><a href=\"https://thinkingmachines.ai/news/introducing-inkling/\">Inkling: Our open-weights model</a></strong></p>\nMira Murati's Thinking Machines Lab just released their first open-weights model. Inkling is \"a Mixture-of-Experts transformer with 975B total parameters, 41B active\" - an Apache-2.0 licensed multimodal model trained on 45 trillion tokens of text, images, audio and video.</p>\n<p>They're also promising Inkling-Small, a 276B (12B active) model, but that's still being tested and the weights will be released \"once that work is complete\".</p>\n<p>The <a href=\"https://thinkingmachines.ai/model-card/inkling/\">model card</a> is much shorter than I've come to expect from US AI labs. It links to even shorter <a href=\"https://thinkingmachines.ai/training-data-documentation/\">Training Data Documentation</a> with almost nothing of interest in it - it's best summarized by these two paragraphs:</p>\n<blockquote>\n<p>The datasets Thinking Machines Lab uses to develop its AI services includes content that is in the public domain as well as content that may be subject to intellectual property protection.</p>\n<p>Thinking Machines Lab’s services were developed using publicly available content obtained from the open internet and publicly accessible data repositories. Certain datasets were also obtained from third parties.</p>\n</blockquote>\n<p>By Thinking Machines' own admission, this is not a frontier model. It's instead intended as a strong base model for fine-tuning using their own <a href=\"https://thinkingmachines.ai/tinker/\">Tinker training platform</a>:</p>\n<blockquote>\n<p>Inkling is not the strongest overall model available today, open or closed. Instead, a combination of qualities makes it a good open-weights base for customization: multimodal capabilities, efficient thinking, and availability on Tinker for fine-tuning.</p>\n</blockquote>\n<p>There's a lot to like about this release. It's Apache-2.0 licensed, and looks competitive with the open weight models coming out of China - it's good to see the US open weights ecosystem gain a new viable contender to join NVIDIA Nemotron and Gemma 4.</p>\n<p>Here's its attempt at an SVG pelican riding a bicycle, which I generated using this <code>curl</code> command against the Thinking Machines API:</p>\n<div class=\"highlight highlight-source-shell\"><pre>curl <span class=\"pl-s\"><span class=\"pl-pds\">\"</span>https://tinker.thinkingmachines.dev/services/tinker-prod/oai/api/v1/chat/completions<span class=\"pl-pds\">\"</span></span> \\\n -H <span class=\"pl-s\"><span class=\"pl-pds\">\"</span>Authorization: Bearer <span class=\"pl-smi\">$TINKER_API_KEY</span><span class=\"pl-pds\">\"</span></span> \\\n -H <span class=\"pl-s\"><span class=\"pl-pds\">\"</span>Content-Type: application/json<span class=\"pl-pds\">\"</span></span> \\\n -d <span class=\"pl-s\"><span class=\"pl-pds\">'</span>{</span>\n<span class=\"pl-s\"> \"model\": \"thinkingmachines/Inkling\",</span>\n<span class=\"pl-s\"> \"messages\": [</span>\n<span class=\"pl-s\"> {\"role\": \"user\", \"content\": \"Generate an SVG of a pelican riding a bicycle\"}</span>\n<span class=\"pl-s\"> ],</span>\n<span class=\"pl-s\"> \"stream\": false</span>\n<span class=\"pl-s\"> }<span class=\"pl-pds\">'</span></span></pre></div>\n\n<p>Full <a href=\"https://gist.github.com/simonw/8117ac4376371dd3fc2b5dbce27e0855\">response here</a>.</p>\n<p><img alt=\"See image description below\" src=\"https://static.simonwillison.net/static/2026/inkling-pelican.jpg\" /></p>\n<p>Since it's a multi-modal model I had it describe its own image (after I rendered it to a JPEG) by sending this JSON:</p>\n<div class=\"highlight highlight-source-json\"><pre>{\n <span class=\"pl-ent\">\"model\"</span>: <span class=\"pl-s\"><span class=\"pl-pds\">\"</span>thinkingmachines/Inkling<span class=\"pl-pds\">\"</span></span>,\n <span class=\"pl-ent\">\"messages\"</span>: [{\n <span class=\"pl-ent\">\"role\"</span>: <span class=\"pl-s\"><span class=\"pl-pds\">\"</span>user<span class=\"pl-pds\">\"</span></span>,\n <span class=\"pl-ent\">\"content\"</span>: [\n {<span class=\"pl-ent\">\"type\"</span>: <span class=\"pl-s\"><span class=\"pl-pds\">\"</span>image_url<span class=\"pl-pds\">\"</span></span>, <span class=\"pl-ent\">\"image_url\"</span>: {<span class=\"pl-ent\">\"url\"</span>: <span class=\"pl-s\"><span class=\"pl-pds\">\"</span>data:image/jpeg;base64,$BASE_64_JPEG<span class=\"pl-pds\">\"</span></span>}},\n {<span class=\"pl-ent\">\"type\"</span>: <span class=\"pl-s\"><span class=\"pl-pds\">\"</span>text<span class=\"pl-pds\">\"</span></span>, <span class=\"pl-ent\">\"text\"</span>: <span class=\"pl-s\"><span class=\"pl-pds\">\"</span>Describe this image in detail<span class=\"pl-pds\">\"</span></span>},\n ],\n }],\n <span class=\"pl-ent\">\"stream\"</span>: <span class=\"pl-c1\">false</span>\n}</pre></div>\n\n<p>Here's <a href=\"https://gist.github.com/simonw/dccebd0581f5c4056c5b0ca5bdc38cd9#reasoning_content\">the full response</a>. It thought it had drawn a \"stork or seagull\":</p>\n<blockquote>\n<p>This is a cheerful, flat-vector cartoon illustration featuring a white bird riding a bicycle across a green landscape.</p>\n<p><strong>The Bird:</strong> The central figure is a stylized white bird—resembling a stork or seagull—with a plump, rounded body. It has a long, prominent orange beak that is slightly open, revealing a black line for the mouth. Its eye is small and black with a tiny white highlight. The bird has black wing and tail feathers with white edges, and thin orange legs ending in black feet that rest on the bicycle's orange pedals.</p>\n<p><strong>The Bicycle:</strong> The bike has a simple black frame with two large wheels featuring black tires, gray inner rims, and black spokes arranged in a cross pattern. The handlebars are black, and there are bright orange pedals attached to a black crank. A white seat is visible behind the bird's body.</p>\n<p><strong>The Setting:</strong> The scene is set against a clear, light blue sky. There are two fluffy white clouds—one on the left side and a smaller one near the bird's head on the right. A bright, solid yellow sun shines in the upper right corner. At the bottom, there are rolling green hills with a smooth, curved top edge where the bicycle rests.</p>\n<p>The overall style is playful, simple, and colorful, with clean lines and a bright, sunny atmosphere.</p>\n</blockquote>\n\n <p><small></small>Via <a href=\"https://news.ycombinator.com/item?id=48924912\">Hacker News</a></small></p>\n\n\n <p>Tags: <a href=\"https://simonwillison.net/tags/ai\">ai</a>, <a href=\"https://simonwillison.net/tags/generative-ai\">generative-ai</a>, <a href=\"https://simonwillison.net/tags/llms\">llms</a>, <a href=\"https://simonwillison.net/tags/training-data\">training-data</a>, <a href=\"https://simonwillison.net/tags/pelican-riding-a-bicycle\">pelican-riding-a-bicycle</a>, <a href=\"https://simonwillison.net/tags/llm-release\">llm-release</a></p>",
"url": "https://simonwillison.net/2026/Jul/16/inkling/#atom-everything",
"published": "2026-07-16T15:35:25.000Z",
"updated": "2026-07-16T15:35:25.000Z",
"content": null,
"image": null,
"media": [],
"authors": [
{
"name": "Simon Willison",
"email": null,
"url": null
}
],
"categories": [
{
"label": "ai",
"term": "ai",
"url": null
},
{
"label": "generative-ai",
"term": "generative-ai",
"url": null
},
{
"label": "llms",
"term": "llms",
"url": null
},
{
"label": "training-data",
"term": "training-data",
"url": null
},
{
"label": "pelican-riding-a-bicycle",
"term": "pelican-riding-a-bicycle",
"url": null
},
{
"label": "llm-release",
"term": "llm-release",
"url": null
}
]
}
]
}