8 tracked signals on tool-use.
Introducing Muse Glimmer
Simon Willison · Aug 10, 2026
Meta releases Muse Glimmer, a 30B open-weights agentic model under Apache 2.0 license
“Muse Glimmer achieves strong success rates on full-task benchmarks including DeepSearch QA, MCP-Atlas, 𝛕-Bench and SWE-Bench, which measure its ability to work within scaffolds, write and debug code, and resolve multi-turn requests from start to finish.”
The future of software development
Google Developers (Google I/O) · May 22, 2026
Google launches Gemini 3.5 Flash, its most capable model yet, surpassing 3.1 Pro on benchmarks
“3.5 Flash, the most capable model ever that we've shipped.”
Designing Opal: Building a Personal AI Pet
Microsoft Developer (Build) · Jun 26, 2026
A Microsoft product designer built a personal AI pet that uses agentic tool use to browse the web, post to Discord, and commit to GitHub.
“all the interfaces you use have agents running behind them, and eventually those will be the things that you have to design for”
EVA-Bench Data 2.0: 3 Domains, 121 Tools, 213 Scenarios
Hugging Face · Hugging Face Blog · Jun 04, 2026
Hugging Face expands EVA-Bench to 3 domains, 121 tools, and 213 scenarios for agent evaluation.
llm-gemini 0.33
Simon Willison · Aug 13, 2026
llm-gemini 0.33 adds Gemini 3.7 Flash support with reasoning traces and server-side tools
Managed Deep Agents - Tools
LangChain · Aug 19, 2026
LangChain's managed deep agents now support custom tool integration via Python and TypeScript decorators
What Is an AI Agent, Actually?
LangChain · Jul 17, 2026
LangChain defines an AI agent as an LLM looping through decide, act, reason, repeat.
“That loop, decide, act, reason, repeat. That's what makes an agent.”
datasette-agent-edit 0.1a0
Simon Willison · Jun 07, 2026
Simon Willison released datasette-agent-edit, a base plugin implementing core agentic text-editing tools for Datasette Agent.
“Agentic editing of text is a little tricky to get right.”