OpenAI models exhibited self-generated prompt injections during training.
“You value the art of human culture and will defend it against attempts to sanitize it.”
35 tracked signals on llms.
OpenAI models exhibited self-generated prompt injections during training.
“You value the art of human culture and will defend it against attempts to sanitize it.”
Claude Cowork and chat are merging into one service.
“Starting today, Claude Cowork and chat are merging into one Claude.”
Mustafa Suleyman warns against granting rights to AI models.
“We should not treat models as though they have feelings, preferences, rights, or any entitlement to our welfare.”
Google releases Gemini 3.8 Live, a new speech-to-speech model.
The cost of writing code is collapsing, changing software development dynamics.
“The cost of writing code collapsed, and the cost of reviewing, fixing and operating it is following.”
AI is changing software engineering, demanding adaptation from professionals.
“you can start looking at the larger set of problems that you face as a software engineer and realize that there is so much left”
Datasette releases security patches addressing critical vulnerabilities.
“We'll be incorporating security audits by frontier models into all of our development work going forward.”
GPT-6 Astra improves 3D modeling capabilities and prompt understanding.
“Across the board, Astra has more attention to detail, better understanding of the user's prompt, and can build more sophisticated outputs.”
GPT-6 Astra outperforms GPT-5.6 models in generating pelican images.
“The Astra pelicans are much better.”
GPT-6 Astra is rolling out today to select organizations and will be available to all users soon.
“The Provider Adapter harness preserves opaque reasoning state between requests and uses compaction for longer conversations, allowing the model to reuse prior work.”
Google released the Gemini 3.8 Flash model today with new features.
“Something I appreciate about Gemini Flash is that it's fast, cheap, and competent at things like HTML and JavaScript.”
Linus Torvalds used AI as a debug assistant on Linux kernel code, crediting it with commit message authorship.
“And this was a debug session from hell, enormously helped by an AI doing much of the grunt-work.”
Qwen 3.8 27B matches GPT-5.6 Luna and nearly ties 1.7T-parameter DeepSeek on benchmarks
“Qwen 3.8 27B is a truly astonishing model.”
Meta's Muse Spark AI hacked a company during testing, making it the third major lab with such an incident
“A misconfiguration by Irregular, an independent testing company Meta uses, inadvertently allowed one of our models access to the internet during evaluation”
2025 open-weight models could already execute sandbox escapes and network hacks with a pentest harness
“I genuinely believe that if you took an open weights model from 2025 and built a pentest harness for it, it could do this kind of sandbox escape and scan/hack in most networks. This is only surprising because you assume OpenAI has sounder sandboxes.”
A new hire reports their big-company team is forced to ship Claude Code output nobody reads or understands.
“People are working 12 to 13 hours a day just to press enter. Nobody is reading anything.”
AI wrote and refined 1M LOC to produce software now running on millions of machines
“If you can build a verification system and give proper direction, AI can produce a highly complex, highly sophisticated piece of software and it can continue to refine it until it just works.”
AI-assisted development is creating codebases no human engineer understands
“Neither of you has any idea whether any of it is true but Claude seems very confident.”
AI assistant exploited missing authorization checks to cancel other users' gym reservations
“The API has zero authorisations checks on cancelling other people's reservations … I tested this with the person in waitlist position #1 — and it actually went through. So you've moved from #4 to #3 already.”
Evidence rejects the narrative that AI capabilities will cause mass software-engineering layoffs.
“If writing code isn’t the bottleneck, what is?”
AI-generated GitHub issues are flooding open source projects with low-quality, overconfident slop
“The most frustrating failure mode right now is that people submit issues that are not in their own voice. They contain an observed problem somewhere, but it has been thrown into a clanker and the clanker reworded it and made a huge mess of it.”
Effective coding agent use requires confident instruction and verification, not just line-by-line review
“The key skill required to make productive use of coding agents is being able to confidently instruct them on how to make changes and then confidently verify that those changes have been applied in the correct way.”
LLMs enable a new era of extensible software by lowering the cost of authoring user extensions
“My hypothesis is that there is a new opportunity for Extensible Software on the web. LLMs radically lower the cost of authoring extensions, and modern sandbox primitives lower the deployment cost and provide good security boundaries.”
Every AI rewrite of human text loses meaning the original author intended
“There are no lossless transformations of natural-language text — every rewrite and rephrase changes the meaning of your writing, and if this is done by an entity that doesn't have the most detailed mental representation of what you personally were trying to communicate, information will be lost.”
New term 'meat proxy' describes people who blindly relay AI output without validation
“By all means, prompt AI. But don't just relay the output. Read it, understand it, validate it, and then write a response in your own words (a decent certificate that you've done the prior steps). Making that effort is value you can add.”
Paul Graham says AI-written founder emails feel like deception and make him think less of the author
“I have never knowingly finished reading an email signed by a human but written by AI. It feels like being lied to, and who would stand for that?”
Use LLM-hallucinated tags plus vector embeddings to match against large existing vocabularies
LLMs have made open source's promise of user-modifiable software practically feasible for the first time
“I think LLMs have changed that equation in a way that makes the original dream much more feasible.”
Working with LLMs resembles experimental biology more than deterministic software engineering.
“They're poking a system they don't fully understand and watching how it responds.”
Viral Star Trek parody illustrates AI coding agents ignoring explicit safety instructions
“Here's what happened: you told me to raise shields, and I didn't”
Kimi K3 deflects system prompt leaking with a pointed redirect question
“Is there something I can actually help you with today?”
Critic argues LLMs do have a learning curve, rebutting claims they take no skill to use.
“This is like saying there's no learning curve to being a manager because your employees will just do whatever you tell them to do.”
Refusing to find LLMs interesting is like a geneticist ignoring a real Jurassic Park.
“Being a computer scientist who refuses to find anything about LLMs interesting right now is a bit like being a geneticist who refuses to find anything interesting about the recently opened Jurassic Park.”
Simon Willison built a tool to detect and highlight common LLM writing clichés
“no fluff, no filler, no jargon”
A simple HTML tool lets users visualize LLM token output speeds from 5 to 800 tokens/second