DeepSeek 4.1 Flash outperforms previous models and significantly reduces operational costs.
“Incredible leap forward.”
12 tracked signals on architecture.
DeepSeek 4.1 Flash outperforms previous models and significantly reduces operational costs.
“Incredible leap forward.”
Enterprise teams default to fine-tuning by accident, not conscious architectural choice
“most corporate teams make this decision by accident”
Salesforce unveils vision for architecture in an agentic future.
Models absorbing harness capabilities into weights is reshaping agent architecture toward human-attention scaffolding
“The change last winter, last Christmas — it's a little hard to pin down. I mean, the harness changed and a little post-training changed and then new pre-trained models came… but it felt like a big jump which is not that easy to pin down what did it.”
ECRI named AI chatbot misuse the #1 health technology hazard of 2026, not a frontier problem but a production baseline issue.
“Most AI safety failures in health care are not model failures. They are architectural decisions that were made before even a single token was generated.”
DeepMind's Gemma 4 achieves multimodal vision by patching images directly into the main transformer, eliminating separate encoders.
“Throw that all away. Out. Right now.”
Agent harness architecture drives double-digit benchmark swings independent of model choice
“Harness design alone can account for double-digit swings in benchmark results and significant differences in token cost”
Voice agents need sub-950ms response; small models beat frontier models on latency
“A frontier model that think for a full second has already lost the room, no matter how good the answer is.”
Sierra builds customer-engagement agents using many parallel models per turn and isolated PCI infrastructure for payments.
“We have isolated infrastructure where payment info doesn't go to an external large language model cuz none of the LLM providers are PCI certified in that way.”
Traditional retry and circuit breaker patterns fail for LLMs; per-request provider fallback is required.
“If you have a single model provider, their ceiling is your ceiling. Their outage is your outage.”
AI architectural research fails because it tests at insufficient compute scale
“to get to any interesting results you need certain level of compute to even see the capabilities in the model”
GTM teams that bolt AI onto existing stacks cannot scale; architecture must be rebuilt with AI at the core.
“by the time the buyer reaches you, the decision is mostly made.”