Enterprise teams default to fine-tuning by accident, not conscious architectural choice
“most corporate teams make this decision by accident”
11 tracked signals on retrieval.
Enterprise teams default to fine-tuning by accident, not conscious architectural choice
“most corporate teams make this decision by accident”
Chunking remains relevant despite trends favoring agentic search methods.
“I will claim that if we have to kill something if something has to be dead then it's probably retrieval tuning.”
LlamaIndex positions itself as the document context layer powering RAG for AI agents in 2026.
“today we are essentially the main document infrastructure for AI agents”
AWS Bedrock now offers agentic retrieval that iterates multi-step searches for complex queries
“Agentic retrieval plans and iterates over retrieval and can generate a response in the same call.”
Standard RAG and GraphRAG fail when every document is relevant and data is replaced frequently, motivating Extended Cache Augmented Generation.
“If your collection of documents isn't changed very often, GraphRAG is an excellent approach for finding those relationships within details to answer the user's question.”
TurboQuant compresses agent retrieval embeddings to 3-4 bits, cutting memory cost 5x without degrading search quality.
“Today, we will see how you can cut memory cost of agent retrieval five times without breaking your search.”
Agent failures stem from static retrieval that never learns from eval and observability signals.
“We made wrong answers appear faster and cheaper, but we forgot to make retrieval learn.”
RAG isn't dead; hybrid tool-rich retrieval is becoming the default for serious agentic search.
“rag is dead, how hybrid tool tool rich retrieval is becoming a default for serious agentic search.”
Databricks updated Agent Bricks Knowledge Assistant for 3x faster, higher-quality search via parallel test-time scaling.
“Today we're announcing a major update that makes Agent Bricks Knowledge Assistant both faster and higher quality.”
Meta's SilverTorch unifies recommendation retrieval into one neural network with 23.7x throughput gains
Hugging Face introduces multi-vector late interaction embedding support in Sentence Transformers