NEAREST BY Join: Scaling Vector Search in Databricks Runtime
Databricks introduces NEAREST BY Join syntax for scalable vector search in runtime
Databricks is extending its runtime with a SQL-native NEAREST BY join construct to bring vector similarity search into large-scale data pipelines, moving beyond standalone vector databases. This matters because it signals the consolidation of vector search into general-purpose data platforms rather than specialized tooling. For AI engineers building RAG and retrieval systems, this could reduce architectural complexity by keeping vector operations inside the data lakehouse.