The Hallway Track

MLOps

16 tracked signals on MLOps.

Your Agents Need a Save Button - Hamza Tahir, ZenML

AI Engineer · Jul 18, 2026

AI agents lack persistent state checkpoints, making debugging and replay impossible today.

“all of that is lost and it is only stamped as a read-only trace by the end, which is sitting in another tool far away from where the actual code is.”
Control How Your GPU Shares Work with Green Contexts

NVIDIA Developer Blog · Oct 06, 2026

NVIDIA introduces Green Contexts for fine-grained GPU resource sharing between concurrent workloads

“Controlling how GPU resources are shared between them remains difficult.”
ModelExpress: Distributing Model Artifacts at the Speed of Light

NVIDIA Developer Blog · Jul 24, 2026

NVIDIA ModelExpress accelerates multi-hundred-GB model weight distribution across GPU clusters

“Every byte moved has a cost. As model checkpoints grow to hundreds of gigabytes or even a terabyte, that cost adds up quickly.”
Build Evals That Actually Matter - Nick Ung, Lyft

AI Engineer · Jul 19, 2026

Lyft applies ML model evaluation rigor to AI agents before production deployment.

“if we are running um offline evaluations for our machine learning model before that goes to productions, I think we should do the same for AI agents as well”
A Developer’s Guide to Managing Models, Cost and Quality in Microsoft Foundry

Microsoft AI Foundry Blog · Jun 02, 2026

Microsoft Foundry positions model operations, not model access, as the core AI challenge

“The hardest part of building AI systems today is no longer getting access to a capable model. It is knowing how to choose, validate, optimize, and operate the right model across the full lifecycle of a real application.”