Tokenomics at scale: How Jamf built real-time spend enforcement for Amazon Bedrock
Jamf built serverless per-user token spend enforcement for Amazon Bedrock without blocking engineers entirely.
“a single engineer running an agentic coding loop against a premium model can burn more tokens in a few hours than a team does in a week”
Jamf open-sourced a production architecture that enforces tiered per-engineer daily spend limits on Amazon Bedrock using IAM policies, Athena cost queries, and scheduled Lambda — restrictions apply within minutes and reset daily. The system blocks expensive models (Claude Opus, then Sonnet) as budgets are hit while always preserving access to a cheap fallback (Claude Haiku). This is a concrete, replicable answer to the AI FinOps problem that most enterprises scaling internal AI access are now facing.