Been building with OpenAI agents for a while and kept running
into the same problem — an agent gets stuck in a loop and burns
through API credits before anyone notices.
Tried a few things: setting account-level limits, adding logging,
wrapping calls in try/except. None of it actually stops the charge
before it fires.
Eventually built a small tool that checks the budget before each
OpenAI call and blocks it if the agent is over limit. Not a
soft warning — the call just doesn’t happen.
Been using it in my own projects and it’s helped. Happy to share
more details if anyone’s dealing with the same thing.
Curious how others are handling this — is there a pattern I’m
missing, or is everyone just setting account caps and hoping?
The pre-call block is the right instinct, but the part I’d want to know more about is where the running total lives. If it’s in-process memory, concurrent agent runs (or a retry that fires two calls close together) can both read the same “under budget” state before either write lands, and you blow past the cap anyway - classic check-then-act race. Doesn’t matter how tight the check is if two calls can pass it at once.
The other edge case I’ve run into: streaming responses. You know the cost of the request tokens before the call, but not the completion cost until the stream actually finishes - so a hard pre-call block only catches costs you could predict in advance, not a response that runs long. Curious whether your tool accounts for either of those, or if it’s mainly for single-threaded scripts where that’s a non-issue.
On the race: you’re right that if the budget check and the write live as two separate steps against an in-process counter, two concurrent calls can both read “under budget” before either commits its spend, and both get through. The fix isn’t a tighter check, it’s not having two steps at all - the reservation itself has to be the atomic operation. Concretely: instead of “read balance, compare, then decrement,” it’s a single atomic decrement-if-sufficient (a SELECT … FOR UPDATE / compare-and-swap against one authoritative counter, or a real ledger hold, depending on your infra). The second call’s decrement simply fails, because the first one already took the balance below the threshold - there’s no window to race in because there’s no read-then-write gap.
On streaming: I don’t try to know the exact cost before the call, I reserve for the worst case and reconcile after. At call time, take a hold for the max plausible cost (usually derived from max_tokens x model pricing), decrement the budget by that hold amount immediately using the same atomic mechanism, then release the difference once the actual completion cost is known. Worst case you’re briefly more conservative mid-run than you needed to be - you never exceed the real budget, because enforcement never trusted a number it didn’t have yet.
Both only work if the hold/decrement happens at the layer that’s actually authoritative for spend, not in app code trying to reconstruct state after the fact. That’s the part I ended up pulling out into its own layer instead of bolting onto the agent loop.
The atomic decrement-if-sufficient point is the key correction - I was still thinking about this as “read fast enough that the race window is small” rather than removing the window entirely. Compare-and-swap against one authoritative counter with no intermediate read makes the race actually impossible instead of just less likely.
The hold-and-reconcile pattern for streaming cost is doing something I didn’t have a clean name for - conservative-then-refunded is exactly right when you can’t know the true cost upfront. Pulling that into its own authoritative layer rather than bolting it onto the agent loop is probably the part most people skip, since it’s more upfront work than an in-process counter, but it’s the difference between a system that’s actually correct under concurrency and one that just happens to work in testing.