For agencies and builders who run agents for clients
See what every client's agents cost you, before the invoice does
Per agent, per client, per job, estimated in dollars at list price, with a monthly report to price and bill from. Give any job a ceiling, and a call whose estimate would cross it gets approved: false before it runs. Your code decides whether to stop, skip, or replan.
Start free · 1,000 preflight calls/mo, no cardFree, no card. Sign up here, connect your code later from your laptop.
next call asks $1.80, $0.08 left
raised in your process
Unattended agents need a job ceiling
Month caps, org caps, and session or window budgets are real. They meter an account, a clock, or one session. A job ceiling meters one job, across processes and providers, with no reset: a call is checked against it only if it asks preflight with the job’s name, and a call whose estimate would cross it gets approved: false. A call that uses more than it estimated can land past by the difference; the next one is refused.
An org cap is monthly: either tonight’s loop fits under it, or every agent in the org is refused until the cap resets or someone raises it.
Estimate a run
What could one unattended job cost you?
Start from our example, then put in your numbers. No signup. The math runs in your browser.
Here every call costs your rate. On a real job, AgentBill prices each call at list price from the tokens it reports.
The inputs need JavaScript. The math is calls in one run × your cost per call.
2,000 calls × $0.05: example inputs until you change them. An estimate, not a measurement.
The 500 approved calls: about $25.00 at your rate. What runs after that is your code’s decision.
Demo · runs in your browser
Watch one job reach its ceiling
A ten-call plan shares job-142 and a ceiling of $5.00. Each call asks for what it costs at list price.
On a ceiling refusal the SDK raises in your process, and the branch you wrote picks one.
Runs in your browser, against no account. Same rule as POST /preflight on a job in dollars, and its dollar fields; the real answer carries a few more.
import os
from agentbill import (
AgentBillClient, TaskCeilingExceededError)
client = AgentBillClient(
api_key=os.environ["AGENTBILL_API_KEY"])
def critique(draft):
try:
r = client.preflight(
agent_id="researcher",
task_ref="job-142",
estimated_usd=1.80)
except TaskCeilingExceededError:
return draft # your code decides
# our quota ran out, not your ceiling
if not r.approved:
return draft
# approved: call your provider, then record
import { preflight, TaskCeilingExceededError }
from 'agentbill'
// Reads AGENTBILL_API_KEY from env.
async function critique(draft) {
try {
const r = await preflight({
agentId: 'researcher',
taskRef: 'job-142',
estimatedUsd: 1.8 })
// our quota ran out, not your ceiling
if (!r.approved) return draft
} catch (e) {
// your code decides
if (e instanceof TaskCeilingExceededError) return draft
throw e
}
// approved: call your provider, then record
}
Three steps, no proxy
1 Give the job a ceiling
Name the job and set its ceiling in dollars in the console. In code, that name is the task_ref, and PUT /tasks/:task_ref/ceiling with ceiling_usd in the body sets the same number.
2 Ask before each call
Call preflight with the same task_ref before your provider call. An approved call reserves its estimate, so two calls racing cannot both take the last room.
3 See what each client cost
wrap() your OpenAI or Anthropic client once: it asks preflight before each call and records the tokens after, at list price, per agent and per customer_id, and the monthly report bills from those. No base URL to change, no provider traffic through us, no provider keys held. If we are unreachable, the SDK raises in your process and your code decides.
Working in Claude, ChatGPT, Cursor or Codex? Connect via MCP → One URL, and preflight is a tool your agent can call.
Running OpenClaw? OpenClaw plugin on ClawHub → One ceiling per session, checked before every model turn and tool call. How it works.
Built for agent loops you own and leave running
- Agent work you bill clients for
- Coding agents you built
- Overnight research jobs
- Scheduled agent runs
- Retry loops you wrote
- Batch pipelines
Examples, not integrations or customers. A call is checked only when your code asks preflight.
~$300 of unintended API usage over one unattended weekend, under a $3,000/month Enterprise limit. The limit is a month. The incident was a weekend.
claude-code issue 64744, opened June 2, 2026. Open, labelled bug and area:cost; a workaround is posted in its comments.
github.com/anthropics/claude-code/issues/64744Pricing
Free to start. Cheap enough to leave on.
Every plan has every feature. The tiers differ in how many preflight calls a month they include, and in who answers when you write in.
Every feature. Direct line to the founder.
Get ScaleWhat AgentBill does not do
- Sit in your request path. No proxy, no base URL to change, no provider keys held. A call that never asks preflight, or a retry buried in a library, is never checked against the ceiling before it runs. If it records, its cost still counts.
- Read your provider bill. No invoice access. A dollar figure appears only as an estimate at public list price, on calls wrap() measured. The estimator above is your rate times your count, run in your browser. A refused call is not money saved: what it would have gone on to cost is unknown.
- Reach into a running job. Preflight answers approved: false, and on a ceiling refusal the SDK raises. Your code decides what happens next.
- Undo what already ran. Calls are refused, not reversed, and the ceiling is keyed on (account_id, task_ref): a loop that opens a new task_ref gets a new ceiling.
Where does the dollar figure come from?
From the model and the tokens your provider reports on each call wrap() measures, priced at public list price from a dated snapshot of the LiteLLM price table. It is an estimate: contract discounts, batch and regional pricing and server-side tool fees are not in it, and a call with no list price is left out of the figure, never counted as $0.
Does AgentBill see my provider bill?
No. It never has access to your OpenAI, Anthropic or cloud account, and it never reads your invoice. The dollars it shows are its own estimate at list price, from the tokens each call reports, so your invoice may differ.
How is a job ceiling different from a monthly spend cap?
A monthly cap meters an organization or a project over the month, and some platforms also cap one session inside their own runtime. A job ceiling is a name your code passes: a call is checked against it only if it asks preflight with that name, from any process and for any provider, and the number has no reset.
Does AgentBill sit in my request path?
No. It is an endpoint your code calls before it calls a provider, not a gateway your traffic routes through. Nothing to point your base URL at and no third party holding your provider keys. If AgentBill is unreachable, preflight raises in your process after its timeout, and your code decides whether to call the provider anyway.
What happens if a job crashes with money still reserved?
Preflight reserves the estimate it approves, so two calls racing cannot both be told there is room for one. A reservation that is never settled expires after 60 minutes and is swept back to the budget every five minutes. Nothing is held forever because a process crashed, and nothing is released before its 60 minutes are up, however slow the process.
What happens when I reach the free tier's 1,000 calls?
Preflight starts answering approved: false with reason free_tier_exceeded (plan_limit_exceeded on a paid plan) and an upgrade_url. The SDK returns that answer instead of raising, so check result.approved. A refused call reserves nothing and is not counted.
Can a ceiling count something other than dollars?
Yes. A job can be opened in tokens, or in units you define and pass, where a unit is whatever you decide it is worth. On a job in dollars, a call with no list price is charged its reservation against the ceiling, never $0.
Give one job a ceiling.
Free tier, no card, 1,000 preflight calls a month.