Get API key →

LangChain agent budget limit, checked in middleware

PythonLangChain 1.4.2

A wrap_model_call middleware that asks AgentBill before each model call your LangChain agent makes, with the job's task_ref. When the call would pass the job's ceiling, the SDK raises, the model call does not go out, and your code decides what the job does next.

Use the call limits LangChain ships

LangChain v1 ships ModelCallLimitMiddleware and ToolCallLimitMiddleware. Each takes a thread_limit and a run_limit, and each counts calls: model calls or tool calls. LangGraph's recursion_limit caps the steps one invocation may take. Use them.

They are bound to a thread or to one invocation, and they count calls. The ceiling below is bound to a task_ref, the name of the job, in units you assign to each call. Every process, worker and agent that passes the same task_ref draws on the same number.

Install

pip install agentbill-sdk "langchain==1.4.2"

Give the job its ceiling before the agent runs, in the console or with PUT /tasks/:task_ref/ceiling, so the code below carries no budget. The samples on this page were run end to end against langchain 1.4.2 and agentbill-sdk 0.6.5, with a stub model in place of a provider.

The middleware

Before each model call it asks with the job's task_ref and what one call is worth to you. A refusal raises TaskCeilingExceededError inside the middleware, so handler(request) is never called and that model call does not go out. After the call, record settles the units preflight reserved.

import os
from dataclasses import dataclass
from agentbill import AgentBillClient
from langchain.agents import create_agent
from langchain.agents.middleware import wrap_model_call

client = AgentBillClient(api_key=os.environ["AGENTBILL_API_KEY"])
UNITS = 12  # what one model call is worth to you, in your own units


@dataclass
class Job:
    task_ref: str  # the job's name, e.g. "job-142"


@wrap_model_call
def agentbill_ceiling(request, handler):
    job = request.runtime.context.task_ref
    # Raises TaskCeilingExceededError when this call would pass the job's
    # ceiling, so handler(request) never runs and no model request goes out.
    client.preflight(agent_id="researcher", task_ref=job, estimated_units=UNITS)
    try:
        response = handler(request)
    except Exception:
        # The model call failed: release the reservation, spend nothing.
        client.record(agent_id="researcher", task_ref=job, units=UNITS, success=False)
        raise
    client.record(agent_id="researcher", task_ref=job, units=UNITS)
    return response


agent = create_agent(model, tools=tools, middleware=[agentbill_ceiling], context_schema=Job)

model is any LangChain chat model and tools your tool list. UNITS is your own estimate for one model call.

Run a job, and decide what a refusal means

from agentbill import TaskCeilingExceededError

try:
    result = agent.invoke(
        {"messages": [{"role": "user", "content": "Research the topic"}]},
        context=Job(task_ref="job-142"),
    )
except TaskCeilingExceededError as e:
    # Your code decides: keep what the agent had, retry smaller, or tell a person.
    print(f"{e.task_ref} was refused at {e.task_used_units}/{e.task_ceiling} units")

The exception carries task_ref, task_ceiling, task_used_units and task_remaining_units, so the handler can say what happened without a second call.

Or let the run finish with an answer

If you would rather the agent end normally, catch the refusal inside the middleware and return a message instead of calling the model. A reply with no tool calls ends the agent's loop, and invoke returns everything the run had until then.

from langchain.messages import AIMessage
from agentbill import TaskCeilingExceededError


@wrap_model_call
def agentbill_ceiling_or_finish(request, handler):
    job = request.runtime.context.task_ref
    try:
        client.preflight(agent_id="researcher", task_ref=job, estimated_units=UNITS)
    except TaskCeilingExceededError as e:
        # No model call. This reply ends the loop with the messages so far.
        return AIMessage(content=f"Refused at {e.task_used_units}/{e.task_ceiling} units for {e.task_ref}.")
    try:
        response = handler(request)
    except Exception:
        client.record(agent_id="researcher", task_ref=job, units=UNITS, success=False)
        raise
    client.record(agent_id="researcher", task_ref=job, units=UNITS)
    return response

Two workers, one job

Run the same job from two processes and pass the same task_ref from both. Each reservation is one conditional UPDATE on the job's row, so when the last units remain and both ask, one is approved and the other is refused. There is no lock to write in your code.

Async agents

The client is synchronous. In an async agent, write the middleware as async def and run the two calls in a thread, so the event loop keeps going while AgentBill answers. Then call agent.ainvoke.

import asyncio


@wrap_model_call
async def agentbill_ceiling_async(request, handler):
    job = request.runtime.context.task_ref
    await asyncio.to_thread(lambda: client.preflight(
        agent_id="researcher", task_ref=job, estimated_units=UNITS))
    try:
        response = await handler(request)
    except Exception:
        await asyncio.to_thread(lambda: client.record(
            agent_id="researcher", task_ref=job, units=UNITS, success=False))
        raise
    await asyncio.to_thread(lambda: client.record(
        agent_id="researcher", task_ref=job, units=UNITS))
    return response

Do not put @client.gate on an async def: the decorator is synchronous, so it would settle before the coroutine runs.

LangGraph

In a graph you build yourself, put the same pair inside the node that calls the model, and read the task_ref from the run's config.

from langchain_core.runnables import RunnableConfig
from langgraph.graph import END, START, MessagesState, StateGraph


def call_model(state: MessagesState, config: RunnableConfig):
    job = config["configurable"]["task_ref"]
    client.preflight(agent_id="researcher", task_ref=job, estimated_units=UNITS)
    try:
        reply = model.invoke(state["messages"])
    except Exception:
        client.record(agent_id="researcher", task_ref=job, units=UNITS, success=False)
        raise
    client.record(agent_id="researcher", task_ref=job, units=UNITS)
    return {"messages": [reply]}


graph = (StateGraph(MessagesState).add_node("model", call_model)
         .add_edge(START, "model").add_edge("model", END).compile())
graph.invoke({"messages": [{"role": "user", "content": "Research the topic"}]},
             config={"configurable": {"task_ref": "job-142"}})

Settling, and what a crash does

record settles the units preflight reserved for that task_ref, in the order they were reserved, so record the number you reserved. A smaller number leaves the rest held until the reservation expires, 60 minutes by default. If the model call raises, the middleware releases the reservation with success=False. If the process dies between the two calls, the reservation expires on its own and its units come back. Until then the job's ceiling is tighter, never looser.

What it does not do

Get API key →