# Understanding Concierge context limits

> 

Source: https://help.kustomer.com/en_us/understanding-concierge-context-limits-BkHCXHjwGe

Last updated: 2026-09-16T16:23:42.656Z

Every procedure, tool, and instruction you add to a Concierge agent adds more context for the model to work with at once. Like a human overloaded with rapid context switching, beyond a certain point, that extra context works against you:

*   The agent can lose track of which procedure to follow or skip steps in it.
*   Tool selection gets less reliable when the agent has many similar tools to choose from.
*   Responses can become slower and less consistent.

Smaller, well-scoped agents — or a team of specialist agents that each handle a distinct intent — produce more accurate and consistent responses. Save-time size limits keep Concierge agents in that shape. They apply the same way on every plan.

When you save or deploy a Concierge agent, Kustomer checks its size against a set of limits. If the agent, a tool, or a procedure is too large, the save or deploy is blocked until you reduce its size.

These limits apply only when a Concierge agent is saved or deployed. They don't apply to a tool sitting in the shared Tools library on its own — a tool is only measured once a Concierge agent references it.

## In this article

*   [How Concierge measures size](#how-concierge-measures-size)
*   [Size limits at a glance](#size-limits-at-a-glance)
*   [What counts toward each limit](#what-counts-toward-each-limit)
*   [Fixing a blocked save or deploy](#fixing-a-blocked-save-or-deploy)

### How Concierge measures size

Concierge estimates size in tokens, using an approximation of roughly one token per four characters. This estimate isn't a precise count of what the underlying model sees, but it gives you a consistent, predictable way to tell how close you are to a limit — and it doesn't change depending on which model powers the agent.

### Size limits at a glance

Limit

Default

What's measured

Agent size

5,000 tokens

The agent's Name, Description, Identity, Custom Guidance, and Tone, plus any unique shared tools @-mentioned in Custom Guidance.

Tools on an agent

20 tools

Unique shared tools referenced from the agent's procedures and Custom Guidance @-mentions. Built-in/system tools and child agents don't count.

Procedures on an agent

10 procedures

The number of procedures on that agent.

Single tool size

5,000 tokens

That tool's name, description, and input schema — only counted if a Concierge agent references it.

Single procedure size

5,000 tokens

That procedure's text, plus any unique shared tools @-mentioned in it.

All procedures combined

25,000 tokens

The text of all procedures on the agent, plus all unique tools referenced across them.

### What counts toward each limit

A few details worth knowing as you build an agent:

*   If a tool is @-mentioned in both Custom Guidance and a procedure, it counts toward **both** agent size and the combined-procedures total.
*   Procedure text and any child (delegated) agent's content aren't counted toward the parent agent's size — procedures have their own separate size caps instead. This is one reason splitting work across a multi-agent team is an effective way to stay under the limits.
*   These save-time limits measure what you configure, not what happens during a live conversation. Knowledge-base results, tool outputs, and conversation history aren't part of this check, and can still grow while a conversation is in progress. Staying under these limits doesn't cap the size of an individual conversation.

**Note:** All calls made by Concierge have a default runtime limit of 64k input tokens.

### Understanding the runtime context limit

Concierge enforces two separate limits on how much context an AI agent can use, and they're checked at different times:

*   The **save-time limit** checks the size of an agent's configuration when you build or edit it. See the section above for details.
*   The **runtime context limit** checks the size of the _assembled context_ — everything sent to the model in a single call — while a conversation is actually running. An agent configuration can pass the save-time check and still hit the runtime limit, because tool output, knowledge base results, conversation history, and scratchpad content all count toward it and none of that exists until the conversation is in progress.

### How the runtime context limit is calculated

The default cap is 64,000 estimated tokens per assembled request. Kustomer estimates token count as non-attachment UTF-8 bytes divided by four; attachments themselves aren't counted toward the cap. A request at exactly 64,000 is allowed — anything above it exceeds the limit.

When context limits are exceeded, the following occurs:

*   Concierge doesn't call the model for that turn.
*   The customer sees a message letting them know their request has been routed to an agent, and the conversation falls back to QnR escalation.
*   On the Traces view in the Automations [Observation sidebar](https://help.kustomer.com/en_us/understand-observability-in-automations-B1CMeXm1ye), the conversation shows a `runtime_context_limit_exceeded` error with details on what was exceeded.

### Fixing a blocked save or deploy

If any limit is exceeded, Kustomer blocks the save or deploy and shows an error identifying which limit was hit, along with the measured value and the allowed value. Each error points to a specific fix:

Error message

What it means

How to fix it

"The agent is too large (measured/allowed tokens)."

Name, Description, Identity, Custom Guidance, and Tone — plus @-mentioned tools — add up to more than the agent size limit.

Shorten Custom Guidance, Identity, or Tone text. Remove tools that are @-mentioned but not actually needed.

"The agent has too many tools attached (measured/allowed)."

The agent references more unique shared tools than the tools-per-agent limit allows. Tool size limits apply only to tools referenced by an AI automation (not those references elsewhere, such as in code procedures).

Remove tools the agent doesn't need. If the agent genuinely needs many distinct tools, consider splitting it into a multi-agent team where each child agent owns a smaller and targeted set.

"The agent has too many procedures (measured/allowed)."

The agent has more procedures than the procedures-per-agent limit allows.

Combine or remove procedures that overlap. If the agent covers several distinct intents, consider splitting it into a multi-agent team with focused proedures — each child agent gets its own procedure count.

"The procedures are too large (measured/allowed tokens)."

All procedures on the agent, plus the tools they reference, add up to more than the combined-procedures limit.

Trim procedure steps down to what's essential, and remove tools @-mentioned in procedures that aren't needed. Splitting procedures across a multi-agent team also reduces what any one agent has to carry.

"The tool _{name}_ is too large (measured/allowed tokens)."

That tool's name, description, and input schema exceed the single-tool size limit.

Narrow the tool's description and simplify its input schema.

"The procedure _{name}_ is too large (measured/allowed tokens)."

That procedure's text, plus its @-mentioned tools, exceed the single-procedure size limit.

Trim that procedure's steps, and remove any @-mentioned tools it doesn't need.

If you see a general message instead — "This configuration exceeds Concierge limits." — one or more of the checks above failed without a more specific message. Start by simplifying Custom Guidance, procedures, and tools, and splitting into a multi-agent team if the agent covers multiple distinct jobs.

A child agent's own content isn't counted against its parent agent's size, which is why splitting work across a multi-agent team is a real way to reduce size — not just a workaround.
