Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 5 additions & 0 deletions _config.yml
Original file line number Diff line number Diff line change
Expand Up @@ -37,5 +37,10 @@ callouts:
# Makes Aux links open in a new tab. Default is false
aux_links_new_tab: true

# Enable mermaid diagrams in fenced ```mermaid code blocks.
# https://just-the-docs.com/docs/ui-components/code/#mermaid-diagram-code-blocks
mermaid:
version: "10.9.0"

kramdown:
syntax_highlighter: coderay
187 changes: 187 additions & 0 deletions technical/guardrails.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,187 @@
---
title: Guardrails
parent: Technical documentation
has_children: false
nav_order: 7
---

# LiteLLM guardrails
Comment thread
cableman marked this conversation as resolved.

LiteLLM guardrails are a way to execute code on and around input and output sent to and from models managed by LiteLLM.

The Open WebUI instance sends requests to LiteLLM, but does not support setting guardrails on a per-request basis. For
this reason, guardrails that should execute on messages from Open WebUI should be set as always on, by setting the
`default_on` setting to `true` when declaring the guardrail (see below).

## Currently, shipped guardrails

| Guardrail | Type | File | What it does |
|----------------------------|----------|------------------------------------------------------------------------------------------------------------------------------------------|---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| `MessageTrimmingGuardrail` | Pre-call | [`message-trimming`](https://github.com/os2ai/helm-deployments/blob/develop/applications/litellm/templates/message-trimming-config.yaml) | Trims oversized message histories to fit the target model's context window, then sanitizes tool-call/tool-response pairings so the trimmed (or otherwise broken) history doesn't crash strict chat templates. |

Pre-call guardrails in [LiteLLM](https://github.com/BerriAI/litellm) proxy applies to inbound chat requests before
forwarding them to the upstream model.

## Usage

The message trimming guardrail can be configured in the litellm
values [file](https://github.com/os2ai/helm-deployments/blob/develop/applications/litellm/litellm-values.yaml#L108)
configuration file in the helm chart.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A note on why we need message trimming would make sense after this paragraph, just very briefly. E.g. what happens if oversized message histories are not trimmed, and how does the guardrail avoid it? In a sentence of two.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added comment about this to the doc


Every model has a fixed context window. If a conversation's message history grows beyond it, the upstream model rejects
the whole request with an error, breaking the conversation. The guardrail prevents this by trimming older messages so the
history always fits within the model's context window before the request is forwarded.

__Note__: that the default configuration contains an option to set max tokens for a named model, which overrides the
global default max tokens value. This is useful for models that have a different context window size than the global
default.

```yaml
model_list:
- model_name: my-model
litellm_params:
model: openai/some-deployed-model
api_base: https://...
api_key: ""
max_tokens: 8192
guardrails:
# attach the guardrail to this model
- message_trimming

guardrails:
- guardrail_name: message_trimming
litellm_params:
guardrail: /app/custom_guardrails/message_overflow.MessageTrimmingGuardrail
mode: pre_call
default_on: true
default_config:
trim_ratio: 0.75
max_output_tokens: 2000
safety_buffer: 500
debug: false
default_max_context_tokens: 8192
max_context_tokens_by_model:
openai/some-deployed-model: 32768
pop_trailing_tool_messages: false
Comment thread
cableman marked this conversation as resolved.
```

For information about the configuration options, see the [configuration reference](#configuration-reference) section.

## How Message Trimming works

Asking the model to generate more tokens than its context window can hold can be fatal for the entire conversation: the
request is rejected outright. To avoid this, the guardrail takes a conservative approach and computes a __safe completion
budget__ — the number of tokens left for the model's reply once the input messages, a configurable safety buffer, and a
headroom factor for tokens the provider may add later have all been subtracted from the context window. It then caps the
request's `max_tokens` at that budget. The trade-off is deliberate: erring on the low side may occasionally give the
model slightly less room to answer than it could technically use, but it guarantees the request fits and the
conversation survives.

`async_pre_call_hook` runs on every chat completion request. The flow:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Before the step-by-step flow, I think a sentence stating what the pros and cons of our approach is would make sense.

E.g. "Sending a too large message to the model can be fatal for the entire conversation, so we take a conservative approach in estimating a safe completion budget for the message" and then explain what a safe completion budget is, why we calculate it as is? I think that would give a lot of good context for evaluating the approach.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Intro section added


1. __Resolve context window__ for the target model (`_resolve_max_context_tokens`):
per-model override map → `litellm.get_max_tokens` → global default. Logs a warning if it falls through to the global
default.
2. __Compute a safe completion budget__ (`_calculate_safe_completion_tokens`) — leaves room for input + safety buffer +
a 25% headroom factor for tokens LiteLLM/the provider may add later.
3. __Update `max_tokens` / `max_completion_tokens`__ in the request so the model can't be asked for more than fits.
4. __Trim input messages__ (`litellm.trim_messages`) if `current_input_tokens > max_input_tokens`, dropping older
messages from the head until it fits.
Comment thread
cableman marked this conversation as resolved.
5. __Sanitize__ (`_sanitize_messages`):
- `_repair_tool_call_pairings` — strip orphan `role: tool` messages and orphan `tool_calls` entries that the trimmer
may have created.
- (Optional, opt-in via `pop_trailing_tool_messages`) pop trailing `role: tool` messages and re-run the repair, then
append `"Please continue"` if the new terminus is an assistant message.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"Please continue" feels like it could skew the output, especially if the language in the context window otherwise isn't English?

6. __Recount and re-budget__ completion tokens once more, since sanitize may have grown or shrunk the message list.

### Why `_repair_tool_call_pairings` exists

LiteLLM's build in `trim_messages` has __no tool-call awareness__ — it drops messages by token count from the head and
freely produces:

- Orphan `role: tool` messages (no surviving `assistant.tool_calls` advertised them).
- Orphan `tool_calls` entries on assistant messages (no surviving `role: tool` answered them).

Both shapes are rejected by strict chat templates (Mistral, vLLM, OpenAI strict mode). The repair pass enforces the
invariant: every surviving `tool_calls[].id` has a later matching `role: tool` message, and every surviving `role: tool`
was advertised by an earlier surviving `assistant.tool_calls` entry. See `_repair_tool_call_pairings` in [`message-trimming`](https://github.com/os2ai/helm-deployments/blob/develop/applications/litellm/templates/message-trimming-config.yaml).

### Why the trailing-tool pop is opt-in

The "normal" agent-loop shape ends on a `role: tool` message:

```mermaid
flowchart LR
U[User] --> A["Assistant{tool_calls}"]
A --> T["Tool{result}"]
T --> C([model is asked to continue here])
```

Here the final `role: tool` message holds the result of the last tool call, and the model is asked to reason from that
result. The `pop_trailing_tool_messages` setting controls whether the guardrail keeps or removes that trailing result
before forwarding the request:

- __Disabled (`false`, the default)__: the trailing `role: tool` message is left in place. Most providers (OpenAI,
Anthropic, Google, Mistral via the official APIs) __accept__ this shape — that's how tool calling works — so the model
receives the tool-call result and reasons from it as intended.
- __Enabled (`true`)__: the guardrail strips the trailing `role: tool` message(s), and if the history then ends on an
assistant message it appends a `"Please continue"` user message in their place. This removes the tool-call result the
model was supposed to reason from, so it has to continue without ever seeing the tool's output.

Because losing the tool-call result degrades the response, the default is __off__. Set `pop_trailing_tool_messages: true`
only for upstream chat templates that explicitly reject `role: tool` messages — notably the strict HuggingFace template
that raises `"Only user and assistant roles are supported!"`. The per-model override map lets you flip it for one model
in a fleet without affecting the others.

### Why both repairs run when pop is enabled

The order is `repair → pop → repair → maybe-append-continue`:

- The first repair cleans up orphans created by `trim_messages`.
- The pop may break a previously-valid `[Assistant{tool_calls=[X]}, Tool X]` pair, leaving the assistant holding orphan
`tool_calls`.
- The second repair restores the invariant — strips the now-orphan `tool_calls`, drops content-empty assistants
entirely.
- *Then* we decide whether to append `"Please continue"`, after seeing the post-repair terminus. (Appending before would
risk leaving a stale "user-continue" line after a now-deleted assistant.)

## Configuration reference

Read from `default_config` of the guardrail entry in `litellm_config.yaml`. All keys optional.

| Key | Type | Default | Purpose |
|---------------------------------------|-------|---------|---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| `trim_ratio` | float | `0.75` | Forwarded to `litellm.trim_messages`. Fraction of `max_tokens` that trimming aims for, leaving headroom for additions later in the pipeline. |
| `max_output_tokens` | int | `2000` | Default completion budget when the request specifies neither `max_tokens` nor `max_completion_tokens`. |
| `safety_buffer` | int | `500` | Reserved tokens carved out of the context window before computing input/output budgets — covers system prompts, function schemas, and other tokens added downstream. |
| `debug` | bool | `false` | When `true`, the guardrail prints `[GUARDRAIL]`-prefixed traces to stdout. Show up in `task compose -- logs -f litellm`. |
| `default_max_context_tokens` | int | `8192` | Fallback context-window size when neither `max_context_tokens_by_model` nor `litellm.get_max_tokens` resolves the model. __Bump this if your fleet's smallest model is bigger than 8k.__ |
| `max_context_tokens_by_model` | dict | `{}` | Per-model overrides keyed by the upstream `model:` value LiteLLM forwards (NOT the friendly `model_name`). Wins over `litellm.get_max_tokens`. Use this for vLLM, Bedrock variants, custom deployments — anything not in [`litellm/model_prices_and_context_window.json`](https://github.com/BerriAI/litellm/blob/main/model_prices_and_context_window.json). |
| `pop_trailing_tool_messages` | bool | `false` | Strip trailing `role: tool` messages before forwarding. __Leave `false` unless the upstream chat template rejects them__ — popping loses tool-call results the model needs to reason from. |
| `pop_trailing_tool_messages_by_model` | dict | `{}` | Per-model override of the flag above, same key shape as `max_context_tokens_by_model`. |

### Resolution order, illustrated

__Context window__ — first hit wins:

```mermaid
flowchart TD
A["max_context_tokens_by_model[model]"] -->|miss| B["litellm.get_max_tokens(model)"]
B -->|raises / 0| C[default_max_context_tokens]
A -. hit .-> H((use value))
B -. hit .-> H
C --> H
```

__Pop trailing tools__ — first hit wins:

```mermaid
flowchart TD
A["pop_trailing_tool_messages_by_model[model]"] -->|miss| B[pop_trailing_tool_messages]
A -. hit .-> H((use value))
B --> H
```

## References

- [LiteLLM custom guardrail docs](https://docs.litellm.ai/docs/proxy/guardrails/custom_guardrail)
Loading