Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 11 additions & 1 deletion docs/observability/tracing/content-capture.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -31,7 +31,7 @@ When content capture is enabled, IORails records content on three types of spans

| Span (Kind) | Attribute or Event | When Recorded |
|-------------|--------------------|---------------|
| `guardrails.request` (SERVER) | `guardrails.request.input`: JSON-encoded list of the caller's input messages | Always, while capturing |
| `guardrails.request` (SERVER) | `guardrails.request.input`: JSON-encoded list of the input messages, after any input-rail rewrite (and, on a check, any output-rail rewrite) | Always, while capturing |
| `guardrails.request` (SERVER) | `guardrails.request.output`: the text actually returned to the caller | When output is produced (a blocked request records the refusal message; an empty stream records nothing) |
| `chat <model>` (CLIENT) | LLM input and output, in the [selected format](#output-format) | Once per LLM call: the main generation call and every rail-action LLM call. On the streaming path, recorded when the stream ends; see [Streaming](#streaming) |
| `guardrails.rail` (INTERNAL) | `guardrails.rail.input`: JSON-encoded `{"messages": ..., "bot_response": ...}` snapshot | Every rail execution, while capturing |
Expand All @@ -44,6 +44,7 @@ Sampling parameters (`gen_ai.request.temperature`, `max_tokens`, and so on) and

A rails-only [`check_async()`](/run-guardrailed-inference/using-python-apis/check-messages) uses the same capture behavior.
The checked messages are recorded in `guardrails.request.input`, and the returned content is recorded in `guardrails.request.output`.
When an input or output rail rewrites the text it checked, `guardrails.request.input` carries the rewritten messages, even when a later rail blocks the check.
Every rail that runs records its own `guardrails.rail` attributes.
A check makes no main-model call, so it produces no CLIENT span for the protected model.
`IORails` captures LLM calls from rail actions as usual.
Expand Down Expand Up @@ -229,6 +230,15 @@ In both cases, if no content was produced, IORails omits the output rather than

- Content capture is **off by default**. Prompts and responses are never written to spans unless you explicitly enable it.
- Captured content can include PII or other sensitive data. Treat your tracing backend as a system that stores user content, and apply the same access controls and retention limits you would apply to any store of conversation data.
- Masking rails limit what the request span records, not what the whole trace records.
`guardrails.request.input` records the turn a masking rail checked as the rail rewrote it, even when a later rail blocks the request.
It records earlier turns of the conversation, which no rail checks, as the caller sent them.
When no mask reaches the request, it records the turn as the caller sent it: the masking rail fails, a tool-result rail blocks before the input rails run, or a stream closes or fails before its input rails finish.
On a `check_async()` call, it records the checked response as the caller sent it when no output rail runs to mask it: another rail blocks before the output rails run, or `rail_types` leaves out `RailType.OUTPUT`.
Each `guardrails.rail` span records the input its rail received, so the masking rail's own span carries the text before the mask.
A rail that calls an LLM also records the prompt it built from that input on the call's own `chat <model>` CLIENT span.
On a `check_async()` call, whose conversation includes the checked response, every rail that runs records that response as the caller sent it.
On `generate_async()` and `stream_async()`, the protected model's `chat <model>` CLIENT span records its response before an output rail masks it.
- Use `OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=false` to force capture off across every service from the environment, regardless of what individual configs request. This is useful as a deployment-wide guardrail for regulated environments.
- The [OpenTelemetry GenAI semantic conventions](https://opentelemetry.io/docs/specs/semconv/gen-ai/) recommend against capturing message content by default because of these privacy risks.

Expand Down
2 changes: 1 addition & 1 deletion docs/observability/tracing/span-reference.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -240,7 +240,7 @@ Every span sets ERROR status and records the exception when one propagates throu
| --------------------------- | ------ | ------------------ | --------------------------------------------------------------------------------------------------------------------- |
| `gen_ai.operation.name` | string | Always | The value `guardrails`, marking the request boundary. |
| `request.id` | string | Always | Request identifier derived from the OpenTelemetry trace ID. |
| `guardrails.request.input` | string | Content capture on | JSON-encoded input messages received from the caller. |
| `guardrails.request.input` | string | Content capture on | JSON-encoded input messages, after any input-rail rewrite (and, on a check, any output-rail rewrite). |
| `guardrails.request.output` | string | Content capture on | The text actually returned to the caller. On a blocked request this is the refusal message, not the raw model output. |

When speculative generation is active, the request span also carries these attributes:
Expand Down
2 changes: 2 additions & 0 deletions docs/run-rails/using-python-apis/check-messages.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -427,6 +427,8 @@ When your application requires that a rail actually ran, assert that the message
A check rejected by a full admission queue increments `guardrails.nonstream.rejections`.
For more information, refer to [Metric Reference](/observability/metrics/reference).
- Content capture records the checked messages in `guardrails.request.input` and the returned content in `guardrails.request.output` on the check's request span.
When an input or output rail rewrites the text it checked, for example to mask personal data, the span records the rewritten messages rather than the ones the caller sent, even when a later rail blocks the check.
An output rail that never runs masks nothing: when another rail blocks before the output rails run, or `rail_types` leaves out `RailType.OUTPUT`, the span records the checked response as the caller sent it.
For more information, refer to [Capturing Prompt and Response Content](/observability/tracing/content-capture).

The synchronous `check()` builds a short-lived engine with tracing and metrics disabled, so it emits no spans and no metrics. Use `check_async()` on instrumented paths.
Expand Down
6 changes: 5 additions & 1 deletion nemoguardrails/guardrails/guardrails_types.py
Original file line number Diff line number Diff line change
Expand Up @@ -103,7 +103,8 @@ class RailResult:
than a second copy that could drift. What this type adds is the aggregation
``RailOutcome`` has no concept of, because it belongs to running *many* rails:
which one blocked (``triggered_rail``), which tool calls or results it blocked
(``tool_violations``) and what every rail did (``records``).
(``tool_violations``), what every rail did (``records``), and what the rails ahead
of a block rewrote the checked text to (``rewrite_before_block``).

``records`` carries the per-rail execution records for every rail that ran in this
check (not just the blocking one), so IORails can synthesize a ``GenerationLog``.
Expand All @@ -119,6 +120,9 @@ class RailResult:
triggered_rail: str | None = None
tool_violations: tuple[ToolViolation, ...] = ()
records: tuple[RailCallRecord, ...] = field(default=(), compare=False)
# Capture data like ``records``, not part of the verdict: the block stands, but a record of
# the request that carried the text a mask removed would undo the mask.
rewrite_before_block: str | None = field(default=None, compare=False)
__hash__ = None

@property
Expand Down
54 changes: 46 additions & 8 deletions nemoguardrails/guardrails/iorails.py
Original file line number Diff line number Diff line change
Expand Up @@ -592,6 +592,30 @@ def _reports_assistant_content(rails_to_run: list[str]) -> bool:
return "tool_call" in rails_to_run and "input" not in rails_to_run


def _rewrite_last_assistant_message(messages: LLMMessages, text: str) -> LLMMessages:
"""Return a copy of *messages* with the last assistant message's content rewritten to *text*."""
for index in range(len(messages) - 1, -1, -1):
if messages[index].get("role") == "assistant":
rewritten = list(messages)
rewritten[index] = {**messages[index], "content": text}
return rewritten
raise ValueError("no assistant message to rewrite")


def _apply_input_rewrite_before_block(messages: LLMMessages, blocked: RailResult) -> LLMMessages:
"""Return *messages* with the user turn as the input rails rewrote it before one of them blocked."""
if blocked.rewrite_before_block is None:
return messages
return rewrite_user_message(messages, blocked.rewrite_before_block)


def _apply_output_rewrite_before_block(messages: LLMMessages, blocked: RailResult) -> LLMMessages:
"""Return *messages* with the checked response as the output rails rewrote it before one of them blocked."""
if blocked.rewrite_before_block is None:
return messages
return _rewrite_last_assistant_message(messages, blocked.rewrite_before_block)


def _blocked_message(result: RailResult) -> str:
"""The text a blocked turn returns, which says whether the rail broke or fired."""
if result.failed:
Expand Down Expand Up @@ -632,7 +656,7 @@ class _TurnConversation:

@dataclass(frozen=True, slots=True)
class _GeneratedTurn:
"""The main model's response and the messages it read, or no response when the rails blocked.
"""The main model's response and the messages it read, or on a block no response and the messages the rails left.

``blocked_by`` carries the verdict that stopped the turn, because the caller renders the
message from it and a rail that broke reads differently from one that fired.
Expand Down Expand Up @@ -1046,8 +1070,8 @@ async def _run_generate(
log.error("[%s] generate_async failed time=%.1fms", req_id, elapsed_ms, exc_info=True)
raise
# Captured at the traced_request boundary, so a future early-return in _do_generate
# is covered. The messages are the ones the model read: a span carrying the text a
# mask removed would defeat the mask.
# is covered. The messages are the ones the rails left, which the model read unless a
# rail blocked: a span carrying the text a mask removed would defeat the mask.
if self._content_capture_enabled:
set_request_content(request_span, conversation.messages, _response_content_for_capture(result))
elapsed_ms = (time.monotonic() - t0) * 1000
Expand Down Expand Up @@ -1132,6 +1156,9 @@ def _blocked_return(blocked_by: RailResult) -> Union[LLMMessage, GenerationRespo
messages, req_id, llm_kwargs, input_enabled=input_enabled, records_out=records
)

# Before the blocked return, so a block behind a mask still records the masked messages.
if conversation is not None:
conversation.messages = turn.messages
if turn.blocked_by is not None:
return _blocked_return(turn.blocked_by)

Expand All @@ -1143,8 +1170,6 @@ def _blocked_return(blocked_by: RailResult) -> Union[LLMMessage, GenerationRespo
# generation bug to the caller as a guardrail decision.
raise RuntimeError("generation returned no response without naming a blocking rail")
messages = turn.messages
if conversation is not None:
conversation.messages = messages

# Log raw content before reasoning extraction and think-token removal
log.debug("[%s] Raw LLM response: %s", req_id, truncate(response.content))
Expand Down Expand Up @@ -1235,7 +1260,8 @@ async def _do_generate_sequential(
log.info("[%s] Input blocked: %s", req_id, display_reason(input_result))
if self._metrics_enabled:
record_request_blocked(RailDirection.INPUT)
return _GeneratedTurn(response=None, messages=messages, blocked_by=input_result)
rewritten_messages = _apply_input_rewrite_before_block(messages, input_result)
return _GeneratedTurn(response=None, messages=rewritten_messages, blocked_by=input_result)

rewritten = _rewritten_user_message(input_result)
if rewritten is not None:
Expand Down Expand Up @@ -1466,16 +1492,19 @@ async def _run_check(
) -> RailsResult:
"""Queue-worker entry for ``check_async``: wrap the rails in a request span."""
tracer = self._tracer if self._tracing_enabled else None
conversation = _TurnConversation(messages=messages)
with traced_request(tracer) as (request_span, req_id):
t0 = time.monotonic()
try:
result = await self._do_check(messages, rail_types, req_id, tools=tools)
result = await self._do_check(messages, rail_types, req_id, tools=tools, conversation=conversation)
except Exception:
elapsed_ms = (time.monotonic() - t0) * 1000
log.error("[%s] check failed time=%.1fms", req_id, elapsed_ms, exc_info=True)
raise
# The messages as the rails masked them, as _run_generate captures them. A check's
# messages carry the checked response too, so an output mask must reach them as well.
if self._content_capture_enabled:
set_request_content(request_span, messages, result.content)
set_request_content(request_span, conversation.messages, result.content)
elapsed_ms = (time.monotonic() - t0) * 1000
log.info(
"[%s] check completed time=%.1fms status=%s",
Expand All @@ -1492,6 +1521,7 @@ async def _do_check(
req_id: str,
*,
tools: Optional[list[dict]] = None,
conversation: "_TurnConversation",
) -> RailsResult:
"""Core check pipeline: run the requested input, output and tool rails on messages."""
log.info("[%s] check called", req_id)
Expand Down Expand Up @@ -1533,11 +1563,13 @@ async def _do_check(
log.info("[%s] Input blocked: %s", req_id, display_reason(input_result))
if self._metrics_enabled:
record_request_blocked(RailDirection.INPUT)
conversation.messages = _apply_input_rewrite_before_block(messages, input_result)
return _blocked_check_result(input_result)
rewritten = _rewritten_user_message(input_result)
if rewritten is not None:
log.info("[%s] Input rails rewrote the user message", req_id)
messages = rewrite_user_message(messages, rewritten)
conversation.messages = messages
if not reports_output:
pass_content = rewritten
else:
Expand All @@ -1559,11 +1591,14 @@ async def _do_check(
log.info("[%s] Output blocked: %s", req_id, display_reason(output_result))
if self._metrics_enabled:
record_request_blocked(RailDirection.OUTPUT)
conversation.messages = _apply_output_rewrite_before_block(messages, output_result)
return _blocked_check_result(output_result)
rewritten = _rewritten_bot_message(output_result)
if rewritten is not None:
log.info("[%s] Output rails rewrote the response", req_id)
pass_content = rewritten
messages = _rewrite_last_assistant_message(messages, rewritten)
conversation.messages = messages
else:
log.info("[%s] Output rails requested but no assistant content to check; skipping", req_id)

Expand Down Expand Up @@ -1734,6 +1769,9 @@ async def _generation_task(request_span):
log.info("[%s] Input blocked: %s", req_id, display_reason(input_result))
if self._metrics_enabled:
record_request_blocked(RailDirection.INPUT)
# The request span records ``messages`` once the stream ends, so a mask applied
# ahead of the block has to reach it before the caller can see the stream close.
messages = _apply_input_rewrite_before_block(messages, input_result)
await streaming_handler.push_chunk(
self._guardrails_violation_payload(
f"Blocked by input rails: {client_reason(input_result)}", "input_rails"
Expand Down
25 changes: 19 additions & 6 deletions nemoguardrails/guardrails/rails_manager.py
Original file line number Diff line number Diff line change
Expand Up @@ -263,6 +263,18 @@ def _result_after_rewrites(
return RailResult(RailOutcome.transform([(_REWRITABLE_TARGET[direction], final_text)]), records=records)


def _blocked_after_rewrites(
blocked: RailResult,
original_text: str,
final_text: str,
records: tuple[RailCallRecord, ...],
) -> RailResult:
"""Keep the block as the verdict, and what the rails ahead of it rewrote, so a record of the request keeps a mask."""
if final_text == original_text:
return replace(blocked, records=records)
return replace(blocked, records=records, rewrite_before_block=final_text)


def _model_free_record(
flow: str, rail_type: str, result: RailResult, tool_name: Optional[str] = None
) -> RailCallRecord:
Expand Down Expand Up @@ -962,26 +974,27 @@ async def _run_rails_sequential(
log.debug("[%s] %s flow %s result %s", req_id, direction.value, flow, result)
if not result.is_safe:
log.info("[%s] %s flow %s blocked", req_id, direction.value, flow)
return replace(result, records=tuple(collected))
return _blocked_after_rewrites(result, original_text, final_text, tuple(collected))
if result.outcome.is_transform:
final_text = _rewritten_text(result.outcome, direction, flow)
rewritten_text = _rewritten_text(result.outcome, direction, flow)
log.info("[%s] %s flow %s rewrote the text it checked", req_id, direction.value, flow)
if direction is RailDirection.INPUT:
try:
messages = rewrite_user_message(messages, final_text)
messages = rewrite_user_message(messages, rewritten_text)
except ValueError:
# Blocking keeps a misbehaving rail inside the fail-closed envelope,
# rather than failing the request as a server error.
log.error(
"[%s] %s flow %s rewrote a turn this request does not have", req_id, direction.value, flow
)
return RailResult.block(
blocked = RailResult.block(
reason="a rail rewrote a message this request does not have",
triggered_rail=_get_flow_name(flow) or flow,
records=tuple(collected),
)
return _blocked_after_rewrites(blocked, original_text, final_text, tuple(collected))
else:
bot_response = final_text
bot_response = rewritten_text
final_text = rewritten_text
return _result_after_rewrites(direction, original_text, final_text, tuple(collected))

async def _run_tool_rails_sequential(
Expand Down
Loading
Loading