This guide explains how to instrument each layer of Fitz with comprehensive observability (logging, metrics, distributed tracing).
Sampling Strategy:
- Hot paths (routing, TLV codec, frame I/O): 0.1% sampling (1 in 1000)
- Critical paths (auth, permission checks, request boundaries): 100% sampling
- Debug paths (family actor workers, internal operations): Debug level (hidden by default)
Metric Types:
- Counters: Increment-only (frames_received, errors_total, operations_total)
- Gauges: Point-in-time values (active_connections, mailbox_depth, pending_operations)
- Histograms: Latency distributions (message_latency_ms, codec_latency_us, operation_latency_ms)
Key Attributes to Record:
message_id,parent_message_id(causation chains)route,domain,realm,area(routing context)session_id,connection_id(session context)operation,error_type(operation context)
Health Endpoints:
/livezreports process liveness only and should drive restart policies./targetzreports scheduling/orchestration health for a non-serving standby and does not require the Midge writer lease./readyzreports native readiness for platforms that support a dedicated readiness probe./healthzmirrors readiness and should drive external traffic admission on platforms that only support one HTTP health check./startupzreports whether Fitz has completed startup and is safe to begin liveness checks.
During planned redeploy drain, /livez remains 200, while /targetz,
/readyz, and /healthz return 503 with accepting_target_traffic or
accepting_traffic set to "draining". Customer traffic admission must use
/healthz; a separate orchestration path may use /targetz to coordinate a
standby. Fitz closes existing ephemeral sessions after
FITZ_DRAIN_GRACE_SECONDS.
While a replacement is waiting for the Midge writer lease, /livez and
/targetz return 200; /startupz, /readyz, and /healthz return 503.
Expected contention retries indefinitely with capped exponential backoff.
Never configure the customer-facing ALB target group to check /targetz: ALB
would route traffic to the waiting task even though its data plane is closed.
A standard one-target-group ECS rolling deployment therefore cannot provide
zero-downtime single-writer handoff. Use separate standby orchestration or a
blue-green/custom lifecycle cutover, or use stop-first deployment and accept
downtime.
SIGTERM and an authenticated runtime drain request are graceful planned shutdowns: an active broker reports draining for its configured grace before cleanup. Ctrl-C and fatal actor or writer-lease-health failures withdraw target and readiness probes and begin cleanup immediately without that delay. A waiting standby skips the active drain grace for every shutdown trigger.
Fitz completes session cleanup and domain joins before releasing storage and joins an in-flight Midge open or recovery attempt before exit. Treat 90 seconds as the operational termination-grace baseline with the default 25-second drain, then increase it for peak session count, domain teardown, and backend/provider latency; it is not a guaranteed upper bound.
Once active, Fitz polls Midge writer-lease renewal health as a lifecycle
invariant rather than deriving behavior from logs. When Midge reports the lease
unhealthy, Fitz withdraws /targetz, /readyz, and /healthz and requests
fatal broker termination; the process does not attempt in-process reacquisition.
Midge independently enforces a monotonic lease-valid-until deadline while a
provider renewal is in flight and treats malformed cloud lease expiration as
indeterminate rather than expired. FITZ_STORAGE_LEASE_TTL_SECS configures the
deadline, defaults to 30 seconds, and cannot be set below 30 seconds. Midge
renews at one third of that TTL. Longer values reduce provider coordination
traffic but extend ungraceful takeover; graceful shutdown conditionally expires
the current lease immediately.
Files: src/api/handlers/tcp_listener.rs, src/api/handlers/tcp_session.rs, src/api/handlers/websocket.rs, src/api/ingress.rs
Goal: Track connection lifecycle, frame I/O, and protocol errors
Template:
// On connection accept (AlwaysSample = 100%)
#[tracing::instrument(skip_all, fields(peer_addr = %peer_addr))]
async fn handle_connection(socket: TcpStream, peer_addr: SocketAddr) {
let span = tracing::Span::current();
METRICS.counter_inc(observability::METRIC_CONNECTIONS_OPENED);
// Inside connection loop:
loop {
let instant = std::time::Instant::now();
// Frame read (hot path, sample at 0.1%)
if should_sample_hot_path() {
let read_span = tracing::debug_span!("frame::read", size = ?frame.len());
read_span.in_scope(|| {
// read frame
});
}
METRICS.histogram_observe_us(
observability::METRIC_FRAME_READ_LATENCY,
instant.elapsed().as_micros() as u64
);
// Wrap frame in span context if sampled
let envelope = parse_frame(frame)?;
}
METRICS.counter_inc(observability::METRIC_CONNECTIONS_CLOSED);
}Key Metrics:
fitz_connections_opened_total(counter)fitz_connections_closed_total(counter)fitz_frames_received_total(counter)fitz_frames_sent_total(counter)fitz_frames_malformed_total(counter)
Files: src/session/session.rs, src/session/permissions.rs, src/api/runtime_ingress.rs
Goal: Track authentication, authorization, and frame processing
Template:
/// Process an inbound frame (AlwaysSample = 100%)
#[tracing::instrument(skip_all, fields(session_id = %session.id(), frame_size = frame.len()))]
pub fn process_frame(session: &Session, frame: &[u8]) -> Result<Message> {
// TLV decode (hot path, sample at 0.1%)
let instant = std::time::Instant::now();
let message = if should_sample_hot_path() {
let decode_span = tracing::debug_span!("tlv::decode");
decode_span.in_scope(|| tlv::decode(frame))
} else {
tlv::decode(frame)
}?;
METRICS.histogram_observe_us(
observability::METRIC_TLV_CODEC_LATENCY,
instant.elapsed().as_micros() as u64
);
// Permission check (AlwaysSample - critical for security)
let perm_instant = std::time::Instant::now();
let _perm_span = tracing::info_span!(
observability::SPAN_PERMISSION_CHECK,
route = %message.route(),
auth_method = %session.auth_method()
);
if !session.has_permission(&message.route()) {
METRICS.counter_inc(observability::METRIC_PERMISSION_DENIALS);
tracing::warn!("Permission denied for route");
return Err(SessionError::PermissionDenied);
}
METRICS.histogram_observe_us(
observability::METRIC_PERMISSION_CHECK_LATENCY,
perm_instant.elapsed().as_micros() as u64
);
Ok(message)
}Key Metrics:
fitz_sessions_created_total(counter)fitz_sessions_closed_total(counter)fitz_tlv_decode_errors_total(counter)fitz_auth_failures_total(counter)fitz_permission_denials_total(counter)fitz_permission_check_latency_us(histogram)
Files: src/runtime/router.rs, src/runtime/family_actor_pool.rs
Goal: Track message routing, delivery, and actor scheduling
Template:
impl Router {
/// Route an envelope (hot path, sample at 0.1%)
pub fn route(&self, envelope: Envelope) -> Result<(), RouteError> {
let dest = envelope.destination().clone();
let start = std::time::Instant::now();
// Only create span if sampled
if should_sample_hot_path() {
let span = tracing::debug_span!(
observability::SPAN_ROUTE_MATCH,
route = %dest,
domain = %dest.domain()
);
span.in_scope(|| self._route_impl(envelope))
} else {
self._route_impl(envelope)
}?;
METRICS.histogram_observe_us(
observability::METRIC_ROUTE_MATCH_LATENCY,
start.elapsed().as_micros() as u64
);
Ok(())
}
fn _route_impl(&self, envelope: Envelope) -> Result<(), RouteError> {
let dest = envelope.destination().clone();
// Try exact route match
let sink = if let Some(sink) = self.registry.get(&dest) {
sink
} else {
// Domain fallback
self.registry.get_by_domain(&dest.domain())
.ok_or_else(|| {
METRICS.counter_inc(observability::METRIC_ROUTE_MISMATCHES);
RouteError::RouteNotFound(dest.clone())
})?
};
match sink.deliver(envelope) {
Ok(()) => {
METRICS.counter_inc(observability::METRIC_MESSAGES_DELIVERED);
Ok(())
}
Err(e) => {
METRICS.counter_inc(observability::METRIC_DELIVERY_FAILURES);
Err(RouteError::DeliveryFailed(dest, e))
}
}
}
}Key Metrics:
fitz_route_mismatches_total(counter)fitz_delivery_failures_total(counter)fitz_messages_pending(gauge - mailbox depth)fitz_route_match_latency_us(histogram)
Files: src/protocol/tlv.rs, src/protocol/{domain}_codec.rs
Goal: Track codec performance and errors
Template:
impl MyDomainCodec {
pub fn encode_message(msg: &MyMessage) -> Result<Vec<u8>> {
let start = std::time::Instant::now();
let result = if should_sample_hot_path() {
let span = tracing::debug_span!("tlv::encode", domain = "my_domain");
span.in_scope(|| Self::_encode_impl(msg))
} else {
Self::_encode_impl(msg)
};
if result.is_ok() {
METRICS.histogram_observe_us(
observability::METRIC_TLV_CODEC_LATENCY,
start.elapsed().as_micros() as u64
);
} else {
METRICS.counter_inc(observability::METRIC_TLV_ENCODE_ERRORS);
}
result
}
}Key Metrics:
fitz_tlv_encode_errors_total(counter)fitz_tlv_decode_errors_total(counter)fitz_tlv_codec_latency_us(histogram)
Files: src/domains/{kv,lease,notice,rpc,queue,stream,schedule}/mod.rs
Goal: Track domain-specific operations and business logic
Template:
impl Actor for MyDomainActor {
fn receive(&mut self, msg: Self::Message, ctx: &mut Context<Self>) -> Self::Response {
let operation_name = format!("{:?}", msg).split('(').next().unwrap_or("unknown");
let start = std::time::Instant::now();
// AlwaysSample domain operations
let _span = tracing::info_span!(
observability::SPAN_DOMAIN_OPERATION,
domain = env!("CARGO_PKG_NAME"),
operation = %operation_name,
realm = %ctx.route().realm(),
actor_id = ?ctx.actor_id()
);
let result = match msg {
MyMessage::Operation1 { param } => {
// Business logic here
self.handle_operation_1(param, ctx)
}
MyMessage::Operation2 { param } => {
self.handle_operation_2(param, ctx)
}
};
// Record latency and success/failure
let latency_ms = start.elapsed().as_millis() as u64;
METRICS.histogram_observe_ms(
observability::METRIC_DOMAIN_OPERATION_LATENCY,
latency_ms
);
if result.is_error() {
METRICS.counter_inc(observability::METRIC_DOMAIN_ERRORS);
} else {
METRICS.counter_inc(observability::METRIC_DOMAIN_OPERATIONS);
}
result
}
}Key Metrics (per domain):
fitz_domain_operations_total(counter, labeled by domain)fitz_domain_errors_total(counter, labeled by domain + error_type)fitz_domain_operation_latency_ms(histogram, labeled by domain)
- Read the code path you want to instrument
- Identify the hot path (small, frequently called), critical path (security/correctness), and debug path
- Add
#[tracing::instrument]or manuallet _span = tracing::info_span!(...)at appropriate places - Add metric counters/histograms for success/failure/latency
- Test with:
FITZ_LOG_LEVEL=debug OTEL_ENABLED=false cargo test
By default, Fitz emits human-readable, colorized text logs. Set FITZ_LOG_FORMAT=json to opt into structured JSON output.
FITZ_LOG_FORMAT=json FITZ_LOG_LEVEL=info ./target/debug/fitz > /tmp/fitz.log 2>&1 &
# Makes a request (or run integration test)
cat /tmp/fitz.log | jq '.' | head -20Expected output:
{
"timestamp": "2026-02-21T12:34:56.789Z",
"level": "INFO",
"message": "Starting Fitz broker",
"target": "fitz::boot",
"module_path": "fitz::boot",
"span": {
"name": "boot"
}
}Raw Prometheus is served by the dedicated metrics listener, not the authenticated main listener:
curl http://localhost:9090/metrics | head -30Expected output:
# HELP fitz_connections_opened_total Fitz counter metrics
# TYPE fitz_connections_opened_total counter
fitz_connections_opened_total 5
# HELP fitz_messages_delivered_total Fitz counter metrics
# TYPE fitz_messages_delivered_total counter
fitz_messages_delivered_total 1024
Instrument layers in this order:
- Runtime/Router (highest impact, hot path)
- Session (security-critical)
- API/Transport (connection lifecycle)
- Protocol/Codecs (performance debugging)
- Domains (per-domain observability)
// Safe to use even if metric export isn't set up yet
METRICS.counter_inc("my_counter"); // No-op if collector not initializedfn should_sample_hot_path() -> bool {
// This could be thread-local or RNG-based
// For now, simple fixed ratio:
std::time::SystemTime::now()
.duration_since(std::time::UNIX_EPOCH)
.map(|d| (d.as_nanos() as u64) % 1000 == 0)
.unwrap_or(false)
}// In envelope creation or routing:
#[tracing::instrument(skip_all, fields(message_id = %msg_id, parent_message_id = ?parent_id))]
fn route_with_causation(parent_id: Option<MessageId>, msg_id: MessageId) {
// Spans are automatically linked by tracing
}let span = tracing::info_span!(
"operation",
route = %route,
realm = %realm,
operation = op,
error_type = tracing::field::Empty, // Fill in later if error
);
span.in_scope(|| {
let result = do_work();
if let Err(e) = &result {
span.record("error_type", &e.kind());
}
result
});See tests/observability_spans.rs and tests/observability_metrics.rs for comprehensive tests.
Quick local test:
# Run a single integration test with JSON logging enabled (opt-in)
FITZ_LOG_FORMAT=json FITZ_LOG_LEVEL=debug cargo test kv_basics -- --nocapture- Implement Layer 1 (API) - Track connections and frames
- Implement Layer 2 (Session) - Track auth and TLV codec
- Implement Layer 3 (Router) - Track routing and delivery (highest ROI)
- Keep the dedicated
/metricslistener private - Export raw Prometheus only fromFITZ_METRICS_BIND_ADDR:FITZ_METRICS_PORT. - Add distributed tracing - Link parent/child messages via OpenTelemetry
- Configure Datadog/Prometheus - Connect external observability platform