Skip to content

Latest commit

 

History

History
1963 lines (1465 loc) · 49 KB

File metadata and controls

1963 lines (1465 loc) · 49 KB

Juniper Canopy API Reference

Version: 1.4.3 Last Updated: September 5, 2026 Base URL: http://127.0.0.1:8050


Table of Contents

  1. Overview
  2. Authentication
  3. REST API Endpoints
  4. Training Control Endpoints
  5. Network Mutation Endpoints
  6. Remote Worker Endpoints
  7. WebSocket Channels
  8. Data Models
  9. Error Handling
  10. Rate Limiting
  11. Code Examples

Overview

Juniper Canopy provides a RESTful HTTP API and WebSocket channels for real-time monitoring of Cascade Correlation neural network training.

API Characteristics

  • Protocol: HTTP/1.1, WebSocket (RFC 6455)
  • Data Format: JSON
  • Encoding: UTF-8
  • CORS: Enabled for localhost origins
  • Rate Limiting: REST limiting is configurable; WebSocket connection caps are enforced

External CasCor Normalization (Service Mode)

When backend_type is service, Canopy normalizes external CasCor ResponseEnvelope payloads before returning API responses. This preserves dashboard contracts across demo and service backends for:

  • status/state fields (training_active/state_machine -> flat status flags)
  • metric naming (loss/accuracy/validation_* -> canonical keys where needed)
  • dataset metadata (input_features/train_samples -> num_* keys)
  • decision boundary key mapping (grid_x/grid_y/predictions -> xx/yy/Z)

Primary codepaths: src/backend/cascor_service_adapter.py, src/backend/service_backend.py, src/backend/state_sync.py.

Base URL

Local Development:

http://127.0.0.1:8050

Custom Port:

export JUNIPER_CANOPY_SERVER__PORT=8051
# Base URL: http://127.0.0.1:8051

API Documentation (Interactive)

When server is running with API-key authentication off -- no API key configured, where an empty or whitespace-only key counts as none (see Authentication) -- visit the URL below. The docs switch is the authentication switch itself, so a real key in CANOPY_API_KEY or in the file named by CANOPY_API_KEY_FILE turns the four docs routes off (404):

http://127.0.0.1:8050/docs

This provides:

  • Interactive API explorer (Swagger UI)
  • Request/response schemas
  • Try-it-out functionality

Authentication

Authentication is configuration-dependent. The key is read from the file named by CANOPY_API_KEY_FILE (its contents, whitespace-stripped) whenever that variable names an existing file, and otherwise from CANOPY_API_KEY:

  • If no key is configured, API-key authentication is disabled for development/demo use.
  • A key that is empty or whitespace-only counts as no key: authentication is disabled exactly as if it were unset, the interactive docs are served, and the dashboard's own requests send no key. Startup logs a WARNING naming which source was blank. This includes a blank file named by CANOPY_API_KEY_FILE while CANOPY_API_KEY holds a real key, because the file takes precedence. With JUNIPER_CANOPY_REQUIRE_AUTH=true a blank key fails the boot exactly as no key does, and the WARNING follows the posture check's CRITICAL, worded for the refused boot.
  • Otherwise keyed callers must send X-API-Key: <value>.
  • A key with leading or trailing whitespace, or a line break, is used exactly as set: it is not stripped, and authentication is enabled on it. No HTTP header carries such a key reliably, so startup logs a WARNING. The dashboard's own requests leave off a key their HTTP client refuses to send (one that starts with whitespace or holds a line break), so they are refused, and the refusal never carries the key. A key file is stripped at both ends, so from CANOPY_API_KEY_FILE only a line break inside the key can do this, and the WARNING then names CANOPY_API_KEY_FILE.
  • A CANOPY_API_KEY_FILE that names no existing file is ignored, exactly as if unset, and startup logs a WARNING naming the variable (never the path it holds).
  • A presented X-API-Key (or WebSocket key or bearer token) holding a non-ASCII character is an ordinary mismatch, refused as any wrong key is: a 401 over HTTP, and on a WebSocket that path's own close code (4001, or 1008 for a bearer token). It used to raise inside the comparison, a 500 whose error report recorded the configured key.
  • Same-origin browser training controls (/api/train/*, /api/csrf, /ws/control) use the browser path: allowed Origin + canopy_session cookie + CSRF token when JUNIPER_CANOPY_BROWSER_CONTROL_AUTH_ENABLED=true (the default). Keyed callers continue to work.

The browser-control path is not a network perimeter by itself. Canopy now enforces that perimeter at startup: JUNIPER_CANOPY_SERVER__HOST must be loopback (127.0.0.0/8, ::1, or localhost) unless a perimeter attestation explicitly declares one is present — JUNIPER_CANOPY_LOOPBACK_PUBLISH_ATTESTED=true (reachable only via a loopback-only host publish) or JUNIPER_CANOPY_AUTH_PROXY_ATTESTED=true (a fronting authenticating proxy terminates access before traffic reaches the app).

Outbound Keys

Canopy sends a key as X-API-Key to three upstreams: juniper-cascor (JUNIPER_CASCOR_API_KEY, falling back to JUNIPER_DATA_API_KEY), juniper-data (JUNIPER_CANOPY_JUNIPER_DATA_API_KEY, then JUNIPER_DATA_API_KEY) and the recurrence service (JUNIPER_CANOPY_RECURRENCE_API_KEY, then JUNIPER_RECURRENCE_API_KEY). Each honours its <NAME>_FILE form first. A key canopy sends must consist of printable ASCII other than space (0x21-0x7E): a value with leading or trailing whitespace, a line break, a space or a non-ASCII character is refused where it is read, because the HTTP clients refuse such a header value and quote it in the error they raise. A refused value is treated exactly like an empty one -- the next variable in the order applies, or no key is sent -- and startup logs one WARNING naming the variable, never its value or a path.


REST API Endpoints

GET /

Description: Root endpoint, redirects to dashboard

Response:

  • Status: 302 Found
  • Location: /dashboard/

Example:

curl -I http://127.0.0.1:8050/

Response Headers:

HTTP/1.1 302 Found
Location: /dashboard/

GET /api/health

Description: Health check endpoint for monitoring and load balancers

Parameters: None

Response Schema:

{
  "status": "healthy",
  "timestamp": 1711459200.123,
  "version": "0.3.0",
  "active_connections": 2,
  "training_active": true,
  "demo_mode": false,
  "juniper_data_available": true
}

Field Descriptions:

  • status (string) - Health status, currently always "healthy"
  • timestamp (number) - Unix timestamp in seconds
  • version (string) - Application version
  • active_connections (integer) - Number of active WebSocket connections
  • training_active (boolean) - Whether training is in progress
  • demo_mode (boolean) - Whether running in demo mode
  • juniper_data_available (boolean) - JuniperData dependency status
  • backend_status (string, service mode, X7 slice 1c) - Cache class: ok / unreachable / indeterminate. Omitted in demo / recurrence.
  • backend_status_stale (boolean, service mode) - true when the last OK payload is older than 5 s or the class is not ok
  • backend_status_age_seconds (number or null, service mode) - Seconds since the last OK payload; null if none

These three extras are additive and appear on /v1/health and inside /v1/health/ready details as well. Status codes are unchanged: an upstream outage stays 200. training_active remains a bool; staleness is reported beside it, not smuggled into it. Landed with #578.

Status Codes:

  • 200 OK - Service healthy

Example Request:

curl http://127.0.0.1:8050/api/health

Use Cases:

  • Docker health checks
  • Load balancer probes
  • Deployment diagnostics

Deprecation: Prefer /v1/health, /v1/health/live, and /v1/health/ready. /api/health and /health remain as aliases and log a warning.

GET /v1/health/live

Description: Liveness probe. Confirms the process is running. Touches no backend — this is the X7 canary: if this stalls, the event loop itself is blocked.

Parameters: None

Response Schema:

{
  "status": "alive"
}

Status Codes:

  • 200 OK - Process is running

Example Request:

curl -s -o /dev/null -w "%{http_code} %{time_total}\n" http://127.0.0.1:8050/v1/health/live

When cascor is unreachable this must still return in milliseconds. See Event-loop I/O discipline (X7).

Related: GET /v1/health (combined status; reaches backend.is_training_active) and GET /v1/health/ready (async dependency probes via probe_dependency).

GET /api/status

Description: Get normalized training status and network information. In service mode, after #578, this is served from the X7 slice-1c status cache — one background poll; every tab costs zero upstream calls. Demo and recurrence still return a live backend.get_status().

Description: Get normalized training status and network information. Reaches backend.get_status() — a synchronous cascor HTTP call. Slice 1a (#567) offloads it with asyncio.to_thread so a slow upstream cannot stall /v1/health/live. This is the T-A2 driver route.

Parameters: None

Response Schema (Demo Backend):

{
  "is_training": true,
  "is_running": true,
  "is_paused": false,
  "completed": false,
  "failed": false,
  "fsm_status": "STARTED",
  "phase": "output",
  "current_epoch": 42,
  "current_loss": 0.234,
  "current_accuracy": 0.876,
  "hidden_units": 3,
  "network_connected": true,
  "monitoring_active": true,
  "input_size": 2,
  "output_size": 1
}

Response Schema (Service Backend, normalized):

{
  "is_training": true,
  "is_running": true,
  "is_paused": false,
  "completed": false,
  "failed": false,
  "fsm_status": "STARTED",
  "phase": "output",
  "current_epoch": 42,
  "hidden_units": 3,
  "network_connected": true,
  "monitoring_active": true,
  "input_size": 2,
  "output_size": 3,
  "learning_rate": 0.01,
  "max_hidden_units": 10,
  "max_epochs": 500,
  "status_class": "ok",
  "stale": false,
  "age_seconds": 0.2
}

Cache envelope (service mode, X7 slice 1c, landed with #578):

Field Meaning
status_class ok / unreachable / indeterminate — what the cache concluded, not the raw payload
stale true when the class is not ok or the last OK payload is older than 5 s
age_seconds Seconds since the last OK payload, or null if none has been seen

A never-OK body omits is_training (C6) and still carries a truthy error so a UI that has not been taught about status_class keeps the PR #340 "Unreachable" branch. A half-dead 200 (dict, no error, not cascor-shaped) is classified unreachable; handing the raw payload to the status bar re-creates that defect as "Stopped".

See AGENTS_REFERENCE.md — Cascor status cache.

Field Descriptions:

  • is_training (boolean) - Training active flag
  • is_running (boolean) - Running state derived from backend FSM
  • is_paused (boolean) - Paused state derived from backend FSM
  • completed (boolean) - Completion state
  • failed (boolean) - Failure state
  • fsm_status (string) - Backend FSM status string
  • phase (string) - Normalized phase (idle, output, candidate, inference, etc.)
  • current_epoch (integer) - Current epoch
  • hidden_units (integer) - Current hidden unit count
  • network_connected (boolean) - Whether a network is loaded/connected
  • monitoring_active (boolean) - Whether training monitoring is active
  • input_size (integer) - Input dimension
  • output_size (integer) - Output dimension
  • learning_rate (number, service mode) - Active learning rate
  • max_hidden_units (integer, service mode) - Max hidden units setting
  • max_epochs (integer, service mode) - Max epochs setting
  • status_class (string, service mode) - Cache verdict (ok / unreachable / indeterminate)
  • stale (boolean, service mode) - Whether the last OK payload is fresh
  • age_seconds (number or null, service mode) - Age of the last OK payload

Status Codes:

  • 200 OK - Status retrieved successfully

Notes:

  • Use phase (not current_phase) as the canonical phase key.
  • In service mode, this endpoint returns normalized fields from nested CasCor status payloads.
  • After #578, service-mode responses are the cache envelope above. Render status_class, not the raw payload. Demo / recurrence are unchanged.

GET /api/state

Description: Get the full training state used by dashboard UI handlers.

Parameters: None

Response Schema (abridged):

{
  "status": "Started",
  "phase": "Output",
  "learning_rate": 0.01,
  "max_hidden_units": 10,
  "current_epoch": 42,
  "grow_iteration": 3,
  "grow_max": 10,
  "phase_started_at": "2026-03-30T12:00:00+00:00",
  "candidate_epoch": 120,
  "candidate_total_epochs": 500
}

Field Groups:

  • Core runtime: status, phase, current_epoch, current_step, timestamp
  • Hyperparameters: learning_rate, max_hidden_units, max_epochs
  • Candidate pool state: candidate_pool_*, pool_metrics, top_candidate_*
  • Progress detail: phase_detail, grow_iteration, grow_max, best_correlation, candidates_*, phase_started_at, candidate_epoch, candidate_total_epochs
  • Service/dashboard compatibility keys: nn_*, cn_* fields may also be present depending on backend mode

Status Codes:

  • 200 OK - State retrieved successfully

Notes:

  • This endpoint is the source for metrics-panel state cards and progress bars.
  • phase_started_at is ISO-8601 and used to compute elapsed phase duration in the dashboard.

GET /api/metrics

Description: Get current training metrics snapshot

Parameters: None

Response Schema (Demo Backend):

{
  "is_running": true,
  "is_paused": false,
  "current_epoch": 42,
  "current_loss": 0.234,
  "current_accuracy": 0.876,
  "hidden_units": 3,
  "metrics_count": 420
}

Response Schema (Service Backend, normalized):

{
  "epoch": 42,
  "train_loss": 0.234,
  "train_accuracy": 0.876,
  "val_loss": 0.251,
  "val_accuracy": 0.861,
  "hidden_units": 3,
  "phase": "output",
  "timestamp": 1711459200.123
}

Status Codes:

  • 200 OK - Metrics retrieved successfully

Notes:

  • /api/metrics is a point-in-time snapshot.
  • For time-series plotting, use /api/metrics/history.

GET /api/metrics/history

Description: Get historical training metrics

Query Parameters:

  • limit (integer, optional): Maximum history entries to return. 0 means "all available" (internally capped).

Response Schema:

{
  "history": [
    {
      "epoch": 1,
      "metrics": {
        "loss": 0.95,
        "accuracy": 0.38,
        "val_loss": 0.99,
        "val_accuracy": 0.35
      },
      "network_topology": {
        "input_units": 2,
        "hidden_units": 0,
        "output_units": 3
      },
      "phase": "output",
      "timestamp": "2026-03-26T18:30:00"
    },
    {
      "epoch": 2,
      "train_loss": 0.82,
      "train_accuracy": 0.51,
      "val_loss": 0.84,
      "val_accuracy": 0.49,
      "hidden_units": 0,
      "phase": "output",
      "timestamp": 1711459201.123
    }
  ]
}

History Entry Shapes:

  • Demo entries use nested metrics + network_topology blocks.
  • Service entries are normalized flat metric objects (train_loss, train_accuracy, val_loss, val_accuracy, etc.).

Status Codes:

  • 200 OK - History retrieved successfully
  • 422 Unprocessable Entity - Invalid limit value type

Example Request:

curl "http://127.0.0.1:8050/api/metrics/history?limit=100"

GET /api/topology

Description: Get current network topology (nodes and connections)

Parameters: None

Response Schema:

{
  "input_units": 2,
  "hidden_units": 3,
  "output_units": 1,
  "nodes": [
    {"id": "input_0", "type": "input", "layer": 0},
    {"id": "hidden_0", "type": "hidden", "label": "H0"},
    {"id": "output_0", "type": "output", "layer": 2}
  ],
  "connections": [
    {"from": "input_0", "to": "hidden_0", "weight": 0.12},
    {"from": "hidden_0", "to": "output_0", "weight": 0.56}
  ]
}

Field Descriptions:

  • input_units (integer) - Input node count
  • hidden_units (integer) - Hidden node count
  • output_units (integer) - Output node count
  • nodes (array) - Topology node list
  • connections (array) - Weighted edges

Status Codes:

  • 200 OK - Topology retrieved successfully
  • 503 Service Unavailable - No topology available

Notes:

  • Node attributes may vary by backend (layer in demo mode, label in service mode).
  • Consumers should rely on id + type + connections as primary contract.
  • Service-mode topology normalization emits a strict 3-layer scheme:
    • input nodes use layer: 0
    • all hidden_* nodes use layer: 1
    • all output_* nodes use layer: 2
  • In service mode, output connection rows are derived by transposing CasCor output_weights from (input+hidden, output) to output-oriented rows.

Network Mutation Endpoints

The network mutation endpoints are thin Canopy proxy routes for the Network Editor tab. They require a live CasCor service backend and forward requests through CascorServiceAdapter to CasCor's /v1/network/... mutation API.

Operational constraints:

  • Demo mode is not supported. Canopy returns 501 Not Implemented when backend_type is not service or no service adapter is attached.
  • CasCor is the source of truth for lifecycle gating and tensor validation. The mutation lifecycle is expected to reject edits outside the Investigating state.
  • Canopy performs request-body shape validation with Pydantic before proxying, but exact parameter shape, NaN/Inf checks, out-of-range hidden unit indexes, and activation validation are enforced by CasCor.
  • Current Canopy proxy handlers wrap adapter failures as 500 responses with a detail message such as patch_weights failed: ...; clients should display the detail rather than infer every upstream failure from the status code alone.

Primary codepaths: src/main.py, src/backend/cascor_service_adapter.py, src/frontend/components/network_editor_panel.py.

PATCH /api/v1/network/weights

Description: Patch one weight or bias group in the loaded network. Used by the Network Editor "Patch weights" form.

Request Schema:

{
  "target": "output_weights",
  "field": "weights",
  "values": [0.1, -0.2, 0.05],
  "hidden_unit_index": null,
  "dtype": "float32"
}

Fields:

  • target (string, required) - Parameter group to patch. The Network Editor emits output_weights, output_bias, hidden_unit_weights, or hidden_unit_bias.
  • field (string, required) - weights or bias. The Network Editor derives this from target.
  • values (array or nested array, required) - New values. Shape must match the selected target; CasCor validates the exact shape.
  • hidden_unit_index (integer, optional) - Required for hidden_unit_* targets.
  • dtype (string, optional) - Defaults to float32.

Status Codes:

  • 200 OK - Patch accepted by CasCor.
  • 422 Unprocessable Entity - Required Canopy request fields are missing or have invalid JSON types.
  • 501 Not Implemented - Canopy is not running with a live CasCor service adapter.
  • 500 Internal Server Error - Adapter or upstream CasCor mutation failed; inspect detail.

Example Request:

curl -X PATCH http://127.0.0.1:8050/api/v1/network/weights \
  -H "Content-Type: application/json" \
  -d '{
    "target": "hidden_unit_weights",
    "field": "weights",
    "values": [0.12, -0.04, 0.31],
    "hidden_unit_index": 2,
    "dtype": "float32"
  }'

POST /api/v1/network/hidden-units

Description: Append a hidden unit at the cascade tail. Used by the Network Editor "Append hidden unit" form.

Request Schema:

{
  "weights": [0.1, -0.2, 0.05],
  "bias": 0.0,
  "activation": "Tanh"
}

Fields:

  • weights (array, required) - Incoming weight vector. The expected length is input_size + existing_hidden_units.
  • bias (number, optional) - Defaults to 0.0.
  • activation (string, optional) - Defaults to Tanh. The Network Editor offers Tanh, Sigmoid, ReLU, and Linear; CasCor validates against its activation registry.

Status Codes:

  • 200 OK - Hidden unit appended by CasCor.
  • 422 Unprocessable Entity - Required Canopy request fields are missing or have invalid JSON types.
  • 501 Not Implemented - Canopy is not running with a live CasCor service adapter.
  • 500 Internal Server Error - Adapter or upstream CasCor mutation failed; inspect detail.

Example Request:

curl -X POST http://127.0.0.1:8050/api/v1/network/hidden-units \
  -H "Content-Type: application/json" \
  -d '{"weights": [0.1, -0.2, 0.05], "bias": 0.0, "activation": "ReLU"}'

DELETE /api/v1/network/hidden-units/{idx}

Description: Remove a hidden unit by zero-based index. CasCor rebuilds the cascade so downstream unit shapes remain valid.

Path Parameters:

  • idx (integer, required) - Hidden unit index to remove.

Status Codes:

  • 200 OK - Hidden unit removed by CasCor.
  • 422 Unprocessable Entity - idx is not an integer.
  • 501 Not Implemented - Canopy is not running with a live CasCor service adapter.
  • 500 Internal Server Error - Adapter or upstream CasCor mutation failed; inspect detail.

Example Request:

curl -X DELETE http://127.0.0.1:8050/api/v1/network/hidden-units/2

GET /api/dataset

Description: Get dataset information

Parameters: None

Response Schema (Demo Backend):

{
  "inputs": [[0.12, 0.34], [-0.56, 0.78]],
  "targets": [0, 1],
  "num_samples": 200,
  "num_features": 2,
  "num_classes": 2
}

Response Schema (Service Backend, normalized):

{
  "num_samples": 1000,
  "num_features": 2,
  "num_classes": 3,
  "loaded": true,
  "train_samples": 800,
  "test_samples": 200,
  "inputs": [[0.1, 0.2], [0.3, 0.4]],
  "targets": [0, 1]
}

Status Codes:

  • 200 OK - Dataset retrieved successfully
  • 503 Service Unavailable - No dataset available

Notes:

  • Demo mode returns full sample arrays.
  • Service mode always returns normalized metadata (num_samples, num_features, num_classes), and may additionally include inputs/targets when available.
  • If service metadata is returned without arrays, Canopy attempts a secondary dataset fetch through the service data endpoint before returning.
  • When arrays are still unavailable, frontend dataset visualizations should treat the payload as metadata-only and avoid assuming inputs/targets exist.

GET /api/decision_boundary

Description: Get decision boundary data for visualization

Query Parameters:

  • resolution (integer, optional): Grid resolution per axis. Values are clamped to [5, 200].

Response Schema:

{
  "xx": [[-1.2, -1.0, -0.8], [-1.2, -1.0, -0.8]],
  "yy": [[-1.2, -1.2, -1.2], [-1.0, -1.0, -1.0]],
  "Z": [[0, 0, 1], [0, 1, 1]],
  "x_min": -1.2,
  "x_max": 1.2,
  "y_min": -1.2,
  "y_max": 1.2,
  "resolution": 100
}

Field Descriptions:

  • xx (array) - X meshgrid (resolution x resolution)
  • yy (array) - Y meshgrid (resolution x resolution)
  • Z (array) - Predicted class grid (resolution x resolution)
  • x_min, x_max, y_min, y_max (number) - Plot bounds
  • resolution (integer) - Applied resolution

Status Codes:

  • 200 OK - Decision boundary computed
  • 503 Service Unavailable - No decision boundary data available

Notes:

  • Service mode maps external CasCor grid_x/grid_y/predictions to xx/yy/Z.
  • Returned Z is class grid data for contour rendering.

GET /api/statistics

Description: Get WebSocket connection statistics

Parameters: None

Response Schema:

{
  "active_connections": 2,
  "total_messages_broadcast": 1523,
  "connections_info": [
    {
      "client_id": "training-client-12345",
      "connected_at": "2025-11-05T10:30:00.123456",
      "messages_sent": 756,
      "last_message_at": "2025-11-05T10:45:23.987654"
    }
  ]
}

Field Descriptions:

  • active_connections (integer) - Current active WebSocket connections
  • total_messages_broadcast (integer) - Total messages broadcast since startup
  • connections_info (array) - Detailed information per connection
    • client_id (string) - Client identifier
    • connected_at (string) - ISO 8601 connection timestamp
    • messages_sent (integer) - Messages sent to this client
    • last_message_at (string) - ISO 8601 timestamp of last message

Status Codes:

  • 200 OK - Statistics retrieved successfully

Example Request:

curl http://127.0.0.1:8050/api/statistics

Example Response:

{
  "active_connections": 1,
  "total_messages_broadcast": 42,
  "connections_info": [
    {
      "client_id": "training-client-123",
      "connected_at": "2025-11-05T10:30:00.000000",
      "messages_sent": 42,
      "last_message_at": "2025-11-05T10:30:42.000000"
    }
  ]
}

Training Control Endpoints

POST /api/train/start

Description: Start training (optionally with reset)

Parameters:

  • reset (boolean, query, optional) - Reset network before starting (default: false)

Response Schema (Demo Backend):

{
  "status": "started",
  "is_running": true,
  "is_paused": false,
  "current_epoch": 0,
  "current_loss": 1.0,
  "current_accuracy": 0.5,
  "hidden_units": 0,
  "metrics_count": 0
}

Response Schema (Service Backend):

{
  "status": "started",
  "ok": true,
  "is_training": true
}

Status Codes:

  • 200 OK - Request accepted

POST /api/train/pause

Description: Request training pause

Parameters: None

Response Schema:

{
  "status": "paused"
}

Status Codes:

  • 200 OK - Request accepted

Notes:

  • Endpoint response is an acknowledgement payload.
  • For authoritative backend state, follow with GET /api/status or WebSocket control responses.

POST /api/train/resume

Description: Request training resume

Parameters: None

Response Schema:

{
  "status": "running"
}

Status Codes:

  • 200 OK - Request accepted

Notes:

  • Endpoint response is an acknowledgement payload.
  • For authoritative backend state, follow with GET /api/status or WebSocket control responses.

POST /api/train/stop

Description: Request training stop

Parameters: None

Response Schema:

{
  "status": "stopped"
}

Status Codes:

  • 200 OK - Request accepted

POST /api/train/reset

Description: Reset training state

Parameters: None

Response Schema (Demo Backend):

{
  "status": "reset",
  "is_running": false,
  "is_paused": false,
  "current_epoch": 0,
  "current_loss": 1.0,
  "current_accuracy": 0.5,
  "hidden_units": 0,
  "metrics_count": 0
}

Response Schema (Service Backend):

{
  "status": "reset",
  "ok": true,
  "data": {
    "message": "reset requested"
  }
}

Status Codes:

  • 200 OK - Request accepted

GET /api/train/status

Description: Get backend-tagged training status

Parameters: None

Response Schema:

{
  "backend": "service",
  "is_training": true,
  "is_running": true,
  "is_paused": false,
  "completed": false,
  "failed": false,
  "fsm_status": "STARTED",
  "phase": "output",
  "current_epoch": 42,
  "hidden_units": 3
}

Status Codes:

  • 200 OK - Status retrieved successfully

Notes:

  • backend is "demo" or "service".
  • Remaining fields mirror GET /api/status for the active backend.

GET /api/selection

Description: The read side of the model/dataset selection: the model canopy has recorded, and the dataset the live backend holds. POST /api/model/select writes the model; this reads it back. The dashboard hydrates both selectors from it once, on page load.

Parameters: None

Response Schema:

{
  "nn_model": "recurrence",
  "backend": "demo",
  "execution": "continuous",
  "status": "live",
  "swapped": false,
  "selected": true,
  "dataset": {"value": "multi_sine", "source": "pending", "generator": "multi_sine"}
}

Status Codes:

  • 200 OK - Selection retrieved. A backend whose status cannot be read still answers 200, with dataset.source "unknown".

Notes:

  • The model fields have exactly the POST /api/model/select response's shape. swapped is always false (a read swaps nothing). selected is false until the first POST /api/model/select, and nn_model is then the model the boot backend serves.
  • dataset.source is the field to branch on:
    • "pending" — a dataset staged for the next start (it wins, because Start consumes it);
    • "loaded" — the dataset the backend holds;
    • "none" — the backend holds nothing;
    • "unknown" — the backend does not report it (a juniper-cascor that predates its current_dataset status field).
  • dataset.value is canopy's dataset-type value ("spirals", not juniper-data's "spiral"). It is null when the backend holds something canopy cannot name: an unseeded generator (generator then names it), or raw inline data.

Remote Worker Endpoints

These endpoints manage distributed training via the RemoteWorkerClient.

Note: Remote worker endpoints require the Cascor backend. They are not available in demo mode.

GET /api/remote/status

Description: Get remote worker connection status

Parameters: None

Response Schema (Connected):

{
  "available": true,
  "connected": true,
  "workers_active": true,
  "num_workers": 4,
  "address": "192.168.1.100:5000"
}

Response Schema (Not Connected):

{
  "available": true,
  "connected": false,
  "workers_active": false
}

Response Schema (No Backend):

{
  "available": false,
  "connected": false,
  "workers_active": false,
  "error": "No backend"
}

Field Descriptions:

  • available (boolean) - Whether remote worker functionality is available
  • connected (boolean) - Whether connected to a remote manager
  • workers_active (boolean) - Whether workers are currently running
  • num_workers (integer, optional) - Number of active workers
  • address (string, optional) - Remote manager address
  • error (string, optional) - Error message if unavailable

Status Codes:

  • 200 OK - Status retrieved successfully

Example Request:

curl http://127.0.0.1:8050/api/remote/status

POST /api/remote/connect

Description: Connect to a remote CandidateTrainingManager

Request Body (JSON):

  • address (string, required) - Remote manager address in host:port format
  • authkey (string, optional) - Authentication key for secure connection

Response Schema (Success):

{
  "status": "connected",
  "address": "192.168.1.100:5000"
}

Status Codes:

  • 200 OK - Connected successfully
  • 500 Internal Server Error - Connection failed
  • 503 Service Unavailable - No backend available

Example Request:

curl -X POST "http://127.0.0.1:8050/api/remote/connect?host=192.168.1.100&port=5000&authkey=secret123"

POST /api/remote/start_workers

Description: Start remote worker processes

Parameters:

  • num_workers (integer, query, optional) - Number of workers to start (default: 1)

Response Schema (Success):

{
  "status": "started",
  "num_workers": 4
}

Status Codes:

  • 200 OK - Workers started successfully
  • 500 Internal Server Error - Failed to start workers
  • 503 Service Unavailable - No backend available

Example Request:

curl -X POST "http://127.0.0.1:8050/api/remote/start_workers?num_workers=4"

POST /api/remote/stop_workers

Description: Stop remote worker processes

Parameters:

  • timeout (integer, query, optional) - Timeout for graceful shutdown in seconds (default: 10)

Response Schema (Success):

{
  "status": "stopped"
}

Status Codes:

  • 200 OK - Workers stopped successfully
  • 500 Internal Server Error - Failed to stop workers
  • 503 Service Unavailable - No backend available

Example Request:

curl -X POST "http://127.0.0.1:8050/api/remote/stop_workers?timeout=30"

POST /api/remote/disconnect

Description: Disconnect from remote manager

Parameters: None

Response Schema (Success):

{
  "status": "disconnected"
}

Status Codes:

  • 200 OK - Disconnected successfully
  • 500 Internal Server Error - Failed to disconnect
  • 503 Service Unavailable - No backend available

Example Request:

curl -X POST http://127.0.0.1:8050/api/remote/disconnect

WebSocket Channels

Connection URL Format

ws://127.0.0.1:8050/ws/<channel>

Common Message Format

Most WebSocket messages use this shape:

{
  "type": "state | metrics | topology | event | control_ack",
  "timestamp": 1711459200.123,
  "data": { }
}

Notes:

  • /ws/training unicasts connection_established, then initial_status, then a state snapshot, in that order, on connect. Broadcast frames are not ordered against that sequence: the socket joins the broadcast set before initial_status is fetched, so a metrics or state broadcast can arrive ahead of it. Treat state as latest-wins, as the dashboard does, rather than assuming initial_status is the first data frame.
  • Some control-channel messages may omit timestamp.
  • Runtime dashboard updates consume metrics, state, topology, and event message types.

Connection Admission and Limits

All WebSocket endpoints (/ws/training, /ws/control, and the legacy /ws) share the same admission policy:

Gate Setting Default Failure
API key / browser auth CANOPY_API_KEY, JUNIPER_CANOPY_BROWSER_CONTROL_AUTH_ENABLED API key unset, browser auth on Close 4001 for missing/invalid keyed auth where required.
Origin allowlist JUNIPER_CANOPY_WEBSOCKET__ALLOWED_ORIGINS localhost origins Close 4003; /ws/control uses opaque reason Policy violation.
Global connection cap JUNIPER_CANOPY_WEBSOCKET__MAX_CONNECTIONS 50 Close 1013 when the stack-wide N+1th socket connects.
Per-IP cap JUNIPER_CANOPY_WEBSOCKET__MAX_CONNECTIONS_PER_IP 5 Close 1013; DoS dampening only, and inert behind NAT where clients share one gateway IP.
Per-session cap JUNIPER_CANOPY_WEBSOCKET__MAX_CONNECTIONS_PER_SESSION 5 Close 1013; keyed on the anonymous canopy_session cookie for best-effort fairness behind NAT.

Cookieless first connections are exempt from the per-session cap and are backstopped by the global cap. These limits are availability controls, not authentication.

WS /ws/training

Description: Stream training state/metrics/topology/event updates.

Connection URL:

ws://127.0.0.1:8050/ws/training

Initial Status Message:

{
  "type": "initial_status",
  "data": {
    "is_training": true,
    "is_running": true,
    "phase": "output",
    "current_epoch": 42
  }
}

State Message:

{
  "type": "state",
  "timestamp": 1711459200.123,
  "data": {
    "status": "Started",
    "phase": "Output",
    "current_epoch": 42
  }
}

Metrics Message:

{
  "type": "metrics",
  "timestamp": 1711459201.123,
  "data": {
    "epoch": 43,
    "metrics": {
      "loss": 0.221,
      "accuracy": 0.885,
      "val_loss": 0.245,
      "val_accuracy": 0.87
    }
  }
}

Topology Message:

{
  "type": "topology",
  "timestamp": 1711459201.456,
  "data": {
    "input_units": 2,
    "hidden_units": 3,
    "output_units": 1,
    "nodes": [],
    "connections": []
  }
}

Event Message:

{
  "type": "event",
  "timestamp": 1711459202.123,
  "data": {
    "event": "training_complete"
  }
}

Ping/Pong:

{"type": "ping"}
{"type": "pong"}

Example JavaScript Client:

const ws = new WebSocket('ws://127.0.0.1:8050/ws/training');

ws.onopen = () => {
  console.log('Connected to training stream');
};

ws.onmessage = (event) => {
  const msg = JSON.parse(event.data);

  switch (msg.type) {
    case 'initial_status':
      console.log('Initial status:', msg.data);
      break;
    case 'state':
      console.log('State:', msg.data);
      break;
    case 'metrics':
      console.log('Epoch:', msg.data.epoch);
      console.log('Loss:', msg.data.metrics?.loss);
      break;
    case 'event':
      console.log('Event:', msg.data);
      break;
    case 'pong':
      console.log('Pong received');
      break;
  }
};

WS /ws/control

Description: Send training-control commands and receive command acknowledgements.

Connection URL:

ws://127.0.0.1:8050/ws/control

Connection Confirmation:

{
  "type": "connection_confirmed",
  "client_id": "control-client-12345"
}

Command Request:

{
  "command": "start",
  "reset": true
}

Command Success Response:

{
  "ok": true,
  "command": "start",
  "state": {
    "is_running": true,
    "current_epoch": 0
  }
}

Command Error Response:

{
  "ok": false,
  "error": "Unknown command: invalid_cmd"
}

Supported Commands: start, stop, pause, resume, reset, set_params

Data Models

TrainingMetrics

interface DemoHistoryMetric {
  epoch: number;
  metrics: {
    loss: number;
    accuracy: number;
    val_loss: number;
    val_accuracy: number;
  };
  network_topology: {
    input_units: number;
    hidden_units: number;
    output_units: number;
  };
  phase: string;
  timestamp: string; // ISO 8601 in demo mode
}

interface ServiceHistoryMetric {
  epoch: number;
  train_loss?: number;
  train_accuracy?: number;
  val_loss?: number;
  val_accuracy?: number;
  hidden_units?: number;
  phase?: string;
  timestamp?: number | string;
}

NetworkTopology

interface NetworkTopology {
  input_units: number;
  hidden_units: number;
  output_units: number;
  nodes: Node[];
  connections: Connection[];
}

interface Node {
  id: string; // e.g., "input_0", "hidden_1", "output_0"
  type: "input" | "hidden" | "output";
  layer?: number; // demo mode
  label?: string; // service mode
}

interface Connection {
  from: string;
  to: string;
  weight: number;
}

Dataset

interface DemoDataset {
  inputs: number[][];
  targets: number[] | number[][];
  num_samples: number;
  num_features: number;
  num_classes: number;
}

interface ServiceDataset {
  num_samples: number;
  num_features: number;
  num_classes: number;
  loaded?: boolean;
  train_samples?: number;
  test_samples?: number;
}

DecisionBoundary

interface DecisionBoundary {
  xx: number[][];  // X meshgrid [resolution, resolution]
  yy: number[][];  // Y meshgrid [resolution, resolution]
  Z: number[][];   // Class grid [resolution, resolution]
  x_min: number;
  x_max: number;
  y_min: number;
  y_max: number;
  resolution: number;
}

TrainingState

interface TrainingState {
  is_training?: boolean;
  is_running: boolean;
  is_paused: boolean;
  completed?: boolean;
  failed?: boolean;
  fsm_status?: string;
  phase?: string;
  current_epoch: number;
  current_loss?: number;
  current_accuracy?: number;
  hidden_units: number;
  metrics_count?: number;
}

Error Handling

HTTP Error Codes

  • 200 OK - Request succeeded
  • 302 Found - Redirect (root endpoint)
  • 400 Bad Request - Invalid parameters
  • 404 Not Found - Resource not available
  • 500 Internal Server Error - Server error
  • 503 Service Unavailable - Service unhealthy

Error Response Format

{
  "error": "Error message description",
  "detail": "Additional error details (optional)",
  "status_code": 400
}

Upstream Failures

When a call canopy makes to juniper-cascor or the recurrence service fails, what a caller reads -- an error or detail field, a control acknowledgement on the WebSocket, the /api/status envelope's error, /api/stream_health's last_disconnect_reason, a failed recurrence fit's completion_reason -- is:

  • the upstream's own answer, when it replied with an HTTP error status (for example cascor's 409 Training data not provided), or
  • the exception's type name alone (for example JuniperCascorConnectionError) for anything else: a refused header value, an unreachable or timed-out upstream, a malformed reply.

The transport text of such a failure -- URLs, socket errors, and a header value a client refused to send -- goes to canopy's logs only.

WebSocket Error Handling

Connection Errors:

  • Network disconnection → Client receives onclose event
  • Server shutdown → server_shutdown message sent before close
  • WebSocket admission cap exceeded → close code 1013
  • Origin/auth policy rejection → close code 4001, 4003, or 1008 depending on the failed gate

Command Errors:

{
  "ok": false,
  "error": "No backend available"
}

Best Practices:

  • Implement exponential backoff for reconnection
  • Handle onclose and onerror events
  • Send ping/pong heartbeats every 30 seconds
  • Gracefully degrade if WebSocket unavailable (fall back to REST polling)

Rate Limiting

REST rate limiting is available but disabled by default (JUNIPER_CANOPY_RATE_LIMIT_ENABLED=false). When enabled, requests are counted by API key or source IP. Canopy's own server-side dashboard self-calls carry an internal per-process header and are exempt so polling does not drain the user's bucket.

WebSocket availability limits are always enforced by WebSocketManager as described in Connection Admission and Limits.

Response Headers:

X-RateLimit-Limit: 100
X-RateLimit-Remaining: 95
X-RateLimit-Reset: 1699200000

429 Too Many Requests Response:

{
  "error": "Rate limit exceeded",
  "retry_after": 60
}

Code Examples

Python Client (REST API)

import requests

BASE_URL = "http://127.0.0.1:8050"

# Health check
response = requests.get(f"{BASE_URL}/api/health")
health = response.json()
print(f"Status: {health['status']}")
print(f"Active connections: {health['active_connections']}")

# Get current status
response = requests.get(f"{BASE_URL}/api/status")
status = response.json()
print(f"Training: {status['is_training']}")
print(f"Epoch: {status['current_epoch']}")
print(f"Loss: {status.get('current_loss', status.get('train_loss'))}")

# Get recent metrics
response = requests.get(f"{BASE_URL}/api/metrics/history?limit=10")
metrics_payload = response.json()
for m in metrics_payload.get("history", []):
    epoch = m.get("epoch", 0)
    if "metrics" in m:
        loss = m["metrics"].get("loss")
    else:
        loss = m.get("train_loss")
    print(f"Epoch {epoch}: Loss={loss}")

# Get topology
response = requests.get(f"{BASE_URL}/api/topology")
topology = response.json()
print(f"Network: {topology['input_units']}-{topology['hidden_units']}-{topology['output_units']}")

# Get dataset
response = requests.get(f"{BASE_URL}/api/dataset")
dataset = response.json()
print(f"Dataset: {dataset['num_samples']} samples, {dataset['num_features']} features")

Python Client (WebSocket)

import asyncio
import json
import websockets

async def training_monitor():
    uri = "ws://127.0.0.1:8050/ws/training"

    async with websockets.connect(uri) as websocket:
        print("Connected to training stream")

        async for message in websocket:
            data = json.loads(message)
            msg_type = data.get('type')

            if msg_type == 'initial_status':
                print('Initial status received')

            elif msg_type == 'state':
                print('State update:', data['data'])

            elif msg_type == 'metrics':
                metrics = data['data'].get('metrics', {})
                epoch = data['data'].get('epoch')
                print(f"Epoch {epoch}: Loss={metrics.get('loss')} Acc={metrics.get('accuracy')}")

            elif msg_type == 'event':
                print('Event:', data['data'])

# Run monitor
asyncio.run(training_monitor())

Python Control Client

import asyncio
import json
import websockets

async def training_controller():
    uri = "ws://127.0.0.1:8050/ws/control"

    async with websockets.connect(uri) as websocket:
        # Start training
        await websocket.send(json.dumps({
            'command': 'start',
            'reset': True
        }))

        response = await websocket.recv()
        data = json.loads(response)

        if data['ok']:
            print("Training started:", data['state'])
        else:
            print("Error:", data['error'])

        # Wait 10 seconds
        await asyncio.sleep(10)

        # Pause training
        await websocket.send(json.dumps({'command': 'pause'}))
        response = await websocket.recv()
        print("Paused:", json.loads(response))

        # Wait 5 seconds
        await asyncio.sleep(5)

        # Resume training
        await websocket.send(json.dumps({'command': 'resume'}))
        response = await websocket.recv()
        print("Resumed:", json.loads(response))

asyncio.run(training_controller())

JavaScript Client (Browser)

// Training monitor
const trainingWs = new WebSocket('ws://127.0.0.1:8050/ws/training');

trainingWs.onmessage = (event) => {
  const data = JSON.parse(event.data);

  switch (data.type) {
    case 'initial_status':
      console.log('Initial status:', data.data);
      break;

    case 'state':
      updateStatusBanner(data.data);
      break;

    case 'metrics':
      const { epoch, metrics } = data.data;
      updateMetricsChart(epoch, metrics.loss, metrics.accuracy);
      break;

    case 'topology':
      updateTopologyGraph(data.data);
      break;

    case 'event':
      console.log('Training event:', data.data);
      break;
  }
};

// Control client
const controlWs = new WebSocket('ws://127.0.0.1:8050/ws/control');

function startTraining() {
  controlWs.send(JSON.stringify({
    command: 'start',
    reset: true
  }));
}

function pauseTraining() {
  controlWs.send(JSON.stringify({ command: 'pause' }));
}

function resumeTraining() {
  controlWs.send(JSON.stringify({ command: 'resume' }));
}

function stopTraining() {
  controlWs.send(JSON.stringify({ command: 'stop' }));
}

controlWs.onmessage = (event) => {
  const response = JSON.parse(event.data);
  if (response.ok) {
    console.log(`Command '${response.command}' succeeded`);
    updateTrainingState(response.state);
  } else {
    console.error('Command failed:', response.error);
  }
};

curl Examples

# Health check
curl http://127.0.0.1:8050/api/health

# Get status
curl http://127.0.0.1:8050/api/status

# Get metrics (last 10)
curl "http://127.0.0.1:8050/api/metrics/history?limit=10"

# Get topology
curl http://127.0.0.1:8050/api/topology

# Get dataset
curl http://127.0.0.1:8050/api/dataset

# Get decision boundary
curl http://127.0.0.1:8050/api/decision_boundary

# Get statistics
curl http://127.0.0.1:8050/api/statistics

# Pretty print JSON
curl -s http://127.0.0.1:8050/api/health | python -m json.tool

Best Practices

REST API

  1. Always check HTTP status codes

    response = requests.get(url)
    if response.status_code == 200:
        data = response.json()
    else:
        print(f"Error: {response.status_code}")
  2. Use appropriate timeouts

    response = requests.get(url, timeout=5)  # 5 second timeout
  3. Handle errors gracefully

    try:
        response = requests.get(url, timeout=5)
        response.raise_for_status()
        data = response.json()
    except requests.exceptions.RequestException as e:
        print(f"Request failed: {e}")
  4. Limit data retrieval

    # Don't retrieve all metrics
    response = requests.get(f"{BASE_URL}/api/metrics/history?limit=100")

WebSocket

  1. Implement reconnection logic

    let reconnectDelay = 1000;
    const maxDelay = 30000;
    
    function connect() {
      const ws = new WebSocket(url);
    
      ws.onclose = () => {
        setTimeout(() => {
          reconnectDelay = Math.min(reconnectDelay * 2, maxDelay);
          connect();
        }, reconnectDelay);
      };
    
      ws.onopen = () => {
        reconnectDelay = 1000; // Reset on successful connection
      };
    }
  2. Send heartbeats

    setInterval(() => {
      if (ws.readyState === WebSocket.OPEN) {
        ws.send(JSON.stringify({ type: 'ping' }));
      }
    }, 30000); // Every 30 seconds
  3. Handle backpressure

    if (ws.bufferedAmount === 0) {
      ws.send(message);  // Safe to send
    } else {
      console.warn('Buffer full, skipping message');
    }
  4. Clean up on disconnect

    window.addEventListener('beforeunload', () => {
      ws.close(1000, 'Page unload');
    });

Support and Contact


End of API Reference