Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
25 changes: 17 additions & 8 deletions cmd/server/main.go
Original file line number Diff line number Diff line change
Expand Up @@ -80,14 +80,23 @@ func main() {
log.Fatalf("Failed to run database migrations: %v", err)
}

// Initialize the instance manager with dependency injection
instanceManager := manager.New(&cfg, db)

// Initialize model manager
modelManager := models.NewManager(cfg.Backends.LlamaCpp.CacheDir, cfg.Backends.LlamaCpp.DownloadTimeout, cfg.Version)

var oidcService *server.OIDCService
if cfg.Auth.OIDC.Enabled() {
oidcService, err = server.NewOIDCService(cfg.Auth)
if err != nil {
// Degrade to key-only auth only
log.Printf("Warning: OIDC login unavailable, continuing without it: %v", err)
oidcService = nil
} else {
log.Printf("OIDC login enabled (issuer: %s)", cfg.Auth.OIDC.IssuerURL)
}
}

// Create a new handler with the instance manager
handler := server.NewHandler(instanceManager, modelManager, cfg, db)
handler := server.NewHandler(instanceManager, modelManager, cfg, db, oidcService)

// Setup the router with the handler
r := server.SetupRouter(handler)
Expand All @@ -102,15 +111,15 @@ func main() {
}

go func() {
fmt.Printf("Llamactl server listening on %s:%d\n", cfg.Server.Host, cfg.Server.Port)
log.Printf("Llamactl server listening on %s:%d", cfg.Server.Host, cfg.Server.Port)
if err := server.ListenAndServe(); err != nil && err != http.ErrServerClosed {
log.Printf("Error starting server: %v\n", err)
}
}()

// Wait for shutdown signal
<-stop
fmt.Println("Shutting down server...")
log.Println("Shutting down server...")

// Create shutdown context with timeout
shutdownCtx, shutdownCancel := context.WithTimeout(context.Background(), 30*time.Second)
Expand All @@ -120,7 +129,7 @@ func main() {
if err := server.Shutdown(shutdownCtx); err != nil {
log.Printf("Error shutting down server: %v\n", err)
} else {
fmt.Println("Server shut down gracefully.")
log.Println("Server shut down gracefully.")
}

// Stop all instances and cleanup
Expand All @@ -133,5 +142,5 @@ func main() {
log.Printf("Error closing database: %v\n", err)
}

fmt.Println("Exiting llamactl.")
log.Println("Exiting llamactl.")
}
73 changes: 61 additions & 12 deletions docs/api-keys.md → docs/authentication.md
Original file line number Diff line number Diff line change
@@ -1,13 +1,16 @@
# API Keys
# Authentication

Llamactl uses two types of API keys to control access:
Llamactl controls access with three mechanisms:

- **Management API Key** — authenticates the web UI and management API (creating/stopping instances, managing keys). See the [Configuration](configuration.md) guide for how to set these.
- **Inference API Key** — authenticates OpenAI-compatible inference requests (`/v1/chat/completions`, `/v1/completions`, etc.). These are created and managed via the web UI or management API and stored in the database.
- **Management API Key** — authenticates the web UI and management API (creating/stopping instances, managing keys). Configured in the config file or via environment variables; see the [Configuration](configuration.md#authentication-configuration) guide.
- **Inference API Key** — authenticates OpenAI-compatible inference requests (`/v1/chat/completions`, `/v1/completions`, etc.). Created and managed via the web UI or management API and stored in the database.
- **OIDC / SSO Session** — optional browser login for the web UI through an OpenID Connect provider; see [OIDC / SSO Login](#oidc-sso-login) below.

This page focuses on **inference API keys** and how their per-instance permissions work.
The rest of this page covers **inference API keys** and their per-instance permissions, followed by **SSO setup**.

## Permission Modes
## Inference API Keys

### Permission Modes

When you create an inference key, you choose a permission mode:

Expand All @@ -16,7 +19,7 @@ When you create an inference key, you choose a permission mode:

Management API keys bypass all permission checks entirely.

## Access Levels
### Access Levels

For each instance granted to a per-instance key, you pick one of three access levels. A key can always **use a running instance** — the level only affects what happens when the instance is **stopped** at request time.

Expand All @@ -31,7 +34,7 @@ For each instance granted to a per-instance key, you pick one of three access le

Both flags default to enabled, so omitting them is equivalent to *Can start and evict others* (backward compatible with keys created before this feature).

## Behavior Reference
### Behavior Reference

How a request is handled depends on the instance's state and the key's access level for it:

Expand All @@ -46,9 +49,9 @@ How a request is handled depends on the instance's state and the key's access le

"Capacity" considers both the instance's [group limit](managing-instances.md#instance-groups) and the global `max_running_instances`. The two 503 responses use distinct error types so clients can tell them apart.

## Managing Keys
### Managing Keys

### Via Web UI
#### Via Web UI

1. Open the web UI and log in with a management API key
2. Navigate to **Settings → API Keys**
Expand All @@ -63,7 +66,7 @@ How a request is handled depends on the instance's state and the key's access le

Expand an existing key in the list to review its per-instance access levels.

### Via API
#### Via API

All key endpoints require a management API key (`<token>` below).

Expand Down Expand Up @@ -114,8 +117,54 @@ curl -X DELETE http://localhost:8080/api/v1/auth/keys/{id} \
-H "Authorization: Bearer <token>"
```

## Use Cases
### Use Cases

- **Production vs. development**: Give production keys *Can start and evict* so requests always succeed; give development keys *Can start on demand* so they never disrupt production workloads.
- **Read-only consumers**: Use *Use running only* for dashboards or monitoring tools that should reuse a warm instance but never spin one up.
- **Shared clusters**: Limit disruptive eviction rights to a small set of trusted keys so noisy neighbors can't evict each other's models.

## OIDC / SSO Login

The web UI can authenticate users through an OpenID Connect provider (Authentik, Keycloak, Authelia, etc.) instead of a management API key. OIDC login is enabled when both `issuer_url` and `client_id` are set; management API keys keep working alongside it, and inference endpoints always use API keys.

```yaml
auth:
oidc:
issuer_url: "https://authentik.example.com/application/o/llamactl/" # IdP issuer URL
client_id: "llamactl" # OAuth2 client ID registered with the IdP
client_secret: "your-client-secret" # OAuth2 client secret
# redirect_url: "" # Optional; derived from the request when empty
# scopes: [openid, profile, email] # Default scopes
# allowed_groups: [] # Restrict login to these groups (empty = all IdP users)
# groups_claim: "groups" # ID-token claim carrying group names
# session_ttl: 12h # Session lifetime (default: 12h)
# secure_cookie: true # Set false only for plain-HTTP LAN deployments
```

Environment variable equivalents (`LLAMACTL_AUTH_OIDC_*`) are listed in the [configuration reference](configuration.md#oidc-sso-login).

### Setting up the IdP client

1. Create a confidential OAuth2/OIDC client in your IdP (Authentik: provider type *OAuth2/OpenID Connect*; Keycloak: client with *Client authentication* enabled)
2. Set the redirect URI to `https://<your-llamactl-host>/api/v1/auth/oidc/callback` — under a subpath proxy, use the full external path
3. Configure the `auth.oidc` block above and restart llamactl. If discovery fails at startup (IdP unreachable, wrong issuer), llamactl logs the error and runs without OIDC until the problem is fixed — management-key login keeps working.

### Restricting logins to specific groups

By default, every user your IdP authenticates can log in. When the IdP serves more users than should have access to llamactl, set `allowed_groups` to require membership in at least one listed group:

```yaml
auth:
oidc:
allowed_groups: ["llamactl-users"]
```

- Groups are read from the ID token's `groups` claim (Authelia, Authentik, Keycloak, and Dex can all emit it — sometimes a mapper must be enabled). If your IdP exposes membership under a different claim, point `groups_claim` at it.
- The check is fail-closed: if the claim is missing or unreadable, login is denied and the reason is logged with the claim names the token actually carried — check the llamactl log if a legitimate user cannot get in.
- Membership is evaluated at login only. Removing a user from the group does not terminate their existing session; access ends when the session expires (`session_ttl`, default 12h) or llamactl restarts. Lower `session_ttl` if you need faster revocation.

### Notes

- Keep `require_management_auth: true` when enabling OIDC — otherwise management endpoints stay unauthenticated and login only adds session support. llamactl logs a warning at startup for this combination.
- Behind a TLS-terminating reverse proxy, either forward `X-Forwarded-Proto`/`X-Forwarded-Host` or set `redirect_url` explicitly. Under a subpath proxy, set `redirect_url` to the full external callback path.
- API keys created while logged in via SSO show the user's email as their owner.
58 changes: 37 additions & 21 deletions docs/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -230,7 +230,7 @@ backends:

### Custom Backends

Custom backends are user-defined named entries under `backends.custom` for managing inference servers without native support. Each entry accepts the same fields as the built-in backends; see the [Custom Backends](custom-backends.md) guide for how they work.
Custom backends are user-defined named entries under `backends.custom` for managing inference servers without native support. Each entry accepts the same fields as the built-in backends, and `args` support the `{port}` and `{model}` placeholders; see the [Custom Backends](custom-backends.md) guide for how they work.

```yaml
backends:
Expand All @@ -248,13 +248,6 @@ backends:
response_headers: {} # Additional response headers to send with responses
```

Configured `args` are prepended to instance arguments and support the `{port}` and `{model}` placeholders; see [Custom Backends](custom-backends.md) for details.

!!! note
- Custom backends are configured in the config file only; there are no environment variable overrides for them.
- On multi-node setups, the command must exist on the node that runs the instance.
- Servers that download models or run long setup on first start may exceed the on-demand start timeout (`instances.on_demand_start_timeout`). Raise the timeout or pre-warm the server's model cache by running the command once manually.

### Data Directory Configuration

```yaml
Expand Down Expand Up @@ -331,10 +324,13 @@ database:

### Authentication Configuration

llamactl supports two types of authentication:
llamactl supports three authentication mechanisms:

- **Management API Keys**: For accessing the web UI and management API (creating/managing instances). These can be configured in the config file or via environment variables.
- **Inference API Keys**: For accessing the OpenAI-compatible inference endpoints. These are managed via the web UI (Settings → API Keys) and stored in the database.
- **OIDC / SSO Sessions**: Optional web UI login through an OpenID Connect provider, configured via the `auth.oidc` block below.

See [Authentication](authentication.md) for inference key permissions and SSO setup.

```yaml
auth:
Expand All @@ -343,24 +339,44 @@ auth:
management_keys: [] # List of valid management API keys
```

**Managing Inference API Keys:**

Inference API keys are managed through the web UI or management API and stored in the database. To create and manage inference keys:

1. Open the web UI and log in with a management API key
2. Navigate to **Settings → API Keys**
3. Click **Create API Key**
4. Configure the key:
- **Name**: A descriptive name for the key
- **Expiration**: Optional expiration date
- **Permissions**: Grant access to all instances or specific instances only, with per-instance access levels controlling auto-start and eviction (see [API Keys](api-keys.md) for details)
5. Copy the generated key - it won't be shown again
Inference API keys are created and managed via the web UI or management API

**Environment Variables:**
- `LLAMACTL_REQUIRE_INFERENCE_AUTH` - Require auth for OpenAI endpoints (true/false)
- `LLAMACTL_REQUIRE_MANAGEMENT_AUTH` - Require auth for management endpoints (true/false)
- `LLAMACTL_MANAGEMENT_KEYS` - Comma-separated management API keys

#### OIDC / SSO Login

The web UI can authenticate users through an OpenID Connect provider (Authentik, Keycloak, Authelia, etc.) instead of a management API key. OIDC login is enabled when both `issuer_url` and `client_id` are set; management API keys keep working alongside it, and inference endpoints always use API keys.

```yaml
auth:
oidc:
issuer_url: "https://authentik.example.com/application/o/llamactl/" # IdP issuer URL
client_id: "llamactl" # OAuth2 client ID registered with the IdP
client_secret: "your-client-secret" # OAuth2 client secret
# redirect_url: "" # Optional; derived from the request when empty
# scopes: [openid, profile, email] # Default scopes
# allowed_groups: [] # Restrict login to these groups (empty = all IdP users)
# groups_claim: "groups" # ID-token claim carrying group names
# session_ttl: 12h # Session lifetime (default: 12h)
# secure_cookie: true # Set false only for plain-HTTP LAN deployments
```

**Environment Variables:**
- `LLAMACTL_AUTH_OIDC_ISSUER_URL` - IdP issuer URL
- `LLAMACTL_AUTH_OIDC_CLIENT_ID` - OAuth2 client ID
- `LLAMACTL_AUTH_OIDC_CLIENT_SECRET` - OAuth2 client secret
- `LLAMACTL_AUTH_OIDC_REDIRECT_URL` - Redirect URL override
- `LLAMACTL_AUTH_OIDC_SCOPES` - Comma-separated scopes
- `LLAMACTL_AUTH_OIDC_ALLOWED_GROUPS` - Comma-separated groups allowed to log in
- `LLAMACTL_AUTH_OIDC_GROUPS_CLAIM` - Name of the ID-token claim carrying groups (default: `groups`)
- `LLAMACTL_AUTH_OIDC_SESSION_TTL` - Session lifetime (Go duration, e.g. `12h`)
- `LLAMACTL_AUTH_OIDC_SECURE_COOKIE` - Mark the session cookie Secure (true/false)

See [Authentication](authentication.md#oidc-sso-login) for IdP client setup.

### Remote Node Configuration

llamactl supports remote node deployments. Configure remote nodes to deploy instances on remote hosts and manage them centrally.
Expand Down
4 changes: 4 additions & 0 deletions docs/custom-backends.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,6 +14,8 @@ The managed server must meet three requirements:

The server learns its port through the `{port}` placeholder in its arguments — llamactl has no other way to communicate it.

On multi-node setups, the command must exist on the node that runs the instance.

## Creating an Instance

Instances of a custom backend are created like any other instance, with `backend_type: "custom"`; see [Managing Instances](managing-instances.md) for the API examples and instance options. In the web UI, each configured name appears in the backend dropdown.
Expand All @@ -23,6 +25,8 @@ The launched command receives arguments from two places: the entry's `args` firs
- `{port}` — the port assigned to the instance. The combined arguments must contain it somewhere, otherwise the server cannot learn its port and instance creation is rejected.
- `{model}` — the model identifier of the instance. Optional; if the arguments use it, the instance must have a model set.

Servers that download models or run long setup on first start may exceed the on-demand start timeout (`instances.on_demand_start_timeout`). Raise the timeout or pre-warm the server's model cache by running the command once manually.

## Health Checks

After launching the command, llamactl polls the entry's `health_path` until it returns a 2xx response, and only then marks the instance ready.
Loading
Loading