Kubernetes-native low-latency distributed workflow orchestrator
Ortrix is a low-latency, Kubernetes-native distributed workflow orchestrator built on partitioned execution, streaming task dispatch, and locality-aware scheduling. It delivers sub-millisecond task dispatch — orders of magnitude faster than poll-based systems.
- Push-based execution — Tasks are streamed to workers instantly via persistent gRPC connections. No polling, no queue consumption delays.
- Partitioned execution model —
hash(workflow_id) → partition → single owner. No distributed locks on the hot path. - Event-sourced WAL — Durable, replayable state with hybrid local/replicated write-ahead log.
- Embedded worker SDK — No separate worker services. Import the SDK into your existing Go services and declare capabilities.
- Capability-based routing — Workers self-declare what they can do. The orchestrator routes dynamically.
- Locality-aware scheduling — Prefer same-pod → same-node → same-zone → any, minimizing network hops.
- Priority queues with fairness — HIGH / MEDIUM / LOW with weighted fair queuing to prevent starvation.
- Saga pattern support — Built-in compensation for multi-step workflows.
- At-least-once execution — Idempotency guarantees for safe retries.
- Canary and blue-green deployments — Version-aware routing for safe rollouts.
- Fast recovery — Snapshot + WAL replay reconstructs state in hundreds of milliseconds.
Client ──▶ Gateway (control plane: auth, routing)
│
▼
Orchestrator (partitioned, in-memory + WAL)
│
gRPC stream (push)
│
▼
Your Service (embedded Worker SDK)
Control Plane (Gateway): Handles bootstrap, authentication, and routing metadata. Not in the execution path.
Data Plane (Orchestrator ↔ Workers): All task dispatch and result collection flows directly over persistent gRPC streams. No intermediate hops, no polling.
Workers: The SDK embeds into your services. Register task handlers, connect to the orchestrator, and receive tasks instantly.
See docs/architecture.md for the full design.
Traditional workflow engines use poll-based task dispatch. Workers repeatedly ask "any work for me?" — adding 500ms+ of latency on every task and wasting resources on empty polls.
Ortrix eliminates this with push-based streaming:
| Metric | Poll-based (e.g., Temporal) | Ortrix (push) |
|---|---|---|
| P50 dispatch | ~500ms | ~1ms |
| P99 dispatch | ~1000ms | ~5ms |
| Idle overhead | Continuous polling | Zero |
Ortrix also eliminates the need for:
- External databases — WAL provides durability without Cassandra/Postgres
- Separate worker infrastructure — SDK embeds into existing services
- Complex deployment — Kubernetes-native from day one
| Dimension | Ortrix | Temporal |
|---|---|---|
| Dispatch model | Push (gRPC streaming) | Pull (long polling) |
| Dispatch latency | ~1ms | ~500ms |
| State storage | In-memory + WAL | External database |
| Workers | Embedded SDK | Separate processes |
| Routing | Capability + locality aware | Task queue based |
| Best for | Low-latency, high-throughput | Rich workflow semantics |
See docs/comparison.md for a detailed analysis.
Ortrix is built on a worker-initiated, zero-exposed-port communication model with intelligent load distribution:
- Worker-initiated connections — Workers open outbound gRPC streams to orchestrators. Orchestrators never dial worker pods. This eliminates inbound attack surface on workers.
- No exposed worker ports — Workers require zero listening ports for orchestration traffic. They are invisible to port scanners and network probes.
- mTLS everywhere — Every connection uses mutual TLS with X.509 service identity. Both sides authenticate on every connection.
- Backpressure-based scheduling — Workers advertise available capacity. The orchestrator respects these limits and never overloads a worker. Tasks queue safely when capacity is exhausted.
- Intelligent load distribution — The orchestrator combines locality scores (same-node, same-zone) with real-time load data (available slots) to select the optimal worker for each task.
Worker ──(outbound mTLS)──▶ Orchestrator
│ │
│ READY(capacity=10) │
│◀──Task──────────────│ (respects capacity)
│──Result─────────────▶│
│◀──Task──────────────│ (load-aware selection)
See docs/security.md and docs/proposals/streaming-protocol.md for full details.
- Go 1.24+
- protoc (for proto generation)
- swag (for Swagger doc generation)
make build# Terminal 1: Start the orchestrator
make run-orchestrator
# Terminal 2: Start the gateway
make run-gatewayOrtrix exposes a REST API alongside its gRPC interface. API documentation is generated using Swagger/OpenAPI via swaggo/swag.
make swaggerThis generates swagger.json, swagger.yaml, and docs.go into docs/swagger/.
make run-gatewayThe gateway starts two servers:
- gRPC on port
8080(configurable viaORTRIX_GATEWAY_PORT) - HTTP/REST on port
8081(gRPC port + 1)
Access the Swagger UI at:
http://localhost:8081/swagger/index.html
Note: Swagger UI is enabled by default in
developmentmode. SetORTRIX_SWAGGER_ENABLED=falseto disable it, orORTRIX_ENVIRONMENT=productionto disable it automatically.
make helppackage main
import (
"context"
"github.com/mayur-tolexo/ortrix/pkg/sdk"
)
func main() {
w := sdk.NewWorker("my-service")
w.RegisterHandler("process_order", func(ctx context.Context, taskID string, payload []byte) ([]byte, error) {
// Your task logic here
return []byte(`{"status": "done"}`), nil
})
w.Start(context.Background(), "localhost:9090")
}# Create a local cluster
make kind-create
# Build Docker images
make docker-all
# Deploy with Helm (coming soon)ortrix/
├── assets/logo/ # Logo and branding assets
├── api/proto/ # gRPC/Protobuf service definitions
├── cmd/
│ ├── gateway/ # Gateway service entry point (gRPC + HTTP/REST)
│ └── orchestrator/ # Orchestrator service entry point
├── docs/
│ └── swagger/ # Generated Swagger/OpenAPI docs
├── internal/
│ ├── config/ # Configuration management
│ ├── logging/ # Structured logging
│ ├── partition/ # Partition ownership and management
│ ├── routing/ # Task routing logic
│ ├── scheduler/ # Priority scheduling
│ └── wal/ # Write-ahead log
├── pkg/sdk/ # Worker SDK (public API)
├── deploy/
│ ├── helm/ # Helm charts
│ └── k8s/ # Kubernetes manifests
└── docs/ # Design documentation
| Document | Description |
|---|---|
| Architecture | System design, control vs data plane, partition model |
| Execution Model | Push-based dispatch, gRPC streaming, task lifecycle |
| State and WAL | Event sourcing, WAL format, snapshots, recovery |
| Scheduling and Routing | Capability routing, locality scheduling, priority queues |
| Failure Handling | Crash recovery, idempotency, saga compensation |
| Partitioning and Scaling | Leases, rebalancing, failover, horizontal scaling |
| Security | mTLS, service identity, authorization |
| Performance | Latency analysis, batching, WAL optimization |
| Comparison | Ortrix vs Temporal |
| Future Work | Research directions: rebalancing, replication, multi-region |
| Streaming Protocol | Protocol design: capacity signaling, flow control, failure handling |
| Proposals | Design proposals for upcoming features |
| Roadmap | Phased development plan with goals and deliverables |
| Testing | Testing strategy, coverage requirements, CI enforcement |
Ortrix follows a proposal-driven development model with phased delivery and strong testing discipline:
-
Proposal-driven development: Every significant feature begins as a written proposal in
docs/proposals/. Proposals define the problem, motivation, design, alternatives, tradeoffs, testing strategy, and rollout plan. This ensures architectural decisions are deliberate, reviewed, and documented before implementation begins. -
Phased delivery: The roadmap defines six development phases, each building on the previous. Each phase delivers a functional, testable system with clear goals, deliverables, and dependencies. This enables incremental progress with confidence.
-
Strong testing discipline: The testing strategy mandates 80%+ code coverage, table-driven unit tests, integration tests for component interactions, and failure tests for crash recovery scenarios. All PRs must include tests, and CI fails if coverage drops below the threshold.
Ortrix is actively evolving. Key areas of upcoming development:
- Locality-aware partition migration — Move partitions closer to their execution zones, reducing cross-zone latency and egress costs
- Load-based rebalancing — Detect hot orchestrator nodes and automatically redistribute partitions based on CPU, queue depth, and latency
- Hot partition mitigation — Sub-partitioning and key-based sharding to break up workflow hotspots
- Partition replication & fast failover — Warm standby nodes with continuous WAL streaming for near-instant promotion
- Multi-region support — Geo-distributed orchestration with cross-region WAL replication and global routing
See docs/future-work.md for the full roadmap with design details, ASCII diagrams, and tradeoff analysis.
Want to contribute? These are excellent areas for new contributors. Check the roadmap and pick an area that interests you.
The Ortrix logo represents:
- Distributed orchestration (graph structure)
- Streaming execution (flow arrows)
- Go-native ecosystem (gopher mascot)
- Central orchestration engine (core cube)
assets/logo/ortrix-logo.pngassets/logo/favicon.png
Use the logo consistently across documentation and UI surfaces.
Note: The logo is optimized for light backgrounds. When using on dark surfaces, consider adding a light backdrop or padding.
The logo assets are reusable across:
- Documentation sites
- Swagger UI branding
- CLI splash screens
We welcome contributions! See CONTRIBUTING.md for:
- How to run locally
- How to add a new capability
- Coding standards
- Testing expectations
- PR guidelines
See LICENSE for details.
