Skip to content

Commit c03eabd

Browse files
Merge upstream master into feat/add-dependabot-govulncheck
2 parents f9b0930 + fdec720 commit c03eabd

81 files changed

Lines changed: 4939 additions & 585 deletions

Some content is hidden

Large Commits have some content hidden by default. Use the searchbox below for content that may be hidden.

‎.github/copilot-instructions.md‎

Lines changed: 115 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,115 @@
1+
# GitHub Copilot Instructions for PipeCD
2+
3+
PipeCD is a GitOps-style continuous delivery platform. It has two main runtime components: **Control Plane** (`cmd/pipecd`) and **Piped agent** (`cmd/piped`/`cmd/pipedv1`), plus a CLI (`cmd/pipectl`) and a React web frontend (`web/`).
4+
5+
## Build, Test, and Lint
6+
7+
### Pre-commit check (run this before submitting a PR)
8+
```bash
9+
make check # runs build + lint + test + generated code check + DCO check
10+
```
11+
12+
### Go
13+
```bash
14+
make build/go # build all Go binaries into .artifacts/
15+
make build/go MOD=pipecd # build a single binary (pipecd, piped, pipectl, launcher)
16+
make test/go # run all Go tests
17+
go test -run TestFooBar ./pkg/foo/... # run a single test
18+
make lint/go # lint via Docker (golangci-lint)
19+
```
20+
21+
To test or build a specific plugin module (each plugin is its own Go module):
22+
```bash
23+
go -C pkg/app/pipedv1/plugin/kubernetes test -race ./...
24+
go -C pkg/app/pipedv1/plugin/ecs test -race ./...
25+
```
26+
27+
### Web (React/TypeScript with Yarn)
28+
```bash
29+
make run/web # start dev server at localhost:9090 with MSW mocks
30+
make test/web # run tests with coverage
31+
yarn --cwd web lint # lint frontend
32+
yarn --cwd web typecheck
33+
```
34+
35+
### Code generation (Protobuf / API)
36+
```bash
37+
make gen/code # regenerate .pb.go and .pb.validate.go (runs via Docker)
38+
```
39+
40+
### Local development environment
41+
```bash
42+
make up/local-cluster # start local kind cluster + registry
43+
make run/pipecd # build and deploy control plane to local cluster
44+
# then: kubectl port-forward -n pipecd svc/pipecd 8080
45+
make run/piped CONFIG_FILE=path/to/piped-config.yaml INSECURE=true
46+
make down/local-cluster # teardown
47+
```
48+
49+
## Architecture
50+
51+
```
52+
cmd/
53+
pipecd/ Control Plane: gRPC server for piped connections, web auth, deployment management
54+
piped/ Legacy piped agent (platform-specific deployment logic built-in)
55+
pipedv1/ Next-gen piped agent (plugin-based architecture)
56+
pipectl/ CLI tool for interacting with control plane
57+
launcher/ Enables remote upgrade of the piped agent
58+
59+
pkg/
60+
model/ Protobuf-defined domain models (.proto → .pb.go + .pb.validate.go)
61+
config/ Application and piped configuration (legacy, for piped v0)
62+
configv1/ Application and piped configuration (for pipedv1/plugin arch)
63+
app/
64+
server/ Control plane application logic; gRPC services in service/
65+
piped/ Legacy piped agent logic
66+
pipedv1/ Next-gen piped agent logic
67+
plugin/ Platform plugins (kubernetes, ecs, cloudrun, lambda, etc.)
68+
rpc/ gRPC server/client utilities, interceptors
69+
plugin/
70+
api/ Plugin gRPC API definitions (protobuf)
71+
sdk/ Plugin SDK (Go)
72+
73+
web/ React + TypeScript frontend
74+
manifests/ Helm charts for all components
75+
```
76+
77+
### Control Plane ↔ Piped communication
78+
The control plane exposes a gRPC API defined in `pkg/app/server/service/pipedservice/service.proto`. Piped agents connect to this API. Web clients use `pkg/app/server/service/webservice/service.proto`. External API consumers use `pkg/app/server/service/apiservice/service.proto`.
79+
80+
### Plugin architecture (pipedv1)
81+
Each plugin (`kubernetes`, `ecs`, `cloudrun`, `lambda`, `scriptrun`, etc.) lives in `pkg/app/pipedv1/plugin/<name>/` as a **separate Go module** with its own `go.mod`. Plugins implement the SDK interfaces for `deployment`, `livestate`, and `planpreview`. At runtime they are separate binaries communicating via gRPC.
82+
83+
## Key Conventions
84+
85+
### License header
86+
Every new Go file must start with this header (the year should be the year first published, not necessarily the current year):
87+
```go
88+
// Copyright 2024 The PipeCD Authors.
89+
//
90+
// Licensed under the Apache License, Version 2.0 (the "License");
91+
// you may not use this file except in compliance with the License.
92+
// You may obtain a copy of the License at
93+
//
94+
// http://www.apache.org/licenses/LICENSE-2.0
95+
//
96+
// Unless required by applicable law or agreed to in writing, software
97+
// distributed under the License is distributed on an "AS IS" BASIS,
98+
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
99+
// See the License for the specific language governing permissions and
100+
// limitations under the License.
101+
```
102+
103+
### Banned imports (enforced by depguard lint)
104+
- `sync/atomic` → use `go.uber.org/atomic` instead
105+
- `io/ioutil` → use `os` or `io` functions instead
106+
- `pipedv1` code must import `github.com/pipe-cd/pipecd/pkg/configv1`, not `pkg/config`
107+
- Plugin code under `pkg/app/pipedv1/plugin/` must NOT import from `github.com/pipe-cd/pipecd` (the main module). Only the `github.com/pipe-cd/piped-plugin-sdk-go` SDK is permitted.
108+
109+
### Protobuf / generated files
110+
Models in `pkg/model/` are defined in `.proto` files and compiled to `.pb.go` and `.pb.validate.go`. Do not manually edit generated files — run `make gen/code` instead.
111+
112+
### Commits
113+
- Sign off every commit: `git commit -s` (DCO required)
114+
- Commit message: single sentence, present tense, capital first letter (e.g., `Add imports to Terraform plan result`)
115+
- PRs target the `master` branch

‎.github/workflows/build.yaml‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -14,7 +14,7 @@ on:
1414

1515
env:
1616
GO_VERSION: 1.25.0
17-
NODE_VERSION: 18.12.0
17+
NODE_VERSION: 20.19.0
1818
HELM_VERSION: 3.8.2
1919

2020
jobs:

‎.github/workflows/lint.yaml‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -15,7 +15,7 @@ concurrency:
1515

1616
env:
1717
GO_VERSION: 1.25.0
18-
NODE_VERSION: 18.12.0
18+
NODE_VERSION: 20.19.0
1919
GOLANGCI_LINT_VERSION: v2.4.0
2020
HELM_VERSION: 3.17.3
2121

‎.github/workflows/publish_site.yaml‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -31,7 +31,7 @@ jobs:
3131
- name: Setup Node
3232
uses: actions/setup-node@v3
3333
with:
34-
node-version: '14'
34+
node-version: '24'
3535

3636
# Build site.
3737
- name: Build site

‎.github/workflows/test.yaml‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -16,7 +16,7 @@ concurrency:
1616
cancel-in-progress: ${{ github.event_name == 'pull_request' }}
1717
env:
1818
GO_VERSION: 1.25.0
19-
NODE_VERSION: 18.12.0
19+
NODE_VERSION: 20.19.0
2020

2121
jobs:
2222
list-go-modules:

‎cmd/pipecd/README.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -4,7 +4,7 @@
44
## Prerequisites
55

66
- [Go 1.24 or later](https://go.dev/)
7-
- [NodeJS v20 or later](https://nodejs.org/en/)
7+
- [NodeJS v20.19.0 or later](https://nodejs.org/en/)
88
- [Docker](https://www.docker.com/)
99
- [kind](https://kind.sigs.k8s.io/docs/user/quick-start/#installation) (If you want to run Control Plane locally)
1010
- [helm 3.8](https://helm.sh/docs/intro/install/) (If you want to run Control Plane locally)
Lines changed: 205 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,205 @@
1+
---
2+
date: 2026-04-10
3+
title: "Building the Kubernetes Multi-Cluster Plugin for PipeCD — LFX Mentorship"
4+
linkTitle: "Building the Kubernetes Multi-Cluster Plugin for PipeCD"
5+
weight: 970
6+
author: Mohammed Firdous ([@mohammedfirdouss](https://github.com/mohammedfirdouss))
7+
categories: ["Contribution"]
8+
tags: ["Kubernetes", "Plugin", "LFX Mentorship"]
9+
---
10+
11+
If you had told me last year that I would be working with Kubernetes and all things clusters, deployments and service meshes, I would have brushed it off. I am truly grateful for the journey thus far.
12+
13+
Earlier last month, I got accepted as an LFX Mentee for Term 1 of this calendar year. For me it is such a big deal, given my background, and how much effort has been put in behind the scenes to get to this stage.
14+
15+
I'm currently a mentee in the LFX Mentorship program working on [PipeCD](https://pipecd.dev), an open-source GitOps continuous delivery platform. For the past four weeks, I've been building out the `kubernetes_multicluster` plugin specifically implementing the deployment pipeline stages that handle canary, primary and baseline deployments across multiple clusters.
16+
17+
---
18+
19+
## What is PipeCD and what is this plugin?
20+
21+
PipeCD is an open-source GitOps CD platform that manages deployments across different infrastructure targets like Kubernetes, ECS, Terraform, Lambda and more. Each target type has a plugin that knows how to deploy to it.
22+
23+
The `kubernetes_multicluster` plugin is for teams running the same application across multiple Kubernetes clusters say US, EU and Asia and needing all of them to stay in sync through a single pipeline. Rolling out a new version across clusters one at a time, manually, with no coordination, is error-prone and slow. The plugin lets you define one pipeline that runs across every cluster at the same time, with canary and baseline checks before anything hits production.
24+
25+
## Progressive Delivery and Why These Stages Exist
26+
27+
Before a new version reaches all users, it goes through stages. A canary sends a small slice of traffic to the new version first. A baseline runs the *current* version at the same scale so you have a fair comparison. Primary is the actual promotion. Clean stages remove the temporary resources when you're done.
28+
29+
This pattern is called progressive delivery, because you roll out gradually, check things look good, then commit. If something looks wrong at the canary stage, you stop there. Nothing has touched production yet.
30+
31+
The `kubernetes_multicluster` plugin runs all of this across every cluster at the same time. One pipeline, every cluster, same stages.
32+
33+
A full pipeline looks like this:
34+
35+
```yaml
36+
stages:
37+
- name: K8S_CANARY_ROLLOUT
38+
- name: K8S_BASELINE_ROLLOUT
39+
- name: K8S_TRAFFIC_ROUTING
40+
- name: K8S_PRIMARY_ROLLOUT
41+
- name: K8S_CANARY_CLEAN
42+
- name: K8S_BASELINE_CLEAN
43+
```
44+
45+
Each of these is a stage I built. The sections below go through what each one does.
46+
47+
## What I Built
48+
49+
### K8S_CANARY_ROLLOUT
50+
51+
The canary stage deploys the new version of your app as a small slice alongside the existing production deployment. If your app normally runs 3 pods, canary might spin up 1 pod (or 20%) of the new version enough to catch problems without affecting most users.
52+
53+
It loads manifests from Git, creates copies of all workloads with a `-canary` suffix, scales them down to the configured replica count, adds a `pipecd.dev/variant=canary` label, and applies them to every target cluster in parallel. The original deployment is never touched this stage only ever adds resources.
54+
55+
![Canary rollout stage log applying manifests to cluster-eu and cluster-us](https://dev-to-uploads.s3.amazonaws.com/uploads/articles/swcu1ppt38ltw87wwbol.png)
56+
57+
![Canary rollout success — deploy targets: cluster-eu + cluster-us](https://dev-to-uploads.s3.amazonaws.com/uploads/articles/0lfhrddbu6r3mt01tlrs.png)
58+
59+
---
60+
61+
### K8S_CANARY_CLEAN
62+
63+
Once the canary window is over, whether you promoted or rolled back, the canary pods are just sitting in every cluster doing nothing. `K8S_CANARY_CLEAN` removes them.
64+
65+
It finds all resources with the label `pipecd.dev/variant=canary` for the application and deletes them in order: Services first, then Deployments, then everything else. The order matters as you don't want to remove the Deployment while the Service is still sending traffic to it.
66+
67+
One thing worth noting: the query is scoped strictly to canary-labelled resources. Even if something goes wrong in the deletion logic, it cannot touch primary resources.
68+
69+
![K8S_CANARY_CLEAN stage log deleting simple-canary resources from both clusters](https://dev-to-uploads.s3.amazonaws.com/uploads/articles/6vzb1fikt47ax3z53bjy.png)
70+
71+
![K8S_CANARY_ROLLOUT → K8S_CANARY_CLEAN pipeline — both stages green on cluster-eu and cluster-us](https://dev-to-uploads.s3.amazonaws.com/uploads/articles/5980a6b46dtiqv9oetz7.png)
72+
73+
---
74+
75+
### K8S_PRIMARY_ROLLOUT
76+
77+
After the canary looks good, you promote the new version to primary, the workload actually serving all your users. This stage takes the manifests from Git, adds the `pipecd.dev/variant=primary` label, and applies them across all clusters in parallel.
78+
79+
It also has a `prune` option: after applying, it checks what's currently running in the cluster against what was just applied, and deletes anything that's no longer in Git. Useful when you remove a resource from your manifests and want the cluster to reflect that.
80+
81+
![K8S_PRIMARY_ROLLOUT success deploy targets: cluster-eu + cluster-us](https://dev-to-uploads.s3.amazonaws.com/uploads/articles/5i4pfqi10ef38i7ltn55.png)
82+
83+
![kubectl confirming simple 2/2 updated in both cluster-eu and cluster-us](https://dev-to-uploads.s3.amazonaws.com/uploads/articles/kbi6iplj1l8g5wbkwgql.png)
84+
85+
---
86+
87+
### K8S_BASELINE_ROLLOUT
88+
89+
This one took me a while to understand and it is the stage I find most interesting to explain as well.
90+
91+
When you're running a canary, the natural thing is to compare it against primary. The issue is that's not a fair comparison primary is handling far more traffic than canary, under different conditions.
92+
93+
Baseline gives you a fairer comparison. You take the *current* version (not the new one) and run it at the same scale as canary. Now your cluster has:
94+
95+
```plaintext
96+
simple 2/2 ← production, current version
97+
simple-canary 1/1 ← new version, being tested
98+
simple-baseline 1/1 ← current version at canary scale
99+
```
100+
101+
You compare canary vs baseline, same number of pods, same traffic conditions. If canary is worse, it's obvious.
102+
103+
The key difference from every other rollout stage is one line of code. Canary and primary load manifests from the new Git commit (`TargetDeploymentSource`). Baseline loads from what's currently running (`RunningDeploymentSource`):
104+
105+
```go
106+
// canary.go — new version
107+
manifests, err := p.loadManifests(ctx, ..., &input.Request.TargetDeploymentSource, ...)
108+
109+
// baseline.go — current version
110+
manifests, err := p.loadManifests(ctx, ..., &input.Request.RunningDeploymentSource, ...)
111+
```
112+
113+
![K8S_BASELINE_ROLLOUT stage log loading manifests from running deployment source](https://dev-to-uploads.s3.amazonaws.com/uploads/articles/a61aeapcwnqquh3v3vdh.png)
114+
115+
![K8S_BASELINE_ROLLOUT](https://dev-to-uploads.s3.amazonaws.com/uploads/articles/0r26i800uc46kfmauo5p.png)
116+
117+
![K8S_BASELINE_ROLLOUT](https://dev-to-uploads.s3.amazonaws.com/uploads/articles/y6lowk38efysvvy0vbmk.png)
118+
119+
![kubectl showing simple, simple-baseline, simple-canary all running in both clusters](https://dev-to-uploads.s3.amazonaws.com/uploads/articles/8n8ailmbo0gae31sf58l.png)
120+
121+
---
122+
123+
### K8S_BASELINE_CLEAN
124+
125+
Once the analysis is done, baseline resources get cleaned up the same way as canary find everything labelled `pipecd.dev/variant=baseline` and delete it in order. No configuration needed. It doesn't matter whether `createService: true` was set during rollout, it finds whatever is there and removes it.
126+
127+
![K8S_BASELINE_CLEAN](https://dev-to-uploads.s3.amazonaws.com/uploads/articles/ddfhlxgi1i40r0x8dgh2.png)
128+
129+
![K8S_BASELINE_CLEAN stage log deleting baseline resources from both clusters](https://dev-to-uploads.s3.amazonaws.com/uploads/articles/fsxadrvt48rbr4hi113m.png)
130+
131+
![K8S_BASELINE_CLEAN stage log deleting baseline resources from both clusters](https://dev-to-uploads.s3.amazonaws.com/uploads/articles/apa1tdqa751bgl2t9g6c.png)
132+
133+
![K8S_BASELINE_CLEAN](https://dev-to-uploads.s3.amazonaws.com/uploads/articles/rder7h0ylhxn0wkaykep.png)
134+
135+
![kubectl confirming no baseline resources remain in cluster-eu or cluster-us](https://dev-to-uploads.s3.amazonaws.com/uploads/articles/8epcyha332ir2xz3jhua.png)
136+
137+
---
138+
139+
### K8S_TRAFFIC_ROUTING
140+
141+
Canary and baseline pods exist in the cluster but get no traffic until this stage runs. Without it, you're analysing pods that nobody is actually hitting. This stage is what sends real user traffic to them.
142+
143+
Two methods are supported:
144+
145+
**PodSelector** (no service mesh needed): changes the Kubernetes Service selector to point at one variant. All-or-nothing 100% to canary or 100% back to primary.
146+
147+
![PodSelector traffic routing full pipeline success across cluster-eu and cluster-us](https://dev-to-uploads.s3.amazonaws.com/uploads/articles/7uqb18dwo6jyhr0aekpw.png)
148+
149+
![PodSelector](https://dev-to-uploads.s3.amazonaws.com/uploads/articles/qddkmg9qs6unwk0bklu3.png)
150+
151+
![PodSelector](https://dev-to-uploads.s3.amazonaws.com/uploads/articles/piwbmpywrxmf4ykbcomp.png)
152+
153+
![PodSelector](https://dev-to-uploads.s3.amazonaws.com/uploads/articles/2u1kqy9jasqhqwtwebbd.png)
154+
155+
**Istio**: updates VirtualService route weights to split traffic across all three variants at once for example, primary 80%, canary 10%, baseline 10%. Also supports `editableRoutes` to limit which named routes the stage is allowed to modify.
156+
157+
One small thing I added on top of the traffic routing stage: per-route logging. When the stage runs, it now logs each route it processes whether it was skipped (because it's not in `editableRoutes`) or updated with new weights. Before this, the log just said "Successfully updated traffic routing" with no detail. Now you can see exactly which routes changed and to what percentages, which is useful when debugging a misconfigured VirtualService.
158+
159+
![Istio traffic routing stage log per-route logging showing which routes were updated in both clusters](https://dev-to-uploads.s3.amazonaws.com/uploads/articles/498uvcytrlppjpxqcq05.png)
160+
161+
![Istio](https://dev-to-uploads.s3.amazonaws.com/uploads/articles/6ql3nmfb19psi5gfu8a2.png)
162+
163+
![Istio](https://dev-to-uploads.s3.amazonaws.com/uploads/articles/oxy1gv5cojotd6ks253m.png)
164+
165+
![Istio](https://dev-to-uploads.s3.amazonaws.com/uploads/articles/tva9uwcmd7prb6qf72dm.png)
166+
167+
![Full Istio pipeline, all 7 stages green on cluster-eu and cluster-us](https://dev-to-uploads.s3.amazonaws.com/uploads/articles/6oj4elm5asfgergkp2se.png)
168+
169+
---
170+
171+
## Something I Found Interesting
172+
173+
The thing that surprised me was how `errgroup` handles running across multiple clusters without much extra code.
174+
175+
Every stage needs to run against N clusters, not one. A simple for-loop would run them one at a time slow, and if cluster 2 fails you don't find out until cluster 1 is already done.
176+
177+
`errgroup` runs all clusters at the same time and returns the first error:
178+
179+
```go
180+
eg, ctx := errgroup.WithContext(ctx)
181+
for _, tc := range targetClusters {
182+
tc := tc
183+
eg.Go(func() error {
184+
return canaryRollout(ctx, tc.deployTarget, ...)
185+
})
186+
}
187+
return eg.Wait()
188+
```
189+
190+
All clusters run in parallel. If any one fails, the stage fails immediately. The same pattern is used across every stage, so adding a new stage is mostly just writing the per-cluster logic the concurrency part is already solved.
191+
192+
## What's Next
193+
194+
The next piece is `DetermineStrategy`, that is the logic that decides what kind of deployment to trigger based on what changed in Git. After that, livestate drift detection so PipeCD can flag when a cluster has drifted from what Git says it should be.
195+
196+
To get involved, check out the PipeCD project and come join us on Slack.
197+
198+
## Links
199+
200+
- [PipeCD repository](https://github.com/pipe-cd/pipecd)
201+
- [LFX Mentorship Program](https://mentorship.lfx.linuxfoundation.org)
202+
- [Issue #6446, kubernetes_multicluster plugin](https://github.com/pipe-cd/pipecd/issues/6446)
203+
- [PR #6629 K8S_TRAFFIC_ROUTING](https://github.com/pipe-cd/pipecd/pull/6629)
204+
- [PR #6648 Per-route logging in K8S_TRAFFIC_ROUTING](https://github.com/pipe-cd/pipecd/pull/6648)
205+
- [Slack #PipeCD](https://app.slack.com/client/T08PSQ7BQ/C01B27F9T0X)

0 commit comments

Comments
 (0)