Flux is installed by the Flux Operator¶
Flux itself is one object: the FluxInstance in
flux/clusters/gewis-prod/flux-system/flux-instance.yaml. The
Flux Operator turns it into the controllers, the
flux-system GitRepository (GEWIS/infra, main) and the root flux-system
Kustomization (flux/clusters/gewis-prod), which applies the
layers.
File in flux/clusters/gewis-prod/flux-system/ |
Holds |
|---|---|
flux-instance.yaml |
Flux version (pinned), components, sync target |
helm-repo.yaml |
The operator's OCI chart repository |
helm-release.yaml |
The operator itself, pinned, plus its Web UI values |
Only source-controller, kustomize-controller and helm-controller run.
notification-controller is left out because nothing here uses an Alert,
Provider or Receiver; add it to components when something does.
Upgrades¶
Both are pinned and bumped by Renovate, so every upgrade is a commit:
- Flux:
distribution.versioninflux-instance.yaml, matched by a regex manager inrenovate.jsonagainstfluxcd/flux2releases. The operator could follow a range like2.8.xon its own, but then upgrades would leave no trace in Git. - The operator: an ordinary pinned HelmRelease, like every other chart.
Bootstrap: tofu starts it, Flux owns it¶
terraform/20_talos-bootstrap/flux.tf uses the upstream
controlplaneio-fluxcd/flux-operator-bootstrap module. It runs a one-shot Job
that installs the operator chart and applies the FluxInstance, both read from
the files above so bootstrap and steady state never disagree. Both are
create-if-missing: once Flux has adopted them the Job leaves them alone, so a
later tofu apply never fights Flux. Changing either file makes the next
20_talos-bootstrap apply re-run the Job, which then finds everything adopted and
does nothing.
The module depends on helm_release.cilium. The Job gets a single attempt
(backoffLimit: 0), and on a fresh cluster it would otherwise start while
Cilium is still rolling out, with no pod network, and fail the apply. Waiting
for the Cilium release guarantees the pod network, not that CoreDNS is ready or
that the first TLS connection to ghcr.io gets through; if the Job still fails,
debug_on_failure = true on the module relays its log into the tofu output.
The module needs the Helm provider 3.x, which is why 20_talos-bootstrap pins
~> 3.1.
Migrating the running cluster¶
The cluster was first set up with flux bootstrap. The operator adopts that
install in place, without downtime: it strips the kustomize.toolkit.fluxcd.io
ownership labels from the controllers and marks them
kustomize.toolkit.fluxcd.io/prune: Disabled, so removing the old gotk-*.yaml
from Git cannot garbage-collect running controllers. Order matters:
tofu applyinterraform/20_talos-bootstrap— installs the operator and theFluxInstance; wait forkubectl -n flux-system get fluxinstance fluxto reportReady.- Only then push the commit that removes
flux-system/gotk-*.yamlandflux-system.yaml. Pushed first, the tree would contain aFluxInstancewhose CRD does not exist yet, and the rootKustomizationwould fail its dry-run. flux trace kustomization flux-systemnow answers object not managed by Flux: the operator owns it, not a Kustomization. The three controllers carryapp.kubernetes.io/managed-by: flux-operator.notification-controlleris still running until the push, because the operator never adopted it: it keeps itskustomize.toolkit.fluxcd.io/name: flux-systemlabel and sits in that Kustomization's inventory, so the first reconcile after the push prunes it along with its CRDs.