Skip to content

[Enhancement]: Allow disabling the Kafka Connect API ClusterIP service #13136

Description

@laughingman7743

Related problem

Every KafkaConnect cluster gets two Services: the already-headless <name>-connect, and <name>-connect-api, which allocates a ClusterIP.

When many KafkaConnect clusters run on a single Kubernetes cluster, those ClusterIPs accumulate. A cluster's Service CIDR is fixed when the cluster is created and generally cannot be resized afterwards, so a busy shared cluster can approach exhaustion. Once the range is full, creation of every new Service and Ingress on that cluster fails with failed to allocate a serviceIP: range is full, which affects all workloads on the cluster, not just Kafka Connect.

There is currently no configuration a user can reach for in that situation: spec.template.apiService is an InternalServiceTemplate, which exposes only metadata, ipFamilyPolicy and ipFamilies.

Suggested solution

Let users opt into creating <name>-connect-api as a headless Service (spec.clusterIP: None) — for example a field on InternalServiceTemplate, or a boolean on the Connect spec.

The Cluster Operator's own reconciliation should be unaffected, because it addresses the REST API by DNS name rather than by ClusterIP: KafkaConnectResources.qualifiedServiceName() returns <name>-connect-api.<namespace>.svc, which KafkaConnectAssemblyOperator uses for connector reconciliation and plugin listing. For a headless Service with a selector, that name resolves to the Connect pod IPs.

One caveat worth noting: a headless name has no A record while no Connect pod is Ready, so clients would get NXDOMAIN rather than a connection error during rollouts. That argues for keeping this opt-in rather than changing the default.

Alternatives

  • Consolidating many small Connect clusters into fewer large ones — this conflicts with isolating connectors per team or per source database.
  • Managing connectors outside of KafkaConnector CRs so that the API Service is not needed — this gives up the CR-based workflow.

Additional context

Flink's Kubernetes integration may be a useful reference, since it solves the same problem and the implementation is small.

It exposes Headless_ClusterIP as one value of kubernetes.rest-service.exposed.type (KubernetesConfigOptions.ServiceExposedType). HeadlessClusterIPService extends ClusterIPService and overrides nothing but the two Service builders, setting spec.clusterIP: None; the Service name, DNS record and port stay the same. Crucially it does not override getRestEndpoint(), which returns the namespaced DNS name — that is what makes the headless variant transparent to callers, and it mirrors how the Cluster Operator already addresses the Connect API here.

The Flink Kubernetes Operator keeps ClusterIP as the default and only switches when the user sets the option (FlinkConfigBuilder.applyFlinkConfiguration()), i.e. it is strictly opt-in — which seems like the right shape here too.

Checked against main: InternalServiceTemplate currently has only the three fields listed above. Happy to open a PR if the approach sounds acceptable.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions