Jenkins and plugins versions report
Environment
Jenkins: 2.541.3
OS: Linux - 6.8.0-63-generic
Java: 21.0.10 - Ubuntu (OpenJDK 64-Bit Server VM)
---
PrioritySorter:936.v2c01c6b_84449
active-directory:2.41
ansible:635.v34b_48e979c86
ansicolor:536.v13fa_b_860c267
ant:520.vd082ecfb_16a_9
antisamy-markup-formatter:173.v680e3a_b_69ff3
apache-httpcomponents-client-4-api:4.5.14-269.vfa_2321039a_83
apache-httpcomponents-client-5-api:5.6-193.vf028a_770a_fa_c
asm-api:9.9.1-189.vb_5ef2964da_91
audit-trail:436.vc0d1e79fc5a_3
authorize-project:2.0.0
bootstrap5-api:5.3.8-895.v4d0d8e47fea_d
bouncycastle-api:2.30.1.84-291.v9f17b_21896e2
branch-api:2.1280.v0d4e5b_b_460ef
build-blocker-plugin:175.vc57a_d7dff5b_4
build-timestamp:1.1.1
built-on-column:1.5
caffeine-api:3.2.3-194.v31a_b_f7a_b_5a_81
checks-api:402.vca_263b_f200e3
cloudbees-folder:6.1100.ve9eed61d16c4
command-launcher:134.v025a_5fcf9dea_
commons-collections4-api:4.5.0-8.va_d5448ef9011
commons-compress-api:1.28.0-3
commons-lang3-api:3.20.0-109.ve43756e2d2b_4
commons-text-api:1.15.0-218.va_61573470393
cppcheck:1.26
credentials:1502.v5c95e620ddfe
credentials-binding:719.v80e905ef14eb_
data-tables-api:2.3.7-1534.v539d4edf109d
display-url-api:2.217.va_6b_de84cc74b_
durable-task:664.v2b_e7a_dfff66c
echarts-api:6.0.0-1247.vf3e35a_c1813f
eddsa-api:0.3.0.1-29.v67e9a_1c969b_b_
email-ext:1933.v45cec755423f
envinject:2.934.vc674e76cf954
envinject-api:1.237.v82803a_511906
extended-choice-parameter:388.ve7b_d0b_920e10
external-monitor-job:223.vb_fddcf42c9b_3
folder-properties:62.v1636b_4a_84608
font-awesome-api:7.2.0-965.ve3840b_696418
git:5.10.1
git-client:6.6.0
git-server:137.ve0060b_432302
github:1.46.0
github-api:1.330-492.v3941a_032db_2a_
github-branch-source:1967.vdea_d580c1a_b_a_
gradle:2.19.1244.v1f9866817fec
gson-api:2.13.2-198.v45e4a_55d9b_a_a_
instance-identity:203.v15e81a_1b_7a_38
ionicons-api:94.vcc3065403257
jackson-annotations2-api:2.21-7.v4777a_f3a_a_d47
jackson2-api:2.21.2-436.v29efdb_7418ff
jackson3-api:3.1.2-73.v3e5485d8b_148
jakarta-activation-api:2.1.4-1
jakarta-mail-api:2.1.5-1
jakarta-xml-bind-api:4.0.6-12.vb_1833c1231d3
javadoc:354.vee1a_660b_4990
javax-activation-api:1.2.0-8
javax-mail-api:1.6.2-11
jaxb:2.3.9-143.v5979df3304e6
jdk-tool:83.v417146707a_3d
jersey2-api:2.48-180.ve47b_264f849b_
jira:3.21
jjwt-api:0.13.0-141.vd58b_a_9592b_6c
jnr-posix-api:3.1.22-204.v925e2b_09a_42c
joda-time-api:2.14.1-187.vdf2def02b_8a_1
jquery:1.12.4-3
jquery3-api:3.7.1-619.vdb_10e002501a_
jsch:0.2.16-95.v3eecb_55fa_b_78
json-api:20251224-185.v0cc18490c62c
json-path-api:3.0.0-218.vcd4dd1355de2
jsoup:1.22.2-95.vc5d00f1eb_42d
junit:1403.vd9d1413fd205
ldap:807.v7d7de30930cf
locale:614.va_6a_5a_1a_f2b_38
lockable-resources:1509.va_6b_5b_5cb_0b_40
mailer:534.v1b_36f5864073
mapdb-api:1.0.9-44.va_1e1310c9118
matrix-auth:3.2.9
matrix-project:870.v9db_fcfc2f45b_
metrics:4.2.37-494.v06f9a_939d33a_
mina-sshd-api-common:2.16.0-167.va_269f38cc024
mina-sshd-api-core:2.16.0-167.va_269f38cc024
oic-auth:4.668.v653c6b_c6cb_f5
oidc-provider:212.v7657c4d7b_29f
okhttp-api:5.3.2-200.vedb_720a_cf1f8
opentelemetry:3.1589.ve81b_b_fa_927d5
opentelemetry-api:1.54.1.102.v276c93ee966d
oss-symbols-api:442.v99039087229b_
p4:1.17.2
pam-auth:1.12
parameterized-trigger:890.vc240a_a_e1217f
pipeline-build-step:584.vdb_a_2cc3a_d07a_
pipeline-graph-analysis:254.v0f63a_a_447dca_
pipeline-graph-view:847.vc7150b_d79f11
pipeline-groovy-lib:797.v90ea_a_9b_e45a_0
pipeline-input-step:551.vdff487c5998c
pipeline-milestone-step:152.v6e22b_8cfc66c
pipeline-model-api:2.2277.v00573e73ddf1
pipeline-model-definition:2.2277.v00573e73ddf1
pipeline-model-extensions:2.2277.v00573e73ddf1
pipeline-stage-step:345.va_96187909426
pipeline-stage-tags-metadata:2.2277.v00573e73ddf1
pipeline-utility-steps:2.20.0
plain-credentials:199.v9f8e1f741799
plugin-util-api:6.1192.v30fe6e2837ff
prism-api:1.30.0-717.vb_f8360844b_53
reverse-proxy-auth-plugin:245.v93f25b_b_b_3102
scm-api:728.vc30dcf7a_0df5
scm-filter-branch-pr:264.v39ce4f34d572
script-security:1399.ve6a_66547f6e1
slack:795.v4b_9705b_e6d47
snakeyaml-api:2.5-149.v72471e9c6371
snakeyaml-engine-api:3.0.1-5.vd98ea_ff3b_92e
ssh-agent:396.vcc7d84e622ec
ssh-credentials:372.va_250881b_08cd
sshd:3.384.vc89b_5e138cf9
structs:362.va_b_695ef4fdf9
subversion:1303.vcfd9679fb_c12
swarm:1254.vf26a_5e188f26
throttle-concurrents:625.vc8b_e469e9a_b_c
timestamper:1.30
token-macro:477.vd4f0dc3cb_cf1
trilead-api:2.284.v1974ea_324382
variant:70.va_d9f17f859e0
woodstox-core-api:7.1.1-1.v4d297985f397
workflow-aggregator:608.v67378e9d3db_1
workflow-api:1413.v2ff1a_5e720fa_
workflow-basic-steps:1098.v808b_fd7f8cf4
workflow-cps:4285.v8df38f05c3c5
workflow-durable-task-step:1475.ved562f6ec8b_3
workflow-job:1571.vb_423c255d6d9
workflow-multibranch:821.vc3b_4ea_780798
workflow-scm-step:466.va_d69e602552b_
workflow-step-api:724.v538c2362b_dfb_
workflow-support:1015.v785e5a_b_b_8b_22
xcode-plugin:2.0.17-565.v1c48051d46ef
What Operating System are you using (both controller, and any agents involved in the problem)?
Linux (Ubuntu) controller and Mac (Tahoe) agents. OS shouldn't matter here.
Reproduction steps
- Jenkins controller with the OpenTelemetry plugin installed.
Confirmed on version 3.1589.ve81b_b_fa_927d5 (latest main was also
verified to have the same unbounded .get() at preOnline:74).
- In Manage Jenkins → System → OpenTelemetry → Configuration Properties,
ensure the line otel.instrumentation.jenkins.agent.enabled=true
(default).
- Connect an inbound agent (JNLP4-connect, Jenkins Swarm Plugin client
in our case) from a network location whose RTT to the controller is
~100 ms or higher. In our environment, the controller is in Helsinki
and agents are in AWS us-west-2 and ap-southeast-1 (~150-250 ms RTT).
Agents in eu-central-1 (low RTT) are unaffected.
- Watch the controller log.
Expected Results
Agent connects and stays online. If agent-side OpenTelemetry configuration
cannot be completed in a reasonable time, the controller should either
(a) succeed with a best-effort partial configuration, or (b) log a warning
and proceed to mark the computer online. It should not terminate the
entire remoting channel.
Actual Results
Agent reaches "Accepted JNLP4-connect connection" on the controller, then 60-90 s later gets terminated with:
java.nio.channels.ClosedChannelException
Caused: hudson.remoting.RequestAbortedException
at hudson.remoting.Request.abort(Request.java:358)
at hudson.remoting.Channel.terminate(Channel.java:1189)
at org.jenkinsci.remoting.protocol.impl.ChannelApplicationLayer.onReadClosed(...)
...
at org.jenkinsci.remoting.protocol.impl.NIONetworkLayer.ready(...)
at org.jenkinsci.remoting.protocol.IOHub$OnReady.run(IOHub.java:808)
Caused: java.util.concurrent.ExecutionException
at hudson.remoting.Request$1.get(Request.java:307)
at hudson.remoting.FutureAdapter.get(FutureAdapter.java:60)
at io.jenkins.plugins.opentelemetry.jenkins.OpenTelemetryConfigurerComputerListener.preOnline(OpenTelemetryConfigurerComputerListener.java:74)
at hudson.slaves.SlaveComputer.setChannel(SlaveComputer.java:724)
at jenkins.slaves.DefaultJnlpSlaveReceiver.afterChannel(DefaultJnlpSlaveReceiver.java:176)
at org.jenkinsci.remoting.engine.JnlpConnectionState.fire(...)
Removing Swarm Node for computer [mac-]
Swarm client retries, connects again, cycle repeats every ~60-90 s.
Agent never becomes usable.
Anything else?
Root cause analysis:
OpenTelemetryConfigurerComputerListener.preOnline() issues a synchronous
remote call via channel.call(...) and does an UNBOUNDED .get() on the
returned Future:
Object result = configureOpenTelemetrySdkOnComputer(
computer, channel, otelSdkProperties,
otelSdkResourceProperties).get(); // no timeout
The sibling method afterConfiguration() in the same class correctly uses
a bounded wait:
... .get(10, TimeUnit.SECONDS);
Because preOnline blocks indefinitely, any subsequent event that closes
the remoting channel (NAT idle timeout, network blip, load balancer
cleanup, or any other transient disturbance) causes every pending
request.get() to abort — including this one. The aborted Future throws
ExecutionException wrapping RequestAbortedException wrapping
ClosedChannelException, which propagates up to SlaveComputer.setChannel
and tears down the computer via SwarmLauncher#afterDisconnect.
On low-RTT agents (sub-50 ms), the OTel SDK configuration RPC usually
completes before any transient disturbance fires, so the agent survives.
On higher-RTT agents, the race leans the other way and the agent is
always kicked within the channel's useful life.
Suggested fix (one-line):
...configureOpenTelemetrySdkOnComputer(...).get(N, TimeUnit.SECONDS);
...matching afterConfiguration's bounded wait, with a catch around
TimeoutException that logs a WARNING and returns without tearing down.
The timeout value could be a config option, defaulting to ~10 s.
Workaround in the meantime: set
otel.instrumentation.jenkins.agent.enabled=false
in the Configuration Properties, which short-circuits the listener.
This eliminates the kick but also disables all agent-side OTel
instrumentation, which is not an acceptable long-term trade.
Related files:
- src/main/java/io/jenkins/plugins/opentelemetry/jenkins/
OpenTelemetryConfigurerComputerListener.java (line 74)
- for comparison: afterConfiguration() in the same class
Are you interested in contributing a fix?
No response
Jenkins and plugins versions report
Environment
What Operating System are you using (both controller, and any agents involved in the problem)?
Linux (Ubuntu) controller and Mac (Tahoe) agents. OS shouldn't matter here.
Reproduction steps
Confirmed on version 3.1589.ve81b_b_fa_927d5 (latest main was also
verified to have the same unbounded .get() at preOnline:74).
ensure the line
otel.instrumentation.jenkins.agent.enabled=true(default).
in our case) from a network location whose RTT to the controller is
~100 ms or higher. In our environment, the controller is in Helsinki
and agents are in AWS us-west-2 and ap-southeast-1 (~150-250 ms RTT).
Agents in eu-central-1 (low RTT) are unaffected.
Expected Results
Agent connects and stays online. If agent-side OpenTelemetry configuration
cannot be completed in a reasonable time, the controller should either
(a) succeed with a best-effort partial configuration, or (b) log a warning
and proceed to mark the computer online. It should not terminate the
entire remoting channel.
Actual Results
Agent reaches "Accepted JNLP4-connect connection" on the controller, then 60-90 s later gets terminated with:
Removing Swarm Node for computer [mac-]
Swarm client retries, connects again, cycle repeats every ~60-90 s.
Agent never becomes usable.
Anything else?
Root cause analysis:
OpenTelemetryConfigurerComputerListener.preOnline() issues a synchronous
remote call via channel.call(...) and does an UNBOUNDED .get() on the
returned Future:
The sibling method afterConfiguration() in the same class correctly uses
a bounded wait:
Because preOnline blocks indefinitely, any subsequent event that closes
the remoting channel (NAT idle timeout, network blip, load balancer
cleanup, or any other transient disturbance) causes every pending
request.get() to abort — including this one. The aborted Future throws
ExecutionException wrapping RequestAbortedException wrapping
ClosedChannelException, which propagates up to SlaveComputer.setChannel
and tears down the computer via SwarmLauncher#afterDisconnect.
On low-RTT agents (sub-50 ms), the OTel SDK configuration RPC usually
completes before any transient disturbance fires, so the agent survives.
On higher-RTT agents, the race leans the other way and the agent is
always kicked within the channel's useful life.
Suggested fix (one-line):
...matching afterConfiguration's bounded wait, with a catch around
TimeoutException that logs a WARNING and returns without tearing down.
The timeout value could be a config option, defaulting to ~10 s.
Workaround in the meantime: set
in the Configuration Properties, which short-circuits the listener.
This eliminates the kick but also disables all agent-side OTel
instrumentation, which is not an acceptable long-term trade.
Related files:
OpenTelemetryConfigurerComputerListener.java (line 74)
Are you interested in contributing a fix?
No response