fix: disable aks node controller and only start at boot and wait for network online target#8911
fix: disable aks node controller and only start at boot and wait for network online target#8911awesomenix wants to merge 1 commit into
Conversation
There was a problem hiding this comment.
Pull request overview
This PR aims to reduce node boot/provisioning latency by removing the aks-node-controller.service dependency on network-online.target (so the unit can be started earlier/in parallel) and instead waiting for network readiness within the provisioning command script (cse_cmd.sh) and (selectively) in the ANC wrapper when hotfix download paths are used.
Changes:
- Make
aks-node-controller.servicea static unit (no[Install]section) and removeAfter=/Wants=fornetwork-online.target. - Add
network-online.targetwait logic tocse_cmd.sh(and the corresponding ANC template), plus conditional waits inaks-node-controller-wrapper.shbefore hotfix network operations. - Update VHD build and validation scripts to keep the ANC unit disabled/static and validate the new state.
Reviewed changes
Copilot reviewed 6 out of 6 changed files in this pull request and generated 4 comments.
Show a summary per file
| File | Description |
|---|---|
vhdbuilder/packer/test/linux-vhd-content-test.sh |
Updates VHD validation to require ANC unit to be static. |
vhdbuilder/packer/pre-install-dependencies.sh |
Disables ANC service during VHD build to keep it from running before provisioning artifacts exist. |
parts/linux/cloud-init/artifacts/cse_cmd.sh |
Adds an explicit wait for network-online.target before running provisioning. |
parts/linux/cloud-init/artifacts/aks-node-controller.service |
Removes network-online.target ordering and makes the unit static. |
parts/linux/cloud-init/artifacts/aks-node-controller-wrapper.sh |
Adds conditional network-online waits for hotfix-related operations. |
aks-node-controller/parser/templates/cse_cmd.sh.gtpl |
Keeps the generated cse_cmd.sh template in sync with the new network-online wait behavior. |
| RemainAfterExit=yes | ||
|
|
||
| [Install] | ||
| WantedBy=basic.target |
There was a problem hiding this comment.
you probably were aware of this too
🔴 High Risk — 🏗️ Architecture / Backward Compatibility
parts/linux/cloud-init/artifacts/aks-node-controller.service removes [Install] and WantedBy, and vhdbuilder/packer/pre-install-dependencies.sh now disables the unit. That drops the long-standing boot fallback path (service auto-start) and makes provisioning depend entirely on boothook start timing.
This aligns with the failing PR E2E signal: multiple scenarios are stuck in provision-wait loop ([ -f /opt/azure/containers/provision.complete ]) and never reach expected CSE output (ADO build 172002687, Run AgentBaker E2E).
Mitigation: restore a fallback start path (or equivalent robust trigger) so ANC still starts when boothook path misses/races.
There was a problem hiding this comment.
The problem is that we drop in the cse-cmd.sh or the aks node config only as part of boothook.
So we want to wait for it to be written before we can startup aks node controller otherwise we will have an issue where the service starts up before that and misses the file completely.
In future when we exclusively move to aks node controller we can add an explicity wait for aks node config before service starts that way everything is synchronized.
ff8ad6a to
0a50baa
Compare
AgentBaker Linux gate detectiveRun: https://msazure.visualstudio.com/CloudNativeCompute/_build/results?buildId=172097539 Detective summary: E2E failed with the recurring broad gotestsum multi-leaf pattern: many ACL/ARM64/FIPS/RCV1P leaves failed, and the task/test-results output is not enough to isolate a single root assertion for the selected failed pair. Likely cause / signature: $aggSig; this pushed the signature above the recurring threshold, so repair item #38796920 was created. Confidence: Medium for grouping; PR author should inspect only targeted subgroups if scenario artifacts show a consistent PR-related assertion. Recommended owner/action: Node Lifecycle E2E owner should inspect scenario artifacts/test-log JSON and split to narrower signatures when concrete assertions recur. Strongest alternative: PR-specific CSE/RCV1P regression; less likely as the top-level gate signature because this run matches the broad recurring aggregate, but targeted artifact review is warranted. Evidence: ADO timeline marks E2E task failed; test run 528965996 contains many failed IDs; gotestsum summary starts with broad ACL failures and no single dominant assertion in the available output. Wiki signature: $aggSig |
| wait_for_network_online() { | ||
| echo "Waiting for network-online.target at $(date -Ins)" | ||
|
|
||
| if timeout 30 sh -c 'until systemctl is-active --quiet network-online.target; do sleep 0.1; done'; then | ||
| echo "network-online.target reached at $(date -Ins)" | ||
| else | ||
| echo "Timed out waiting for network-online.target at $(date -Ins)" >&2 | ||
| return 1 | ||
| fi | ||
| } |
| wait_for_network_online() { | ||
| echo "Waiting for network-online.target at $(date -Ins)" | ||
|
|
||
| if timeout 30 sh -c 'until systemctl is-active --quiet network-online.target; do sleep 0.1; done'; then | ||
| echo "network-online.target reached at $(date -Ins)" | ||
| else | ||
| echo "Timed out waiting for network-online.target at $(date -Ins)" >&2 | ||
| return 1 | ||
| fi | ||
| } |
| wait_for_network_online() { | ||
| echo "Waiting for network-online.target at $(date -Ins)" | ||
|
|
||
| if timeout 30 sh -c 'until systemctl is-active --quiet network-online.target; do sleep 0.1; done'; then | ||
| echo "network-online.target reached at $(date -Ins)" | ||
| else | ||
| echo "Timed out waiting for network-online.target at $(date -Ins)" >&2 | ||
| return 1 | ||
| fi | ||
| } | ||
|
|
||
| wait_for_network_online || exit 124 | ||
|
|
AgentBaker Linux gate detectiveRun: https://msazure.visualstudio.com/CloudNativeCompute/_build/results?buildId=172464976 Detective summary: The E2E task ended with 17 failed leaves. The visible leaf is Test_ACL_Scriptless/default timing out around scenario setup, with RCV1P/ACL slow-readiness warnings in the same run; this does not isolate to the CSE boot/network-online change. Likely cause/signature: �2e-gotestsum-multi-leaf-no-error-body — broad E2E session/readiness degradation with insufficient leaf body, rather than a targeted product regression. Confidence: Medium; the evidence shows aggregate multi-leaf readiness failures, but not a single deterministic leaf root cause. Recommended owner/action: E2E/platform owners should inspect the run's scenario logs and RCV1P/shared-cluster readiness; if recurring, improve leaf error surfacing and route through the existing repair item. Strongest alternative: PR #8911's CSE boot sequencing change could affect provisioning timing. Less likely because the failure shape is broad multi-leaf readiness/context-deadline noise, not a stable CSE boot assertion. Evidence: Timeline log 562, ADO test run 532204889, visible ACL_Scriptless context deadline, RCV1P slow readiness warnings. |
Devinwong
left a comment
There was a problem hiding this comment.
LGTM. Just the PR title seems conflicting with description.
26a5835 to
d828d22
Compare
d828d22 to
fe66bac
Compare
fe66bac to
e06a1e7
Compare
| wait_for_network_online() { | ||
| echo "Waiting for network-online.target at $(date -Ins)" | ||
|
|
||
| if timeout 30 sh -c 'until systemctl is-active --quiet network-online.target; do sleep 0.1; done'; then | ||
| echo "network-online.target reached at $(date -Ins)" | ||
| else | ||
| echo "Timed out waiting for network-online.target at $(date -Ins)" >&2 | ||
| return 1 | ||
| fi | ||
| } | ||
|
|
||
| wait_for_network_online || exit 124 | ||
|
|
| wait_for_network_online() { | ||
| echo "Waiting for network-online.target at $(date -Ins)" | ||
|
|
||
| if timeout 30 sh -c 'until systemctl is-active --quiet network-online.target; do sleep 0.1; done'; then | ||
| echo "network-online.target reached at $(date -Ins)" | ||
| else | ||
| echo "Timed out waiting for network-online.target at $(date -Ins)" >&2 | ||
| return 1 | ||
| fi | ||
| } | ||
|
|
||
| wait_for_network_online || exit 124 | ||
|
|
| if [ "${ENABLE_PROVISIONING_HOTFIX:-}" = "true" ]; then | ||
| wait_for_network_online || return 124 | ||
| log "ENABLE_PROVISIONING_HOTFIX=true; running check-hotfix to refresh hotfix pointer" | ||
| if "$BIN_PATH" check-hotfix; then | ||
| log "ANC check-hotfix completed; hotfix pointer refresh attempted" |
| if [ -f "$HOTFIX_JSON" ]; then | ||
| if jq -e 'has("version")' "$HOTFIX_JSON" >/dev/null 2>&1; then | ||
| wait_for_network_online || return 124 | ||
| fi | ||
| log "Found ANC hotfix config at ${HOTFIX_JSON}; running download-hotfix" |
| flatcarTemplate = `{ | ||
| "ignition": { "version": "3.4.0" }, | ||
| "systemd": { | ||
| "units": [{ | ||
| "name": "aks-node-controller.service", | ||
| "enabled": true | ||
| }] | ||
| }, | ||
| "storage": { | ||
| "files": [%s] | ||
| "files": [%s], | ||
| "links": [{ | ||
| "path": "/etc/systemd/system/basic.target.wants/aks-node-controller.service", | ||
| "target": "/etc/systemd/system/aks-node-controller.service", | ||
| "overwrite": true | ||
| }] | ||
| } |
e06a1e7 to
5314760
Compare
| wait_for_network_online() { | ||
| echo "Waiting for network-online.target at $(date -Ins)" | ||
|
|
||
| if timeout 30 sh -c 'until systemctl is-active --quiet network-online.target; do sleep 0.1; done'; then | ||
| echo "network-online.target reached at $(date -Ins)" | ||
| else | ||
| echo "Timed out waiting for network-online.target at $(date -Ins)" >&2 | ||
| return 1 | ||
| fi | ||
| } |
| if [ "${ENABLE_PROVISIONING_HOTFIX:-}" = "true" ]; then | ||
| wait_for_network_online || return 124 | ||
| log "ENABLE_PROVISIONING_HOTFIX=true; running check-hotfix to refresh hotfix pointer" | ||
| if "$BIN_PATH" check-hotfix; then | ||
| log "ANC check-hotfix completed; hotfix pointer refresh attempted" | ||
| else | ||
| log "ANC check-hotfix failed; continuing (fail-open)" | ||
| fi | ||
| fi |
| if jq -e 'has("version")' "$HOTFIX_JSON" >/dev/null 2>&1; then | ||
| wait_for_network_online || return 124 | ||
| fi |
| wait_for_network_online() { | ||
| echo "Waiting for network-online.target at $(date -Ins)" | ||
|
|
||
| if timeout 30 sh -c 'until systemctl is-active --quiet network-online.target; do sleep 0.1; done'; then | ||
| echo "network-online.target reached at $(date -Ins)" | ||
| else | ||
| echo "Timed out waiting for network-online.target at $(date -Ins)" >&2 | ||
| return 1 | ||
| fi | ||
| } |
| wait_for_network_online() { | ||
| echo "Waiting for network-online.target at $(date -Ins)" | ||
|
|
||
| if timeout 30 sh -c 'until systemctl is-active --quiet network-online.target; do sleep 0.1; done'; then | ||
| echo "network-online.target reached at $(date -Ins)" | ||
| else | ||
| echo "Timed out waiting for network-online.target at $(date -Ins)" >&2 | ||
| return 1 | ||
| fi | ||
| } |
5314760 to
d62c318
Compare
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 11 out of 11 changed files in this pull request and generated 5 comments.
Comments suppressed due to low confidence (1)
parts/linux/cloud-init/artifacts/aks-node-controller.service:8
- With the service being enabled at boot (e.g., via Flatcar Ignition),
RemainAfterExit=yescan prevent the boothook’s latersystemctl startfrom re-running the oneshot after provisioning files are written, because the unit may already be in an active (exited) state from an earlier no-op run. Removing RemainAfterExit avoids this “started too early once” latch.
Type=oneshot
ExecStart=/opt/azure/containers/aks-node-controller-wrapper.sh
RemainAfterExit=yes
| "systemd": { | ||
| "units": [{ | ||
| "name": "aks-node-controller.service", | ||
| "enabled": true | ||
| }] | ||
| }, | ||
| "storage": { | ||
| "files": [%s] | ||
| "files": [%s], | ||
| "links": [{ | ||
| "path": "/etc/systemd/system/basic.target.wants/aks-node-controller.service", | ||
| "target": "/etc/systemd/system/aks-node-controller.service", | ||
| "overwrite": true | ||
| }] | ||
| } |
| echo "Waiting for network-online.target at $(date -Ins)" | ||
|
|
||
| if timeout 30 sh -c 'until systemctl is-active --quiet network-online.target; do sleep 0.1; done'; then |
| echo "Waiting for network-online.target at $(date -Ins)" | ||
|
|
||
| if timeout 30 sh -c 'until systemctl is-active --quiet network-online.target; do sleep 0.1; done'; then |
| echo "Waiting for network-online.target at $(date -Ins)" | ||
|
|
||
| if timeout 30 sh -c 'until systemctl is-active --quiet network-online.target; do sleep 0.1; done'; then |
| # The cloud-init boothook starts ANC after writing its provisioning files. | ||
| # Keep the unit disabled in the VHD image so it cannot run before those files are ready. | ||
| systemctl disable aks-node-controller.service || exit 1 |
awesomenix
left a comment
There was a problem hiding this comment.
will hold off on this change, feels a bit high risk
|
this PR is actually useless because netowrk is always online before boothook |
dont wait for network-online.target start in parallel, this saves about 4s