Skip to content

Investigation: Terraform upgrade would not fix Azure state-backend DNS failures (no code changes) - #5178

Open
Jay W (JayDoubleu) with Copilot wants to merge 2 commits into
mainfrom
copilot/fix-dns-lookup-failure
Open

Jay W (JayDoubleu) with Copilot wants to merge 2 commits into
mainfrom
copilot/fix-dns-lookup-failure

Conversation

Copilot AI commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

What is being addressed

Core deployment intermittently fails during terraform apply. The error is dial tcp: lookup <account>.blob.core.windows.net … no such host, raised while the azurerm backend reads or writes state. The question was whether bumping the pinned Terraform version or the provider would fix it.

This PR contains no code changes. The evidence shows that upgrading would not affect the failure.

How is this addressed

  • Provider version doesn't matter. backend "azurerm" {} (core/terraform/main.tf) is implemented inside Terraform itself. The hashicorp/azurerm provider version has no effect on state requests.

  • Newer Terraform releases behave the same. Terraform v1.14.3 (pinned), v1.15.9, v1.16.5 (latest stable) and v1.17.0-rc1 all pin go-azure-sdk/sdk v0.20250131.1134653 and giovanni v0.28.0. The internal/backend/remote-state/azure/* files are identical across these versions apart from the copyright header. None adds its own retry.

  • The SDK skips retries for these errors on purpose. In sdk/client/client.go, the retry function does not retry errors matching dial tcp on hosts other than management.azure.com. Blob storage requests fall into that group. The same logic is on go-azure-sdk main, so a future Terraform release that picks up a newer SDK would still not retry:

    if !isResourceManagerHost(req) {
        return extendedRetryPolicy(r, err) // `dial tcp` → return false (no retry)
    }
  • Suggested follow-up work (not in this PR):

    • Remove the likely cause: create the management-storage private endpoint (core/terraform/resource_processor/vmss_porter/main.tf) in an earlier step, rather than in the same apply that uses that storage for state.
    • Add limited recovery in the deployment script: on this specific DNS error, capture DNS answers from 168.63.129.16 and the blob lease state, then re-plan a limited number of times after a delay. No unconditional apply retry and no automatic lease break.
    • Optionally, report it upstream in hashicorp/terraform (Azure backend retry on DNS errors) or hashicorp/go-azure-sdk.
  • No documentation, CHANGELOG or template version updates are needed.

@github-actions github-actions Bot added the external PR from an external contributor label Oct 9, 2026
Copilot AI changed the title [WIP] Fix intermittent deployment failure due to DNS lookup error Investigation: Terraform upgrade would not fix Azure state-backend DNS failures (no code changes) Oct 9, 2026
)

Co-authored-by: JayDoubleu <40270505+JayDoubleu@users.noreply.github.com>
@JayDoubleu
Jay W (JayDoubleu) marked this pull request as ready for review October 9, 2026 20:26
@JayDoubleu
Jay W (JayDoubleu) requested a review from a team as a code owner October 9, 2026 20:27
Copilot AI balanced review requested due to automatic review settings October 9, 2026 20:27

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

The DNS gate can accept stale answers and is skipped entirely on recovery reruns.

3 open findings
What changed in this PR

Introduces a core deployment mitigation for management-storage private-endpoint DNS failures.

Changes:

  • Adds a targeted private-endpoint apply and DNS wait.
  • Adds mocked deployment-script tests and CI routing.
  • Bumps the core version and changelog.
File Description
core/​terraform/​mgmt_storage_private_endpoint.sh Adds endpoint creation and DNS readiness logic.
core/​terraform/​deploy.sh Runs the new step before the main apply.
devops/​tests/​test_mgmt_storage_private_endpoint.py Tests deployment-script paths.
.github/​workflows/​build_validation_develop.yml Routes script changes through validation.
core/​version.txt Bumps core to 0.18.12.
CHANGELOG.md Records the mitigation.

🧠 Review effort: Balanced


echo "Waiting for ${host} to resolve ${DNS_REQUIRED_SUCCESSES} consecutive times"
for ((attempt=1; attempt<=DNS_MAX_ATTEMPTS; attempt++)); do
if answer=$(getent ahosts "${host}" 2>&1) && [[ -n "${answer}" ]]; then
Comment on lines +70 to +77
state_addresses=$(terraform state list)
if grep -qxF "${PE_ADDRESS}" <<< "${state_addresses}"; then
echo "Management storage private endpoint already in state; skipping targeted apply"
exit 0
fi

host=$(state_blob_host)
echo "$(timestamp) Creating management storage private endpoint before the main core apply"
Comment thread core/terraform/deploy.sh
-k "${TRE_ID}" \
-l "${LOG_FILE}" \
-c "terraform plan --parallelism=25 -out ${PLAN_FILE} && \
-c "./mgmt_storage_private_endpoint.sh && \

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

external PR from an external contributor

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Core deployment intermittently fails on Terraform state-storage DNS lookup during apply

3 participants