Repository navigation
Add Helm chart to deploy Scope app onto the Bicep-provisioned AKS cluster - #1481
Open
Josh Duffney (duffney) wants to merge 10 commits into
Open
Josh Duffney (duffney) wants to merge 10 commits into
Josh Duffney (duffney) wants to merge 10 commits into
Conversation
Umbrella chart at deploy/helm/scope/ with Chart.yaml, values.yaml, shared helpers, workload-identity ServiceAccount, Secrets Store CSI SecretProviderClass, common-env ConfigMap, and the api service (deployment, service, ingress) as the reference implementation for remaining services. Pairs with the Azure infrastructure provisioned by deploy/azure/ (PR #1477): reads its outputs (acrLoginServer, keyVaultName, workloadIdentityClientId, workloadIdentityServiceAccountName, storageAccountName) as values, uses CSI-synced secrets for static connection strings/API keys, and workload identity for direct Blob/Queue/Key Vault SDK access. Remaining services (portal, judge, token-manager, scheduler, workers, jobs) are configured in values.yaml but not yet templated — tracked as follow-up work. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 094bb93a-ec65-4f80-8896-0dd041aedeff
… phase - Add per-image registry override to scope.image helper so core services (api, portal, judge, token-manager, scheduler, report-generator) default to GHCR while the ACP coding-agent workers and db-migrate (builder stage) point at the Bicep-provisioned ACR instead. - Template the remaining core services: judge service, token-manager, scheduler, portal (deployment/service/ingress), and the report-generator background worker. - Add the db-migrate pre-install/pre-upgrade Job, running against a dedicated :builder-tagged api image since the production image only ships compiled db-migrations output (no tsx/src). - Fix Redis secret wiring: the Bicep template's KV secret is redis-primary-key (a bare access key), not a connection string; wire discrete REDIS_HOST/PORT/PASSWORD/TLS env vars to match what every service actually reads. - Add AZURE_KEYVAULT_URI to the common-env ConfigMap for token-manager's direct Key Vault SDK client. - Add .github/workflows/publish-images.yml to publish the 6 core images (+ an api :builder tag) to GHCR on push to main. - Document the 3-phase flow (Azure infra -> helm install core app from GHCR -> build/push ACP workers to ACR + register secrets via Key Vault or the token-manager admin API) in the chart README and NOTES.txt. - Defer ACP worker Deployment manifests (coder-acp-copilot, coder-acp-claude-code) entirely; they need the Kubedock/MCP-gateway sidecar pattern from docs/architecture/kubedock.md, not yet templated. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 094bb93a-ec65-4f80-8896-0dd041aedeff
Found and fixed while running `helm install` against a real Bicep-provisioned test AKS cluster: 1. db-migrate Job hook-ordering bug: pre-install/pre-upgrade hooks run before any regular (non-hook) resource exists, so the Job's ServiceAccount and SecretProviderClass weren't present yet. Switched to post-install/ post-upgrade and added the serviceAccountName, workload-identity label, and kv-secrets CSI volume (mirrors api Deployment) so the Job can trigger the CSI driver's secret sync. 2. Portal nginx upstream resolution: apps/portal/nginx.conf hardcodes `proxy_pass http://api:80` at image-build time with no templating hook. The chart was naming the api Service "<release>-api", which nginx could never resolve. Renamed the Service to the literal "api" and updated the few other templates (ingress backend, SCOPE_MT_API_URL, CRITERIA_API_URL, NOTES.txt) that referenced the old componentName. 3. SecretProviderClass all-or-nothing secret sync: the Azure Key Vault CSI provider fails the *entire* volume mount if any single listed secret is missing, even for pods that don't need it. Made each secrets.keyVaultSecretNames.<logical> entry skippable via an empty string, so optional secrets (githubToken, anthropicApiKey — needed only by Phase 3 workers / portal AI features) can be omitted until populated. Verified end-to-end: `helm install` reaches STATUS: deployed, all 6 core Deployments reach 1/1 Running, the post-install db-migrate Job completes, and api confirms a live MongoDB (Cosmos) connection. Also relayed two infra-level bugs discovered during this validation to the Deploy-to-Azure Bicep session (PR #1477): a missing dependsOn causing redis-primary-key to never be written to Key Vault, and a misnamed Cosmos private DNS zone (privatelink.mongo.cosmos.azure.net instead of .com) that caused Cosmos DNS to resolve to the public IP and get firewalled. Both are being fixed on that branch, not this one. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 094bb93a-ec65-4f80-8896-0dd041aedeff
The chart previously defaulted portal.authEnabled to true, matching the portal image's own 'secure by default' fail-safe. But the publicly published GHCR portal image has no Entra app registration baked in (VITE_AUTH_* are build-time-only values — see docs/architecture/auth-rbac.md §8), so a fresh `helm install` with no overrides landed every user on RequireAuth.tsx's dead-end "Authentication not configured" page with no way to sign in or disable it short of a values override. Flipped the chart default to false so a vanilla install is usable out of the box; set portal.authEnabled=true (and rebuild/republish the portal image with real VITE_AUTH_* values) once a real IdP is wired up. Verified live: helm upgrade on the test AKS env rolled the portal pod, config.js now reports authEnabled: false, and the portal UI loads (HTTP 200) instead of the auth-not-configured screen. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 094bb93a-ec65-4f80-8896-0dd041aedeff
package.json has referenced this file (the 'get-copilot-token' pnpm script) since before this repo was split from the internal scope-core-internal repo, but the file itself was never copied over, leaving `pnpm get-copilot-token` / `pnpm -s get-copilot-token` broken with ERR_MODULE_NOT_FOUND. Restored verbatim from growth-ecosystems/scope-core-internal. Reviewed for anything that shouldn't be public: none — it only uses VS Code's well-known public OAuth client_id for the GitHub device-code flow and calls public github.com OAuth endpoints, no secrets or internal URLs. Verified: `pnpm -s get-copilot-token` now starts the device-code flow correctly instead of failing to resolve the module. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 094bb93a-ec65-4f80-8896-0dd041aedeff
Adds templates/workers/{coder-acp-copilot,coder-acp-claude-code}/deployment.yaml,
gated on workers.<name>.enabled (default false). Mirrors the existing
report-generator pattern: pure background Storage Queue consumers (no
HTTP port/probes), common-env ConfigMap + scope-secrets envFrom, Key
Vault CSI volume mount, shared workload identity ServiceAccount.
New values.yaml fields: workers.<name>.queueName and .agentVersion,
defaulted to each worker's agent-version.dev.yaml contents for chart
smoke-testing; real deployments should register a non-dev version
manifest and point agentVersion at it.
Explicitly deferred (documented via code comments + README): MCPJungle
MCP-gateway sidecar, AI Gateway/DevProxy HAR-capture sidecar, and
Kubedock in-task Docker access — none are templated yet.
Credential provisioning for these workers needs no new chart plumbing:
both workers acquire GitHub/Anthropic credentials dynamically from
Token Manager at runtime (POST /api/v1/keys via the Portal's "Create
Token" page or direct API call, with the static scope-secrets env vars
as an optional fallback only.
Updated README's Phase 3 section to describe the now-templated worker
Deployments and the agentVersion registration contract.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 094bb93a-ec65-4f80-8896-0dd041aedeff
EOF
)
Wraps building/pushing both ACP worker images, prompting for and registering GitHub/Anthropic credentials via the Token Manager API, helm upgrade to enable the worker Deployments, and agent registration into a single idempotent command, eliminating the prior multi-step manual az acr build + helm upgrade --set + register-agent.sh flow. Validated end-to-end against the live test AKS cluster: both worker images built via az acr build, helm upgrade enabled both Deployments, both pods came up Running and connected to Cosmos DB/Redis/Storage queues, and both agent types registered successfully via the API. Updated deploy/helm/scope/README.md's Phase 3 section to recommend the script as the primary path, keeping the underlying helm upgrade mechanics as a "what it does under the hood" reference. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 094bb93a-ec65-4f80-8896-0dd041aedeff
…ario quickstart `scope criteria import` and `scope prompt-feature import` never sent the projectId the API's /criteria/seed and /prompt-features/seed routes require (ProjectIdQuerySchema), so both commands always failed with a 400 against any real deployment. Confirmed live against the test cluster before and after the fix. Added `--project <id>` following the same requireProjectId pattern already used by every other criteria/prompt-feature subcommand in these files, and verified by importing all 14 config/criteria samples into the test env's "Initial Project" end to end. Also documents, in deploy/helm/scope/README.md, how to try a sample scenario after install — explicitly NOT as a Helm hook/Job: criteria are project-scoped business data, not infrastructure, and no project exists until db-migrate's "Initial Project" migration has run, so baking a project assumption into the chart would be wrong for anyone bringing their own projects. Scenarios/personas aren't server-side data at all (read straight from config/ by `run submit`), so there's nothing to import for those either way. Fixed a stale doc bug in passing: README said db-migrate is a pre-install,pre-upgrade hook; the template (and its own inline comment) say post-install,post-upgrade. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 094bb93a-ec65-4f80-8896-0dd041aedeff
…e-values in bootstrap-workers.sh BlobStorage: add a narrow withContainerRetry() wrapper (via the existing withRetry utility) around every snapshots-container write, retrying only on AuthorizationPermissionMismatch / 403 "not authorized" errors (never on other 4xx) with capped exponential backoff. This guards against transient RBAC-propagation races (e.g. right after createIfNotExists auto-creates a container for the first time). Note: this does NOT fix the deterministic "Snapshot upload failed" AuthorizationPermissionMismatch seen on uploadWorkspaceSnapshot / uploadSnapshotFromDirectory. Root-caused separately on the test AKS cluster: those two methods set Blob Index Tags (x-ms-tags) on upload, which requires the blobs/tags/write data action that Storage Blob Data Contributor does not grant. Confirmed via live reproduction (plain PUT succeeds, same PUT with x-ms-tags fails with 403 every time). Fix is a Bicep role upgrade (Storage Blob Data Contributor -> Storage Blob Data Owner for the workload identity) tracked in the deploy-azure-button-infra PR, not addressed by this commit. bootstrap-workers.sh: --values-file <path> lets helm upgrade re-merge the chart's current values.yaml instead of freezing a historical --reuse-values snapshot, which was silently dropping new default keys (agentVersion, queueName) from reaching deployed workers.
The README had grown to cover install steps, secrets architecture, worker bootstrapping internals, and a sample-scenario walkthrough all in one file. Split it so the README focuses purely on the steps needed to get the app running, and links out to detail docs for anyone who wants to go deeper: - README.md: three-phase overview, prerequisites, install steps, verify (including the port-forward commands needed to reach the API/portal without an ingress), and links to the detail docs. - CONFIGURATION.md (new): values.yaml conventions, image registry strategy, and the two bring-your-own secrets (GitHub token, Anthropic API key) - trimmed to just what's actionable for getting the app running, not the full three-pattern secrets architecture. - WORKERS.md (new): bootstrap-workers.sh details, how the chart templates worker Deployments, agent-version rules, and the Token Manager API as an alternative way to register credentials. - SAMPLE-SCENARIO.md (new): the criteria-seed + sample-run walkthrough. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 094bb93a-ec65-4f80-8896-0dd041aedeff
Josh Duffney (duffney)
requested review from
Cedric Vidal (cedricvidal) and
Wassim Chegham (manekinekko)
as code owners
October 9, 2026 19:12
Test Results (Node.js 22)test: Run #226
🎉 All tests passed! |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds a Helm chart (
deploy/helm/scope/) that deploys the Scope application — api, portal, judge, token-manager, scheduler, report-generator, and the two ACP coding-agent workers (coder-acp-copilot,coder-acp-claude-code) — onto the AKS cluster provisioned by the "Deploy to Azure" Bicep template (#1477). The Bicep template provisions infrastructure only; this chart is the app-install piece it calls out as a separate, not-yet-built step.Covers three phases:
deploy/azure/(Bicep), unchanged by this PR..github/workflows/publish-images.yml), so no registry build step is needed.scripts/bootstrap-workers.sh, a single entry point that builds/pushes both worker images to ACR, registers their GitHub/Anthropic credentials through the Token Manager API, enables their Deployments viahelm upgrade, and registers both agent types with the API.Secrets flow through the AKS Key Vault Secrets Provider CSI driver (workload identity, no secrets in values files) for infra-provisioned values (Cosmos DB connection string, Redis key), with two bring-your-own secrets (GitHub token, Anthropic API key) settable via
az keyvault secret setor through the Token Manager API.Also includes two fixes found while validating the chart end-to-end on a live test AKS cluster:
scope criteria import/scope prompt-feature importnever sent theprojectIdthe API's seed routes require, so both commands always 400'd against a real deployment.BlobStoragesnapshot/tool-call/log uploads now retry on transientAuthorizationPermissionMismatcherrors (e.g. right after a container is first auto-created), via the existingwithRetryutility.Demos
N/A — this is infrastructure/chart tooling and CLI/backend fixes, not a Portal UI change.
Testing
bash -n scripts/bootstrap-workers.sh/scripts/register-agent.sh— passed.npx vitest run packages/shared/src/storage/blob-storage.test.ts— 19/19 passed.pnpm --filter shared exec tsc --noEmit -p .— passed, no errors.oss-scope-jduffney/aks-oss-scope-test3-csqp7e3kner4s):helm install/helm upgradeof the core chart, confirmed all 6 core services come up healthy with CSI-synced secrets and workload identity.scripts/bootstrap-workers.shend-to-end: built/pushed both worker images to ACR, registered credentials, enabled worker Deployments, registered both agent types.config/criteriasamples into the seeded "Initial Project" viascope criteria import --project <id>(confirms the projectId fix).hello-world-express-v2.yamlend to end (scope run submit) — reachedstatus: done, including successful workspace-snapshot uploads to Blob Storage with index tags (confirms the companion storage-RBAC fix landing in Add Deploy-to-Azure Bicep template #1477, which upgrades the workload identity's storage role toStorage Blob Data Owner).Documentation and compatibility
deploy/helm/scope/README.md— quickstart: prerequisites, install steps, bringing workers online, and verifying the deployment.deploy/helm/scope/CONFIGURATION.md— values.yaml conventions, image registry strategy, bring-your-own secrets.deploy/helm/scope/WORKERS.md—bootstrap-workers.shdetails and how the chart templates worker Deployments.deploy/helm/scope/SAMPLE-SCENARIO.md— seeding criteria and submitting a sample run.No breaking changes to existing APIs/CLI behavior, aside from the
criteria import/prompt-feature importbug fix (these commands previously always failed with a 400 against a real deployment;--project <id>now works as documented). No DB migration changes. Requires the Bicep RBAC fix in #1477 (storage role upgrade toStorage Blob Data Owner) for workspace snapshot uploads with index tags to succeed.Checklist
up()/down()and keep it CosmosDB-compatible. — N/A, no DB schema changes.NOTICE/NOTICE-REVIEW.txtwithpnpm noticeas needed. — N/A, no dependency changes.Co-authored-by: Copilot App 223556219+Copilot@users.noreply.github.com