Skip to content

Add Helm chart to deploy Scope app onto the Bicep-provisioned AKS cluster - #1481

Open
Josh Duffney (duffney) wants to merge 10 commits into
mainfrom
duffney-helm-chart-app-deployment
Open

Josh Duffney (duffney) wants to merge 10 commits into
mainfrom
duffney-helm-chart-app-deployment

Conversation

@duffney

Copy link
Copy Markdown
Collaborator

Summary

Adds a Helm chart (deploy/helm/scope/) that deploys the Scope application — api, portal, judge, token-manager, scheduler, report-generator, and the two ACP coding-agent workers (coder-acp-copilot, coder-acp-claude-code) — onto the AKS cluster provisioned by the "Deploy to Azure" Bicep template (#1477). The Bicep template provisions infrastructure only; this chart is the app-install piece it calls out as a separate, not-yet-built step.

Covers three phases:

  1. Infrastructure — deploy/azure/ (Bicep), unchanged by this PR.
  2. Core app — this chart. Images pull from GHCR by default (public images published by .github/workflows/publish-images.yml), so no registry build step is needed.
  3. ACP coding-agent workers — scripts/bootstrap-workers.sh, a single entry point that builds/pushes both worker images to ACR, registers their GitHub/Anthropic credentials through the Token Manager API, enables their Deployments via helm upgrade, and registers both agent types with the API.

Secrets flow through the AKS Key Vault Secrets Provider CSI driver (workload identity, no secrets in values files) for infra-provisioned values (Cosmos DB connection string, Redis key), with two bring-your-own secrets (GitHub token, Anthropic API key) settable via az keyvault secret set or through the Token Manager API.

Also includes two fixes found while validating the chart end-to-end on a live test AKS cluster:

  • scope criteria import / scope prompt-feature import never sent the projectId the API's seed routes require, so both commands always 400'd against a real deployment.
  • BlobStorage snapshot/tool-call/log uploads now retry on transient AuthorizationPermissionMismatch errors (e.g. right after a container is first auto-created), via the existing withRetry utility.

Demos

N/A — this is infrastructure/chart tooling and CLI/backend fixes, not a Portal UI change.

Testing

  • bash -n scripts/bootstrap-workers.sh / scripts/register-agent.sh — passed.
  • npx vitest run packages/shared/src/storage/blob-storage.test.ts — 19/19 passed.
  • pnpm --filter shared exec tsc --noEmit -p . — passed, no errors.
  • Live validation on a real AKS cluster provisioned by the Bicep template (oss-scope-jduffney / aks-oss-scope-test3-csqp7e3kner4s):
    • helm install/helm upgrade of the core chart, confirmed all 6 core services come up healthy with CSI-synced secrets and workload identity.
    • scripts/bootstrap-workers.sh end-to-end: built/pushed both worker images to ACR, registered credentials, enabled worker Deployments, registered both agent types.
    • Imported all 14 config/criteria samples into the seeded "Initial Project" via scope criteria import --project <id> (confirms the projectId fix).
    • Submitted hello-world-express-v2.yaml end to end (scope run submit) — reached status: done, including successful workspace-snapshot uploads to Blob Storage with index tags (confirms the companion storage-RBAC fix landing in Add Deploy-to-Azure Bicep template #1477, which upgrades the workload identity's storage role to Storage Blob Data Owner).

Documentation and compatibility

  • deploy/helm/scope/README.md — quickstart: prerequisites, install steps, bringing workers online, and verifying the deployment.
  • deploy/helm/scope/CONFIGURATION.md — values.yaml conventions, image registry strategy, bring-your-own secrets.
  • deploy/helm/scope/WORKERS.md — bootstrap-workers.sh details and how the chart templates worker Deployments.
  • deploy/helm/scope/SAMPLE-SCENARIO.md — seeding criteria and submitting a sample run.

No breaking changes to existing APIs/CLI behavior, aside from the criteria import/prompt-feature import bug fix (these commands previously always failed with a 400 against a real deployment; --project <id> now works as documented). No DB migration changes. Requires the Bicep RBAC fix in #1477 (storage role upgrade to Storage Blob Data Owner) for workspace snapshot uploads with index tags to succeed.

Checklist

  • If Portal features changed, keep CLI capabilities in sync. — N/A, no Portal changes.
  • If Portal components changed, update their Storybook stories. — N/A, no Portal components changed.
  • If database changes require a migration, include up() / down() and keep it CosmosDB-compatible. — N/A, no DB schema changes.
  • If dependencies changed, update the lockfile and regenerate NOTICE / NOTICE-REVIEW.txt with pnpm notice as needed. — N/A, no dependency changes.
  • Video showing the behavior before the suggested change — N/A, not a user-visible Portal change.
  • Video showing the behavior after the suggested change — N/A, not a user-visible Portal change.

Co-authored-by: Copilot App 223556219+Copilot@users.noreply.github.com

Josh Duffney (duffney) and others added 10 commits October 8, 2026 14:05
Umbrella chart at deploy/helm/scope/ with Chart.yaml, values.yaml, shared
helpers, workload-identity ServiceAccount, Secrets Store CSI
SecretProviderClass, common-env ConfigMap, and the api service (deployment,
service, ingress) as the reference implementation for remaining services.

Pairs with the Azure infrastructure provisioned by deploy/azure/ (PR #1477):
reads its outputs (acrLoginServer, keyVaultName, workloadIdentityClientId,
workloadIdentityServiceAccountName, storageAccountName) as values, uses
CSI-synced secrets for static connection strings/API keys, and workload
identity for direct Blob/Queue/Key Vault SDK access.

Remaining services (portal, judge, token-manager, scheduler, workers, jobs)
are configured in values.yaml but not yet templated — tracked as follow-up
work.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 094bb93a-ec65-4f80-8896-0dd041aedeff
… phase

- Add per-image registry override to scope.image helper so core services
  (api, portal, judge, token-manager, scheduler, report-generator) default
  to GHCR while the ACP coding-agent workers and db-migrate (builder stage)
  point at the Bicep-provisioned ACR instead.
- Template the remaining core services: judge service, token-manager,
  scheduler, portal (deployment/service/ingress), and the report-generator
  background worker.
- Add the db-migrate pre-install/pre-upgrade Job, running against a
  dedicated :builder-tagged api image since the production image only
  ships compiled db-migrations output (no tsx/src).
- Fix Redis secret wiring: the Bicep template's KV secret is
  redis-primary-key (a bare access key), not a connection string; wire
  discrete REDIS_HOST/PORT/PASSWORD/TLS env vars to match what every
  service actually reads.
- Add AZURE_KEYVAULT_URI to the common-env ConfigMap for token-manager's
  direct Key Vault SDK client.
- Add .github/workflows/publish-images.yml to publish the 6 core images
  (+ an api :builder tag) to GHCR on push to main.
- Document the 3-phase flow (Azure infra -> helm install core app from
  GHCR -> build/push ACP workers to ACR + register secrets via Key Vault
  or the token-manager admin API) in the chart README and NOTES.txt.
- Defer ACP worker Deployment manifests (coder-acp-copilot,
  coder-acp-claude-code) entirely; they need the Kubedock/MCP-gateway
  sidecar pattern from docs/architecture/kubedock.md, not yet templated.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 094bb93a-ec65-4f80-8896-0dd041aedeff
Found and fixed while running `helm install` against a real Bicep-provisioned
test AKS cluster:

1. db-migrate Job hook-ordering bug: pre-install/pre-upgrade hooks run before
   any regular (non-hook) resource exists, so the Job's ServiceAccount and
   SecretProviderClass weren't present yet. Switched to post-install/
   post-upgrade and added the serviceAccountName, workload-identity label,
   and kv-secrets CSI volume (mirrors api Deployment) so the Job can trigger
   the CSI driver's secret sync.

2. Portal nginx upstream resolution: apps/portal/nginx.conf hardcodes
   `proxy_pass http://api:80` at image-build time with no templating hook.
   The chart was naming the api Service "<release>-api", which nginx could
   never resolve. Renamed the Service to the literal "api" and updated the
   few other templates (ingress backend, SCOPE_MT_API_URL,
   CRITERIA_API_URL, NOTES.txt) that referenced the old componentName.

3. SecretProviderClass all-or-nothing secret sync: the Azure Key Vault CSI
   provider fails the *entire* volume mount if any single listed secret is
   missing, even for pods that don't need it. Made each
   secrets.keyVaultSecretNames.<logical> entry skippable via an empty
   string, so optional secrets (githubToken, anthropicApiKey — needed only
   by Phase 3 workers / portal AI features) can be omitted until populated.

Verified end-to-end: `helm install` reaches STATUS: deployed, all 6 core
Deployments reach 1/1 Running, the post-install db-migrate Job completes,
and api confirms a live MongoDB (Cosmos) connection.

Also relayed two infra-level bugs discovered during this validation to the
Deploy-to-Azure Bicep session (PR #1477): a missing dependsOn causing
redis-primary-key to never be written to Key Vault, and a misnamed Cosmos
private DNS zone (privatelink.mongo.cosmos.azure.net instead of .com)
that caused Cosmos DNS to resolve to the public IP and get firewalled.
Both are being fixed on that branch, not this one.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 094bb93a-ec65-4f80-8896-0dd041aedeff
The chart previously defaulted portal.authEnabled to true, matching the
portal image's own 'secure by default' fail-safe. But the publicly
published GHCR portal image has no Entra app registration baked in
(VITE_AUTH_* are build-time-only values — see docs/architecture/auth-rbac.md
§8), so a fresh `helm install` with no overrides landed every user on
RequireAuth.tsx's dead-end "Authentication not configured" page with no
way to sign in or disable it short of a values override.

Flipped the chart default to false so a vanilla install is usable
out of the box; set portal.authEnabled=true (and rebuild/republish the
portal image with real VITE_AUTH_* values) once a real IdP is wired up.

Verified live: helm upgrade on the test AKS env rolled the portal pod,
config.js now reports authEnabled: false, and the portal UI loads
(HTTP 200) instead of the auth-not-configured screen.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 094bb93a-ec65-4f80-8896-0dd041aedeff
package.json has referenced this file (the 'get-copilot-token' pnpm
script) since before this repo was split from the internal
scope-core-internal repo, but the file itself was never copied over,
leaving `pnpm get-copilot-token` / `pnpm -s get-copilot-token` broken
with ERR_MODULE_NOT_FOUND.

Restored verbatim from growth-ecosystems/scope-core-internal. Reviewed
for anything that shouldn't be public: none — it only uses VS Code's
well-known public OAuth client_id for the GitHub device-code flow and
calls public github.com OAuth endpoints, no secrets or internal URLs.

Verified: `pnpm -s get-copilot-token` now starts the device-code flow
correctly instead of failing to resolve the module.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 094bb93a-ec65-4f80-8896-0dd041aedeff
Adds templates/workers/{coder-acp-copilot,coder-acp-claude-code}/deployment.yaml,
gated on workers.<name>.enabled (default false). Mirrors the existing
report-generator pattern: pure background Storage Queue consumers (no
HTTP port/probes), common-env ConfigMap + scope-secrets envFrom, Key
Vault CSI volume mount, shared workload identity ServiceAccount.

New values.yaml fields: workers.<name>.queueName and .agentVersion,
defaulted to each worker's agent-version.dev.yaml contents for chart
smoke-testing; real deployments should register a non-dev version
manifest and point agentVersion at it.

Explicitly deferred (documented via code comments + README): MCPJungle
MCP-gateway sidecar, AI Gateway/DevProxy HAR-capture sidecar, and
Kubedock in-task Docker access — none are templated yet.

Credential provisioning for these workers needs no new chart plumbing:
both workers acquire GitHub/Anthropic credentials dynamically from
Token Manager at runtime (POST /api/v1/keys via the Portal's "Create
Token" page or direct API call, with the static scope-secrets env vars
as an optional fallback only.

Updated README's Phase 3 section to describe the now-templated worker
Deployments and the agentVersion registration contract.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 094bb93a-ec65-4f80-8896-0dd041aedeff
EOF
)
Wraps building/pushing both ACP worker images, prompting for and
registering GitHub/Anthropic credentials via the Token Manager API,
helm upgrade to enable the worker Deployments, and agent registration
into a single idempotent command, eliminating the prior multi-step
manual az acr build + helm upgrade --set + register-agent.sh flow.

Validated end-to-end against the live test AKS cluster: both worker
images built via az acr build, helm upgrade enabled both Deployments,
both pods came up Running and connected to Cosmos DB/Redis/Storage
queues, and both agent types registered successfully via the API.

Updated deploy/helm/scope/README.md's Phase 3 section to recommend
the script as the primary path, keeping the underlying helm upgrade
mechanics as a "what it does under the hood" reference.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 094bb93a-ec65-4f80-8896-0dd041aedeff
…ario quickstart

`scope criteria import` and `scope prompt-feature import` never sent the
projectId the API's /criteria/seed and /prompt-features/seed routes
require (ProjectIdQuerySchema), so both commands always failed with a 400
against any real deployment. Confirmed live against the test cluster
before and after the fix. Added `--project <id>` following the same
requireProjectId pattern already used by every other criteria/prompt-feature
subcommand in these files, and verified by importing all 14 config/criteria
samples into the test env's "Initial Project" end to end.

Also documents, in deploy/helm/scope/README.md, how to try a sample
scenario after install — explicitly NOT as a Helm hook/Job: criteria are
project-scoped business data, not infrastructure, and no project exists
until db-migrate's "Initial Project" migration has run, so baking a
project assumption into the chart would be wrong for anyone bringing their
own projects. Scenarios/personas aren't server-side data at all (read
straight from config/ by `run submit`), so there's nothing to import for
those either way.

Fixed a stale doc bug in passing: README said db-migrate is a
pre-install,pre-upgrade hook; the template (and its own inline comment)
say post-install,post-upgrade.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 094bb93a-ec65-4f80-8896-0dd041aedeff
…e-values in bootstrap-workers.sh

BlobStorage: add a narrow withContainerRetry() wrapper (via the existing
withRetry utility) around every snapshots-container write, retrying only
on AuthorizationPermissionMismatch / 403 "not authorized" errors (never on
other 4xx) with capped exponential backoff. This guards against transient
RBAC-propagation races (e.g. right after createIfNotExists auto-creates a
container for the first time).

Note: this does NOT fix the deterministic "Snapshot upload failed"
AuthorizationPermissionMismatch seen on uploadWorkspaceSnapshot /
uploadSnapshotFromDirectory. Root-caused separately on the test AKS
cluster: those two methods set Blob Index Tags (x-ms-tags) on upload,
which requires the blobs/tags/write data action that Storage Blob Data
Contributor does not grant. Confirmed via live reproduction (plain PUT
succeeds, same PUT with x-ms-tags fails with 403 every time). Fix is a
Bicep role upgrade (Storage Blob Data Contributor -> Storage Blob Data
Owner for the workload identity) tracked in the deploy-azure-button-infra
PR, not addressed by this commit.

bootstrap-workers.sh: --values-file <path> lets helm upgrade re-merge the
chart's current values.yaml instead of freezing a historical --reuse-values
snapshot, which was silently dropping new default keys (agentVersion,
queueName) from reaching deployed workers.
The README had grown to cover install steps, secrets architecture, worker
bootstrapping internals, and a sample-scenario walkthrough all in one file.
Split it so the README focuses purely on the steps needed to get the app
running, and links out to detail docs for anyone who wants to go deeper:

- README.md: three-phase overview, prerequisites, install steps, verify
  (including the port-forward commands needed to reach the API/portal
  without an ingress), and links to the detail docs.
- CONFIGURATION.md (new): values.yaml conventions, image registry
  strategy, and the two bring-your-own secrets (GitHub token, Anthropic
  API key) - trimmed to just what's actionable for getting the app
  running, not the full three-pattern secrets architecture.
- WORKERS.md (new): bootstrap-workers.sh details, how the chart templates
  worker Deployments, agent-version rules, and the Token Manager API as an
  alternative way to register credentials.
- SAMPLE-SCENARIO.md (new): the criteria-seed + sample-run walkthrough.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 094bb93a-ec65-4f80-8896-0dd041aedeff
@github-actions github-actions Bot added type: documentation Documentation additions, corrections, and improvements. area: cicd Build, test, release, and deployment pipelines. area: cli Scope command-line interface and terminal workflows. language: javascript Work involving JavaScript code, tooling, or dependencies. topic: testing Test coverage, test infrastructure, and validation quality. labels Oct 9, 2026
@github-actions

github-actions Bot commented Oct 9, 2026

Copy link
Copy Markdown

Test Results (Node.js 22)

test: Run #226

Tests 📝 Passed ✅ Failed ❌ Skipped ⏭️ Pending ⏳ Other ❓ Flaky 🍂 Duration ⏱️
3202 3202 0 0 0 0 0 1m27s

🎉 All tests passed!

Github Test Reporter

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area: cicd Build, test, release, and deployment pipelines. area: cli Scope command-line interface and terminal workflows. language: javascript Work involving JavaScript code, tooling, or dependencies. topic: testing Test coverage, test infrastructure, and validation quality. type: documentation Documentation additions, corrections, and improvements.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant