Repository navigation
Conversation
Provisions a 3-node b3-16 OVH Managed Kubernetes cluster (private network + router for node egress) with an Octavia load balancer and floating IP for ingress, and a *.eoapi-workshop.ds.io wildcard A record in Route53 pointing at it. Local state by default (gitignored); S3 backend stubbed for later remote-state migration. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The proxy's OIDC URL used the namespace-qualified Service DNS, pinning the release to namespace `eoapi`. The short Service name resolves in any namespace; only the release name must stay `eoapi`. Bring the proxy in line with docker-compose on main: v1.2.0 (what compose's :latest ran at the workshop), PRIVATE_ENDPOINTS requiring the stac:write scope, and the chapter 7 row-level TenantFilter. The filter module is mounted from a ConfigMap built from docs/workshop_filters.py through a symlink, so compose and the chart share one copy. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
On a cluster that already has a ClusterIssuer of that name, TLS=1 rendered a second one, which helm refuses to adopt. Had it been adopted, teardown would delete an issuer other tenants use. Reuse it and render none. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A Lab per entry in jupyter.participants meant a list edit per participant, and nothing used the freedom to name them. jupyter.count: N renders lab-01..lab-NN; deploy.sh writes their tokens to a jupyter.tokens map and still reads the old list format, so existing Lab URLs survive the switch. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The event needs more capacity than a standing cluster, for a day. A `workshop` node pool can now be added and removed with Terraform: workshop_node_count on the main stack for a cluster it created, or the standalone workshop-pool/ root for any existing OVH cluster. OVH labels the pool's nodes nodepool=workshop; the chart's new jupyter.nodeSelector puts the Labs there, so they can burst without crowding other tenants on the standing nodes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Chart README: lab-01..lab-NN from jupyter.count and the tokens map, nodeSelector, reuse of an existing ClusterIssuer, what deploy.sh installs on a shared cluster, the write scope and row-level filter, notebooks 00-08, and sizing for ~20 participants. Terraform README: the workshop pool in the overview, file table and cost note. Root README: point at the Kubernetes path beside the AWS one. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every deploy re-ran `helm upgrade --install pgo`, pulling the chart and waiting on the rollout each time. Skip it when a PGO Deployment exists in any namespace. Checks the Deployment, not the CRDs, which outlive `helm uninstall pgo`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Move the chart's eoapi dependency from 0.13.1 to 0.16.3 and keep docker-compose in lockstep with what it renders: stac-fastapi-pgstac 7.0.0, titiler-pgstac 3.2.0, tipg 1.6.1, stac-browser 5.1.0, stac-auth-proxy v1.2.0 and the mock-oidc digest the chart pins. stac-fastapi-pgstac 7.0.0 drops the POSTGRES_* settings and the module-level `app`, so compose now passes PG* variables and starts `create_app --factory`, as the chart does. pgstac goes to 0.9.11 in both environments. The chart's default pgstac-pypgstac:v0.10.0 is a stray build of the 0.9.10 commit, and 0.9.12 has no ghcr.io/stac-utils/pgstac image for compose, so 0.9.11 is the newest both can run. The loader and the labs' pypgstac follow. The browser and mock-oidc image overrides are dropped: upstream now ships the same vanilla stac-browser and pins mock-oidc by digest, which left the old repository/tag keys unused. eoapi >= 0.14 requires Kubernetes >= 1.32, so the Terraform default moves from 1.31 to 1.35, the newest OVH MKS lists. Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(docs): notebooks follow a per-participant URL contract workshop_setup.py gains endpoints(), to_browser(), show_links() and collection_id(): server-side and browser-facing URLs come from env vars, replacing the per-notebook .replace() heuristics. Notebooks use per-user collection ids, and 9 notebook-02 points with no Sentinel-2 items are dropped. Compose and the chart set the browser-facing variables. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(spike): one participant's eoAPI stack in one network namespace Compose layout of a participant pod: the Lab owns the netns, every other service joins it and binds 127.0.0.1. jupyter-server-proxy exposes each service same-origin under /stac /raster /vector /browser /oidc /manager, behind the Lab password. The DB image bakes the ecoregions and glad data. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(spike): Helm chart with one pod per participant One Deployment per participant from the compose containers (DB as a native sidecar), a credentials Secret kept across upgrades, one Ingress for the Lab hosts and a NetworkPolicy admitting only the ingress controller. No cluster-scoped objects. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(spike): JupyterHub front door for comparison zero-to-jupyterhub 4.4.2 running the same participant stack as singleuser extraContainers, with a static per-user password authenticator. Kept to compare against the own chart. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test(spike): checks and evidence for the per-participant stack Reproducible check scripts per topic (build, apis, auth, browser apps, notebooks, footprint, both front doors, verify) and the evidence they produced, including the skeptic re-run in evidence/verify.md. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * docs: bring notebook 00 up to date for per-participant stacks The infrastructure section still described the shared CDK endpoints (workshop-*.eoapi.dev), Binder and DB credentials handed out on the day. It now describes the participant's own stack behind the Lab login, the work/ folder that survives a restart, and the links cell at the end. The outline lists chapters 6-8; also fixes the 4.3 heading and typos. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * chore(spike): drop the JupyterHub front door The own chart is the front door: it fits one password per Lab URL and laptop API access, for the same amount of code. The z2jh variant stays in commit 10aac1d; evidence/frontdoor-hub.md keeps the comparison. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * chore(spike): keep only the footprint tables; mark pre-fix evidence superseded Each run keeps results/<label>/analysis.txt, the numbers the sizing cites. The raw samples stay in commit e3c0479 and run.sh regenerates them. fix.md and verify.md describe the current stack. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(spike): compare image architectures with the Docker host stac-browser and stac-manager always printed PASS, and the other images had to be arm64. Now each image must match the host, except those two, which are only published for amd64. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * style(spike): pass ruff lint and format CI lint failed with 114 errors under spike/. The auth checks are fragments run after common.py, so F821 is ignored there; the rest are renamed variables, two lambdas turned into defs, two unused imports. Notebook 03's filter line now fits on one line. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(spike): persist participant data, restrict egress, pin images - A PVC per participant for the DB (PGDATA one level down, past lost+found) and one for /home/jovyan/work. A pod replacement no longer wipes the participant; removing them or the release deletes both. - The NetworkPolicy now also limits egress: DNS, and 80/443 on public addresses. Lab terminals can no longer reach the API server, other namespaces or node IPs. - The Lab and DB images come from <registry>/eoapi-workshop-{lab,db}:<tag> (`local` on kind); nodeSelector/tolerations pin the pods to the workshop pool, and prepull adds a DaemonSet that pulls every image. The kind checks probe egress both ways (policy present, then deleted) and write their persistence files under work/. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * ci: publish the participant Lab and DB images for amd64 Pushes ghcr.io/<owner>/eoapi-workshop-{lab,db}:sha-<commit> on main and feat/k8s, the tag spike/chart pins. Pull requests build without pushing. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(spike): deploy.sh for the participant stacks Context, namespace, image tag and the whole participant list are always explicit; up uses --reset-values (a bare helm upgrade reuses the last --set) and refuses to drop participants without REMOVE=1. creds prints the CSV for the handout slips. down needs CONFIRM=<namespace> and only uninstalls the release, never the namespace. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * docs(spike): evidence for the hardening runs Kind: frontdoor-own 61/0/0 (egress probed with and without the policy), verify/own.sh 4/0/0 (data survives pod replacement), deploy.sh's guards. Compose checks unchanged from verify.md except build's fixed arch check. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * chore(spike): drop loc.py, the front-door line counter It measured the own chart against z2jh; that choice is made. The evidence keeps its numbers and points at commit e3c0479 for the script. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * chore(spike): move the investigation to an evidence branch The PR keeps what the workshop runs on (chart, deploy.sh, images, compose, CI, notebook fixes) and the chart's kind tests. The per-topic checks, write-ups and screenshots (~11,800 lines) stay on branch spike/per-user-stacks-evidence (commit d50c0ce); the README links it and now carries the kind setup steps and the sizing numbers. The persistence probe deletes its collection first: with PVCs, a previous run's copy survives, and POST answered 409. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Check out this pull request on See visual diffs & provide feedback on Jupyter Notebooks. Powered by ReviewNB |
The workshop runs one eoAPI stack per participant (spike/), so the shared-stack chart in infrastructure/charts/eoapi-workshop goes, and publish-workshop-image.yml with it: its GHCR image was the chart's jupyter.image, and nothing else pulls it (compose and spike/lab build Dockerfile.local themselves). CI's Helm job now lints and renders spike/chart. The README, the Terraform docs and comments, and three spike/ comments that cited the old chart now point at spike/. The workshop pool sizing follows spike/README.md "Sizing": 6x b3-16 for 20 participants. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- tls: a namespaced cert-manager Issuer and one Certificate for a fixed list of Lab hosts (spares included, not the participant list, so a participant change never triggers a new ACME order). - values-labs.yaml: hosts on *.eoapi-workshop.ds.io, staging TLS for rehearsals, the pool settings to uncomment once the pool exists. - warm.sh: fetches every stack's first world-view tile (110 s cold, 1 s warm on kind). - README: the real-cluster steps and the rehearsal checklist. The default render is unchanged; the Issuer and Certificate pass a server-side dry run against cert-manager's CRDs. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Records what the design chose, what each choice costs and what to do on the day: one pod per participant, the Lab as the only way in, the single STAC DB connection (a large request blocks that participant's STAC API for about a minute, seen on labs), memory sizing, what survives a restart, the release freeze, cold tiles, network isolation, TLS and pinned images. Also what the 2026-10-02 rehearsal on labs verified. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
With the rehearsal kit's tls.yaml, the chart renders a namespaced Issuer and Certificate when tls.acmeServer is set; the README said nothing did. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The folder is the workshop's deployment now, not a spike: one eoAPI stack per participant (compose locally, a chart on Kubernetes). Local names follow: compose project and images eoapi-participant-*, kind cluster eoapi-participant, the kind tests' release and namespace participants, test hosts lab-uNN.participant.local. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
From a skeptic-checked review of PR #35's non-notebook files: - fix cross-file inconsistencies: the pool label (nodepool, as Terraform sets it), cert-manager comments (the chart now issues its own certificate), cold-tile timings, CI now renders values-labs.yaml so the Issuer and Certificate are checked; - drop investigation wording (verify stage, the spike, test-run dates, PR self-references) and comments that repeat the code or the README; - say each fact once in participant/README.md (the numbers live in Sizing, the trade-offs in Design and trade-offs); - deduplicate test code (one Collection helper in ingress.py, one hashing helper in own.sh) and drop an unused Terraform output. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
From the notebook content review (wf_8473515d-ebb), only the findings that are wrong on every deployment (shared AWS, per-participant stacks, local compose): typos, the 'April 4' search date, 01's minimal STAC item (geometry, links), 04's asset order and colormap/map.html references, 05's ecoregions table and filter field, 06's printed detail, 07's policy table, 08's collection names, the README link in 00 (it downloaded a raw file on the book site), chapters 06-08 in the myst toc with unique cell ids in 08, and cql2 in environment.yml for workshop_filters.py. Deployment-dependent findings stay open. Notebooks 00-08 executed on the local per-participant stack: 0 errored cells. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- values-labs.yaml: the rehearsal found labs' API server on a private address (10.3.0.1:6443), which the egress policy already blocks; the general egressExcept guidance stays in values.yaml and the README. - frontdoor-own/run.sh: pod.arch expected an arm64 host (stac-fastapi=aarch64) and failed on amd64; broken emulation already fails the 9/9-ready check. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The repository is stac-api-extensions/transaction (singular). The plural URL returns 404, which fails the Docs job's --check-links --strict build. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Hi @hrodmn and all! Before marking this PR ready I want to know your thought about the wording in the notebooks. They currently need to run on three deployments: a) the shared AWS stack, b) one stack per participant on Kubernetes (this PR), and c) local docker compose. I ran a review of notebooks 00–08 with Claude and it found 22 places where the text is right for one deployment and wrong for another. What was wrong everywhere is already fixed in 4e63e81. My question is: Should the notebooks describe all three deployments? or only the one we run at the workshop? I put a proposal that keeps the text correct on all three, for discussion: Some of the most visible difference are maybe:
|
STAC Browser draws an item's `preview` COG, the asset with the `overview` role. Earth Search's preview declares no no-data value, so the empty part of a partially covered tile came out black (97% of it on S2B_T39QWD_20250413T071904_L2A). - Notebook 02 marks 0 as the preview's no-data value before loading the items. ol-stac passes it to OpenLayers, which draws those pixels transparent. - STAC Browser gets SB_displayOverviewsForChildren=true in the chart and in both compose files. Collection and search pages, such as the one notebook 02 embeds, then draw the preview COG instead of the JPEG thumbnail, which can't be transparent. Verified against stac-browser 5.1.0 in Chrome and Firefox: the black share of the no-data area dropped from 0.97 to 0. The field survives pypgstac upsert, stac-fastapi and stac-auth-proxy unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…f listing them participants: N renders u01..uNN, and deploy.sh up takes TAG N. tls.count does the same for the certificate's names, a second count on purpose: it is set once with spares, so stacks can be added on the day without a new ACME order (5 per week per set of names, 50 per week across ds.io). The chart fails if participants exceeds tls.count. values-labs.yaml moves to production Let's Encrypt with 25 names. deploy.sh reads a release's old list as its length, so the running u01..u03 become 3 and keep their credentials. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ks 03 and 04 On labs, notebook 03's `limit=1000` items request raised ReadTimeout and notebook 04's external preview raised RemoteProtocolError. - stac-auth-proxy buffers and rewrites the whole response: a 1,000-item page (18.7 MB) peaked at 365 MiB against a 384 MiB limit and took 5.3 s, past httpx's 5 s default. Limit 384 -> 768 MiB; the cell gets timeout=60. - titiler held ~600 MiB after one map's mosaic-tile burst and was OOM-killed at 768 MiB, taking the in-flight preview with it. MALLOC_ARENA_MAX=2 and GDAL_CACHEMAX=64 cut the burst peak from 494 to 355 MiB; limit 768 MiB -> 1 GiB. GDAL now retries S3 503 SlowDown. - Notebook 04 cells 25 and 27 render the same 2048 px preview as cell 23 but used the 5 s default timeout; they now match cell 23. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01G22jbyiWkxrVUp3PSMLFTK
Notebook 04's colormap painted value 0 (no forest loss) opaque black, so every loaded Hansen tile showed as a black block on the world map (44% of a z1 tile). 0 -> [0, 0, 0, 0] leaves only the loss pixels (under 1% opaque). Cell 26 explains the fourth value. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01G22jbyiWkxrVUp3PSMLFTK
The DB image baked the MAAP collection's first 100 items: a 50-80N strip with gaps, under a global extent, so notebook 04's map opened on an almost empty world. - glad_to_sql.py takes bboxes and follows pagination. BBOXES covers North America and Europe with the UK (152 tiles). The collection extent becomes their union, [-180, 0, 50, 80], so map.html opens on them at z2. - warm.sh warms those six z2 tiles in parallel (49 s cold on a test stack) instead of the world tile. - Notebook 04 no longer promises a global view on AWS Lambda. Existing stacks keep their old glad items: init SQL only runs on an empty volume. To update one: delete_collection + this SQL in one transaction, delete the glad rows from pgstac.searches (their metadata keeps the old bounds), restart titiler, then run warm.sh. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01G22jbyiWkxrVUp3PSMLFTK
Zoomed-out Sentinel-2 mosaic tiles are bigger than the participant's 4-degree box, never count as filled, and read the 100-item cap one COG at a time from us-west-2 (~1.5 s each): 2.5-3 min per view. items_limit=20 on the three Sentinel-2 maps (cells 5, 14, 17): z5 view 170 -> 33 s, z6 158 -> 29 s, z7 95 -> 64 s, z8+ unchanged; about 3x less titiler CPU and slightly less memory (local benchmark). Cost: zoomed-out tiles draw fewer scenes (z5 fill 22% -> 13%, z6 44% -> 33%). The GLAD map keeps the default. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01G22jbyiWkxrVUp3PSMLFTK
Deploys the workshop to Kubernetes as one full eoAPI stack per participant, plus the infrastructure under it.
Per-participant stacks (
participant/, from #64)./stac,/raster,/vector,/browser,/managerand/oidc, behind one password per participant.participant/chart):work/on volumes;publish-participant-images.ymlbuilds for amd64;values-labs.yamlholds the labs settings).participant/deploy.shinstalls exactly the participants it is given, prints their credentials as CSV, and tears down without touching the namespace.participant/warm.shwarms each stack's first map tiles.docs/workshop_setup.py).participant/README.md, including the single STAC database connection per stack, the release freeze during the event, and cold tiles.infrastructure/charts/eoapi-workshop) has been removed from this branch.Terraform (
infrastructure/terraform). An OVH managed cluster, ingress-nginx, cert-manager, a Route53 wildcard record, and an optionalworkshopnode pool. The pool comes fromworkshop_node_count, or from the standaloneworkshop-pool/root for an existing cluster.CI. Helm lint and render of
participant/chart(including the labs values), and ruff. The participant images publish on pushes tomainandfeat/k8s.Tested
Securelogin cookie;work/surviving a pod restart on Cinder volumes;Not done
b3-16(or 3–4×b3-32) plus a spare. Theworkshop-pool/root still defaults to 3 nodes.kube_version = "1.31"in the Terraform is likely out of date for OVH.🤖 Generated with Claude Code