Rewrite AKS zone-down sample-app tutorial for async signals and hard placement (#5240)
* Rewrite AKS zone-down sample-app tutorial for async signals and hard placement
- Replace the fixed "~5 minute" Kubernetes reschedule claim with observable,
monitor-driven guidance; downtime duration is no longer asserted as a fact.
- Add explicit Cloud Shell bootstrap (az account set, az aks get-credentials,
kubelogin convert-kubeconfig) before the first kubectl call.
- Pin the front end to a single zone as an explicit, labeled demo
anti-pattern instead of relying on the scheduler's default placement,
so run 1's failure is deterministic instead of easy to miss.
- Add a browser-based monitor (monitor.py) as the primary demo surface,
showing storefront HTTP status, target-node readiness, pod placement,
and a transition history; Metrics and kubectl become supporting signals.
- Call out that the storefront HTTP signal and node NotReady status update
asynchronously, and that the HTTP signal is primary.
- Set the Compute Zone Down scenario duration to 5 minutes for this demo,
with a note that other scenario types need their own duration guidance.
- Replace the ScheduleAnyway topology spread constraint (previously
described as guaranteeing per-zone placement, which it doesn't) with a
DoNotSchedule hard constraint, plus an explicit verify-fix.sh check that
every zone has a Ready store-front replica before run 2.
- Add a troubleshooting section for when run 1's impact isn't visible.
- Reframe run 2's outcome as sustained availability through the outage,
not zero dropped requests, noting brief load-balancer convergence is
still possible.
Mirrors the final contract from the companion microsoft/chaos-studio
samples/aks-zone-down-demo update.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
* Strengthen monitor error visibility and run-2 wording
- Monitor: call out that a failed kubectl/API poll surfaces as an explicit
error state per signal, instead of hanging on a "checking" placeholder.
- Verification: clarify that verify-fix.sh retries transient kubectl/API/
JSON failures until its own timeout before failing, rather than exiting
on the first transient error.
- Run 2: reframe explicitly as sustained availability with a possible brief
Azure Load Balancer convergence blip, not an unbroken "stays green
throughout" state.
Interim update while the companion microsoft/chaos-studio sample branch
finishes its correction pass; not final until the revised contract lands.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
* Reconcile zone-pin, monitor, and verify steps with the immutable f9573c9 sample
- Derive PIN_ZONE dynamically from the pod's actual node/zone label
(kubectl get pods -> node -> zone label), instead of asking the
reader to pick a bare zone number.
- Move the deliberate-anti-pattern annotation to the Deployment's
top-level metadata.annotations (matches deploy.sh); use the exact
annotation key chaos-demo.aks-zone-down-demo/deliberate-anti-pattern.
- Add the rollout restart/status wait and a self-check that the pod
landed back in $PIN_ZONE, mirroring deploy.sh's own safety check.
- Pin the monitor.py / verify-fix.sh download URLs to the immutable
commit f9573c943694e88cf3aacbca38debe84fa91a62c instead of a 404ing
main branch link; note the move to a tagged ref once merged.
- Pass $PIN_ZONE verbatim (full zone label) to monitor.py --target-zone
instead of a hardcoded eastus2- prefix.
- Tighten the monitor error-state wording to match the real red banner
behavior, and the verify-fix.sh wording to match its rollout-wait +
per-zone Ready check.
- Point the scenario zone-number entry and troubleshooting steps at
$PIN_ZONE instead of an ad hoc 'the zone you pinned' reference.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
* Remove unsupported fixed recovery-time claim in run 2 wording
Replace 'recovery within a few seconds' with environment-dependent
wording: a brief Azure Load Balancer convergence blip is possible with
no fixed time bound, and the monitor's transition history is what
distinguishes a transient blip from a sustained outage.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
* Apply Learn Authoring Assistant style suggestions
Accepts 10 of 11 flagged suggestions: splits four semicolon-joined
independent clauses into separate sentences, replaces 'may' with
'might' per style guide, removes a dangling 'below' reference, adds
a noun after the demonstrative 'these', and replaces a dash-joined
clause with a full sentence introducing a cross-reference.
Rejects the suggestion to drop the trailing colon on 'Remove the zone
pin you added earlier:' (line 256) because it would break the
established colon-before-code-block pattern used by every other
numbered step in this article that introduces a fenced code block
(for example, the 'Create a resource group and the cluster:' and
'Deploy the application:' steps).
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
---------
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>