feat(ariadne): allowlist clear_stuck_agent_pods
A build whose agent never started is a distinct failure from one that lost a connection mid-run, and a plain retry queues behind the same stuck pods. The remediation clears pods that have already succeeded or failed - Ariadne's existing scheduled cleanup - and only then rebuilds. Mapping jenkins_agent_provisioning_failure to it keeps the one-classification- one-action rule: a diagnosis asking for this action under any other classification is still refused before anything runs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
649e6f1e37
commit
c73a79a243
@ -465,13 +465,13 @@ spec:
|
||||
- name: ARIADNE_HERMES_AUTOREMEDIATION_ENABLED
|
||||
value: "true"
|
||||
- name: ARIADNE_HERMES_ALLOWED_ACTIONS
|
||||
value: repair_demo_fixture,retry_transient_infra,reclaim_workspace_storage
|
||||
value: repair_demo_fixture,retry_transient_infra,reclaim_workspace_storage,clear_stuck_agent_pods
|
||||
# Each classification maps to exactly one action id. A diagnosis
|
||||
# whose classification is absent here can never reach an action,
|
||||
# and a diagnosis asking for an action that is not its
|
||||
# classification's own is refused before anything runs.
|
||||
- name: ARIADNE_HERMES_ACTION_CLASSIFICATIONS
|
||||
value: known_demo_fixture_failure=repair_demo_fixture,transient_infra_failure=retry_transient_infra,workspace_storage_exhausted=reclaim_workspace_storage
|
||||
value: known_demo_fixture_failure=repair_demo_fixture,transient_infra_failure=retry_transient_infra,workspace_storage_exhausted=reclaim_workspace_storage,jenkins_agent_provisioning_failure=clear_stuck_agent_pods
|
||||
# Adds a service's own namespace to log evidence alongside
|
||||
# jenkins, so a build failure that coincides with the service being
|
||||
# unhealthy carries some trace of the service. Only mapped jobs are
|
||||
|
||||
Loading…
x
Reference in New Issue
Block a user