fix(skill): retry registry fetches and tolerate blocked backup cleanup - #185
Merged
Conversation
- bl skill init: raise index timeout 10s→30s and retry transient network failures (3 attempts); advisor sync silent channel stays fail-fast - make post-swap backup deletion best-effort so host safe-delete guards cannot fail a completed install (postinstall.js mirror included) - add unit coverage for guard-blocked cleanup and registry retry policy
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Fixes two issues in
bl skill init:INDEX_TIMEOUT_MSraised from 10s to 30s;fetchSkillsIndex/downloadSkillAssetnow go throughwithRetry(3 attempts by default; network/timeout/5xx/404 errors are retried)attemptsparameter so the silent advisor background channel passes 1 and stays fail-fast, without stallingbl advisor recommendbailian-docs-llm-wikihas 1583 files. AfteratomicSwapcompletes, deleting the.old-*backup trips the host guard (">500 files deleted per turn requires confirmation"), and the error bubbles up: an install that already succeeded is reported as failed and the lock entry is never updated.old-*dirs are inert (status scans already ignore them). The duplicated logic inpostinstall.jsis fixed the same wayTest plan
skills-installer.test.ts: guard-blocked backup deletion still installs, cleanup failure does not mask the original error, index fetch retries then succeeds,attempts=1disables retries, asset download retries transient HTTP errors; existing 404 case now asserts the 3 retries are exhaustedvp checkpasses (0 errors)Known issues
advisor-sync (1 case) and skills-agents (8 cases) fail on this machine — pre-existing environment issues (
/etc/codexand other real agent dirs on this host break detection expectations), reproduced on unmodified code, unrelated to this change.