skill-creator: description optimization loop silently measures nothing on Windows (select() on pipe, then cp1252 report crash)

Status Open
Reported on v2.1.119
Maintainer reply None cached
Activity 0 comments · opened Aug 25, 2026

Summary

The skill-creator description-optimization loop (scripts/run_loop.pyscripts/run_eval.py)
is unusable on Windows. Two independent defects: every eval query fails silently, and the report
writer then crashes. The combination is bad because the loop still exits 0 and prints a plausible
result — recall=0% for every candidate description — which reads as "your skill never triggers"
rather than "nothing was measured".

Environment: Windows 11, Python 3.13.1, Claude Code 2.1.119, run from Git Bash.

Bug 1 — select.select() on a subprocess pipe (Windows: sockets only)

scripts/run_eval.py polls the claude -p subprocess with:

ready, _, _ = select.select([process.stdout], [], [], 1.0)
...
chunk = os.read(process.stdout.fileno(), 8192)

On Windows select.select() accepts sockets only, so this raises
OSError: [WinError 10038] An operation was attempted on something that is not a socket.

The exception is swallowed and reported as Warning: query failed: ..., once per query per run.
With 20 queries × 3 runs that is 60 identical warning lines, easily mistaken for network flakiness.

Impact: no query is ever evaluated. Every candidate description scores identically
(precision=100% recall=0%), so the "best" description is chosen by an arbitrary tie-break. The
run reports success.

Repro: run python -m scripts.run_loop --eval-set <set> --skill-path <skill> --model <model>
on Windows. claude -p itself works fine when invoked directly, which confirms the fault is in the
polling, not the CLI.

Suggested fix: replace the select-based polling with a reader thread feeding a queue —
portable and behaviourally identical:

class _LineReader:
    def __init__(self, stream):
        self.q, self.eof = queue.Queue(), False
        threading.Thread(target=self._pump, args=(stream,), daemon=True).start()

    def _pump(self, stream):
        try:
            for raw in iter(stream.readline, b""):
                self.q.put(raw)
        finally:
            self.q.put(None)

    def read(self, timeout):
        try:
            item = self.q.get(timeout=timeout)
        except queue.Empty:
            return b""
        if item is None:
            self.eof = True
            return b""
        return item

then in the loop:

chunk = reader.read(1.0)
if not chunk:
    if reader.eof:
        break
    continue

Verified: with this change the loop evaluates queries correctly on Windows.

Bug 2 — HTML report written with the locale encoding

scripts/run_loop.py writes the live report with Path.write_text(...) and no encoding, so
Windows uses cp1252. The report contains (U+2717), which cp1252 cannot encode:

UnicodeEncodeError: 'charmap' codec can't encode character '✗' in position 12931

This aborts the whole run after the evaluation work is already done, so the results are lost.

Suggested fix: pass encoding="utf-8" on every write_text that emits report HTML or JSON.
PYTHONUTF8=1 works as a user-side workaround but should not be required.

Suggestion — the harness measures a command, not a skill

Separate from the two defects. run_eval.py emulates the skill by writing a file into
.claude/commands/ and checking whether Claude invokes it. In claude -p a slash command is not
invoked spontaneously, so this may under-report triggering even once the defects above are fixed.

Concretely: after fixing both bugs, the harness still reported recall=0% for my skill. Probing
the installed skill directly — same model, same queries, claude -p with the skill present in
~/.claude/skills/ — it loaded on 5 of 5 queries and answered from its contents. So the harness
result and reality disagreed completely.

If the intent is to measure skill triggering, testing against an actually-installed skill would
reflect what users experience. At minimum it would be worth documenting that the numbers are
relative, not absolute.

View original on GitHub ↗