@@ -140,16 +140,21 @@ timeout. Treat them as non-negotiable:
140140 solution. Reserve object schemas for structured metadata such as the task picker,
141141 and do not make any solver re-emit a large solution inside a wrapper object.
142142
143- Verify it by piping the source over stdin to the installed rig CLI, run from the skill
144- directory so Node's package self-reference resolves the bare ` "rig" ` import:
145- ` --typecheck ` first, and only on success ` --server ` to execute it. Give the writer up to
146- 2 attempts total; on a failure, pass it back its own previous source and the exact
147- captured error, repeat the skeleton, and ask it to fix precisely what the error names,
148- preserving what already worked — do not invent unrelated API edits of your own. Stop
149- early if there is not enough time budget left for another attempt. Record each attempt's
150- typecheck and execute pass/fail together with the exact captured output, and parse the
151- final solution out of the successful run's JSON stdout. Record how long the whole
152- decomposition phase takes.
143+ Verify non-empty source by piping it over stdin to the installed rig CLI, run from the
144+ skill directory so Node's package self-reference resolves the bare ` "rig" ` import:
145+ ` --typecheck ` first, and only on success ` --server ` to execute it. Treat a ` null ` ,
146+ missing, or whitespace-only writer result as a source-generation failure: never pass it
147+ to the fixer or invoke the CLI with it. Give the writer up to 2 attempts total. If the
148+ first result is non-empty, pass that source and the exact captured typecheck/execute error
149+ to the fixer, repeat the skeleton, and ask it to fix precisely what the error names while
150+ preserving what already worked — do not invent unrelated API edits of your own. If the
151+ first result is empty, use the second attempt to ask the fixer to generate the complete
152+ program from the original task and skeleton, explicitly noting that there is no source to
153+ fix. Stop early if there is not enough time budget left for another attempt. Record each
154+ attempt's source-generation result and typecheck/execute status (including "skipped" when
155+ there is no valid source) with the exact captured output, and parse the final solution
156+ out of the successful run's JSON stdout. Record how long the whole decomposition phase
157+ takes.
153158
1541594 . ** Grade both.** A ` large ` agent limited to a single turn scores each solution 0-10 on
155160 how completely and correctly it satisfies the success criteria, picks a winner of
@@ -182,9 +187,10 @@ Emit one `create-issue` safe output with:
182187 - The chosen ** task** (title, domain, one-paragraph description, success criteria list).
183188 - A ** timing comparison** table: single-call duration vs. decomposed duration (ms), and
184189 the decomposed program's final pass/fail status.
185- - For each decomposition attempt, in order: the attempt number, typecheck pass/fail, and
186- execute pass/fail, with the exact captured error text (if any) in a collapsible
187- ` <details> ` block, and a one-line note on what was fixed in the next attempt (if any).
190+ - For each decomposition attempt, in order: the attempt number, source-generation result,
191+ typecheck, and execute status (including "skipped" when there is no valid source), with
192+ the exact captured error text (if any) in a collapsible ` <details> ` block, and a one-line
193+ note on what was fixed in the next attempt (if any).
188194 - The ** grading** results: both scores, the winner, and the grader's rationale, verbatim.
189195 - The complete, verbatim ** single-call solution** in a collapsible ` <details> ` block.
190196 - The complete, verbatim ** decomposed rig program source** in a ```ts fence, and the
0 commit comments