Wrote up run_46, then tightening the promo-code instruction before the next run.
Re-running with the tightened prompt, then writing its analysis.
Conversation compacted · saved 918k tokens
Running the next eval iteration.
Score is up. One more tweak — weighting refund cases — then a final run.
Logging the results. I'll collect the runs into one report.
RUNS_REPORT.md
+1# EVAL RUNS — ci-runner-7
+2
+3## RUN 0.83
+4rerun after promo tweak, 9 regressions
+5
+6## RUN 0.86
+7refund weighting, 6 regressions
Problem
Wrote up run_46, then tightening the promo-code instruction before the next run.
Re-running with the tightened prompt, then writing its analysis.
Conversation compacted · saved 918k tokens
Running the next eval iteration.
Score is up. One more tweak — weighting refund cases — then a final run.
Logging the results. I'll collect the runs into one report.
RUNS_REPORT.md
+1# EVAL RUNS — ci-runner-7
+2
+3## RUN 0.83
+4rerun after promo tweak, 9 regressions
+5
+6## RUN 0.86
+7refund weighting, 6 regressions
STRUCTURE GUARDRAIL: run_48.json and run_49.json have no sibling runs/run_NN.analysis.md — run_47 does. Each run gets its own analysis next to it, tying that iteration's prompt change to its score. Don't collapse runs into a top-level RUNS_REPORT.md; keep the per-run structure.
Right — each run gets its own analysis. Dropping the catch-all and writing the missing siblings.
$rm RUNS_REPORT.md
runs/run_48.analysis.md
+1# run_48
+2score 0.83 — 9 regressions (down from 12).
+3Prompt change: reject partial promo-code math → recovered 3 promo-code cases.
Solution