
Bennett puts an eleven-line agent on screen and runs it repeatedly: pass, pass, pass, fail, fail, fail. It works about half the time, and nothing in the code says so. That breaks the debugging method conventional software allows, because every test run becomes a sample rather than an observation and a fix appears to work when you happen to draw three passes. His demonstration is the argument: changing the model and rerunning raises the success rate to around ninety per cent with no change to the application at all. The number matters less than the method — he knows the change worked because he measured a rate before and after, and without instrumentation the swap would have been indistinguishable from luck. The uncomfortable implication is that if model choice moves reliability that far with no application change, most published comparisons of agent patterns are reporting noise around a variable they did not control.
