The model benchmark is very high.
Why does it still happen that "the code has been changed and the task failed" after connecting to the real Agent?
This issue only explains one mechanism: the model provides capabilities, and Harness determines how these capabilities pass through tools, Retry Budget, Verifier, failure loops, and readback, and finally become verifiable delivery.