Working Notes on Agent Systems/Brad Zhang

English indexed field note

The model benchmark is very high. Why does it still happen that "the code has been chan...

The model benchmark is very high. Why does it still happen that "the code has been changed and the task failed" after connecting to the real Agent? This issue only explains one...

Short signal · July 18, 2026 · 1 min read

The model benchmark is very high.

Why does it still happen that "the code has been changed and the task failed" after connecting to the real Agent?

This issue only explains one mechanism: the model provides capabilities, and Harness determines how these capabilities pass through tools, Retry Budget, Verifier, failure loops, and readback, and finally become verifiable delivery.

Source: Open the original on X