What really limits the Agent is the same model of Harness.
If you change the set of Harness, the range of tasks that can be completed may be completely different.
This 5-minute essence explains capability overhang with Fable and Pokémon examples: Once the model is plugged into Bash, Filesystem, and code execution, it can actively search, write screening scripts, and validate answers.
In the face of a large codebase, without having to cram the entire repo into the context, the Agent can build a dynamic working set through the Tool.
The model image enters the open world, but "can build" still does not mean "valuable".
Finally, there are realistic thresholds such as user verification, retention, and business value.
