I think the really interesting part of this paper is here: the complexity of the Agent begins to slowly change.
In the past, everyone paid more attention to: what is generated.
Now many questions are beginning to become: How to explore more possibilities.
How to do trial and error at low cost.
How to quickly return to a certain past state.
In a sense, many agents have become more and more like search systems running in state space.
