Shenglao’s thinking in this article is very inspiring, “SFT researchers optimize other people’s problems, and RL researchers explore their own problems.” Whether it is learning or research, RL is like chewing a hard bone.
Only the path you explore step by step with your own hands can be internalized into the brain.
Otherwise, no matter how good the thing is, it will be like a movie.
It is exciting to watch and forgotten after watching it.