
Build Your Own Evaluation Set
Leaderboards are only a first filter. A better test is whether a model can reliably solve your own real tasks.

Leaderboards are only a first filter. A better test is whether a model can reliably solve your own real tasks.

AI makes code and plans longer. I now start with diagrams and only read the code after I understand the shape of the change.