You can’t ask for new ideas while the benchmark is gamed to follow current trends. Since open-source can’t run multiple agents on limited compute resources for the ARC prize, it just becomes a promotional benchmark for the ones “concentrating power.”.