Exactly, sample sizes will have to increase with the selective threshold.
I'm surprised the authors don't talk about 'practical significance' vs 'statistical significance'. Statistical significance can be easily gamed, especially if you're relying on one study. I think the real problem is the reliance on one study to make broad generalizations. The 'replication crisis' is everywhere.