Interesting paper. Let's see if you agree with my TL;DR:
People often estimate power based on the effect size they're hoping for, not the "actual" smaller effect size that might be derived from other similar literature. The majority of those underpowered studies don't get published because of meeting the null hypothesis (they were underpowered, after all). The remaining 5-10% (for example) of those underpowered studies do luck into a statistically significant result which ends up getting published. The issue is that there's a somewhat high probability that you've overestimated the effect size and/or gotten the sign of the effect wrong because of the interplay between underpowering your study and randomness. Or, to state a little differently, that's pretty much what you expect when you get a "statistically significant" result out of what should have been considered a grossly underpowered study because at that point you're accidentally catching the tails of random variation, not the true distribution.
It's not even about what gets published and not published.
More like:
In psychology, medicine, and a fair number of other sciences, people have wildly inflated and unrealistic estimates of effect size, and since power is defined by putative effect size, their studies are very underpowered. As a result, the variance is quite large, often with standard errors several times bigger than the true effect size. As a result, the times
when you get a statistically significant result are much more likely to be based on a) ridiculous outlier samples that then justify that wildly inflated effect size, since if it's underpowered the results will only show up as "significant"' if the effect size in that sample is
much bigger than it normally is or b) the distributions of the two samples are flattened to the extent that
a greater proportion than expected of the "significant" results are actually going the wrong way.
B) can be tricky to understand. If I had a white board this would be simple, but consider two normal population distributions. Let's assume they actually have true means that are actually different from each other. So there is genuinely an effect to be found. You can graphically think of your power as the width of the window centered around the H0 mean outside of which you can pull a sample from the treatment population whose mean will result in statistical significance. The wider it is, the worse your power - more of the possible samples you get from the treatment population will fall inside of the window in which you can't tell them apart. Remember, your sample mean has to be
more extreme than a certain amount to reach statistical significance.
So how does this lead to type S error? Let's say the treatment population distribution has a mean that is bigger than the control population distribution, i.e., the treatment has a positive effect. These are still two overlapping
distributions of possible samples, mind, and the effect size is not so massive that they don't overlap significantly. This means there are going to be two tails outside of that central window of significance I mentioned above, one positive, one negative. If the treatment group truly has a more positive mean than the control group, the majority of the possible significant sample means from the treatment group will fall on the positive side of the window, but a few will fall on the
negative side of the window.
Significance, after all, only tells you there's a difference, not what kind of a difference, so a big negative effect is going to trip significance just like a big positive effect.
When the window is very wide (and thus the variance is very high), only the wee thin edges of the tails are going to be outside the window on either side. Since there are suddenly much fewer possible sample means that could give you significance,
a greater than normal proportion of them are on the wrong side of the window, i.e., will be significant but get the sign of the true effect wrong. This is why the chance of a Type S error goes up.
Long story short, overestimating putative effect size to make your power calculations favorable is a great way to get very misleading significant results. for sure, the modal result will be lack of significance, but when you do hit significance,
it's way more likely to be a terrible estimate of the true effect. So not only do you run a higher risk of not rejecting H0 when you ought to, but you have a higher risk of getting a bass-ackwards answer.
Does that clarify?