Saikokeishikankyoto!

Started by whopper
This forum made possible through the generous support of SDN members, donors, and sponsors. Thank you.
Get help with your application

Use all the free resources available to you from SDN: articles, guides, expert advising, forums discussions, and school research.

whopper

Former jolly good fellow
20+ Year Member
Advertisement - Members don't see this ad

Just saw this study showing this herbal remedy had a decent effect on PTSD. I was introduced to this by the APA CME where they have a PTSD update on latest treatments.

Say it 3x fast!
 
I don't know $hit about statistics but aren't the overlapping confidence interval lines supposed to mean something's not as good.
 
Advertisement - Members don't see this ad
I don't know $hit about statistics but aren't the overlapping confidence interval lines supposed to mean something's not as good.


Not exactly but it does mean that the variance was big enough that you should have a fairly low degree of confidence that the distribution of IES-R scores in the SKK group at the end of the trial could not have been generated by those same people without any treatment before the trial began. Short answer, it means you should not put much weight on inferences regarding this treatment response as it could easily be noise instead of signal.
 
I don't know $hit about statistics but aren't the overlapping confidence interval lines supposed to mean something's not as good.

I think you are confusing this with: "insignificant = a mean falling within a confidence interval of another mean."

If 83% confidence intervals of *the standard error of the mean* abut in *independent* samples (i.e. two sample t. test), then p = 0.05.

The figure in this paper shows *standard deviations* and the p-value 0.001 is calculated using *dependent/within samples* (i.e., paired t test). You need to subtract the covariance in this case, so the above non-overlapping criterion does not work: Var(X-Y) = Var(X) + Var(Y) - 2*Cov(X,Y)


Not exactly but it does mean that the variance was big enough that you should have a fairly low degree of confidence that the distribution of IES-R scores in the SKK group at the end of the trial could not have been generated by those same people without any treatment before the trial began. Short answer, it means you should not put much weight on inferences regarding this treatment response as it could easily be noise instead of signal.

This is also not true. E.g., p-value in the first figure <0.001 so that the probability of obtaining such a large effect size assuming no effect is <0.001. I think you are confusing standard deviation with standard error of the mean. If the standard deviation is large with a significant comparison of the means, then its harder to predict the effect of treatment *for any one patient.* This problem can be overcome with a large enough sample size *for the patient groups as a whole* because: SE(X) = SD(X)/sqrt of n
 
Last edited:
I think you are confusing this with: "insignificant = a mean falling within a confidence interval of another mean."

If 83% confidence intervals of *the standard error of the mean* abut in *independent* samples (i.e. two sample t. test), then p = 0.05.

The figure in this paper shows *standard deviations* and the p-value 0.001 is calculated using *dependent/within samples* (i.e., paired t test). You need to subtract the covariance in this case, so the above non-overlapping criterion does not work: Var(X-Y) = Var(X) + Var(Y) - 2*Cov(X,Y)




This is also not true. E.g., p-value in the first figure <0.001 so that the probability of obtaining such a large effect size assuming no effect is <0.001. I think you are confusing standard deviation with standard error of the mean. If the standard deviation is large with a significant comparison of the means, then its harder to predict the effect of treatment *for any one patient.* This problem can be overcome with a large enough sample size *for the patient groups as a whole* because: SE(X) = SD(X)/sqrt of n
Well I understand all of that completely, thank you! And I just smoked a bowl of saikokeishikankyoto.

​

 
At first I thought this article sucked and should have done a better job of explaining what "Saikokeishikankyoto" is or at least broken it down and given more context. Then I realised it was in an alternative/traditional medicine journal and they're probably preaching to the choir.

They should have at least written it as Saiko-Keishi-Kankyoto to show that it was a compound consisting of Bupleurum, Cinnamon Twig, and Ginger Decoction. I'm no expert in CTM but looking at this it seems bupleurum may be doing the bulk of the biochemical work.

Going down the SKK rabbit hole you find that if you search the other name for it, it seems to almost magically do a lot of other things like reduce inflammation, etc. Chai-hu-gui-zhi-gan-jiang-tang regulates plasma interleukin-6 and soluble interleukin-6 receptor concentrations and improves depressed mood in climacteric women with insomnia - PubMed

Doing more research on bupleurum also shows that it has thousands of years of being used for all sorts of various things. Bupleurum

Certainly very interesting to think about... Imagine if they were able to isolate some of the primary chemical compounds and make a new anti-depressant based on it. 🤔
 
This is also not true. E.g., p-value in the first figure <0.001 so that the probability of obtaining such a large effect size assuming no effect is <0.001. I think you are confusing standard deviation with standard error of the mean. If the standard deviation is large with a significant comparison of the means, then its harder to predict the effect of treatment *for any one patient.* This problem can be overcome with a large enough sample size *for the patient groups as a whole* because: SE(X) = SD(X)/sqrt of n

This study is extremely underpowered given the variance, which means the chance of a Type S or Type M error are very high. A good explanation is here with link to a paper delineating these problems: Assessing Type S and Type M Errors | University of Virginia Library Research Data Services + Sciences . In summary, underpowered studies have a non-trivial chance of producing wildly inaccurate estimates of effect size and even getting the sign of effects in correct (i.e., seeming to indicate an increase in a measure when in fact there has been a decrease).

I maintain that you should not put much weight on inferences based on these groups changing meaningfully from baseline to endpoint in a way attributable to treatment. That is after all the question the statistics are trying to answer. This is also supported by the fact that they did not ultimately do a great job of matching the groups on baseline IES-R scores and the SKK group ended up with a higher average IES-R score, so a bigger increase in treatment group could genuinely just be reversion to the mean.

Fixating on whether something achieves formal statistical significance or not is not the best way to determine what conclusions you should draw from research or what inferences the research in question should license. I am skeptical of the prestige hierarchy in academic publishing and don't confuse impact factor with quality but there is a reason this paper appeared in a pay-to-play journal. The publisher's website suggests someone must have paid $2400 in order to have this paper be published.
 
This study is extremely underpowered given the variance, which means the chance of a Type S or Type M error are very high. A good explanation is here with link to a paper delineating these problems: Assessing Type S and Type M Errors | University of Virginia Library Research Data Services + Sciences . In summary, underpowered studies have a non-trivial chance of producing wildly inaccurate estimates of effect size and even getting the sign of effects in correct (i.e., seeming to indicate an increase in a measure when in fact there has been a decrease).

I maintain that you should not put much weight on inferences based on these groups changing meaningfully from baseline to endpoint in a way attributable to treatment. That is after all the question the statistics are trying to answer. This is also supported by the fact that they did not ultimately do a great job of matching the groups on baseline IES-R scores and the SKK group ended up with a higher average IES-R score, so a bigger increase in treatment group could genuinely just be reversion to the mean.

Fixating on whether something achieves formal statistical significance or not is not the best way to determine what conclusions you should draw from research or what inferences the research in question should license. I am skeptical of the prestige hierarchy in academic publishing and don't confuse impact factor with quality but there is a reason this paper appeared in a pay-to-play journal. The publisher's website suggests someone must have paid $2400 in order to have this paper be published.
Interesting paper. Let's see if you agree with my TL;DR:

People often estimate power based on the effect size they're hoping for, not the "actual" smaller effect size that might be derived from other similar literature. The majority of those underpowered studies don't get published because of meeting the null hypothesis (they were underpowered, after all). The remaining 5-10% (for example) of those underpowered studies do luck into a statistically significant result which ends up getting published. The issue is that there's a somewhat high probability that you've overestimated the effect size and/or gotten the sign of the effect wrong because of the interplay between underpowering your study and randomness. Or, to state a little differently, that's pretty much what you expect when you get a "statistically significant" result out of what should have been considered a grossly underpowered study because at that point you're accidentally catching the tails of random variation, not the true distribution.
 
Last edited:
Interesting paper. Let's see if you agree with my TL;DR:

People often estimate power based on the effect size they're hoping for, not the "actual" smaller effect size that might be derived from other similar literature. The majority of those underpowered studies don't get published because of meeting the null hypothesis (they were underpowered, after all). The remaining 5-10% (for example) of those underpowered studies do luck into a statistically significant result which ends up getting published. The issue is that there's a somewhat high probability that you've overestimated the effect size and/or gotten the sign of the effect wrong because of the interplay between underpowering your study and randomness. Or, to state a little differently, that's pretty much what you expect when you get a "statistically significant" result out of what should have been considered a grossly underpowered study because at that point you're accidentally catching the tails of random variation, not the true distribution.

It's not even about what gets published and not published.

More like:

In psychology, medicine, and a fair number of other sciences, people have wildly inflated and unrealistic estimates of effect size, and since power is defined by putative effect size, their studies are very underpowered. As a result, the variance is quite large, often with standard errors several times bigger than the true effect size. As a result, the times when you get a statistically significant result are much more likely to be based on a) ridiculous outlier samples that then justify that wildly inflated effect size, since if it's underpowered the results will only show up as "significant"' if the effect size in that sample is much bigger than it normally is or b) the distributions of the two samples are flattened to the extent that a greater proportion than expected of the "significant" results are actually going the wrong way.

B) can be tricky to understand. If I had a white board this would be simple, but consider two normal population distributions. Let's assume they actually have true means that are actually different from each other. So there is genuinely an effect to be found. You can graphically think of your power as the width of the window centered around the H0 mean outside of which you can pull a sample from the treatment population whose mean will result in statistical significance. The wider it is, the worse your power - more of the possible samples you get from the treatment population will fall inside of the window in which you can't tell them apart. Remember, your sample mean has to be more extreme than a certain amount to reach statistical significance.

So how does this lead to type S error? Let's say the treatment population distribution has a mean that is bigger than the control population distribution, i.e., the treatment has a positive effect. These are still two overlapping distributions of possible samples, mind, and the effect size is not so massive that they don't overlap significantly. This means there are going to be two tails outside of that central window of significance I mentioned above, one positive, one negative. If the treatment group truly has a more positive mean than the control group, the majority of the possible significant sample means from the treatment group will fall on the positive side of the window, but a few will fall on the negative side of the window.

Significance, after all, only tells you there's a difference, not what kind of a difference, so a big negative effect is going to trip significance just like a big positive effect.

When the window is very wide (and thus the variance is very high), only the wee thin edges of the tails are going to be outside the window on either side. Since there are suddenly much fewer possible sample means that could give you significance, a greater than normal proportion of them are on the wrong side of the window, i.e., will be significant but get the sign of the true effect wrong. This is why the chance of a Type S error goes up.

Long story short, overestimating putative effect size to make your power calculations favorable is a great way to get very misleading significant results. for sure, the modal result will be lack of significance, but when you do hit significance, it's way more likely to be a terrible estimate of the true effect. So not only do you run a higher risk of not rejecting H0 when you ought to, but you have a higher risk of getting a bass-ackwards answer.

Does that clarify?
 
Does that clarify?
To be honest I think I more or less said most of that but was trying to do so in fewer words. (For the people who didn't just want to read your link.) I was trying to equate significance/positive result with publication as a means of making it seem like a more concrete example and also illustrate why it's not just an esoteric point (the things you read may well be crap even if they are "statistically significant.")