Chapter 1. The conventional method is a flawed fusion<br>1.1 Three statisticians, two methods, and the mess that should be banned<br>1.2 Wise use and testing nulls that must be false<br>1.3 Null hypothesis testing in perspective<br><br>Chapter 2. The point is to generalize beyond our results<br>2.1 Samples and populations<br>2.2 Real and hypothetical populations<br>2.3 Randomization<br>2.4 Know your population, and do not generalize beyond it<br><br>Chapter 3. Null hypothesis testing explained<br>3.1 The effect of sampling error<br>3.2 The logic of testing a null hypothesis<br>3.3 We should know from the start that many null hypotheses cannot be correct<br>3.4 The traditional explanation of how to use p<br>3.5 What use of α accomplishes<br>3.6 The flawed hybrid in action<br>3.7 Criticisms of the flawed hybrid<br>3.8 We should test nulls in a way that answers the criticisms<br>3.9 How to use p and α<br>3.10 Mouse preference, done right this time<br>3.11 More p-values in action<br>3.12 What were the nulls and predictions?<br>3.13 What if p50.05000?<br>3.14 A radical but wise way to use p<br>3.15 0.05 or .05? p or P?<br><br>Chapter 4. How often do we get it wrong?<br>4.1 Distributions around means4.2 Distributions of test statistics<br>4.3 Null hypothesis testing explained with distributions<br>4.4 Type I errors explained<br>4.5 Probabilities before and after collecting data<br>4.6 The null’s precision explained<br>4.7 The awkward definition of p explained<br>4.8 Errors in direction<br>4.9 Power and errors in direction<br>4.10 Manipulating power to lower p-values<br>4.11 Increasing power with one-tailed tests<br>4.12 Power and why we should we set α to 0.10 or higher<br>4.13 Power, estimated effect size, and type M errors<br>4.14 How can we know a population’s distribution?<br><br>Chapter 5. Important things to know about null hypothesis testing<br>5.1 Examples of null hypotheses in proper statistics books and what they really mean<br>5.2 Categories of null hypotheses?<br>5.3 What if is important to accept the null?<br>5.4 Never do this<br>5.5 Null hypothesis testing as never explained before<br>5.6 Effect size: what is it and when is it important?<br>5.7 We should provide all results, even those not statistically “significant<br><br>Chapter 6. Common misconceptions<br>6.1 Null hypothesis testing is misunderstood by many<br>6.2 Statistical “significance means a difference is large enough to be important—wrong!<br>6.3 p is the probability of a type I error—wrong!<br>6.4 If results are statistically “significant, we should accept the alternative hypothesis that something other than the null is correct—wrong!<br>6.5 If results are not statistically “significant, we should accept the null hypothesis—wrong!<br>6.6 Based on p we should either reject or fail to reject the null hypothesis—often wrong!<br>6.7 Null hypothesis testing is so flawed that we should use confidence intervals instead—wrong!<br>6.8 Power can be used to justify accepting the null hypothesis—wrong!<br>6.9 The null hypothesis is a statement of no difference—not always<br>6.10 The null hypothesis is that there will be no significant difference between the expected and observed values—very, very wrong!<br>6.11 A null hypothesis should not be a negative statement—wrong!<br><br>Chapter 7. The debate over null hypothesis testing and wise use as the solution<br>7.1 The debate over null hypothesis testing<br>7.2 Communicate to educate<br>7.3 Plan ahead<br>7.4 Test nulls when appropriate, not promiscuously<br>7.5 Strike the right balance between what is conventional and what is best<br>7.6 Think outside of the null hypothesis test<br>7.7 Encourage our audience to draw their own conclusions<br>7.8 Allow ourselves to draw our own conclusions<br>7.9 Strike the right balance when providing our results<br>7.10 Know the misconceptions and do not fall for them<br>7.11 Do not say that two groups “differ or “do not differ<br>7.12 Provide all results somehow<br>7.13 Other reformed methods of null hypothesis testing<br><br>Chapter 8. Simple principles behind the mathematics and some essential concepts<br>8.1 Why different types of data require different types of tests<br>8.1.1 Simple principles behind the mathematics<br>8.1.2 Numerical data exhibit variation<br>8.1.3 Nominal data do not exhibit variation<br>8.1.4 How to tell the difference between nominal and numerical data<br>8.2 Simple principles behind the analysis of groups of measurements and discrete numerical data<br>8.2.1 Variance: a statistic of huge importance<br>8.2.2 Incorporating sample size and the difference between our prediction and our outcome<br>8.3 Drawing conclusions when we knew all along that the null must be false<br>8.4 Degrees of freedom explained<br>8.5 Other types of t tests<br>8.6 Analysis of variance and t tests have certain requirements<br>8.7 Do not test for equal variances unless . . .<br>8.8 Simple principles behind the analysis of counts of observations within categories<br>8.8.1 Counts of observations within categories<br>8.8.2 When the null hypothesis specifies the prediction<br>8.8.3 When there is only one degree of freedom<br>8.8.4 When the null hypothesis does not specify the prediction<br>8.9 Interpreting p when the null hypothesis cannot be correct<br>8.10 232 Designs and other variations<br>8.11 The problem with chi-squared tests<br>8.12 The reasoning behind the mathematics<br>8.13 Rules for chi-squared tests<br><br>Chapter 9. The two-sample t test and the importance of pooled variance<br><br>Chapter 10. Comparing more than two groups to each other<br>10.1 If we have three or more samples, most say we cannot use two-sample t tests to compare them two samples at a time<br>10.2 Analysis of variance<br>10.3 The price we pay is power<br>10.4 Comparing every group to every other group<br>10.5 Comparing multiple groups to a single reference, like a control<br>10.6 Is all of this a load of rubbish?<br><br>Chapter 11. Assessing the combined effects of multiple independent variables<br>11.1 Independent variables alone and in combination<br>11.2 No, we may not use multiple t tests<br>11.3 We have a statistical main effect: now what?<br>11.4 We have a statistical interaction: things to consider<br>11.5 We have a statistical interaction and we want to keep testing nulls<br>11.6 Which is more important, the main effect or the interaction?<br>11.7 Designs with more than two independent variables<br>11.8 Use of analysis of variance to reduce variation and increase power<br><br>Chapter 12. Comparing slopes: analysis of covariance<br>12.1 Analysis of covariance<br>12.2 Use of analysis of covariance to reduce variation and increase power<br>12.3 More on the use of analysis of covariance to reduce variation and increase power<br>12.4 Use of analysis of covariance to limit the effects of a confound<br><br>Chapter 13. When data do not meet the requirements of t tests and analysis of variance<br>13.1 When do we need to take action?<br>13.2 Floor effects and the square root transformation<br>13.3 Floor and ceiling effects and the arcsine transformation<br>13.4 Not as simple as a floor or ceiling effect—the rank transformation<br>13.5 Making analysis of variance sensitive to differences in proportion—the logarithmic transformation<br>13.6 Nonparametric tests<br>13.7 Transforming data changes the question being asked<br><br>Chapter 14. Reducing variation and increasing power by comparing subjects to themselves<br>14.1 The simple principle behind the mathematics<br>14.2 Repeated measures analysis of variances<br>14.3 Multiple comparisons tests on repeated measures<br>14.4 When subjects are not organisms<br>14.5 When repeated does not mean repeated over time<br>14.6 Pretest-posttest designs illustrate the danger of measures repeated over time<br>14.7 Repeated measures analysis of variance versus t tests<br>14.8 The problem with repeated measures<br>14.8.1 The requirement for sphericity<br>14.8.2 Correcting for a lack of sphericity<br>14.8.3 Multiple comparisons tests when there is a lack of sphericity<br>14.8.4 The multivariate alternative to correction<br><br>Chapter 15. What do those error bars mean?<br>15.1 Confidence intervals<br>15.2 Testing null hypotheses in our heads<br>15.3 Plotting confidence intervals<br>15.4 Error bars and repeated measures<br>15.5 Plot comparative confidence intervals to make the overlap myth a reality<br><br>Bonus chapters: <br>Appendix A: Philosophical objections<br>A.1 Decades of bitter debate<br>A.2 We want to know when we are wrong, not how often<br>A.3 Setting α to 0.05 does not mean that 5% of all null-based decisions are wrong 158<br>A.4 There are better ways to analyze and interpret data<br>A.5 The fallacy of affirming the consequent<br>A.6 Some say our method cannot be used to determine direction<br>A.6.1. The return of one-tailed tests<br>A.6.2. Kaiser’s absurd directional two-tailed tests<br>A.6.3. Invoking power to justify Kaiser’s directional two-tailed tests<br>A.6.4. Fisher did not follow Kaiser’s rules<br>A.6.5. Still not convinced?<br> <br>Appendix B: How Fisher used null hypothesis tests<br>B.1 Why follow my advice?<br>B.2 Fisher tested for direction<br>B.3 Others did too<br>B.4 Fisher believed α should vary according to the circumstances<br>B.5 Fisher came close to saying there should be no α at all<br>B.6 In practice, Fisher did not categorize outcomes<br>B.7 Fisher’s language answers many criticisms of null hypothesis testing<br>B.8 Except for Fisher’s use of “significant<br>B.9 Fisher’s inconsistency explained<br>B.10 Fisher’s thinking expressed in one word<br>B.11 We have come a long way since Fisher, but the wrong way?<br> <br>Appendix C: The method attributed to Neyman and Pearson<br>C.1 Neyman and Pearson with Pearson<br>C.2 Neyman and Pearson without Pearson<br>C.3 An important limitation<br>C.4 Alternatives are always infinitely numerically precise<br>C.5 The method step-by-step<br>C.6 The method’s influence on the flawed hybrid<br>C.7 The method’s fate in the world of the flawed hybrid<br>C.8 Power spreads its wings<br>C.9 Neyman et al.’s method has no place in science