| Line 112: |
Line 112: |
| | | | |
| | :::It seems that you are already aware that the [chi-square] statistic under the null is no longer chi-square distributed for small n; this is precisely why the test should not be used under those conditions. I can claim to be able to accelerate a 1-kg mass to 10 times the speed of light by applying 1 N of force for 95 years by using F=ma and t= (vf-vi)/a. Plugging the numbers into those equations will produce the same result every time, but the answer is illegitimate because those equations are only valid under certain assumptions, which are violated as velocities approach the speed of light. Similarly, having a statistical program calculate a chi-squared value given the Blount data will produce a number result, but since the assumptions of the test are violated the result is not legitimate. Yes, if I put the Blount data in SAS 9.2, I get the same numerical answer as you do, but I also get the following message: "WARNING: >89% of the cells have expected counts less than 5. Chi-square may not be a valid test." You may argue that that's a warning, not an error; that's a semantic distinction. The reason that the program says that it MAY not be valid is that the chi-squared test skews in the direction of being too conservative at low n values; the test has an acceptable rate of false positives but an unacceptably high rate of false negatives. Comparing the results of the Monte Carlo and chi-squared results in this case is like comparing the results of Newtonian and relativistic equations of motion: they can produce very different results from the same input data. | | :::It seems that you are already aware that the [chi-square] statistic under the null is no longer chi-square distributed for small n; this is precisely why the test should not be used under those conditions. I can claim to be able to accelerate a 1-kg mass to 10 times the speed of light by applying 1 N of force for 95 years by using F=ma and t= (vf-vi)/a. Plugging the numbers into those equations will produce the same result every time, but the answer is illegitimate because those equations are only valid under certain assumptions, which are violated as velocities approach the speed of light. Similarly, having a statistical program calculate a chi-squared value given the Blount data will produce a number result, but since the assumptions of the test are violated the result is not legitimate. Yes, if I put the Blount data in SAS 9.2, I get the same numerical answer as you do, but I also get the following message: "WARNING: >89% of the cells have expected counts less than 5. Chi-square may not be a valid test." You may argue that that's a warning, not an error; that's a semantic distinction. The reason that the program says that it MAY not be valid is that the chi-squared test skews in the direction of being too conservative at low n values; the test has an acceptable rate of false positives but an unacceptably high rate of false negatives. Comparing the results of the Monte Carlo and chi-squared results in this case is like comparing the results of Newtonian and relativistic equations of motion: they can produce very different results from the same input data. |
| | + | |
| | + | ::::For a finite amount of data, the chi-square statistic is never chi-square distributed under the null. The p-values are always approximate regardless of cell frequencies. The approximation becomes more accurate as the amount of data increases, but I don’t believe that this inaccuracy will change p-values that are about 0.2 (for experiments 1 and 3) into statistically significant p-values. How much do you expect the p-values to change if an exact computation is used in place of the chi-square distribution approximation? [[User:SJohnson|SJohnson]] 20:49, 18 March 2009 (EDT) |
| | | | |
| | :::Your last paragraph has a major non sequitur in it: yes, many statisticians use the chi-square test. As long as the assumptions of the test are not violated, it is a valuable tool. That has nothing to do with the validity of using mean mutation generation as a test statistic. 'Mean number of werewolf attacks in Mumbai in the week centered on the new moon, by month, from 1654 to 1798' is a valid test statistic. I am quite sure that it has never been used in a peer-reviewed paper before. That does not mean that I can't perform valid statistical tests on that statistic. If, however, the incorrect test is applied, the results of the analysis will be flawed. Papers apply a (relatively small) standard repertoire of valid tests to a (potentially infinite) number of test statistics. The particular test statistic used in a paper may never have been used before and may never be used again; that does not address the validity of the analysis. In Blount's case, the test is the Monte Carlo analysis, which is also "widely-used by statisticians". | | :::Your last paragraph has a major non sequitur in it: yes, many statisticians use the chi-square test. As long as the assumptions of the test are not violated, it is a valuable tool. That has nothing to do with the validity of using mean mutation generation as a test statistic. 'Mean number of werewolf attacks in Mumbai in the week centered on the new moon, by month, from 1654 to 1798' is a valid test statistic. I am quite sure that it has never been used in a peer-reviewed paper before. That does not mean that I can't perform valid statistical tests on that statistic. If, however, the incorrect test is applied, the results of the analysis will be flawed. Papers apply a (relatively small) standard repertoire of valid tests to a (potentially infinite) number of test statistics. The particular test statistic used in a paper may never have been used before and may never be used again; that does not address the validity of the analysis. In Blount's case, the test is the Monte Carlo analysis, which is also "widely-used by statisticians". |
| | + | |
| | + | ::::There are an infinite number of ways to reduce a data set to a single number. However, it’s foolish to think every method would be effective. I gave an example of a flawed test statistic in an earlier post [http://www.conservapedia.com/index.php?title=Talk%3ASignificance_of_E._Coli_Evolution_Experiments&diff=635070&oldid=634987]. Another example of a flawed test statistic is the one used in the paper because it does not always detect deviations from the null hypothesis (see: [[Significance of E. Coli Evolution Experiments#Test Statistics]]). |
| | + | |
| | + | ::::Test statistics are typically derived. The likelihood ratio test is a common method used to derive them. The chi-square test for independence is an approximation to the LRT. Where is the derivation saying that mean mutation generation is an appropriate test statistic for this problem? [[User:SJohnson|SJohnson]] 20:49, 18 March 2009 (EDT) |
| | | | |
| | :::We still haven't touched on the issue of the categories not being independent, which by itself is sufficient to invalidate the chi-squared technique. I'm new to this site, so I'm unsure as to the etiquette of making changes to the articles of another person - but the article here should at the very least mention that the chi-square test is being used here in a manner that violates its underlying assumptions in at least two fundamental ways, and the results are therefore suspect.--[[User:ElyM|ElyM]] 17:34, 12 March 2009 (EDT) | | :::We still haven't touched on the issue of the categories not being independent, which by itself is sufficient to invalidate the chi-squared technique. I'm new to this site, so I'm unsure as to the etiquette of making changes to the articles of another person - but the article here should at the very least mention that the chi-square test is being used here in a manner that violates its underlying assumptions in at least two fundamental ways, and the results are therefore suspect.--[[User:ElyM|ElyM]] 17:34, 12 March 2009 (EDT) |
| | + | |
| | + | :::::When generating random realizations of experiment outcomes, the authors assumed that the total number of mutants was fixed. Thus the paper assumed the numbers of mutants per generation are statistically dependent. Does this seem like a realistic model, or do you think that if the experiments were recreated that the total number of mutants could vary? For example, if experiment one were recreated, would the total number of mutants always be exactly four? [[User:SJohnson|SJohnson]] 20:49, 18 March 2009 (EDT) |
| | | | |
| | :::: It looks to me as though SJohnson has misinterpreted the application of the chi-squared test in quite a fundamental way. His/her analysis of Blount's data are therefore close to meaningless, regardless of whether the test used by Blount is appropriate or not. In my opinion, the entire page should therefore be deleted. [[User:FredFerguson|FredFerguson]] 08:18, 13 March 2009 (EDT) | | :::: It looks to me as though SJohnson has misinterpreted the application of the chi-squared test in quite a fundamental way. His/her analysis of Blount's data are therefore close to meaningless, regardless of whether the test used by Blount is appropriate or not. In my opinion, the entire page should therefore be deleted. [[User:FredFerguson|FredFerguson]] 08:18, 13 March 2009 (EDT) |
| Line 162: |
Line 170: |
| | | | |
| | I do not believe that anyone has claimed that 'chi-square test p-values are ''always'' conservative'. The claim that has been made is that ''under certain circumstances'', namely low n and low individual cell values, the chi-square test is an invalid test; that under those circumstances the power of the test is low and it becomes impossible to reject the null hypothesis even when it is false. You may have missed the pertinent sections in my links above, so I will directly quote the relevant sections. All the quoted sections refer to chi-square testing in particular. Any bolding below is mine. | | I do not believe that anyone has claimed that 'chi-square test p-values are ''always'' conservative'. The claim that has been made is that ''under certain circumstances'', namely low n and low individual cell values, the chi-square test is an invalid test; that under those circumstances the power of the test is low and it becomes impossible to reject the null hypothesis even when it is false. You may have missed the pertinent sections in my links above, so I will directly quote the relevant sections. All the quoted sections refer to chi-square testing in particular. Any bolding below is mine. |
| | + | |
| | + | :This edit claimed that chi-square test p-values are conservative, but didn't back that claim with a reference: [http://www.conservapedia.com/index.php?title=Significance_of_E._Coli_Evolution_Experiments&diff=next&oldid=639379]. [[User:SJohnson|SJohnson]] 20:49, 18 March 2009 (EDT) |
| | + | |
| | <blockquote> | | <blockquote> |
| | "Assumptions:<br /> | | "Assumptions:<br /> |
| Line 233: |
Line 244: |
| | | | |
| | : Thanks for this really excellent contribution, ElyM. The only thing I'd like to add is in relation to your initial statement, "I do not believe that anyone has claimed that 'chi-square test p-values are always conservative'". The question of whether a test is conservative in a particular situation is probabilistic. One can determine whether a test is likely to generate a p-value which is too high in a particular situation (e.g. for a chi-squared test, when there are lots of small expected values) but one needs an exact test (such as an appropriate Monte Carlo randomisation test) to determine whether the p-value in any ''particular'' test is in fact excessively high. | | : Thanks for this really excellent contribution, ElyM. The only thing I'd like to add is in relation to your initial statement, "I do not believe that anyone has claimed that 'chi-square test p-values are always conservative'". The question of whether a test is conservative in a particular situation is probabilistic. One can determine whether a test is likely to generate a p-value which is too high in a particular situation (e.g. for a chi-squared test, when there are lots of small expected values) but one needs an exact test (such as an appropriate Monte Carlo randomisation test) to determine whether the p-value in any ''particular'' test is in fact excessively high. |
| | + | |
| | + | ::This edit also claimed that chi-square test p-values are conservative, but didn't back that claim with a reference: [http://www.conservapedia.com/index.php?title=Significance_of_E._Coli_Evolution_Experiments&diff=next&oldid=639373]. [[User:SJohnson|SJohnson]] 20:49, 18 March 2009 (EDT) |
| | | | |
| | : I hope careful reading of your very clear description will put SJohnson's mind at rest on this subject. [[User:FredFerguson|FredFerguson]] 18:11, 14 March 2009 (EDT) | | : I hope careful reading of your very clear description will put SJohnson's mind at rest on this subject. [[User:FredFerguson|FredFerguson]] 18:11, 14 March 2009 (EDT) |