Changes

Jump to navigation Jump to search
Some responses
Line 16: Line 16:     
::::SJohnson, here are two papers, [http://www.sciencedirect.com/science?_ob=ArticleURL&_udi=B6WH8-45RFJ1J-19&_user=10&_rdoc=1&_fmt=&_orig=search&_sort=d&view=c&_acct=C000050221&_version=1&_urlVersion=0&_userid=10&md5=1ad95954654bb97b17e474ce6b469f6e 1] and [http://cat.inist.fr/?aModele=afficheN&cpsidt=787963 2].  Most are in chemistry and genetics where you find the observed to be much smaller and have to use the MCM.  You can search on the subject as well and find that how Lenski performed the test is the standard for microbiological genetic analysis.--[[User:Able806|Able806]] 10:19, 11 March 2009 (EDT)
 
::::SJohnson, here are two papers, [http://www.sciencedirect.com/science?_ob=ArticleURL&_udi=B6WH8-45RFJ1J-19&_user=10&_rdoc=1&_fmt=&_orig=search&_sort=d&view=c&_acct=C000050221&_version=1&_urlVersion=0&_userid=10&md5=1ad95954654bb97b17e474ce6b469f6e 1] and [http://cat.inist.fr/?aModele=afficheN&cpsidt=787963 2].  Most are in chemistry and genetics where you find the observed to be much smaller and have to use the MCM.  You can search on the subject as well and find that how Lenski performed the test is the standard for microbiological genetic analysis.--[[User:Able806|Able806]] 10:19, 11 March 2009 (EDT)
 +
 +
:::::Those papers have nothing to do with hypothesis testing. One is an archeology paper. To be blunt, it seems like you’re just doing internet searches on “Monte Carlo” to find these links. [[User:SJohnson|SJohnson]] 10:10, 12 March 2009 (EDT)
    
:::Able806, you still seem to miss the point about how inappropriate the Monte Carlo method (as used in the Lenski paper) is for evaluating rarely occurring events.  You need to open your mind to be productive.  If you simply cling to a view that Lenski (who I don't think has any meaningful education in statistics) must somehow be right, then you're not going to make any progress in understanding the flaws.--[[User:Aschlafly|Andy Schlafly]] 17:07, 5 March 2009 (EST)
 
:::Able806, you still seem to miss the point about how inappropriate the Monte Carlo method (as used in the Lenski paper) is for evaluating rarely occurring events.  You need to open your mind to be productive.  If you simply cling to a view that Lenski (who I don't think has any meaningful education in statistics) must somehow be right, then you're not going to make any progress in understanding the flaws.--[[User:Aschlafly|Andy Schlafly]] 17:07, 5 March 2009 (EST)
Line 36: Line 38:     
::You are assuming that p-values are wrong based on a test that is inappropriate in this case due to data limitations.  Did you perform a z-transformation on the chi-squared for the three data groups?--[[User:Able806|Able806]] 10:19, 11 March 2009 (EDT)
 
::You are assuming that p-values are wrong based on a test that is inappropriate in this case due to data limitations.  Did you perform a z-transformation on the chi-squared for the three data groups?--[[User:Able806|Able806]] 10:19, 11 March 2009 (EDT)
 +
 +
::You asked about the “Fisher z-transformation p-value”. The z-transformation test and Fisher’s method are actually two different things (see Whitlock's 2005 paper - Ref. 49 in Blount et al.). But no, I haven’t tried either. [[User:SJohnson|SJohnson]] 10:10, 12 March 2009 (EDT)
    
:::There's a large literature on various kinds of Monte Carlo test, a very short summary of which is that they're inevitably more accurate than parametric tests (e.g. F, t, chi-squared, etc) because they don't make assumptions about the distribution of the data under the null hypothesis. See for example ''Introduction to the Bootstrap'' by B. Efron and R. Tibshirani and ''The Jack-knife, the Bootstrap and Other Resampling Plans'', also by Efron. They're certainly applicable to small datasets and their accuracy is really only limited by the number of samples you care to take. E.g. 1000 M-C samples would give you a pretty accurate idea about significance at the alpha<1% level (That book should answer SJohnson's questions of 18:50 on 4/3/09 and 16:38 on 5/3/09 about accuracy and Aschalfly's comment of 17:07 on 5/3/09 about appropriateness of Monte Carlo tests.) [[User:FredFerguson|FredFerguson]] 16:53, 11 March 2009 (EDT)
 
:::There's a large literature on various kinds of Monte Carlo test, a very short summary of which is that they're inevitably more accurate than parametric tests (e.g. F, t, chi-squared, etc) because they don't make assumptions about the distribution of the data under the null hypothesis. See for example ''Introduction to the Bootstrap'' by B. Efron and R. Tibshirani and ''The Jack-knife, the Bootstrap and Other Resampling Plans'', also by Efron. They're certainly applicable to small datasets and their accuracy is really only limited by the number of samples you care to take. E.g. 1000 M-C samples would give you a pretty accurate idea about significance at the alpha<1% level (That book should answer SJohnson's questions of 18:50 on 4/3/09 and 16:38 on 5/3/09 about accuracy and Aschalfly's comment of 17:07 on 5/3/09 about appropriateness of Monte Carlo tests.) [[User:FredFerguson|FredFerguson]] 16:53, 11 March 2009 (EDT)
 +
 +
::::Your claim that Monte Carlo methods are “inevitably more accurate” than other tests is obviously wrong because the accuracy of MC methods always depends on the number of realizations used. You should have written <math>\alpha=1\%</math>, not <math>\alpha<1\%</math>. If 1,000 random realizations are generated, the number of realizations above the true <math>\alpha=1\%</math> level is binomial with mean 10 and variance about 10. Thus, the standard deviation of the MC estimate is >0.003. In this example, a Monte Carlo p-value could be off by 30% and still be within a standard deviation. Is that really “pretty accurate”?
 +
 +
::::Using one million MC realizations (as done in the paper) at the <math>\alpha=0.001</math> level means the standard deviation is about 10%. The paper reported a p-value of less than 0.001 (experiment two). It wouldn’t surprise me to find out that the experiment two p-value for the flawed test is off because only one million realizations were used. My original statement, “When p-values are small, Monte Carlo methods are notoriously inaccurate unless the number of realizations generated is enormous” is correct. [[User:SJohnson|SJohnson]] 10:10, 12 March 2009 (EDT)
 +
 +
There’s still confusion about the difference between test statistics and Monte Carlo methods. Before you find a Monte Carlo estimate of a p-value, you need to select a test statistic to reduce the data set to a scalar. I am interested in hearing which test statistic you believe should be used in place of the chi-square test and why. [[User:SJohnson|SJohnson]] 10:10, 12 March 2009 (EDT)
    
----
 
----
Line 65: Line 75:  
</math>
 
</math>
 
:where <math>n_{i,j}</math> is the observed value and <math>E\left[n_{i,j}\right]</math> is the expected null hypothesis value. So if you have MS Excel, another way to arrive at the p-value of 0.19 is to type “=CHIDIST(14.82,11)” into a cell. Cheers! [[User:SJohnson|SJohnson]] 16:38, 5 March 2009 (EST)
 
:where <math>n_{i,j}</math> is the observed value and <math>E\left[n_{i,j}\right]</math> is the expected null hypothesis value. So if you have MS Excel, another way to arrive at the p-value of 0.19 is to type “=CHIDIST(14.82,11)” into a cell. Cheers! [[User:SJohnson|SJohnson]] 16:38, 5 March 2009 (EST)
 +
 
::OK, thanks for the info. From what I'd calculated and looked up in tables, the numbers seemed close to a df=11 for a chi-square of ~14. (Aside: With terms having 17/3 in the denominator in the figures above, were you using the test of independence? I was using Pearson's test for [http://en.wikipedia.org/wiki/Pearson%27s_chi-square_test#Test_for_fit_of_a_distribution fit of a distribution] which returns a chi-squared value of 14 and roughly matched the p-values you reported, assuming the df was 11).
 
::OK, thanks for the info. From what I'd calculated and looked up in tables, the numbers seemed close to a df=11 for a chi-square of ~14. (Aside: With terms having 17/3 in the denominator in the figures above, were you using the test of independence? I was using Pearson's test for [http://en.wikipedia.org/wiki/Pearson%27s_chi-square_test#Test_for_fit_of_a_distribution fit of a distribution] which returns a chi-squared value of 14 and roughly matched the p-values you reported, assuming the df was 11).
   Line 88: Line 99:     
:There are other significant problems with the use of the chi-squared test in this circumstance, but they can wait until you address these first major problems.--[[User:ElyM|ElyM]] 12:18, 11 March 2009 (EDT)   
 
:There are other significant problems with the use of the chi-squared test in this circumstance, but they can wait until you address these first major problems.--[[User:ElyM|ElyM]] 12:18, 11 March 2009 (EDT)   
 +
 +
::Wackerly et al. says in general it’s assumed that the cell frequencies are above five so that the chi-square statistic (under the null) is approximately chi-square distributed (see p. 703). That book does not say chi-square test results are invalid if frequencies are five or less. Your example of a chi-square test warning message (it said "warning" not "error" as you stated) in Minitab [http://www.minitab.com/support/answers/answer.aspx?log=0&id=2236] said “approximation probably invalid” referring to the chi-square distribution approximation to the chi-square test statistic’s distribution. Your example did not say “chi-square test invalid”. I agree that when cell frequencies are low, the chi-square test statistic’s distribution starts to deviate from the chi-square distribution. I maintain that this deviation is not enough to explain the >2.5x and >20x differences in the chi-square test p-values and the p-values from the paper.
 +
 +
::As the numerous links in your post proved, the chi-square test is widely-used by statisticians. Can you give examples of statisticians using mean mutation generation as a test statistic? Also, did your software agree with the chi-square test p-values I presented? Thanks. [[User:SJohnson|SJohnson]] 10:10, 12 March 2009 (EDT)
    
== Misinterpretation of test ==
 
== Misinterpretation of test ==
37

edits

Navigation menu