Changes

Jump to navigation Jump to search
response to ASchlafly
Line 263: Line 263:     
::: You say, "I'm not sure why you expect me to address Blout's use of Monte Carlo."  The reason is obvious:  the title of the content page is the "Significance of E. Coli Evolution Experiments."  You haven't addressed the inappropriateness of using Monte Carlo simulations for assessing the significance rarely occurring events, which was central to Lenski's statistical claims.  I suggest you address this flaw if you want to be taken seriously.--[[User:Aschlafly|Andy Schlafly]] 23:12, 18 March 2009 (EDT)
 
::: You say, "I'm not sure why you expect me to address Blout's use of Monte Carlo."  The reason is obvious:  the title of the content page is the "Significance of E. Coli Evolution Experiments."  You haven't addressed the inappropriateness of using Monte Carlo simulations for assessing the significance rarely occurring events, which was central to Lenski's statistical claims.  I suggest you address this flaw if you want to be taken seriously.--[[User:Aschlafly|Andy Schlafly]] 23:12, 18 March 2009 (EDT)
 +
 +
::::ASchlafly, again per your request, and based on your statement regarding the "inappropriateness of using Monte Carlo simulations for assessing the significance of rarely occurring events",  I have spent the last several days reviewing the literature available to me on Monte Carlo and other resampling techniques, looking for ways in which Blount may have made a methodological error of the sort that SJohnson has made. I have been unable to find any examples of authors suggesting that Monte Carlo be avoided for low ''n'', or for events with low probability regardless of'' n'', much less providing specific cutoff numbers as are seen in the references that I provided for the chi-square test. Similarly, the technique that Blount used does not require/assume that categories are unrelated, as the chi-square test does.  Of course, the absence of evidence is not evidence of absence, and I may have misinterpreted the basis of your objection.  At this point I'll need you to explain your objection in more detail if you wish me to find the appropriate literature addressing your concerns. Do you believe that the number of resamplings was too low in Blount's paper? That the analysis should have been performed with a software package other than Statistics101? Some other procedural issue? Some issue of interpretation?
 +
 +
::::The statistical problem that Blount must address is straightforward: given a distribution of mutant cultures that ''appears'' to be skewed toward the higher generations, what is the probability that this same amount of skew (or a greater degree) could arise by chance, given the null hypothesis that every generation is equally likely to produce a mutant? Interestingly, in the case of the first replay experiment, the total number of ways to randomly select (equal probability, no replacement) four cultures from seventy-two is 72x71x70x69, or 24,690,960. This number is small enough that a program can brute-force-calculate the 'mean generation number' of ''all possible'' combinations of four cultures in a reasonable amount of time. An experimentally-derived 'mean generation number' can be checked against this exhaustive list, and the number of means equal to or larger than the experimental mean can be found exactly. Converting this number to a percentage of 24,690,960 provides an exact p-value for any given experimental 'mean generation number'. This exhaustive approach is different than the Monte Carlo technique, in that ''all possible'' outcomes are examined, rather than a ''random subset'' of all possible outcomes. For the first replay experiment, it provides a way to independently check Blount's Monte Carlo results. This approach is not possible for the second and third replay experiments, in which the total number of possible combinations becomes impractically large: 340!/335! = 4.41 x10^12 and 2800!/2792! = 3.74 x 10^27, respectively.
 +
 +
::::I asked a colleague to run just such a brute-force program for me on the first replay data. I also ran several Monte Carlo simulations ('''not''' using Statsistics101) with Blount's data, using twenty-five million, one hundred million, and 493,819,200 resamplings - note that this last is twenty times the number of all possible combinations of 4 samples drawn without replacement from 72. The p-values from the 25M, 100M, and 493M Monte Carlo resamplings (0.00844, 0.00846, and 0.00846, respectively) compare favorably with Blount's 1M value of 0.0085 and the non-Monte-Carlo brute-force exact calculation, which provides a p-value of 0.008457. Thus it appears that Blount's statistical results are confirmed by a ''non-Monte Carlo'' technique, at least for the first replay experiment.
 +
 +
::::My intention is not to get caught up in a digression about Monte Carlo, though - I'd rather keep the focus on the fact that the main article should acknowledge that SJohnson is using chi-square in a way that violates accepted guidelines; this remains true whether Blount's analysis is valid or not.--[[User:ElyM|ElyM]] 12:13, 23 March 2009 (EDT)
    
== References ==
 
== References ==
 
{{reflist}}
 
{{reflist}}
17

edits

Navigation menu