| Line 85: |
Line 85: |
| | | | |
| | :::As I said it would be interesting to read how these estimates were arrived at. I have no idea what the percentages should be and would love to see the data but I know that since this is an essay that a cite is not required and my not be possible to get in any case. --[[User:WillB|WillB]] 18:41, 14 December 2008 (EST) | | :::As I said it would be interesting to read how these estimates were arrived at. I have no idea what the percentages should be and would love to see the data but I know that since this is an essay that a cite is not required and my not be possible to get in any case. --[[User:WillB|WillB]] 18:41, 14 December 2008 (EST) |
| | + | |
| | + | |
| | + | Aschlafly asks "Do you doubt its [the statistics'] truth?" To which I respond, yes, I do doubt it, but more importantly I doubt very much that these "statistics" should be referred to as such. I, further, dispute WillB's claim that because this is an essay Aschlafly is not obligated to provide citations or data to support his claim. |
| | + | |
| | + | My basic issue here is that absent any sort of supporting documentation these "estimates" are nothing of the sort and are, instead, simply guesses. A guess is not the same as a statistical estimate. To understand why, consider ASchlafly's first cited reason, "Did not hear about conservative principles...", which I will refer to as Claim A. Now, in theory Claim A is true of some portion of the population (A "population" being defined as the set of things we're interested in studying. In this case, for the sake of argument, we'll assume the population is American citizens or legal residents). This value is known as a population parameter and we will refer to it as "Mu", following the common statistical practice of referring to true population values with Greek symbols. Now, the difficulty is that the population is too large for us to assess as a whole. To deal with this, we take a sample of the population- for example, a random sample. Now, the composition of a true random sample is determined by its parent population. So, for example, if Claim A is true of 40% of the population, as ASchlafly asserts, then a true random sample will- in theory- contain the same proportion of persons for whom Claim A is true. The quantity in the sample for whom Claim A is true will be referred to as X-bar, a symbol used to refer to the arithmetic mean of a sample, and is not written in Greek as it refers to an estimate or "statistic" rather than to a true population parameter. So, if our random sample is perfect Mu should equal x-bar. |
| | + | |
| | + | Now, this is true in theory, but in practice things get tricky. While a true random sample is the ideal, it is also the case that larger samples provide better estimates (i.e. more accurate values of x-bar) than smaller samples. Thus, it is almost always the case that Mu does not equal x-bar but rather is only approximately equal to x-bar. The relation between a particular value of x-bar and the value of Mu is determined using two things: the estimated [[standard deviation]] of a distribution of means and a [[confidence interval]]. The distribution of means is a mathematical construct that indicates how often we ought to see the x-bar of a sample of a given size, n, vary by a certain amount from the value of Mu. For our purposes we would estimate it using sample information by dividing the sample variance by n (i.e the sample size) and then taking the square root. We will refer to this as SDm (i.e. the standard deviation of the distribution of means). We would then use this value to construct a confidence interval. A confidence interval is a range of scores within which we are certain to a specific probability that we will find Mu. Put differently, it is like saying, "The true population value is x-bar plus or minus y amount, with a certainty of 95%." Confidence intervals are almost always included with point estimates (i.e. estimates of a specific value, such as x-bar) because statisticians are well aware that x-bar almost never precisely equals Mu. We would construct our confidence interval by taking x-bar and then adding, and subtracting, the value of the product of SDm and a t-value corresponding to our desired level of certainty and degrees of freedom (i.e. x-bar + (SDm)(t-value) and x-bar - (SDm)(t-value)). A t-value is a score taken from the t-distribution, which is an approximation of the [[normal distribution]] used when a smaller sample size produces non-normality. The t-distribution asymptotically approximates the normal distribution, so with large sample sizes you can essentially use the normal distribution instead. The degrees of freedom in this case are equal to n-1. So, if we drew a sample of 101 people and wanted to be 99% sure that our confidence interval included Mu then we would use a t-value of 2.626. So, in summation, using the sample information we can compute not simply an estimate of the value of Mu (i.e. x-bar) but also an interval within which we can be confident (e.g. 99% certain) that the true value of Mu lies. Obviously, the smaller the confidence interval, the more exact our estimate of Mu is likely to be and, if we assume that the sample x-bar is derived from is a perfect random sample, then the quality of our estimate is based entirely on sample size. Further, the relationships discussed above are well-documented empirically and have been proven out by mathematicians and statisticians since about the turn of the century. |
| | + | |
| | + | The reality, of course, is that samples are rarely if ever perfectly random. Sampling error inevitably creeps in and, as a result, confidence intervals often have to be adjusted for this added error. If we cannot determine the extent of the sampling error's influence on our statistics then there is no mathematical adjustment possible and we, instead, have to assess the robustness of our estimates against the probable size of the error. And what all this means is that something is a statistical estimate rather than an offhand guess precisely because it includes not only a point estimate (e.g. 40% of the population subscribes to Claim A) but also a confidence interval around that point estimate (e.g. plus or minus 10%). Moreover, the point estimates as well as the confidence intervals are produced using a set of established procedures that are rooted in the mathematical characteristics of both population/sample relationships and the estimators (i.e. the computations used to produce the point estimates). Given that this is the case, point estimates are almost always provided with confidence intervals or standard errors and, additionally, information must be provided in order to allow others to assess the likely accuracy of the estimates. |
| | + | |
| | + | ASchlafly has provided a set of point estimates generated in an unknown manner from unknown data. He has included no confidence intervals and no indication of the degree of accuracy in these estimates. As a result they are not statistical estimates and in no way should be referred to as the product of "statistical analysis". They are, to the contrary, nothing more or less than offhand guesses and should be referred to as such. On the other hand, if ASchlafly has some basis for these point estimates he should indicate where the data derive from and provide- at a bare minimum- the confidence intervals around these estimates as well as an outline indicating how he produced those intervals. |
| | + | |
| | + | The point of statistics is not simply to compute an answer (e.g. a point estimate) but to produce along with it an estimate of the degree or error contained in that estimate. No human, after all, is perfect and it is in the nature of statistics to honestly quantify that imperfection. Absent confidence intervals or standard errors, and some understanding of how the data were gathered the appropriate response to the question "Do you doubt its truth" from anyone who is marginally competent in statistics must be "yes." Moreover, if ASchlafly wishes to label these assertions as "statistical estimates" then it is his responsibility to justify them in the manner accepted by statisticians. Otherwise, they should be (correctly) labeled as guesses. |
| | + | |
| | + | Understand, the substance of my objection has nothing to do with whether or not ASchlafly's claims are correct but, rather, with the appropriateness of their presentation and treatment. -Drek |
| | | | |
| | == My political alignment == | | == My political alignment == |