Changes

Jump to navigation Jump to search
Line 2,234: Line 2,234:  
::::You both raise valid points, which hopefully will be answered soon. Andy and August both mention that the x^2 curve isn't a perfect fit - it certainly isn't, and I'm pretty sure a 2^x curve with appropriate constants will fit better; I'm planning to do that tonight. August mentions different ways of representing the data - I'll happily produce a histogram of the data if that'd be interesting, but the reason for plotting it as I have is to produce a curve that I can use for my more grandiose scheme, of which more later. My background is not so much in statistics - although I've done a fair bit of that - but in purer maths, so my thinking is mostly based around the relationships between smooth(-ish) functions. That may not be the best way to deal with these data qua data, but to extract patterns for further, more abstract work, it's ideal. [[User:Jcw|Jcw]] 12:43, 20 June 2011 (EDT)
 
::::You both raise valid points, which hopefully will be answered soon. Andy and August both mention that the x^2 curve isn't a perfect fit - it certainly isn't, and I'm pretty sure a 2^x curve with appropriate constants will fit better; I'm planning to do that tonight. August mentions different ways of representing the data - I'll happily produce a histogram of the data if that'd be interesting, but the reason for plotting it as I have is to produce a curve that I can use for my more grandiose scheme, of which more later. My background is not so much in statistics - although I've done a fair bit of that - but in purer maths, so my thinking is mostly based around the relationships between smooth(-ish) functions. That may not be the best way to deal with these data qua data, but to extract patterns for further, more abstract work, it's ideal. [[User:Jcw|Jcw]] 12:43, 20 June 2011 (EDT)
 
::::[http://static.inky.ws/image/412/image.jpg Et voila], a better fit. This is an exponential curve fitted to the same data. Note that it fits much better in the region with the most words, but is a bit out for the earlier period where there are fewer words in the list. This is because we can more easily find suitable words from more recent periods, so naturally the pattern is most exact there. No doubt if we could go through a large, representative corpus and extract words uniformly, it would fit nicely all the way along. [[User:Jcw|Jcw]] 16:25, 20 June 2011 (EDT)
 
::::[http://static.inky.ws/image/412/image.jpg Et voila], a better fit. This is an exponential curve fitted to the same data. Note that it fits much better in the region with the most words, but is a bit out for the earlier period where there are fewer words in the list. This is because we can more easily find suitable words from more recent periods, so naturally the pattern is most exact there. No doubt if we could go through a large, representative corpus and extract words uniformly, it would fit nicely all the way along. [[User:Jcw|Jcw]] 16:25, 20 June 2011 (EDT)
 +
 +
:::::'''@Aschlafly: ''' Could you recount the words? My count gave me 26-51-103-210-18 (Sum: 408) instead of 26-52-103-208-18 (Sum: 408). Perhaps a fourth column for the century (or even better, the decade) could be added? That would make it much easier to keep track of the numbers!
 +
 +
:::::'''@Jcw: ''' I don't think that your ''better fit'' is the function which Aschlafly has in mind: it should be <math>F_{theo}(t) = \frac{\#words }{15}</math><math>(2^{\frac{t-1599}{100}}-1)</math>, where ''#words'' is the number of words created before 2000, i.e., 390. This function touches/intersects the empirical cdf at the turn of each century, a fact which betrays the biased method of looking for these words.
 +
 +
:::::[[User:AugustO|AugustO]] 11:12, 21 June 2011 (EDT)
    
== Americanadians ==
 
== Americanadians ==
Block, SkipCaptcha, Automoderated users, edit, rollback
5,023

edits

Navigation menu