Changes

Jump to navigation Jump to search
2,659 bytes added ,  14:24, February 18, 2008
New page: The well known entity in the theory of information, seemingly founded by Claude Shannon, is the information entropy defined as H = - sum { p(i) log [p(i) ] } where the index i...
The well known entity in the theory of information, seemingly founded by [[Claude Shannon]], is the [[information entropy]] defined as

H = - sum { p(i) log [p(i) ] }

where the index i runs through all positive integers up to n, for instance (i = 1, 2, …, n). This entropy may calculated for any statistical frequency function, such that the sum { p(i) } = 1. The maximum in H (= log(n)) is obtained when all p(i) are equal. Mathematically, H is also equivalent to entropy and disorder in physics.

H may also be interpreted as an [[average information]] based on a certain axiom that when an event occurs with probability p(i), then we get the [[self-information]] - log [p(i) ]. This information is only about probability and has no mental meaning.

Shannon derived the formula from axioms, which in a natural way may be interpreted as information. When we toss a coin we may have two different messages, each with probability 0.5 which may be denoted 0 and 1. If the coin is tossed two times in a sequence we may have the following four possible messages ‘00’, ‘01’, ‘10’ and ‘11’, each with probability 0.25, which may be seen as the information from a pair of tosses.

A reasonable postulate is that the sum of information from any two different messages should equal the total information in the joined messages. And one way to accomplish this is to let the self-information equal the negative logarithm of the probability. Thus, the information in a pair of tosses becomes

-log(0.5) - log(0.5) = -log(0.25) or log(2) + log(2) = log(4)

Another reasonable postulate is that when a message increases in length, then its information increases. For instance, a triple of tosses gives messages with probability = 0.125 and a quadruple with probability = 0.0625. and

-log(0.0625) > -log(0.125).

And an event occurring with probability = 1 is a certain event, which will carry no information. Consequently -log(1) = 0.

If all these properties should be maintained for all possible divisions of messages written in different alphabets, then the negative logarithm of the probability is a very good candidate for being a measure of information.

More generally - the world “we” may be followed by “are”, but hardly by “am” – some part Y of a message may depend on another part X, making the average information H(Y|X) less than H(XY), i. e. the average information when X and Y are statistically independent. We have

H(Y|X) = H(XY) – H(X)

and the only functions satisfying this formula are those proportional to the entropy H, Åslund, 1961. Then, entropy may be seen as [[average information]].
39

edits

Navigation menu