Difference between revisions of "UTF-8"

From Conservapedia
Jump to navigation Jump to search
(a way of encoding Unicode characters)
 
(a bit more about the encoding scheme)
Line 1: Line 1:
−
The '''UTF-8''' standard is a way of encoding [[Unicode]] characters, invented in a single evening by two computer programmers in New Jersey. [http://www.cl.cam.ac.uk/~mgk25/ucs/utf-8-history.txt UTF-8 history
+
The '''UTF-8''' standard is a way of encoding [[Unicode]] characters, invented in a single evening by two computer programmers in New Jersey. [http://www.cl.cam.ac.uk/~mgk25/ucs/utf-8-history.txt UTF-8 history]
−
]
+
Each character is formatted into 8-bit bytes, with simple [[ASCII]] characters requiring only one byte, and more complex character sets taking up to 6 bytes. The idea is to pack as many bits as possible into each byte, while also indicating how many (more) bytes are needed to complete the encoding.
 +
 
 +
Each byte sequence starts with a byte whose initial bits indicates how many total bytes it takes to encode the character. If the initial bit is 0, then it's a 7-bit ASCII character which fits in one byte. Nearly all text written in American English will fit in this encoding.

Revision as of 17:43, February 20, 2009

The UTF-8 standard is a way of encoding Unicode characters, invented in a single evening by two computer programmers in New Jersey. UTF-8 history Each character is formatted into 8-bit bytes, with simple ASCII characters requiring only one byte, and more complex character sets taking up to 6 bytes. The idea is to pack as many bits as possible into each byte, while also indicating how many (more) bytes are needed to complete the encoding.

Each byte sequence starts with a byte whose initial bits indicates how many total bytes it takes to encode the character. If the initial bit is 0, then it's a 7-bit ASCII character which fits in one byte. Nearly all text written in American English will fit in this encoding.