A survey of Unicode compression
Doug Ewell · 2004
The Unicode (ISO/IEC 10646) coded character set is the largest of its kind.1 Almost a million code positions are available in Unicode for formal character encoding, with more than 137,000 additional code positions reserved for private-use characters. This is quite a change from the 128 or 256 characters available in 8-bit “legacy” code pages, or even the thousands available in East Asian double-byte character sets (DBCS).