CurrentINFORMATIONALINDEPENDENT stream

RFC 4042: UTF-9 and UTF-18 Efficient Transformation Formats of Unicode

In plain English — editorial summary, not part of the RFC

ISO-10646 defines a large character set called the Universal Character Set (UCS), which encompasses most of the world's writing systems. The same set of codepoints is defined by Unicode, which further defines additional character properties and other implementation details. By policy of the relevant standardization committees, changes to Unicode and amendments and additions to ISO/IEC 10646 track each other, so that the character repertoires and code point assignments remain in synchronization. The current representation formats for Unicode (UTF-7, UTF-8, UTF-16) are not storage and computation efficient on platforms that utilize the 9 bit nonet as a natural storage unit instead of the 8 bit octet. This document describes a transformation format of Unicode that takes advantage of the nonet so that the format will be storage and computation efficient. This memo provides information for the Internet community.

Document record

Document ID
RFC4042
Published
April 2005
Authors
M. Crispin
Status
INFORMATIONAL
Stream
INDEPENDENT
Area
Pages
9
Also known as

Topics

Related documents

Ranked automatically by shared keywords, IETF area and stream — not by editorial selection.

Also filed under

About this page

The document record above — title, authors, date, status, stream, area, relationships, DOI and errata — is imported verbatim from the public RFC Editor index. The “in plain English” section is editorial: written by The metasystema editorial team, not part of the RFC. Where the two differ, the RFC text governs.

Last checked against the RFC Editor index on . RFCs are never revised after publication; changes are issued as new documents.

Data sources · Editorial policy · Report a correction · What is an RFC?

canonical URL: /rfc/4042-utf-9-and-utf-18-efficient-transformation-formats-of-unicode