Plane (Unicode)
In the Unicode standard, a plane is a continuous group of 65,536 code points. There are 17 planes, identified by the numbers 0 to 16, which corresponds with the possible values 00–1016 of the first two positions in six position hexadecimal format. Plane 0 is the Basic Multilingual Plane, which contains most commonly used characters. The higher planes 1 through 16 are called "supplementary planes". The very last code point in Unicode is the last code point in plane 16, U+10FFFF. As of Unicode version 13.0, seven of the planes have assigned code points, and five are named.
The limit of 17 planes is due to UTF-16, which can encode 220 code points as pairs of words, plus the BMP as a single word. UTF-8 was designed with a much larger limit of 231 code points, and can encode 221 code points even under the current limit of 4 bytes.
The 17 planes can accommodate 1,114,112 code points. Of these, 2,048 are surrogates, 66 are non-characters, and 137,468 are reserved for private use, leaving 974,530 for public assignment.
Planes are further subdivided into Unicode blocks, which, unlike planes, do not have a fixed size. The 308 blocks defined in Unicode 13.0 cover 26% of the possible code point space, and range in size from a minimum of 16 code points to a maximum of 65,536 code points. For future usage, ranges of characters have been tentatively mapped out for most known current and ancient writing systems.
Overview
Plane | Allocated code points | Assigned characters |
0 BMP | 65,472 | 55,503 |
1 SMP | 24,704 | 22,279 |
2 SIP | 60,912 | 60,866 |
3 TIP | 4,944 | 4,939 |
14 SSP | 368 | 337 |
15 SPUA-A | 65,536 | |
16 SPUA-B | 65,536 | |
Totals | 287,472 | 143,924 |
Basic Multilingual Plane
The first plane, plane 0, the Basic Multilingual Plane contains characters for almost all modern languages, and a large number of symbols. A primary objective for the BMP is to support the unification of prior character sets as well as characters for writing. Most of the assigned code points in the BMP are used to encode Chinese, Japanese, and Korean characters.The High Surrogate and Low Surrogate codes are reserved for encoding non-BMP characters in UTF-16 by using a pair of 16-bit codes: one High Surrogate and one Low Surrogate. A single surrogate code point will never be assigned a character.
65,472 of the 65,536 code points in this plane have been allocated to a Unicode block, leaving just 64 code points in unallocated ranges.
, the BMP comprises the following 163 blocks:
- Basic Latin
- Latin-1 Supplement
- Latin Extended-A
- Latin Extended-B
- IPA Extensions
- Spacing Modifier Letters
- Combining Diacritical Marks
- Greek and Coptic
- Cyrillic
- Cyrillic Supplement
- Armenian
- Aramaic Scripts:
- * Hebrew
- * Arabic
- * Syriac
- * Arabic Supplement
- * Thaana
- * N'Ko
- * Samaritan
- * Mandaic
- * Syriac Supplement
- * Arabic Extended-A
- Brahmic scripts:
- * Devanagari
- * Bengali
- * Gurmukhi
- * Gujarati
- * Oriya
- * Tamil
- * Telugu
- * Kannada
- * Malayalam
- * Sinhala
- * Thai
- * Lao
- * Tibetan
- * Myanmar
- Georgian
- Hangul Jamo
- Ethiopic
- Ethiopic Supplement
- Cherokee
- Unified Canadian Aboriginal Syllabics
- Ogham
- Runic
- Philippine scripts:
- * Tagalog
- * Hanunoo
- * Buhid
- * Tagbanwa
- Khmer
- Mongolian
- Unified Canadian Aboriginal Syllabics Extended
- Limbu
- Tai scripts:
- * Tai Le
- * New Tai Lue
- * Khmer Symbols
- * Buginese
- * Tai Tham
- Combining Diacritical Marks Extended
- Balinese
- Sundanese
- Batak
- Lepcha
- Ol Chiki
- Cyrillic Extended-C
- Georgian Extended
- Sundanese Supplement
- Vedic Extensions
- Latin-2 supplement:
- * Phonetic Extensions
- * Phonetic Extensions Supplement
- * Combining Diacritical Marks Supplement
- * Latin Extended Additional
- Greek Extended
- Symbols:
- * General Punctuation
- * Superscripts and Subscripts
- * Currency Symbols
- * Combining Diacritical Marks for Symbols
- * Letterlike Symbols
- * Number Forms
- * Arrows
- * Mathematical Operators
- * Miscellaneous Technical
- * Control Pictures
- * Optical Character Recognition
- * Enclosed Alphanumerics
- * Box Drawing
- * Block Elements
- * Geometric Shapes
- * Miscellaneous Symbols
- * Dingbats
- * Miscellaneous Mathematical Symbols-A
- * Supplemental Arrows-A
- * Braille Patterns
- * Supplemental Arrows-B
- * Miscellaneous Mathematical Symbols-B
- * Supplemental Mathematical Operators
- * Miscellaneous Symbols and Arrows
- Glagolitic
- Latin Extended-C
- Coptic
- Georgian Supplement
- Tifinagh
- Ethiopic Extended
- Cyrillic Extended-A
- Supplemental Punctuation
- CJK scripts and symbols:
- * CJK Radicals Supplement
- * Kangxi Radicals
- * Ideographic Description Characters
- * CJK Symbols and Punctuation
- * Hiragana
- * Katakana
- * Bopomofo
- * Hangul Compatibility Jamo
- * Kanbun
- * Bopomofo Extended
- * CJK Strokes
- * Katakana Phonetic Extensions
- * Enclosed CJK Letters and Months
- * CJK Compatibility
- * CJK Unified Ideographs Extension A
- * Yijing Hexagram Symbols
- * CJK Unified Ideographs
- Yi Syllables
- Yi Radicals
- Lisu
- Vai
- Cyrillic Extended-B
- Bamum
- Modifier Tone Letters
- Latin Extended-D
- Syloti Nagri
- Common Indic Number Forms
- Phags-pa
- Saurashtra
- Devanagari Extended
- Kayah Li
- Rejang
- Hangul Jamo Extended-A
- Javanese
- Myanmar Extended-B
- Cham
- Myanmar Extended-A
- Tai Viet
- Meetei Mayek Extensions
- Ethiopic Extended-A
- Latin Extended-E
- Cherokee Supplement
- Meetei Mayek
- Hangul Syllables
- Hangul Jamo Extended-B
- Surrogates:
- * High Surrogates
- * High Private Use Surrogates
- * Low Surrogates
- Private Use Area
- CJK Compatibility Ideographs
- Alphabetic Presentation Forms
- Arabic Presentation Forms-A
- Variation Selectors
- Vertical Forms
- Combining Half Marks
- CJK Compatibility Forms
- Small Form Variants
- Arabic Presentation Forms-B
- Halfwidth and Fullwidth Forms
- Specials
Supplementary Multilingual Plane
, the SMP comprises the following 134 blocks:
- Archaic Greek and Other Left-to-right scripts:
- * Linear B Syllabary
- * Linear B Ideograms
- * Aegean Numbers
- * Ancient Greek Numbers
- * Ancient Symbols
- * Phaistos Disc
- * Lycian
- * Carian
- * Coptic Epact Numbers
- * Old Italic
- * Gothic
- * Old Permic
- * Ugaritic
- * Old Persian
- * Deseret
- * Shavian
- * Osmanya
- * Osage
- * Elbasan
- * Caucasian Albanian
- * Linear A
- Right-to-left scripts:
- * Cypriot Syllabary
- * Imperial Aramaic
- * Palmyrene
- * Nabataean
- * Hatran
- * Phoenician
- * Lydian
- * Meroitic Hieroglyphs
- * Meroitic Cursive
- * Kharoshthi
- * Old South Arabian
- * Old North Arabian
- * Manichaean
- * Avestan
- * Inscriptional Parthian
- * Inscriptional Pahlavi
- * Psalter Pahlavi
- * Old Turkic
- * Old Hungarian
- * Hanifi Rohingya
- * Rumi Numeral Symbols
- * Yezidi
- * Old Sogdian
- * Sogdian
- * Chorasmian
- * Elymaic
- Brahmic scripts:
- * Brahmi
- * Kaithi
- * Sora Sompeng
- * Chakma
- * Mahajani
- * Sharada
- * Sinhala Archaic Numbers
- * Khojki
- * Multani
- * Khudawadi
- * Grantha
- * Newa
- * Tirhuta
- * Siddham
- * Modi
- * Mongolian Supplement
- * Takri
- * Ahom
- * Dogra
- * Warang Citi
- * Dives Akuru
- * Nandinagari
- * Zanabazar Square
- * Soyombo
- * Pau Cin Hau
- * Bhaiksuki
- * Marchen
- * Masaram Gondi
- * Gunjala Gondi
- * Makasar
- Lisu Supplement
- Tamil Supplement
- Cuneiform
- Cuneiform Numbers and Punctuation
- Early Dynastic Cuneiform
- Egyptian Hieroglyphs
- Egyptian Hieroglyph Format Controls
- Anatolian Hieroglyphs
- Bamum Supplement
- Mro
- Bassa Vah
- Pahawh Hmong
- Medefaidrin
- Miao
- Ideographic Symbols and Punctuation
- Tangut
- Tangut Components
- Khitan Small Script
- Tangut Supplement
- Kana Supplement
- Kana Extended-A
- Small Kana Extension
- Nushu
- Duployan
- Shorthand Format Controls
- Supplementary symbols:
- * Musical notation:
- ** Byzantine Musical Symbols
- ** Musical Symbols
- ** Ancient Greek Musical Notation
- * Mayan Numerals
- * Mathematical symbols:
- ** Tai Xuan Jing Symbols
- ** Counting Rod Numerals
- ** Mathematical Alphanumeric Symbols
- * Sutton SignWriting
- Glagolitic Supplement
- Nyiakeng Puachue Hmong
- Wancho
- Mende Kikakui
- Adlam
- Indic Siyaq Numbers
- Ottoman Siyaq Numbers
- Arabic Mathematical Alphabetic Symbols
- Game tiles and cards:
- * Mahjong Tiles
- * Domino Tiles
- * Playing Cards
- Enclosed Alphanumeric Supplement
- Enclosed Ideographic Supplement
- Miscellaneous Symbols and Pictographs
- Emoticons
- Ornamental Dingbats
- Transport and Map Symbols
- Alchemical Symbols
- Geometric Shapes Extended
- Supplemental Arrows-C
- Supplemental Symbols and Pictographs
- Chess Symbols
- Symbols and Pictographs Extended-A
- Symbols for Legacy Computing
Supplementary Ideographic Plane
, the SIP comprises the following six blocks:
- CJK Unified Ideographs Extension B
- CJK Unified Ideographs Extension C
- CJK Unified Ideographs Extension D
- CJK Unified Ideographs Extension E
- CJK Unified Ideographs Extension F
- CJK Compatibility Ideographs Supplement
Tertiary Ideographic Plane
, the TIP comprises the following block:
- CJK Unified Ideographs Extension G
Unassigned planes
Supplementary Special-purpose Plane
Plane 14, the Supplementary Special-purpose Plane. comprising the following two blocks :- Tags
- Variation Selectors Supplement – used to indicate alternate glyphs for characters.
Private Use Area planes