Width¶
Status: Convention. No standard defines terminal cell width; wcwidth and
Unicode Annex #11 are the shared references.
Width is the number of cells a character advances the cursor. Terminals, libc, and application libraries each compute it independently. When they disagree, the application believes the cursor is somewhere it is not.
The wcwidth rule¶
POSIX wcwidth(3) returns 0, 1, 2, or -1 for a code point. The conventional
mapping, from Markus Kuhn's reference implementation, is:
| Class | Width |
|---|---|
| Control characters | -1 (undefined) |
Combining marks (Mn, Me), most format characters (Cf) |
0 |
East Asian Wide (W) and Fullwidth (F) |
2 |
| Everything else | 1 |
The East Asian Width property comes from Unicode Annex #11. Categories N,
Na, and H are one cell. Category A, ambiguous, is one cell by default
and two cells in CJK locales; emulators expose this as a setting:
| Emulator | Setting |
|---|---|
| xterm | cjkWidth resource, -cjk_width option1 |
| VTE | VTE_CJK_WIDTH environment variable, terminal "ambiguous width" option2 |
| kitty | Follows the wide list; no ambiguous toggle documented |
| WezTerm | treat_east_asian_ambiguous_width_as_wide3 |
| foot | tweak.ambiguous-width? (?) |
| Windows Terminal | Follows the locale (?) |
Zero-width characters¶
Combining marks attach to the preceding cell. Zero Width Joiner (U+200D),
Zero Width Non-Joiner (U+200C), and variation selectors are format characters
with width 0 in wcwidth terms, yet each may change the width of the unit
they belong to. That is the seam between width and
grapheme clusters.
Emoji and variation selectors¶
Emoji presentation is the largest source of drift. The rules that most emulators converge on:
- Code points with
Emoji_Presentation=Yesare width 2. - Text-default emoji such as U+2764 HEART are width 1, but U+2764 U+FE0F is width 2 in emulators that honor VS16.
- U+FE0E (VS15) requests text presentation; a width-2 emoji followed by VS15 may become width 1.
- Skin-tone modifiers (U+1F3FB–U+1F3FF) and ZWJ sequences form one width-2 unit only when the emulator clusters; otherwise each part is measured alone.
wcwidth implementations that predate a given Unicode version return 1 for
emoji they do not know, so the same string measures differently on an old
libc and a new emulator.
Unicode version skew¶
Three components carry their own Unicode tables:
- the emulator, compiled against a Unicode version;
- libc, whose
wcwidththe shell and many C programs use; - the application's own library, such as
wcwidthfor Python,unicode-widthfor Rust, orutf8proc.
Any mismatch produces drift after the first affected character. There is no in-band negotiation of Unicode version; the text sizing and grapheme protocols exist partly to reduce dependence on it.
Some emulators ship generated tables from Unicode's EastAsianWidth.txt and
emoji-data.txt (kitty, foot, Ghostty, WezTerm); others call libc or a
bundled utf8proc. The difference is visible with recently added emoji.
Box drawing and symbols¶
U+2500–U+257F box drawing, U+2580–U+259F block elements, and Powerline
private-use glyphs are width 1 by wcwidth. Emulators often synthesize these
glyphs instead of using the font so the lines meet; see
Rendering. U+2000–U+200A spaces are width 1 except the
zero-width ones.
Wide character at the last column¶
If a width-2 character is written when only one cell remains, the terminal must either wrap it to the next line, leaving a blank spacer cell, or clip it. DEC terminals did not have wide characters; xterm wraps and leaves the last column blank, and other emulators follow. An application computing wrap positions must reproduce this rule.
Probe¶
Print a character, ask for the cursor position, and compare the column. This measures the emulator's opinion directly and is the only portable method.
stty -echo -icanon min 0 time 5
for s in 'A' '中' '❤' '❤️' '👍🏽' '🇺🇸' '👨👩👧'; do
printf '\r\033[K%s\033[6n' "$s"
reply=$(dd bs=64 count=1 2>/dev/null | tr -d '\033')
col=${reply##*;}; col=${col%R}
printf ' width=%s\n' "$((col - 1))"
done
stty sane
Compare with python3 -c 'import unicodedata;print(unicodedata.east_asian_width("中"))'
and with the application's own width library.
Sources¶
- Unicode Standard Annex #11, East Asian Width
- Markus Kuhn's wcwidth.c
- Unicode Technical Standard #51, Emoji
- XTerm Control Sequences
- utf8proc
Compatibility¶
Choose terminals
| Feature / stable ID | Alacritty | Apple Terminal | ConPTY (conhost) | Contour | DomTerm | foot | Ghostty | iTerm2 | kitty | Konsole | mintty | mlterm | PuTTY | Revenant | RLogin | st | tmux | rxvt-unicode | VTE | wayst | WezTerm | Windows Terminal | xterm | xterm.js |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
Ambiguous-width optiontext-ambiguous-width |
Unknown | Unknown | Unknown | Unknown | Unknown | Unknown | Unknown | SupportedImported · unverified | Unknown | Unknown | Unknown | Unknown | Unknown | Unknown | Unknown | Unknown | Unknown | Unknown | SupportedImported · unverified | Unknown | SupportedImported · unverified | Unknown | SupportedImported · unverified | Unknown |
VS16 widenstext-vs16-width |
Unknown | Unknown | Unknown | Unknown | Unknown | SupportedImported · unverified | SupportedImported · unverified | Unknown | SupportedImported · unverified | Unknown | Unknown | Unknown | Unknown | Unknown | Unknown | Unknown | Unknown | Unknown | SupportedImported · unverified | Unknown | SupportedImported · unverified | Unknown | PartialImported · unverified | Unknown |
Bundled Unicode tablestext-unicode-tables |
SupportedImported · unverified | Unknown | Unknown | Unknown | Unknown | SupportedImported · unverified | SupportedImported · unverified | Unknown | SupportedImported · unverified | Unknown | Unknown | Unknown | Unknown | Unknown | Unknown | Unknown | Unknown | Unknown | PartialImported · unverified | Unknown | SupportedImported · unverified | Unknown | SupportedImported · unverified | SupportedImported · unverified |
-
xterm manual,
cjkWidth,mkWidth. ↩