Hi, Matthias. A few questions/comments on P4235R0 (Removing Digraphs).
First, thank you for a thorough paper; I love seeing all the history presented!
Have you reached out to Keld Simonsen for comment regarding the history and the original intent? I can put you in touch with him if you don't have contact information.
The introduction states, "Digraphs are a complicated solution to a very old problem ...". I think it would be helpful to briefly describe the problems they solve. The history section suggests that the problem solved is source code written in character encodings that lack support for (some of) the members of the basic character set. That is one of the problems, but not the only one. Digraphs also allow source code to be written in a subset of characters that allow it to be correctly interpreted in multiple incompatible character encodings. I think this is best exemplified by relating the digraphs to the variant members of EBCDIC. According to IBM documentation, the variant EBCDIC members that are also members of the basic character set include:
| Unicode character name | Character | Digraph | Comment |
| U+0021 EXCLAMATION MARK | ! | ||
| U+0023 NUMBER SIGN | # | %: | with %:%: for ## |
| U+0024 DOLLAR SIGN | $ | Added in C++26 | |
| U+0040 COMMERCIAL AT | @ | Added in C++26 | |
| U+007C VERTICAL LINE | | | ||
| U+005B LEFT SQUARE BRACKET | [ | <: | |
| U+005C REVERSE SOLIDUS | \ | ||
| U+005D RIGHT SQUARE BRACKET | ] | :> | |
| U+005E CIRCUMFLEX ACCENT | ^ | ||
| U+0060 GRAVE ACCENT | ` | Added in C++26 | |
| U+007B LEFT CURLY BRACKET | { | <% | |
| U+007D RIGHT CURLY BRACKET | } | %> | |
| U+007E TILDE | ~ |
It is perhaps worth noting that digraphs are lacking for some members of the basic character set that do not have invariant encoding across all EBCDIC code pages.
The paper could also explicitly mention that universal-character-names provide a modern alternative to digraphs.
Tom.