C++ Logo

sg16

Advanced search

Re: [isocpp-sg16] Thoughts on P4235R0: Removing Digraphs

From: Matthias Wippich <mfwippich_at_[hidden]>
Date: Thu, 24 Sep 2026 06:19:23 +0200
Thanks for the feedback, I'll reply inline.

> First, thank you for a thorough paper; I love seeing all the history presented!

Happy to hear that. It was an immense pain to reconstruct. Out of the
17091 X3J11, X3J16, WG14 and WG21 documents hosted on open-std.org
(15391 after removing working drafts, editors reports etc), 293
mentioned "trigraph", 169 "digraph" for a total of 351 documents.
Unfortunately some of the early references use neither of those terms.
I hope I didn't make any unfortunate mistakes or misrepresented
anything - I've read through almost all of those documents to figure
out what matters, but y'know.. the query might've been bad and I have
no idea if there was any undetected conversion failures (I OCR'd all
those pdf papers that had no better source.. that took a weekend).
I've been meaning to ask Keld about the correctness of the early
history.

There was a few interesting (albeit irrelevant for this paper) ones.
For example, John Skaller published WG21/N0259 in 1993, which proposed
`addrof` as alternative to `bitand`. The motivation is largely
```
int* a = bitand b;
```
looking rather strange.

> Have you reached out to Keld Simonsen for comment regarding the history and the original intent? I can put you in touch with him if you don't have contact information.

I haven't yet, but will. Peter provided me with Keld's current email address.

> The introduction states, "Digraphs are a complicated solution to a very old problem ...". I think it would be helpful to briefly describe the problems they solve. The history section suggests that the problem solved is source code written in character encodings that lack support for (some of) the members of the basic character set. That is one of the problems, but not the only one. Digraphs also allow source code to be written in a subset of characters that allow it to be correctly interpreted in multiple incompatible character encodings. I think this is best exemplified by relating the digraphs to the variant members of EBCDIC. According to IBM documentation, the variant EBCDIC members that are also members of the basic character set include:
> [...]
> It is perhaps worth noting that digraphs are lacking for some members of the basic character set that do not have invariant encoding across all EBCDIC code pages.

Yeah that's a good point. I tried alluding to that fact with the
linked Google spreadsheet (Affected encodings). I'll try making that
point more explicit for the next revision :)

> The paper could also explicitly mention that universal-character-names provide a modern alternative to digraphs.

I really like the idea of using UCNs as a generic way of spelling
various characters that have meaning in C++, especially if it also
allows named universal characters to be used. I'd be happy to consider
working on a proposal for this if there's appetite.

However, I do not think it truly provides an alternative to digraphs.
If you look at the spreadsheet from the paper - this would only be a
usable alternative for BS_4730, EBCDIC-ES, EBCDIC-ES-S, IBM275,
LATIN-GREEK, EBCDIC-DK-NO, EBCDIC-UK, EBCDIC-US and IBM880 (using the
iconv primary spellings). For the other 39 affected encodings that are
otherwise usable, \ is missing - so you can't spell UCNs either. We
currently have no alternative token for that.

For the much nicer named universal characters you additionally run
into the issue of requiring {}. Considering that \ is also required,
this narrows the list down to BS_4730, EBCDIC-ES, EBCDIC-ES-S and
LATIN-GREEK. That does not seem like a viable replacement unless we
also introduce something for \.


Thanks,
Matthias

Received on 2026-09-24 04:19:40