Date: Thu, 24 Sep 2026 16:20:04 -0400
On 9/24/26 2:04 AM, Corentin Jabot via SG16 wrote:
>
>
> On Thu, Sep 24, 2026, 06:19 Matthias Wippich via SG16
> <sg16_at_[hidden]> wrote:
>
> Thanks for the feedback, I'll reply inline.
>
> > First, thank you for a thorough paper; I love seeing all the
> history presented!
>
> Happy to hear that. It was an immense pain to reconstruct. Out of the
> 17091 X3J11, X3J16, WG14 and WG21 documents hosted on open-std.org
> <http://open-std.org>
> (15391 after removing working drafts, editors reports etc), 293
> mentioned "trigraph", 169 "digraph" for a total of 351 documents.
> Unfortunately some of the early references use neither of those terms.
> I hope I didn't make any unfortunate mistakes or misrepresented
> anything - I've read through almost all of those documents to figure
> out what matters, but y'know.. the query might've been bad and I have
> no idea if there was any undetected conversion failures (I OCR'd all
> those pdf papers that had no better source.. that took a weekend).
> I've been meaning to ask Keld about the correctness of the early
> history.
>
> There was a few interesting (albeit irrelevant for this paper) ones.
> For example, John Skaller published WG21/N0259 in 1993, which proposed
> `addrof` as alternative to `bitand`. The motivation is largely
> ```
> int* a = bitand b;
> ```
> looking rather strange.
>
> > Have you reached out to Keld Simonsen for comment regarding the
> history and the original intent? I can put you in touch with him
> if you don't have contact information.
>
> I haven't yet, but will. Peter provided me with Keld's current
> email address.
>
> > The introduction states, "Digraphs are a complicated solution to
> a very old problem ...". I think it would be helpful to briefly
> describe the problems they solve. The history section suggests
> that the problem solved is source code written in character
> encodings that lack support for (some of) the members of the basic
> character set. That is one of the problems, but not the only one.
> Digraphs also allow source code to be written in a subset of
> characters that allow it to be correctly interpreted in multiple
> incompatible character encodings. I think this is best exemplified
> by relating the digraphs to the variant members of EBCDIC.
> According to IBM documentation, the variant EBCDIC members that
> are also members of the basic character set include:
> > [...]
> > It is perhaps worth noting that digraphs are lacking for some
> members of the basic character set that do not have invariant
> encoding across all EBCDIC code pages.
>
> Yeah that's a good point. I tried alluding to that fact with the
> linked Google spreadsheet (Affected encodings). I'll try making that
> point more explicit for the next revision :)
>
Ah! I had missed the spreadsheet. That's a nice collection of encodings! :)
If possible, it would be helpful to include that content as an appendix
in the paper itself.
>
> > The paper could also explicitly mention that
> universal-character-names provide a modern alternative to digraphs.
>
> I really like the idea of using UCNs as a generic way of spelling
> various characters that have meaning in C++, especially if it also
> allows named universal characters to be used. I'd be happy to consider
> working on a proposal for this if there's appetite.
>
>
> We went down this road once and decided we didn't want all of the cost
> and complexity associated with UCNs identifiers and that the current
> limitations that they do not represent basic character set elements
> can't reasonably be removed.
Yes, and I think that is the right decision in general. But for cases
where we already recognize digraphs (and trigraphs outside of character
and string literals) as equivalent to their corresponding character,
allowing a UCN as well would not have much impact on existing
implementations.
>
>
>
> However, I do not think it truly provides an alternative to digraphs.
> If you look at the spreadsheet from the paper - this would only be a
> usable alternative for BS_4730, EBCDIC-ES, EBCDIC-ES-S, IBM275,
> LATIN-GREEK, EBCDIC-DK-NO, EBCDIC-UK, EBCDIC-US and IBM880 (using the
> iconv primary spellings). For the other 39 affected encodings that are
> otherwise usable, \ is missing - so you can't spell UCNs either. We
> currently have no alternative token for that.
>
> For the much nicer named universal characters you additionally run
> into the issue of requiring {}. Considering that \ is also required,
> this narrows the list down to BS_4730, EBCDIC-ES, EBCDIC-ES-S and
> LATIN-GREEK. That does not seem like a viable replacement unless we
> also introduce something for \.
>
Yeah, adding an alternative for \ would require some invention unless we
want to bring back the ??/ trigraph (but limit it to use outside of
character and string literals this time).
I would have to defer to our friends at IBM regarding how important the
code pages that don't support \ are for C++ code.
>
>
> I think the questions that need to be answered, by IBM is:
> - can they just keep using ??= pragma
> - can they use _Pragma
> - how serious about using digraphs are they, there does seem to be
> evidence they are currently doing that
> - why do they need a solution that impacts things unrelated to pragma ?
>
>
> From my own research, digraphs were not initially intended or suitable
> to support EBCDIC encoding and whether that changed is unclear.
They certainly aren't sufficient to support all EBCDIC code pages, but
it might be that they have been sufficient to support the EBCDIC code
pages people actually write C++ code in.
Tom.
>
>
>
>
> Thanks,
> Matthias
> --
> SG16 mailing list
> SG16_at_[hidden]
> https://lists.isocpp.org/mailman/listinfo.cgi/sg16
> Link to this post: http://lists.isocpp.org/sg16/2026/09/4859.php
>
>
>
>
> On Thu, Sep 24, 2026, 06:19 Matthias Wippich via SG16
> <sg16_at_[hidden]> wrote:
>
> Thanks for the feedback, I'll reply inline.
>
> > First, thank you for a thorough paper; I love seeing all the
> history presented!
>
> Happy to hear that. It was an immense pain to reconstruct. Out of the
> 17091 X3J11, X3J16, WG14 and WG21 documents hosted on open-std.org
> <http://open-std.org>
> (15391 after removing working drafts, editors reports etc), 293
> mentioned "trigraph", 169 "digraph" for a total of 351 documents.
> Unfortunately some of the early references use neither of those terms.
> I hope I didn't make any unfortunate mistakes or misrepresented
> anything - I've read through almost all of those documents to figure
> out what matters, but y'know.. the query might've been bad and I have
> no idea if there was any undetected conversion failures (I OCR'd all
> those pdf papers that had no better source.. that took a weekend).
> I've been meaning to ask Keld about the correctness of the early
> history.
>
> There was a few interesting (albeit irrelevant for this paper) ones.
> For example, John Skaller published WG21/N0259 in 1993, which proposed
> `addrof` as alternative to `bitand`. The motivation is largely
> ```
> int* a = bitand b;
> ```
> looking rather strange.
>
> > Have you reached out to Keld Simonsen for comment regarding the
> history and the original intent? I can put you in touch with him
> if you don't have contact information.
>
> I haven't yet, but will. Peter provided me with Keld's current
> email address.
>
> > The introduction states, "Digraphs are a complicated solution to
> a very old problem ...". I think it would be helpful to briefly
> describe the problems they solve. The history section suggests
> that the problem solved is source code written in character
> encodings that lack support for (some of) the members of the basic
> character set. That is one of the problems, but not the only one.
> Digraphs also allow source code to be written in a subset of
> characters that allow it to be correctly interpreted in multiple
> incompatible character encodings. I think this is best exemplified
> by relating the digraphs to the variant members of EBCDIC.
> According to IBM documentation, the variant EBCDIC members that
> are also members of the basic character set include:
> > [...]
> > It is perhaps worth noting that digraphs are lacking for some
> members of the basic character set that do not have invariant
> encoding across all EBCDIC code pages.
>
> Yeah that's a good point. I tried alluding to that fact with the
> linked Google spreadsheet (Affected encodings). I'll try making that
> point more explicit for the next revision :)
>
Ah! I had missed the spreadsheet. That's a nice collection of encodings! :)
If possible, it would be helpful to include that content as an appendix
in the paper itself.
>
> > The paper could also explicitly mention that
> universal-character-names provide a modern alternative to digraphs.
>
> I really like the idea of using UCNs as a generic way of spelling
> various characters that have meaning in C++, especially if it also
> allows named universal characters to be used. I'd be happy to consider
> working on a proposal for this if there's appetite.
>
>
> We went down this road once and decided we didn't want all of the cost
> and complexity associated with UCNs identifiers and that the current
> limitations that they do not represent basic character set elements
> can't reasonably be removed.
Yes, and I think that is the right decision in general. But for cases
where we already recognize digraphs (and trigraphs outside of character
and string literals) as equivalent to their corresponding character,
allowing a UCN as well would not have much impact on existing
implementations.
>
>
>
> However, I do not think it truly provides an alternative to digraphs.
> If you look at the spreadsheet from the paper - this would only be a
> usable alternative for BS_4730, EBCDIC-ES, EBCDIC-ES-S, IBM275,
> LATIN-GREEK, EBCDIC-DK-NO, EBCDIC-UK, EBCDIC-US and IBM880 (using the
> iconv primary spellings). For the other 39 affected encodings that are
> otherwise usable, \ is missing - so you can't spell UCNs either. We
> currently have no alternative token for that.
>
> For the much nicer named universal characters you additionally run
> into the issue of requiring {}. Considering that \ is also required,
> this narrows the list down to BS_4730, EBCDIC-ES, EBCDIC-ES-S and
> LATIN-GREEK. That does not seem like a viable replacement unless we
> also introduce something for \.
>
Yeah, adding an alternative for \ would require some invention unless we
want to bring back the ??/ trigraph (but limit it to use outside of
character and string literals this time).
I would have to defer to our friends at IBM regarding how important the
code pages that don't support \ are for C++ code.
>
>
> I think the questions that need to be answered, by IBM is:
> - can they just keep using ??= pragma
> - can they use _Pragma
> - how serious about using digraphs are they, there does seem to be
> evidence they are currently doing that
> - why do they need a solution that impacts things unrelated to pragma ?
>
>
> From my own research, digraphs were not initially intended or suitable
> to support EBCDIC encoding and whether that changed is unclear.
They certainly aren't sufficient to support all EBCDIC code pages, but
it might be that they have been sufficient to support the EBCDIC code
pages people actually write C++ code in.
Tom.
>
>
>
>
> Thanks,
> Matthias
> --
> SG16 mailing list
> SG16_at_[hidden]
> https://lists.isocpp.org/mailman/listinfo.cgi/sg16
> Link to this post: http://lists.isocpp.org/sg16/2026/09/4859.php
>
>
Received on 2026-09-24 20:20:11
