Date: Tue, 29 Sep 2026 17:35:57 -0400
On 9/28/26 5:09 PM, Corentin Jabot wrote:
>
>
> On Mon, Sep 28, 2026, 23:01 Tom Honermann <tom_at_[hidden]> wrote:
>
> On 9/28/26 4:04 AM, Corentin Jabot wrote:
>>
>>
>> On Mon, Sep 28, 2026, 05:14 Yongwei Wu via SG16
>> <sg16_at_[hidden]> wrote:
>>
>> First, I think the paper is a fantastic summary.
>>
>> Second, the encoding list in the Google Sheet may contain
>> some false positives. As a Chinese, I noticed the weird GBGBK
>> immediately. I have no idea what it is, and googling does not
>> reveal interesting results, and GBGBK does not exist on macOS
>> (Sequoia) or Ubuntu (24.04 LTS). Claude Sonnet says it might
>> be an internal conversion module, but it does not give much
>> detail. Currently I do not think it is worthwhile to dig out
>> what it really is.
>>
>> I think there are fewer in-use encodings that require
>> digraphs than the spreadsheet implies. From what I can
>> recognize, GBBIG5 may be similar to GBGBK. What is ISO646 (as
>> versus ISO646-something), which does not really make sense?
>> On macOS there is an ISO646-BASIC:1983 (I don't see
>> an equivalent on Ubuntu), but it behaves differently than the
>> spreadsheet too.
>>
>>
>> That table is missing some critical information: how many if
>> these encodings are in common use, and used to write C++ code (or
>> code in general)
>>
>> Or rather, the question of whether a given encoding needs
>> digraphs is sort of irrelevant, the question is whether digraphs
>> are used.
> A difficult question to answer, unfortunately. Maybe our friends
> at IBM can share some insight.
>>
>> EBCDIC-derived encodings use trigraphs and we should keep
>> accepting the magical phase one mapping of trigraphs to be a
>> conforming extension.
>
> I don't think trigraphs can ever be a conforming extension. The
> following example has well-defined behavior in ISO C++, but
> behaves differently when trigraphs are enabled.
> https://godbolt.org/z/o3xznsrGY.
>
> The model is that we can magically map the trigraphs that would appear
> outside of string literal in later phases. Which is even implementable
> given the fact no one actually implement distinct phases. It's a
> reasonable-ish justification to support trigraphs without warnings.
Thanks to Jens for reminding me to think of trigraphs as a phase 1
translation.
Corentin, I'm not sure I understand your suggestion. It sounds like the
digraph model (since character and string literals aren't affected), but
using the character sequences from trigraphs.
Tom.
>
> #include <cstdio>
> int main() {
> std::printf("It's like magic: '??!'\n");
> }
>
>> And if we really want to do encoding agnostic pragma without
>> trigraphs (which there is no evidence of anyone doing today), we
>> don't need digraphs.
>>
>> As Matthias's research shows, digraphs were a solution to a
>> problem that no longer existed by the time digraphs where shoved
>> in the standard, and because of internalization, internet,
>> standardization of keyboards, OSes, programming languages,
>> communication protocols and so forth, products that don't support
>> the range of ASCII characters would not have been viable, since
>> before the 00s.
>>
>> Shift JIS is the only popular encoding that is not ASCII
>> compatible but we are all happy to pretend that ¥ and \ are the
>> same character (same for ~/overline). Which is bonkers, but it is
>> what it is. (Tom table also makes that mistake).
>
> I just copied the data from the spreadsheet; I didn't attempt any
> validation of it.
>
> Tom.
>
>>
>>
>>
>> My 2 cents.
>>
>> On Sun, 27 Sept 2026 at 11:21, Tom Honermann via SG16
>> <sg16_at_[hidden]> wrote:
>>
>> On 9/24/26 8:27 PM, Matthias Wippich wrote:
>> > On Thu, Sep 24, 2026 at 10:20 PM Tom Honermann via SG16
>> > <sg16_at_[hidden]> wrote:
>> >> If possible, it would be helpful to include that
>> content as an appendix in the paper itself.
>> > I'll try to do that, but I'm worried that it'll be
>> impossible to
>> > format correctly. Maybe a link to a CSV export of that
>> table (hosted
>> > on github/gist) is sufficient? The google link might
>> indeed not be
>> > very stable.
>>
>> Any original content used to justify the motivation or
>> design choices
>> really should be in the paper for the historical record.
>>
>> Attached is a quick hack job I did that I would
>> personally call good
>> enough if you want to use it. The HTML isn't pretty and I
>> didn't try to
>> style it in any way, so there is plenty of opportunity to
>> make it prettier.
>>
>> Tom.
>> --
>> SG16 mailing list
>> SG16_at_[hidden]
>> https://lists.isocpp.org/mailman/listinfo.cgi/sg16
>> Link to this post:
>> http://lists.isocpp.org/sg16/2026/09/4865.php
>>
>>
>>
>> --
>> Yongwei Wu
>> URL: http://wyw.dcweb.cn/
>> --
>> SG16 mailing list
>> SG16_at_[hidden]
>> https://lists.isocpp.org/mailman/listinfo.cgi/sg16
>> Link to this post: http://lists.isocpp.org/sg16/2026/09/4866.php
>>
>
>
> On Mon, Sep 28, 2026, 23:01 Tom Honermann <tom_at_[hidden]> wrote:
>
> On 9/28/26 4:04 AM, Corentin Jabot wrote:
>>
>>
>> On Mon, Sep 28, 2026, 05:14 Yongwei Wu via SG16
>> <sg16_at_[hidden]> wrote:
>>
>> First, I think the paper is a fantastic summary.
>>
>> Second, the encoding list in the Google Sheet may contain
>> some false positives. As a Chinese, I noticed the weird GBGBK
>> immediately. I have no idea what it is, and googling does not
>> reveal interesting results, and GBGBK does not exist on macOS
>> (Sequoia) or Ubuntu (24.04 LTS). Claude Sonnet says it might
>> be an internal conversion module, but it does not give much
>> detail. Currently I do not think it is worthwhile to dig out
>> what it really is.
>>
>> I think there are fewer in-use encodings that require
>> digraphs than the spreadsheet implies. From what I can
>> recognize, GBBIG5 may be similar to GBGBK. What is ISO646 (as
>> versus ISO646-something), which does not really make sense?
>> On macOS there is an ISO646-BASIC:1983 (I don't see
>> an equivalent on Ubuntu), but it behaves differently than the
>> spreadsheet too.
>>
>>
>> That table is missing some critical information: how many if
>> these encodings are in common use, and used to write C++ code (or
>> code in general)
>>
>> Or rather, the question of whether a given encoding needs
>> digraphs is sort of irrelevant, the question is whether digraphs
>> are used.
> A difficult question to answer, unfortunately. Maybe our friends
> at IBM can share some insight.
>>
>> EBCDIC-derived encodings use trigraphs and we should keep
>> accepting the magical phase one mapping of trigraphs to be a
>> conforming extension.
>
> I don't think trigraphs can ever be a conforming extension. The
> following example has well-defined behavior in ISO C++, but
> behaves differently when trigraphs are enabled.
> https://godbolt.org/z/o3xznsrGY.
>
> The model is that we can magically map the trigraphs that would appear
> outside of string literal in later phases. Which is even implementable
> given the fact no one actually implement distinct phases. It's a
> reasonable-ish justification to support trigraphs without warnings.
Thanks to Jens for reminding me to think of trigraphs as a phase 1
translation.
Corentin, I'm not sure I understand your suggestion. It sounds like the
digraph model (since character and string literals aren't affected), but
using the character sequences from trigraphs.
Tom.
>
> #include <cstdio>
> int main() {
> std::printf("It's like magic: '??!'\n");
> }
>
>> And if we really want to do encoding agnostic pragma without
>> trigraphs (which there is no evidence of anyone doing today), we
>> don't need digraphs.
>>
>> As Matthias's research shows, digraphs were a solution to a
>> problem that no longer existed by the time digraphs where shoved
>> in the standard, and because of internalization, internet,
>> standardization of keyboards, OSes, programming languages,
>> communication protocols and so forth, products that don't support
>> the range of ASCII characters would not have been viable, since
>> before the 00s.
>>
>> Shift JIS is the only popular encoding that is not ASCII
>> compatible but we are all happy to pretend that ¥ and \ are the
>> same character (same for ~/overline). Which is bonkers, but it is
>> what it is. (Tom table also makes that mistake).
>
> I just copied the data from the spreadsheet; I didn't attempt any
> validation of it.
>
> Tom.
>
>>
>>
>>
>> My 2 cents.
>>
>> On Sun, 27 Sept 2026 at 11:21, Tom Honermann via SG16
>> <sg16_at_[hidden]> wrote:
>>
>> On 9/24/26 8:27 PM, Matthias Wippich wrote:
>> > On Thu, Sep 24, 2026 at 10:20 PM Tom Honermann via SG16
>> > <sg16_at_[hidden]> wrote:
>> >> If possible, it would be helpful to include that
>> content as an appendix in the paper itself.
>> > I'll try to do that, but I'm worried that it'll be
>> impossible to
>> > format correctly. Maybe a link to a CSV export of that
>> table (hosted
>> > on github/gist) is sufficient? The google link might
>> indeed not be
>> > very stable.
>>
>> Any original content used to justify the motivation or
>> design choices
>> really should be in the paper for the historical record.
>>
>> Attached is a quick hack job I did that I would
>> personally call good
>> enough if you want to use it. The HTML isn't pretty and I
>> didn't try to
>> style it in any way, so there is plenty of opportunity to
>> make it prettier.
>>
>> Tom.
>> --
>> SG16 mailing list
>> SG16_at_[hidden]
>> https://lists.isocpp.org/mailman/listinfo.cgi/sg16
>> Link to this post:
>> http://lists.isocpp.org/sg16/2026/09/4865.php
>>
>>
>>
>> --
>> Yongwei Wu
>> URL: http://wyw.dcweb.cn/
>> --
>> SG16 mailing list
>> SG16_at_[hidden]
>> https://lists.isocpp.org/mailman/listinfo.cgi/sg16
>> Link to this post: http://lists.isocpp.org/sg16/2026/09/4866.php
>>
Received on 2026-09-29 21:36:02
