Date: Mon, 28 Sep 2026 08:29:09 +0000
On Monday, September 28th, 2026 at 10:04 AM, Corentin Jabot via SG16 <sg16_at_[hidden]> wrote:
> On Mon, Sep 28, 2026, 05:14 Yongwei Wu via SG16 <sg16_at_[hidden]> wrote:
>
>> First, I think the paper is a fantastic summary.
>>
>> Second, the encoding list in the Google Sheet may contain some false positives. As a Chinese, I noticed the weird GBGBK immediately. I have no idea what it is, and googling does not reveal interesting results, and GBGBK does not exist on macOS (Sequoia) or Ubuntu (24.04 LTS). Claude Sonnet says it might be an internal conversion module, but it does not give much detail. Currently I do not think it is worthwhile to dig out what it really is.
My first guess was that GBGBK was just GBK, but GBK is listed lower down (and indeed doesn't need digraphs).
> That table is missing some critical information: how many if these encodings are in common use, and used to write C++ code (or code in general)
>
> Or rather, the question of whether a given encoding needs digraphs is sort of irrelevant, the question is whether digraphs are used.
>
> EBCDIC-derived encodings use trigraphs and we should keep accepting the magical phase one mapping of trigraphs to be a conforming extension.
> And if we really want to do encoding agnostic pragma without trigraphs (which there is no evidence of anyone doing today), we don't need digraphs.
>
> As Matthias's research shows, digraphs were a solution to a problem that no longer existed by the time digraphs where shoved in the standard, and because of internalization, internet, standardization of keyboards, OSes, programming languages, communication protocols and so forth, products that don't support the range of ASCII characters would not have been viable, since before the 00s.
The table has lots of information but leaves me without the relevant outcomes. We want to find the relevance of digraphs in the current world. That means that we want to know
- Which encodings are used to write C++ code (potentially in places we cannot see)
- That need digraphs to write C++ code in the full language (ie, no leaving out parts)
- That have all the characters that you need to be able to write the digraphs in the first place
If the encoding doesn't have the characters to write the digraphs, it's not usable (so do not contribute to digraph relevance). If it doesn't need the digraphs to write C++, it doesn't contribute to digraph relevance. And if an encoding definitely is not used to write C++, it doesn't contribute to relevance.
Leaving out all of these - or at least grouping them to reduce the size of the table, for example grouping all "ascii-containing encodings" - leaves us with a much smaller table that provides information on the relevance of digraphs. From that point on, we can try to get a quantitative estimate on the amount of affected C++ code, by looking at proliferation of a particular encoding, or by seeing which compilers support the relevant encoding to start with. Candidate encodings without any compiler support are again not contributing to relevance.
The conclusion of that is an estimate of the amount of affected code, and should be usable to inform us of how much future code we should expect to be written in C++29 or later - ie, the versions we'd affect with the removal - as well as the amount of existing code that would no longer be valid C++29.
Maybe Tom and others are able to find some of those numbers from closed ecosystems that use these encodings that can inform us of the importance of digraphs? That'd be great - both for the encoding they use and for having a bigger factual basis behind extrapolation.
And for a final thing; can you add references for the important encodings in question where you found them, what links point to them etc. They are almost by definition obscure and old, so it'd help us greatly to start with your information anchors.
Regards,
Peter Bindels
> On Mon, Sep 28, 2026, 05:14 Yongwei Wu via SG16 <sg16_at_[hidden]> wrote:
>
>> First, I think the paper is a fantastic summary.
>>
>> Second, the encoding list in the Google Sheet may contain some false positives. As a Chinese, I noticed the weird GBGBK immediately. I have no idea what it is, and googling does not reveal interesting results, and GBGBK does not exist on macOS (Sequoia) or Ubuntu (24.04 LTS). Claude Sonnet says it might be an internal conversion module, but it does not give much detail. Currently I do not think it is worthwhile to dig out what it really is.
My first guess was that GBGBK was just GBK, but GBK is listed lower down (and indeed doesn't need digraphs).
> That table is missing some critical information: how many if these encodings are in common use, and used to write C++ code (or code in general)
>
> Or rather, the question of whether a given encoding needs digraphs is sort of irrelevant, the question is whether digraphs are used.
>
> EBCDIC-derived encodings use trigraphs and we should keep accepting the magical phase one mapping of trigraphs to be a conforming extension.
> And if we really want to do encoding agnostic pragma without trigraphs (which there is no evidence of anyone doing today), we don't need digraphs.
>
> As Matthias's research shows, digraphs were a solution to a problem that no longer existed by the time digraphs where shoved in the standard, and because of internalization, internet, standardization of keyboards, OSes, programming languages, communication protocols and so forth, products that don't support the range of ASCII characters would not have been viable, since before the 00s.
The table has lots of information but leaves me without the relevant outcomes. We want to find the relevance of digraphs in the current world. That means that we want to know
- Which encodings are used to write C++ code (potentially in places we cannot see)
- That need digraphs to write C++ code in the full language (ie, no leaving out parts)
- That have all the characters that you need to be able to write the digraphs in the first place
If the encoding doesn't have the characters to write the digraphs, it's not usable (so do not contribute to digraph relevance). If it doesn't need the digraphs to write C++, it doesn't contribute to digraph relevance. And if an encoding definitely is not used to write C++, it doesn't contribute to relevance.
Leaving out all of these - or at least grouping them to reduce the size of the table, for example grouping all "ascii-containing encodings" - leaves us with a much smaller table that provides information on the relevance of digraphs. From that point on, we can try to get a quantitative estimate on the amount of affected C++ code, by looking at proliferation of a particular encoding, or by seeing which compilers support the relevant encoding to start with. Candidate encodings without any compiler support are again not contributing to relevance.
The conclusion of that is an estimate of the amount of affected code, and should be usable to inform us of how much future code we should expect to be written in C++29 or later - ie, the versions we'd affect with the removal - as well as the amount of existing code that would no longer be valid C++29.
Maybe Tom and others are able to find some of those numbers from closed ecosystems that use these encodings that can inform us of the importance of digraphs? That'd be great - both for the encoding they use and for having a bigger factual basis behind extrapolation.
And for a final thing; can you add references for the important encodings in question where you found them, what links point to them etc. They are almost by definition obscure and old, so it'd help us greatly to start with your information anchors.
Regards,
Peter Bindels
Received on 2026-09-28 08:29:18
