C++ Logo

sg16

Advanced search

Re: [isocpp-sg16] Thoughts on P4235R0: Removing Digraphs

From: Corentin Jabot <corentinjabot_at_[hidden]>
Date: Wed, 30 Sep 2026 11:25:24 +0200
On Tue, 29 Sept 2026 at 23:35, Tom Honermann <tom_at_[hidden]> wrote:

> On 9/28/26 5:09 PM, Corentin Jabot wrote:
>
>
>
> On Mon, Sep 28, 2026, 23:01 Tom Honermann <tom_at_[hidden]> wrote:
>
>> On 9/28/26 4:04 AM, Corentin Jabot wrote:
>>
>>
>>
>> On Mon, Sep 28, 2026, 05:14 Yongwei Wu via SG16 <sg16_at_[hidden]>
>> wrote:
>>
>>> First, I think the paper is a fantastic summary.
>>>
>>> Second, the encoding list in the Google Sheet may contain some false
>>> positives. As a Chinese, I noticed the weird GBGBK immediately. I have no
>>> idea what it is, and googling does not reveal interesting results, and
>>> GBGBK does not exist on macOS (Sequoia) or Ubuntu (24.04 LTS). Claude
>>> Sonnet says it might be an internal conversion module, but it does not give
>>> much detail. Currently I do not think it is worthwhile to dig out what it
>>> really is.
>>>
>>> I think there are fewer in-use encodings that require digraphs than the
>>> spreadsheet implies. From what I can recognize, GBBIG5 may be similar to
>>> GBGBK. What is ISO646 (as versus ISO646-something), which does not really
>>> make sense? On macOS there is an ISO646-BASIC:1983 (I don't see
>>> an equivalent on Ubuntu), but it behaves differently than the spreadsheet
>>> too.
>>>
>>
>> That table is missing some critical information: how many if these
>> encodings are in common use, and used to write C++ code (or code in general)
>>
>> Or rather, the question of whether a given encoding needs digraphs is
>> sort of irrelevant, the question is whether digraphs are used.
>>
>> A difficult question to answer, unfortunately. Maybe our friends at IBM
>> can share some insight.
>>
>>
>> EBCDIC-derived encodings use trigraphs and we should keep accepting the
>> magical phase one mapping of trigraphs to be a conforming extension.
>>
>> I don't think trigraphs can ever be a conforming extension. The following
>> example has well-defined behavior in ISO C++, but behaves differently when
>> trigraphs are enabled. https://godbolt.org/z/o3xznsrGY.
>>
> The model is that we can magically map the trigraphs that would appear
> outside of string literal in later phases. Which is even implementable
> given the fact no one actually implement distinct phases. It's a
> reasonable-ish justification to support trigraphs without warnings.
>
> Thanks to Jens for reminding me to think of trigraphs as a phase 1
> translation.
>
> Corentin, I'm not sure I understand your suggestion. It sounds like the
> digraph model (since character and string literals aren't affected), but
> using the character sequences from trigraphs.
>

I'm not suggesting anything.
I'm saying the phase 1 wording gives enough flexibility to perform or not
trigraphs replacement, and that an implementation can choose to do that
replacement or not in string, raw strings etc.


To get us back on topic, I'm arguing that we should do what the paper is
proposing.
The status quo is that

  - EBCDIC encodings can use trigraphs as a conforming extension
  - Encodings for which digraphs were standardized are no longer in use -
they were obsoleted by the time digraphs were standardized.
  - The standard librarie prevent the existence of a conforming
implementation that would not support ascii-like encoded source files
  - There is a basic expectation that code is written in an ASCII superset.
I imagine IBM might want to write code in e.g, rust, and I doubt rust has
any appetite for digraphs. I cannot imagine anyone being keen on spreading
digraphs all over source code.
  - We know that ??=pragma is the one important use case. i.e we should
assume new code in an EBCDIC environment ought to be written in IBM47 and
if the compiler assume another encoding, switching seems like a reasonable
feature, which leaves us to ponder
    - is _Pragma a reasonable substitute
    - can an implementation just treat `??=pragma` as a magic marker as a
conforming extension - i don't think anything needs standardizing
    - Given that to this day, %:pragma is presumably not something IBM
uses, why should we suddenly try to preserve that use case?




> Tom.
>
>
> #include <cstdio>
>> int main() {
>> std::printf("It's like magic: '??!'\n");
>> }
>>
>> And if we really want to do encoding agnostic pragma without trigraphs
>> (which there is no evidence of anyone doing today), we don't need digraphs.
>>
>> As Matthias's research shows, digraphs were a solution to a problem that
>> no longer existed by the time digraphs where shoved in the standard, and
>> because of internalization, internet, standardization of keyboards, OSes,
>> programming languages, communication protocols and so forth, products that
>> don't support the range of ASCII characters would not have been viable,
>> since before the 00s.
>>
>> Shift JIS is the only popular encoding that is not ASCII compatible but
>> we are all happy to pretend that ¥ and \ are the same character (same for
>> ~/overline). Which is bonkers, but it is what it is. (Tom table also makes
>> that mistake).
>>
>> I just copied the data from the spreadsheet; I didn't attempt any
>> validation of it.
>>
>> Tom.
>>
>>
>>
>>
>>> My 2 cents.
>>>
>>> On Sun, 27 Sept 2026 at 11:21, Tom Honermann via SG16 <
>>> sg16_at_[hidden]> wrote:
>>>
>>>> On 9/24/26 8:27 PM, Matthias Wippich wrote:
>>>> > On Thu, Sep 24, 2026 at 10:20 PM Tom Honermann via SG16
>>>> > <sg16_at_[hidden]> wrote:
>>>> >> If possible, it would be helpful to include that content as an
>>>> appendix in the paper itself.
>>>> > I'll try to do that, but I'm worried that it'll be impossible to
>>>> > format correctly. Maybe a link to a CSV export of that table (hosted
>>>> > on github/gist) is sufficient? The google link might indeed not be
>>>> > very stable.
>>>>
>>>> Any original content used to justify the motivation or design choices
>>>> really should be in the paper for the historical record.
>>>>
>>>> Attached is a quick hack job I did that I would personally call good
>>>> enough if you want to use it. The HTML isn't pretty and I didn't try to
>>>> style it in any way, so there is plenty of opportunity to make it
>>>> prettier.
>>>>
>>>> Tom.
>>>> --
>>>> SG16 mailing list
>>>> SG16_at_[hidden]
>>>> https://lists.isocpp.org/mailman/listinfo.cgi/sg16
>>>> Link to this post: http://lists.isocpp.org/sg16/2026/09/4865.php
>>>>
>>>
>>>
>>> --
>>> Yongwei Wu
>>> URL: http://wyw.dcweb.cn/
>>> --
>>> SG16 mailing list
>>> SG16_at_[hidden]
>>> https://lists.isocpp.org/mailman/listinfo.cgi/sg16
>>> Link to this post: http://lists.isocpp.org/sg16/2026/09/4866.php
>>>
>>

Received on 2026-09-30 09:25:50