Date: Tue, 29 Sep 2026 18:20:21 -0400
On 9/29/26 5:59 PM, Robin Leroy wrote:
> Le mar. 29 sept. 2026 à 10:31, Jens Maurer via SG16
> <sg16_at_[hidden]> a écrit :
>
> The model is that trigraphs are mapped in phase 1 (which is
> implementation-defined),
> and the implementation is simply considered to use a (rather
> strange) source file
> encoding. Which is conforming.
>
> Whether that kind of allowance jeopardizes the portability
> promises of C++
> is another question.
>
>
> Le lun. 28 sept. 2026 à 23:09, Corentin Jabot via SG16
> <sg16_at_[hidden]> a écrit :
>
> The model is that we can magically map the trigraphs that would
> appear outside of string literal in later phases. Which is even
> implementable given the fact no one actually implement distinct
> phases. It's a reasonable-ish justification to support trigraphs
> without warnings.
>
>
> Le mar. 29 sept. 2026 à 23:36, Tom Honermann via SG16
> <sg16_at_[hidden]> a écrit :
>
> It sounds like the digraph model (since character and string
> literals aren't affected), but using the character sequences from
> trigraphs.
>
>
> Presumably the wording of phase 1 allows for a stateful encoding (a
> compiler could allow for ISO 2022 or SCSU source files).
>
> But then that should allow for an even stranger encoding where the
> byte corresponding to " in ASCII (or EBCDIC), as well as some other
> byte sequences, are simultaneously printing characters and shift
> functions, controlling a mode D, and where the byte sequence
> corresponding to %: in ASCII (or EBCDIC) represents U+0023 in mode D,
> and represents the sequence U+0025, U+003A outside of mode D.
>
> Admittedly raw string literals make that hypothetical encoding
> extremely weird.
So, basically, a stateful encoding that shifts state based on C++
tokenization rules? I love it! 😂 I might just have to ask Claude to do
something for me...
Tom.
>
> Best regards,
>
> Robin Leroy
> Le mar. 29 sept. 2026 à 10:31, Jens Maurer via SG16
> <sg16_at_[hidden]> a écrit :
>
> The model is that trigraphs are mapped in phase 1 (which is
> implementation-defined),
> and the implementation is simply considered to use a (rather
> strange) source file
> encoding. Which is conforming.
>
> Whether that kind of allowance jeopardizes the portability
> promises of C++
> is another question.
>
>
> Le lun. 28 sept. 2026 à 23:09, Corentin Jabot via SG16
> <sg16_at_[hidden]> a écrit :
>
> The model is that we can magically map the trigraphs that would
> appear outside of string literal in later phases. Which is even
> implementable given the fact no one actually implement distinct
> phases. It's a reasonable-ish justification to support trigraphs
> without warnings.
>
>
> Le mar. 29 sept. 2026 à 23:36, Tom Honermann via SG16
> <sg16_at_[hidden]> a écrit :
>
> It sounds like the digraph model (since character and string
> literals aren't affected), but using the character sequences from
> trigraphs.
>
>
> Presumably the wording of phase 1 allows for a stateful encoding (a
> compiler could allow for ISO 2022 or SCSU source files).
>
> But then that should allow for an even stranger encoding where the
> byte corresponding to " in ASCII (or EBCDIC), as well as some other
> byte sequences, are simultaneously printing characters and shift
> functions, controlling a mode D, and where the byte sequence
> corresponding to %: in ASCII (or EBCDIC) represents U+0023 in mode D,
> and represents the sequence U+0025, U+003A outside of mode D.
>
> Admittedly raw string literals make that hypothetical encoding
> extremely weird.
So, basically, a stateful encoding that shifts state based on C++
tokenization rules? I love it! 😂 I might just have to ask Claude to do
something for me...
Tom.
>
> Best regards,
>
> Robin Leroy
Received on 2026-09-29 22:20:28
