Le mar. 29 sept. 2026 à 10:31, Jens Maurer via SG16 <sg16@lists.isocpp.org> a écrit :
The model is that trigraphs are mapped in phase 1 (which is implementation-defined),
and the implementation is simply considered to use a (rather strange) source file
encoding. Which is conforming.
Whether that kind of allowance jeopardizes the portability promises of C++
is another question.
Le lun. 28 sept. 2026 à 23:09, Corentin Jabot via SG16 <sg16@lists.isocpp.org> a écrit :The model is that we can magically map the trigraphs that would appear outside of string literal in later phases. Which is even implementable given the fact no one actually implement distinct phases. It's a reasonable-ish justification to support trigraphs without warnings.
Le mar. 29 sept. 2026 à 23:36, Tom Honermann via SG16 <sg16@lists.isocpp.org> a écrit :
It sounds like the digraph model (since character and string literals aren't affected), but using the character sequences from trigraphs.
Presumably the wording of phase 1 allows for a stateful encoding (a compiler could allow for ISO 2022 or SCSU source files).
But then that should allow for an even stranger encoding where the byte corresponding to " in ASCII (or EBCDIC), as well as some other byte sequences, are simultaneously printing characters and shift functions, controlling a mode D, and where the byte sequence corresponding to %: in ASCII (or EBCDIC) represents U+0023 in mode D, and represents the sequence U+0025, U+003A outside of mode D.
Admittedly raw string literals make that hypothetical encoding extremely weird.
So, basically, a stateful encoding that shifts state based on C++ tokenization rules? I love it! 😂 I might just have to ask Claude to do something for me...
Tom.
Best regards,
Robin Leroy