Date: Tue, 29 Sep 2026 01:41:49 -0400
Historically, trigraphs in string literals were exactly that kind of a problem, and those issues would be solved in the way that we all had to work around such trigraph issues, possibly using the ??? trigraph for the first ?.
I believe that kind of encoding is perfectly reasonable, and is similar to the rationale that GCC used to ignore salient whitespace following a ‘\’ line continuation before we established that as the conforming behavior (and I had more programs bitten by this than by trigraphs).
Digraphs never ran into this problem as they are alternative spellings, not text transformations, so never had an impact in string literals, header names, etc.
Bringing myself up to speed, I am very much in favor of papers like this that simplify the language from a legacy that is no longer needed, or used. My preference is always to deprecate before removal, but we opted to directly remove trigraphs without a period of deprecation so we have not just precedent, but relevant precedent.
As a heads up I am also working on a paper to reclassify the alternate textual spelling of operators as “alternative operators” rather than “alternative tokens”, i.e., to consider the alternative spellings only where an operator would be valid, and not universally so that ‘and’ is never a valid spelling to denote an rvalue reference. That would seem to be a natural complement to this paper, handling the remaining alternative tokens in the table. I do *not* expect to have that paper published in time for the October mailing, but might be able to share a preview with this group if interested in giving early feedback on direction.
AlisdairM
อลิสแดร์ม
Sent from my iPhone
On Sep 28, 2026, at 21:43, Fraser Gordon via SG16 <sg16_at_[hidden]> wrote:
--On Mon, 28 Sept 2026 at 17:09, Corentin Jabot via SG16 <sg16_at_[hidden]> wrote:On Mon, Sep 28, 2026, 23:01 Tom Honermann <tom_at_[hidden]> wrote:On 9/28/26 4:04 AM, Corentin Jabot wrote:EBCDIC-derived encodings use trigraphs and we should keep accepting the magical phase one mapping of trigraphs to be a conforming extension.I don't think trigraphs can ever be a conforming extension. The following example has well-defined behavior in ISO C++, but behaves differently when trigraphs are enabled. https://godbolt.org/z/o3xznsrGY.
The model is that we can magically map the trigraphs that would appear outside of string literal in later phases. Which is even implementable given the fact no one actually implement distinct phases. It's a reasonable-ish justification to support trigraphs without warnings.Would it be a conforming extension to declare an implementation-defined input encoding where e.g. the sequence '??!' happens to be the encoding for '|'? And similarly '%:' for '#' etc? It'd be an encoding with some very strange properties (in particular, needing lookahead!) but it might give the needed license where vendors need to support it. (The flaw that comes to mind for me is that the trigraph sequences become difficult to spell, but more indirection might overcome that: the encoding for '??!' in a string -- which I'm assuming is the only place it would be observable in a valid C++ program? -- could be encoded as the sequence ??""! ).
SG16 mailing list
SG16_at_[hidden]
https://lists.isocpp.org/mailman/listinfo.cgi/sg16
Link to this post: http://lists.isocpp.org/sg16/2026/09/4871.php
Received on 2026-09-29 05:42:08
