Skip to content

HTML comment inside inline raw HTML creates malformed output #1643

Description

@posita

Originally mis-filed against Zensical as zensical/zensical#981. @squidfunk graciously deduced this project as the root cause. The repro repo illustrates via Zensical, but let me know if I should provide a more compact version that doesn't rely on that as the vehicle.

Context

Discovered while investigating zensical/zensical#975.

Bug description

This (in docs/symptom.md) ...

<picture>
  <!-- TODO(@posita): https://github.com/zensical/zensical/issues/975 - What the heck? 👉 -->
  <source media="(prefers-color-scheme: dark)" srcset="../images/example.svg">
  <img alt="Example" src="../images/example.svg">
</picture>

... renders as this (in site/symptom/index.html) ...

<p><picture>
  <!-- TODO(@posita): https://github.com/zensical/zensical/issues/975 - What the heck? 👉 --></p>
<p><source media="(prefers-color-scheme: dark)" srcset="../images/example.svg">
  <img alt="Example" src="../../images/example.svg">
</picture></p>

Note the insertion of </p> ... <p> in the middle of the <picture> ... </picture> block immediately after the comment.

Reproduction / Steps to reproduce

git clone https://github.com/posita/zensical-picture-repro.git
cd zensical-picture-repro
git checkout posita/0/comment-in-picture-destroys-picture  # branch pertaining to this issue
zensical build --clean
grep --after-context=4 '<picture>' site/symptom/index.html

Environment

% uname -prism
Linux 7.1.5-76070105-generic x86_64 x86_64 x86_64
% python3 --version
Python 3.11.15
% zensical --version
0.0.65

Related links

Repro: https://github.com/posita/zensical-picture-repro/tree/posita/0/comment-in-picture-destroys-picture
Zip file: https://github.com/posita/zensical-picture-repro/archive/refs/heads/posita/0/comment-in-picture-destroys-picture.zip

Activity

  1. facelessuser commented on Sep 28, 2026

    @facelessuser
    Collaborator

    A picture is inherently an inline element, so Python Markdown treats it that way. Python Markdown also treats HTML comments as their own block. So the comment breaks your content into two paragraphs.

  2. posita commented on Sep 29, 2026

    @posita
    Author

    A picture is inherently an inline element, so Python Markdown treats it that way. Python Markdown also treats HTML comments as their own block. So the comment breaks your content into two paragraphs.

    Sorry if I'm confused. Does that mean my example is invalid HTML? Apologies if that's the case.

  3. facelessuser commented on Sep 29, 2026

    @facelessuser
    Collaborator

    It's not necessarily invalid, but this happens if you try to do this with any inline element in Python Markdown.

    If <picture> was treated as a block, the Python Markdown parser would handle it differently, but it is treated as inline. This would happen if you did this with any inline element in Python Markdown.

    It should be noted that browsers treat <picture> as inline by default, so it being inline is not a bug.

  4. waylan commented on Sep 29, 2026

    @waylan
    Member

    this happens if you try to do this with any inline element in Python Markdown.

    The key is the above quote. This is not unique to <picture> but all inline elements. Therefore, a fix which only addresses <picture> is not the correct fix (and that is why I closed #1645). If we were to change anything it would be to recognize when a raw HTML comment is within a raw inline HTML element and not treat that comment as block-level. Of course, in all other contexts, comments should continue to be treated as block-level. Whether we should make that change however, could depend on the rules and how the reference implementation behaves. I will need to do some investigating.

  5. waylan commented on Sep 29, 2026

    @waylan
    Member

    Python-Markdown's behavior is currently as follows:

    Input

    <span><!-- comment --></span>

    Output

    <p><span><!-- comment --></span></p>

    This first example behaves as expected. But then when a comment is not on a line by itself, Markdown has always treated it differently. This matches the reference implementation and does not need to change.

    Input

    <span>
    <!-- comment -->
    </span>

    Output

    <p><span></p>
    <!-- comment -->
    <p></span></p>

    This is the issue in its most basic form. Interestingly, the reference implementation (markdown.pl) gets this correct. We should be outputting the same:

    Expected Output

    <p><span>
    <!-- comment -->
    </span></p>

    This is confirmed as a bug which needs to be fixed.

    For completeness, I also checked this input:

    Input

    <span>
    
    <!-- comment -->
    
    </span>

    and the reference implementation outputs the following:

    Output of markdown.pl

    <p><span></p>
    
    <!-- comment -->
    
    <p></span></p>

    I would be okay with calling that a bug in the reference implementation. We should match the behavior of the previous example (presumably while preserving whitespace).

    Expected Output

    <p><span>
    
    <!-- comment -->
    
    </span></p>
  6. added
    bugBug report.
    coreRelated to the core parser code.
    confirmedConfirmed bug report or approved feature request.
    on Sep 29, 2026
  7. changed the title [-]HTML comment inside `<picture>...</picture>` in Markdown creates malformed HTML[/-] [+]HTML comment inside inline raw HTML creates malformed output[/+] on Sep 29, 2026
  8. facelessuser commented on Sep 29, 2026

    @facelessuser
    Collaborator

    Based on how block handling is done throughout Python Markdown, I'm not sure if we can reasonably handle it...well, not without some effort, at least.

    <span>
    
    <!-- comment -->
    
    </span>

    I think this is certainly possible:

    <span>
    <!-- comment -->
    </span>

    What about things like below? How does the reference implementation handle this?

    text <span>
    
    <!-- comment -->
    
    </span>
  9. waylan commented on Sep 29, 2026

    @waylan
    Member

    Thinking further about the example with blank lines, I think we need to follow the refernce implementation. raw inline HTML does not change how blocks are parsed. Specifically, consider this:

    input

    <span>
    
    </span>

    That is (correctly) rendered as

    <p><span></p>
    <p></span></p>

    Therefore, we shouldn't special case it when there is a comment inserted in the middle.

    Conversely, this:

    <span>
    </span>

    renders as

    <p><span>
    </span></p>

    So inserting a comment in the middle should not change that.

    A few more examples.

    Input

    <span><!-- comment -->
    </span>

    Output (correct)

    <p><span><!-- comment -->
    </span></p>

    Input

    <span>
    <!-- comment --></span>

    Output (incorrect)

    <p><span></p>
    <!-- comment -->
    <p></span></p>

    Expected Output (what markdown.pl outputs)

    <p><span>
    <!-- comment --></span></p>
  10. facelessuser commented on Sep 29, 2026

    @facelessuser
    Collaborator

    Okay, this sounds like it requires a change in our HTML parser, then. It should take context as to whether a comment is its own block or not, I assume. I haven't looked at the code yet. Just thinking out loud.

  11. added a commit that references this issue on Oct 8, 2026
    0bf535b
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugBug report.confirmedConfirmed bug report or approved feature request.coreRelated to the core parser code.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions