Markdown to HTML: raw HTML, safety, GFM differences and heading ids
Converting Markdown to HTML source is easy until the HTML gets pasted somewhere else, such as a CMS, a template or an email. Then the questions are about what the converter did with the HTML already inside the Markdown, which flavour it assumed, and whether the anchors still work.
Raw HTML inside Markdown: three choices
Markdown lets you write HTML directly, and a converter has to decide what to do with it. The Markdown to HTML tool on this site exposes the choice as a setting:
| Setting | <b onclick="x()">hi</b> becomes | Use when |
|---|---|---|
| Show it as text (default) | <b onclick=...> visible on the page | the Markdown is untrusted, or you are documenting HTML |
| Keep it as HTML | copied verbatim, onclick included | you wrote the Markdown yourself |
| Remove it | nothing | you want plain Markdown output only |
Keep mode is a straight copy. Anything written as raw HTML reaches the output as written, scripts and event handlers included, and the tool shows a note under the options when it is selected. The preview pane is unaffected either way, because it renders inside a sandboxed frame with scripts disabled. That protects the preview and not the HTML you copy out of it.
Markdown is not a sanitizer
It is a common assumption that converting Markdown makes content safe. A Markdown parser’s job is
translating syntax, and most of them, including the marked library used here, pass raw HTML through by
design. Safety is a separate step, and where it belongs depends on who wrote the text.
The tool does apply one protection regardless of mode: links and images written in Markdown syntax with a
scheme such as javascript: are neutralised. A [click](javascript:alert(1)) link comes out as plain text,
and the status line counts how many were removed. Web, mailto:, tel:, anchor and relative addresses
pass, and images may also use data:image/ for common formats. That check covers only Markdown-syntax
links. An <a href="javascript:..."> written as raw HTML is raw HTML, so it is escaped in the default mode but copied
unchanged in keep mode.
If you render untrusted Markdown on your own site, do not keep raw HTML from the converter and rely on
it. Run the final HTML through a sanitizer designed for the job, such as DOMPurify in the browser or
sanitize-html on a Node server, using an allowlist of the tags and attributes you accept. Do this at the
point where the HTML is put into the page, since that is the only place that knows the context.
Why List<String> vanishes
Writing List<String> in prose and finding List in the output is a Markdown trap, not a converter bug.
<String> is a valid HTML open tag as far as CommonMark is concerned, so it is treated as raw HTML. In the
default mode the converter shows it as text and nothing is lost, but in keep mode the browser reads it as an
unknown element and hides it. The same goes for <T>, <div> and <br> typed in a sentence. Put code
in backticks (`List<String>`) or fenced blocks, where the characters are escaped for you.
A related surprise: Markdown inside an HTML block is not processed unless a blank line separates it.
<div>
*hello*
</div>
stays as the literal text *hello* inside the div in CommonMark, while adding blank lines around the
*hello* line makes it an <em>. If the Markdown you are converting mixes HTML wrappers with Markdown
content, check the result with those blank lines in mind.
GitHub-flavoured or plain CommonMark
By default the tool runs GitHub-flavoured Markdown: tables, task lists (- [x]), ~~strikethrough~~ and
bare URLs turned into links. Switch the first option off for plain CommonMark, and tables become paragraphs
of pipe-separated text. This matters when the destination renderer is strict. If the place you will paste the
output is going to re-render Markdown rather than accept HTML, a table that only exists in GFM will not
survive.
The output is plain semantic HTML with no classes or inline styles, apart from language-xxx on the <code> inside a
fenced block. Syntax-highlighting libraries such as highlight.js and Prism look for that class. The result
takes on the styling of whatever page it lands in. An email client is the exception: support for
<style> blocks varies, and inline styles are the dependable route there, so expect to add them yourself.
Line breaks: one newline or two spaces
In CommonMark a single newline inside a paragraph is a soft break and renders as a space. A hard break
needs two trailing spaces or a backslash before the newline. Renderers disagree about the default: some
treat every newline as a <br>, which is why text copied between a chat or comment box and a README can
change shape. The option “Treat a single line break as <br>” turns that behaviour on. Leave it off
for documents, where hard-wrapped source lines should flow together, and turn it on for notes and chat
text that were typed with a newline for each line.
Heading ids and anchors that break
Links like page.html#what-changed depend on the id a renderer gives the heading. With Add id attributes
on, the tool lowercases the heading text, removes punctuation, joins words with hyphens and appends -1,
-2 to repeats. ## What changed becomes what-changed, ## Q&A: Setup becomes qa-setup, and a second
## Setup becomes setup-1. Letters in any language are kept, so ## 本次变化 gets the id 本次变化.
Other generators follow their own rules, and some drop non-ASCII letters entirely, which leaves a Chinese
heading with no usable anchor. A table of contents or cross-reference written against one generator’s
ids can break on another. After converting, open the Preview or read the id= values in the source. If a
link must never break, set the id explicitly in your own template instead of deriving it from the title.
Fragment or complete document
By default the output is a fragment, suitable for a CMS body field. Wrap in a complete HTML document adds
the doctype, a viewport meta tag and a <title> taken from the first h1, and the optional small stylesheet
makes the standalone file readable. Open links in a new tab adds target="_blank" with
rel="noopener noreferrer", to web links only. Minify removes the whitespace between tags and leaves
<pre> blocks alone. Conversion runs in the page, so pasting private documents is fine.
If your goal is a PDF or an image rather than HTML source, the Markdown viewer and exporter on this site is the better fit.