Markdown to HTML: raw HTML, safety, GFM differences and heading ids

Updated 5 min read

Converting Markdown to HTML source is easy until the HTML gets pasted somewhere else, such as a CMS, a template or an email. Then the questions are about what the converter did with the HTML already inside the Markdown, which flavour it assumed, and whether the anchors still work.

Raw HTML inside Markdown: three choices

Markdown lets you write HTML directly, and a converter has to decide what to do with it. The Markdown to HTML tool on this site exposes the choice as a setting:

Setting<b onclick="x()">hi</b> becomesUse when
Show it as text (default)&lt;b onclick=...&gt; visible on the pagethe Markdown is untrusted, or you are documenting HTML
Keep it as HTMLcopied verbatim, onclick includedyou wrote the Markdown yourself
Remove itnothingyou want plain Markdown output only

Keep mode is a straight copy. Anything written as raw HTML reaches the output as written, scripts and event handlers included, and the tool shows a note under the options when it is selected. The preview pane is unaffected either way, because it renders inside a sandboxed frame with scripts disabled. That protects the preview and not the HTML you copy out of it.

Markdown is not a sanitizer

It is a common assumption that converting Markdown makes content safe. A Markdown parser’s job is translating syntax, and most of them, including the marked library used here, pass raw HTML through by design. Safety is a separate step, and where it belongs depends on who wrote the text.

The tool does apply one protection regardless of mode: links and images written in Markdown syntax with a scheme such as javascript: are neutralised. A [click](javascript:alert(1)) link comes out as plain text, and the status line counts how many were removed. Web, mailto:, tel:, anchor and relative addresses pass, and images may also use data:image/ for common formats. That check covers only Markdown-syntax links. An <a href="javascript:..."> written as raw HTML is raw HTML, so it is escaped in the default mode but copied unchanged in keep mode.

If you render untrusted Markdown on your own site, do not keep raw HTML from the converter and rely on it. Run the final HTML through a sanitizer designed for the job, such as DOMPurify in the browser or sanitize-html on a Node server, using an allowlist of the tags and attributes you accept. Do this at the point where the HTML is put into the page, since that is the only place that knows the context.

Why List<String> vanishes

Writing List<String> in prose and finding List in the output is a Markdown trap, not a converter bug. <String> is a valid HTML open tag as far as CommonMark is concerned, so it is treated as raw HTML. In the default mode the converter shows it as text and nothing is lost, but in keep mode the browser reads it as an unknown element and hides it. The same goes for <T>, <div> and <br> typed in a sentence. Put code in backticks (`List<String>`) or fenced blocks, where the characters are escaped for you.

A related surprise: Markdown inside an HTML block is not processed unless a blank line separates it.

<div>
*hello*
</div>

stays as the literal text *hello* inside the div in CommonMark, while adding blank lines around the *hello* line makes it an <em>. If the Markdown you are converting mixes HTML wrappers with Markdown content, check the result with those blank lines in mind.

GitHub-flavoured or plain CommonMark

By default the tool runs GitHub-flavoured Markdown: tables, task lists (- [x]), ~~strikethrough~~ and bare URLs turned into links. Switch the first option off for plain CommonMark, and tables become paragraphs of pipe-separated text. This matters when the destination renderer is strict. If the place you will paste the output is going to re-render Markdown rather than accept HTML, a table that only exists in GFM will not survive.

The output is plain semantic HTML with no classes or inline styles, apart from language-xxx on the <code> inside a fenced block. Syntax-highlighting libraries such as highlight.js and Prism look for that class. The result takes on the styling of whatever page it lands in. An email client is the exception: support for <style> blocks varies, and inline styles are the dependable route there, so expect to add them yourself.

Line breaks: one newline or two spaces

In CommonMark a single newline inside a paragraph is a soft break and renders as a space. A hard break needs two trailing spaces or a backslash before the newline. Renderers disagree about the default: some treat every newline as a <br>, which is why text copied between a chat or comment box and a README can change shape. The option “Treat a single line break as <br>” turns that behaviour on. Leave it off for documents, where hard-wrapped source lines should flow together, and turn it on for notes and chat text that were typed with a newline for each line.

Heading ids and anchors that break

Links like page.html#what-changed depend on the id a renderer gives the heading. With Add id attributes on, the tool lowercases the heading text, removes punctuation, joins words with hyphens and appends -1, -2 to repeats. ## What changed becomes what-changed, ## Q&A: Setup becomes qa-setup, and a second ## Setup becomes setup-1. Letters in any language are kept, so ## 本次变化 gets the id 本次变化.

Other generators follow their own rules, and some drop non-ASCII letters entirely, which leaves a Chinese heading with no usable anchor. A table of contents or cross-reference written against one generator’s ids can break on another. After converting, open the Preview or read the id= values in the source. If a link must never break, set the id explicitly in your own template instead of deriving it from the title.

Fragment or complete document

By default the output is a fragment, suitable for a CMS body field. Wrap in a complete HTML document adds the doctype, a viewport meta tag and a <title> taken from the first h1, and the optional small stylesheet makes the standalone file readable. Open links in a new tab adds target="_blank" with rel="noopener noreferrer", to web links only. Minify removes the whitespace between tags and leaves <pre> blocks alone. Conversion runs in the page, so pasting private documents is fine.

If your goal is a PDF or an image rather than HTML source, the Markdown viewer and exporter on this site is the better fit.

Open the tool: Markdown to HTML Converter

Back to guides

More guides