Markdown has become the lingua franca of documentation: GitHub READMEs, static-site blogs, wikis, and AI chat interfaces all prefer it. But a huge amount of the world’s content still lives as HTML — legacy CMS exports, copied web pages, email newsletters, and WYSIWYG output. Moving that content into a Markdown workflow means converting it, and doing it by hand is tedious and error-prone.
An HTML to Markdown converter automates the mechanical part: headings become # markers, lists become - and 1. lines, links collapse to [text](url), and code blocks get their fences. What would take an hour of careful find-and-replace becomes a single paste.
This free converter handles the tags you actually encounter — headings, paragraphs, bold, italic, links, inline code, code blocks, ordered and unordered lists, blockquotes, and horizontal rules — while stripping scripts, styles, and other non-content tags. Conversion runs locally in your browser, so private intranet pages stay private.
How to use the HTML to Markdown converter
- Paste your HTML — a full page, a snippet, or CMS export. Click “Load sample” to try it first.
- Click “Convert to Markdown” and read through the result.
- Check links and code blocks — the two areas where source HTML varies most.
- Copy the Markdown into your README, docs folder, or static-site content directory.
- Going the other way? Our Markdown to HTML converter reverses the process.
Key features & benefits
- Comprehensive tag coverage. Headings h1–h6, paragraphs, emphasis, links, both list types, code, blockquotes, and rules.
- Code blocks preserved.
<pre><code>sections become fenced blocks with their content intact. - Junk stripped automatically. Scripts, styles, and presentational tags are removed rather than converted into noise.
- Readable output. Excess blank lines collapse and trailing whitespace is trimmed, so the Markdown is clean on arrival.
- Entity decoding built in.
<and friends become real characters in the Markdown output. - Private conversion. Everything runs in your browser — safe for content behind logins or under NDA.
Getting the best conversion results
The converter is faithful, which means messy HTML produces messy Markdown. A little source hygiene goes a long way.
Simplify before converting
If your HTML comes from Word or a page builder, expect nested spans and inline styles — these are stripped, but deeply nested lists may need a manual tidy afterwards. Converting the smallest meaningful chunk (the article body, not the whole page chrome) gives the cleanest result.
Verify links and images
Relative links like /about convert as-is and will break outside their original site — rewrite them as absolute URLs when migrating content. If your source mixes encoded entities into link text, our HTML entity encoder/decoder can pre-clean the HTML.
Check code samples
Inline <code> becomes backticks and blocks become fences, but syntax highlighting classes are dropped. If you publish to a platform with language-tagged fences, add the language name (e.g. ```js) manually — it is a ten-second job per block. For prose migrated this way, our word counter confirms nothing was lost in transit.
Frequently asked questions
Will it handle a full HTML page?
Yes, but you will get the best results pasting just the content region. Full pages include navigation, footers, and sidebars that convert into Markdown noise — select the article body in your browser’s inspector for a clean extraction.
What happens to images?
Standard <img> tags are converted to Markdown image syntax with their alt text and URL. Complex responsive markup (picture, srcset) is simplified to the primary source, so verify image URLs after conversion.
Are tables converted?
Basic tables are flattened to readable text rather than Markdown table syntax, because real-world HTML tables are often used for layout rather than data. For genuine data tables, converting via our CSV to JSON converter after exporting to CSV is usually more reliable.
Is the conversion lossless?
Content is preserved; presentation is not. Colors, fonts, exact spacing, and CSS classes have no Markdown equivalent and are intentionally dropped. If you need pixel fidelity, keep the HTML — Markdown is for content, not design.
Can I convert Markdown back to HTML later?
Absolutely — that is the standard round trip. Our Markdown to HTML converter takes the output of this tool straight back to clean HTML whenever you need it.
How It Works: Under the Hood
HTML-to-Markdown conversion is a tree transformation: the converter parses HTML into a DOM tree, then walks it depth-first, emitting Markdown syntax for each node type. Headings (h1–h6) become # prefixes, strong/em become **/*, links become [text](url), and lists become indented – or 1. items. Block elements (paragraphs, divs, blockquotes) get blank-line separation; inline elements stay in flow.
The hard cases are where converters earn their keep. Tables map to pipe syntax, but colspan/rowspan have no Markdown equivalent and get flattened or dropped. Nested lists need precise indentation (2–4 spaces per level, depending on flavor). Images become  — but base64 data-URI images bloat the Markdown enormously, so good converters warn about them. Anything without a Markdown equivalent (iframes, complex divs, custom widgets) falls back to embedded raw HTML, which most Markdown renderers pass through untouched.
Flavor matters: GitHub Flavored Markdown supports tables, strikethrough, and task lists; vanilla Markdown doesn’t. A converter targeting the wrong flavor produces syntax the destination won’t render.
Real-World Use Cases
- Migrating a blog to a static site: a developer exports 200 WordPress posts as HTML and converts them to Markdown for a Hugo/Jekyll site — turning a week of manual reformatting into an afternoon of cleanup.
- Writing GitHub READMEs from docs: a technical writer drafts in Google Docs, exports HTML, converts to Markdown, and pastes into README.md with formatting intact.
- Archiving web articles for notes: a researcher saves the readable HTML of articles, converts to Markdown, and files them in Obsidian — where Markdown is the native format and HTML would be clutter.
- Cleaning CMS exports: a content manager pulls product descriptions from a legacy CMS (div-soup HTML) and converts to clean Markdown before importing into a modern headless CMS.
- Preparing LLM training data: an ML engineer converts scraped web pages to Markdown because it’s dramatically more token-efficient than raw HTML for the same readable content.
Advanced Tips
- Strip the chrome before converting: remove nav bars, sidebars, footers, and ad containers from the HTML first. Converters faithfully convert everything, and cleaning 200 lines of nav links out of Markdown is painful.
- Check tables manually: always eyeball converted tables — merged cells, nested tables, and wide tables are the #1 source of mangled output. Fix them in the Markdown directly; it’s faster than fighting the converter.
- Normalize to one flavor: decide your target (GFM, CommonMark, etc.) before converting a batch. Mixed-flavor Markdown across a documentation site causes subtle rendering inconsistencies.
- Handle images deliberately: decide upfront whether to keep remote image URLs, download and re-host, or drop images. A 300-post migration with hotlinked images that later 404 is a slow-motion disaster.
Common Mistakes to Avoid
- Converting styled HTML expecting styled Markdown: Markdown has no colors, fonts, or alignment. If the source relies on inline styles for meaning (red text = warnings), that information is lost — add callout syntax manually.
- Ignoring relative URLs: converted links and images keep their original href/src values. Relative paths like /images/logo.png break the moment the Markdown lives somewhere else — rewrite them to absolute URLs.
- Trusting nested-list indentation blindly: different renderers want 2, 3, or 4 spaces per nesting level. Verify nested lists render correctly in your target renderer, not just the converter’s preview.
- Forgetting that Markdown can’t do everything: collapsible sections, tabs, and embedded videos need raw HTML fallbacks. Don’t waste time forcing them into “pure” Markdown — hybrid documents are normal and fine.
From CMS Export to Clean Docs in One Paste
Legacy CMS exports are the classic use case: thousands of words trapped in div-soup with inline styles. Paste the export, convert, and you get Markdown ready for a modern docs site or wiki — headings, lists, links, and code fences intact, while scripts, styles, and layout markup are stripped automatically. Review links and images afterward (the two things source HTML varies most on), then commit the result. Going the other direction later? The Markdown to HTML converter reverses the trip.
What Survives Conversion — and What Doesn’t
Structural content converts faithfully: headings, paragraphs, bold, italic, links, images, lists, blockquotes, code blocks, and tables become clean Markdown. What’s deliberately dropped: scripts, stylesheets, forms, and layout divs — presentation, not content. Embedded videos and iframes usually reduce to their fallback links, so check those spots manually. Heavily styled or JavaScript-rendered pages convert best after simplifying the source. For entity-encoded content from migrations, the HTML entity encoder decoder cleans it first.
Frequently Asked Questions
Does it preserve code syntax highlighting?
Code blocks keep their content and language fences, but highlighting colors aren’t part of Markdown — your renderer or docs theme applies those.
What happens to embedded videos or iframes?
They’re typically reduced to fallback links or removed, since Markdown has no native video syntax. Re-embed videos manually in your target platform.
Can I convert a whole website at once?
No — it converts one pasted document at a time. For multi-page migrations, convert each page separately and assemble the files afterward.