Skip to content

htmlToMarkdown

Useful and high-quality HTML to Markdown converter. Internally, it just calls htmlToMdast and mdastToMarkdown in that order.

Usage

import { htmlToMarkdown } from "webforai";
 
const html = "<h1>Hello, world!</h1>";
const markdown = htmlToMarkdown(html);
=> "# Hello, world!"

Returns

string

The converted Markdown string.

htmlToMarkdownWithMetadata

Takes the same arguments and returns the Markdown together with what the conversion learned about the page:

import { htmlToMarkdownWithMetadata } from "webforai";
 
const html = "<html><head><title>Hello</title></head><body><article><p>Hello, world!</p></article></body></html>";
const { markdown, metadata, extraction } = htmlToMarkdownWithMetadata(html, { url: "https://example.com/" });
// metadata: { title, description, author, published, modified, siteName, canonicalUrl, lang, image, type }
// extraction: { extractor: "kiwame" | "takumi" | "adapter", confidence?: number }

extraction.confidence (0–1, kiwame only) estimates how well the selection matches the page's main content; see How it works.

Parameters

htmlOrHast

type: string | Hast

The HTML string or HAST tree to convert.

const markdown = htmlToMarkdown("<h1>Hello, world!</h1>");
// => "# Hello, world!"

options.baseUrl

type: string

The base URL to use for replacing relative links.

const markdown = htmlToMarkdown("<a href='/foo'>bar</a>", {
	baseUrl: "https://example.com",
});
// => "[bar](https://example.com/foo)"

options.url

type: string

The page's URL. Selects a site adapter (GitHub, MDN, …) and anchors the page title; read from the canonical link when absent.

options.frontmatter

type: boolean (default false)

Prepend a YAML front-matter block built from the page's metadata (JSON-LD, Open Graph, meta tags).

options.title

type: boolean (default true)

Prepend the page title as a # heading when the extracted content does not already open with it. Set to false to keep only what the extractor found.

options.extractors

type: ExtractorSelectors

The extractor, or an array of extractors run in order, that select the content to convert. Omitted, autoExtractor runs; false converts the whole document. The built-in presets:

ExtractorPreset nameSelects
autoExtractor / readabilityExtractorauto / readabilityA site adapter when one applies, otherwise kiwameExtractor (the default).
agentExtractoragentThe same main content, then reader comments and the page's other links grouped by role. createAgentExtractor({ roles, comments }) picks the groups.
kiwameExtractorkiwameThe learned block classifier alone. createKiwameExtractor({ threshold, … }) tunes it.
takumiExtractortakumiThe heuristic, Readability-style extractor.
minimalFilterminimalRemoves only what is clearly not content.

presetExtractors(name) returns the option for a preset name, and EXTRACTOR_PRESETS lists the names (the CLI's --extractor and the platform's convert.extractor take the same ones):

import { htmlToMarkdown, presetExtractors } from "webforai";
 
const html = "<article><p>Hello, world!</p></article>";
const markdown = htmlToMarkdown(html, { extractors: presetExtractors("agent") });

You can define your own functions in addition to the presets:

import { htmlToMarkdown, type Extractor, takumiExtractor } from "webforai"
 
const yourCustomExtractor: Extractor = (params) => {
	const { hast, url } = params
	// ... your logic ...
	return hast
};
 
const html = "<h1>Hello, world!</h1>"
const markdown = htmlToMarkdown(html, {
	extractors: [yourCustomExtractor, takumiExtractor]
});
// => "# Hello, world!"

options.formatting

type: Omit<MdastToMarkdownOptions, "baseUrl">

Formatting options passed to mdast-util-to-markdown.

const markdown = htmlToMarkdown("<h1>Hello, world!</h1>", {
	formatting: {
		bullet: "*",
	},
});
// => "* Hello, world!"

options.linkAsText

type: boolean

Whether to convert links to plain text.

const markdown = htmlToMarkdown("<a href='/foo'>bar</a>", {
	linkAsText: true,
});
// => "bar"

options.tableAsText

type: boolean

Whether to convert tables to plain text.

options.hideImage

type: boolean

Whether to hide images.

options.lang

type: string

The language of the HTML.