# htmlToMarkdown

Useful and high-quality HTML to Markdown converter. Internally, it just calls [htmlToMdast](https://webforai.dev/docs/html-to-mdast) and [mdastToMarkdown](https://webforai.dev/docs/mdast-to-markdown) in that order.

## Usage

```ts
import { htmlToMarkdown } from "webforai";
 
const html = "<h1>Hello, world!</h1>";
const markdown = htmlToMarkdown(html);
=> "# Hello, world!"
```

## Returns

`string`

The converted Markdown string.

## htmlToMarkdownWithMetadata

Takes the same arguments and returns the Markdown together with what the conversion learned about the page:

```ts
import { htmlToMarkdownWithMetadata } from "webforai";
 
const html = "<html><head><title>Hello</title></head><body><article><p>Hello, world!</p></article></body></html>";
const { markdown, metadata, extraction } = htmlToMarkdownWithMetadata(html, { url: "https://example.com/" });
// metadata: { title, description, author, published, modified, siteName, canonicalUrl, lang, image, type }
// extraction: { extractor: "kiwame" | "takumi" | "adapter", confidence?: number }
```

`extraction.confidence` (0–1, kiwame only) estimates how well the selection matches the page's main content; see [How it works](https://webforai.dev/how-it-works#extraction-report-and-confidence).

## Parameters

### htmlOrHast

type: `string | Hast`

The HTML string or HAST tree to convert.

```plain
const markdown = htmlToMarkdown("<h1>Hello, world!</h1>");
// => "# Hello, world!"
```

### options.baseUrl

type: `string`

The base URL to use for replacing relative links.

```uri
const markdown = htmlToMarkdown("<a href='/foo'>bar</a>", {
	baseUrl: "https://example.com",
});
// => "[bar](https://example.com/foo)"
```

### options.url

type: `string`

The page's URL. Selects a site adapter (GitHub, MDN, …) and anchors the page title; read from the canonical link when absent.

### options.frontmatter

type: `boolean` (default `false`)

Prepend a YAML front-matter block built from the page's metadata (JSON-LD, Open Graph, meta tags).

### options.title

type: `boolean` (default `true`)

Prepend the page title as a `#` heading when the extracted content does not already open with it. Set to `false` to keep only what the extractor found.

### options.extractors

type: `ExtractorSelectors`

The extractor, or an array of extractors run in order, that select the content to convert. Omitted, `autoExtractor` runs; `false` converts the whole document. The built-in presets:

| Extractor                                | Preset name            | Selects                                                                                                                                               |
| ---------------------------------------- | ---------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------- |
| `autoExtractor` / `readabilityExtractor` | `auto` / `readability` | A site adapter when one applies, otherwise `kiwameExtractor` (the default).                                                                           |
| `agentExtractor`                         | `agent`                | The same main content, then reader comments and the page's other links grouped by role. `createAgentExtractor({ roles, comments })` picks the groups. |
| `kiwameExtractor`                        | `kiwame`               | The learned block classifier alone. `createKiwameExtractor({ threshold, … })` tunes it.                                                               |
| `takumiExtractor`                        | `takumi`               | The heuristic, Readability-style extractor.                                                                                                           |
| `minimalFilter`                          | `minimal`              | Removes only what is clearly not content.                                                                                                             |

`presetExtractors(name)` returns the option for a preset name, and `EXTRACTOR_PRESETS` lists the names (the CLI's `--extractor` and the platform's `convert.extractor` take the same ones):

```ts
import { htmlToMarkdown, presetExtractors } from "webforai";
 
const html = "<article><p>Hello, world!</p></article>";
const markdown = htmlToMarkdown(html, { extractors: presetExtractors("agent") });
```

You can define your own functions in addition to the presets:

```ts
import { htmlToMarkdown, type Extractor, takumiExtractor } from "webforai"
 
const yourCustomExtractor: Extractor = (params) => {
	const { hast, url } = params
	// ... your logic ...
	return hast
};
 
const html = "<h1>Hello, world!</h1>"
const markdown = htmlToMarkdown(html, {
	extractors: [yourCustomExtractor, takumiExtractor]
});
// => "# Hello, world!"
```

### options.formatting

type: `Omit<MdastToMarkdownOptions, "baseUrl">`

Formatting options passed to [mdast-util-to-markdown](https://github.com/syntax-tree/mdast-util-to-markdown).

```plain
const markdown = htmlToMarkdown("<h1>Hello, world!</h1>", {
	formatting: {
		bullet: "*",
	},
});
// => "* Hello, world!"
```

### options.linkAsText

type: `boolean`

Whether to convert links to plain text.

```plain
const markdown = htmlToMarkdown("<a href='/foo'>bar</a>", {
	linkAsText: true,
});
// => "bar"
```

### options.tableAsText

type: `boolean`

Whether to convert tables to plain text.

### options.hideImage

type: `boolean`

Whether to hide images.

### options.lang

type: `string`

The language of the HTML.
