# CLI

`npx webforai` converts a URL or a local HTML file to clean Markdown. It is **non-interactive by default**: the Markdown goes to stdout and every log goes to stderr, so the command composes with pipes, scripts and AI agents. The guided wizard is still there behind `-i`.

```uri
# Markdown to stdout
npx webforai@latest https://example.com/article
 
# write to a file instead
npx webforai@latest https://example.com/article -o article.md
 
# convert a local HTML file
npx webforai@latest ./page.html
 
# machine-readable output
npx webforai@latest https://example.com --json | jq -r .markdown
```

## Flags

| flag                    | meaning                                                                                                                                                                                                                                                        |
| ----------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `-o, --output <path>`   | Write Markdown to a file instead of stdout.                                                                                                                                                                                                                    |
| `-l, --loader <loader>` | `fetch` (default for URLs), `playwright`, or `platform`. Local files are detected automatically.                                                                                                                                                               |
| `-m, --mode <mode>`     | `default` or `ai` — `ai` strips links, tables and images down to plain text (local conversion only).                                                                                                                                                           |
| `--extractor <name>`    | Main-content extraction preset: `auto` (default: site adapters, then the learned `kiwame` classifier), `kiwame`, `takumi` (the heuristic, Readability-style extractor), `minimal`, or `none` for the whole page. The platform loader accepts all but `kiwame`. |
| `--frontmatter`         | Prepend YAML front matter built from the page metadata (title, author, canonical URL, …).                                                                                                                                                                      |
| `--json`                | Print a JSON envelope instead of raw Markdown (below).                                                                                                                                                                                                         |
| `--engine <engine>`     | Platform engine: `auto` (renders JavaScript only when the page needs it), `fetch`, `browser`, `proxy-fetch`, `proxy-browser` (proxy engines need a subscription). Implies `-l platform`.                                                                       |
| `--region <region>`     | Platform proxy egress region: `auto` or `jp` (`jp` runs on the proxy tier, so it needs a subscription). Implies `-l platform`.                                                                                                                                 |
| `--screenshot`          | Platform browser engines only: also capture a screenshot (expiring URL).                                                                                                                                                                                       |
| `--respect-robots`      | Platform loader only: honor the target site's robots.txt rules (off by default). A disallowed URL fails with `robots_disallowed` and is not billed. See [robots.txt](https://webforai.dev/platform/api-reference#robotstxt).                                   |
| `--api-key <key>`       | Platform API key; defaults to `$WEBFORAI_API_KEY`.                                                                                                                                                                                                             |
| `--platform-url <url>`  | Platform base URL; defaults to `$WEBFORAI_PLATFORM_URL` or the hosted instance.                                                                                                                                                                                |
| `-i, --interactive`     | Guided prompt flow.                                                                                                                                                                                                                                            |
| `-d, --debug`           | Verbose logs on stderr.                                                                                                                                                                                                                                        |

Exit codes: `0` success, `1` runtime failure, `2` usage error. API failures include the platform [error code](https://webforai.dev/platform/api-reference#errors) and `retryAfter` (seconds, when supplied) on stderr; `invalid_api_key`, `payment_required` and `spend_cap_reached` name the platform [dashboard](https://platform.webforai.dev/dashboard) URL (your `--platform-url` when set), as does the usage error for a missing API key. Explicit local loaders cannot be combined with `--engine` or `--region`; this exits with code `2` before making a request. Interactive mode applies the same flag validation and environment defaults.

## Loaders

- **fetch** — plain HTTP fetch. Fast, no JavaScript execution. Follows `<meta http-equiv="refresh">` redirect stubs (an HTTP 200 "Redirecting…" page) like ordinary redirects, up to 3 hops.
- **playwright** — renders the page in local headless Chromium first. Install the browser once with `npx playwright install chromium`.
- **platform** — the [hosted webforai platform](https://webforai.dev/platform) fetches and converts server-side. Useful for JavaScript-heavy pages without a local browser, bot-protected sites (proxy engines), region-pinned fetching, and screenshots.

```uri
export WEBFORAI_API_KEY=wfa_...   # from https://platform.webforai.dev/dashboard
npx webforai@latest https://spa.example.com --engine auto
npx webforai@latest https://blocked.example.com --engine proxy-fetch --region jp   # subscription
```

## Batch and crawl

Two subcommands run asynchronous [platform jobs](https://webforai.dev/platform/api-reference#post-v1batch) end to end: submit, wait (a live progress line on a terminal, quiet when piped), page through the results and write **one Markdown file per page** into a directory.

```uri
export WEBFORAI_API_KEY=wfa_...
 
# crawl a docs site (same-origin), plus llms.txt and llms-full.txt
npx webforai@latest crawl https://docs.example.com -o docs --limit 100 --include '^/docs' --llms-txt
 
# convert a list of URLs (arguments, or one per line with --file; "-" reads stdin)
npx webforai@latest batch https://a.example/post https://b.example/news -o pages
cat urls.txt | npx webforai@latest batch --file - -o pages --json
```

| flag                                                                                        | meaning                                                                                                                                                                                                           |
| ------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `-o, --output <dir>`                                                                        | Required. Files mirror URL paths: `/docs/intro` → `docs/intro.md`, `/` → `index.md`; batch prefixes the host (`a.example/post.md`). A query string adds a short hash; collisions get `_1`, `_2`.                  |
| `--max-depth <n>`                                                                           | crawl: link depth from the seed, `0`–`5` (default 2).                                                                                                                                                             |
| `--limit <n>`                                                                               | crawl: maximum pages, `1`–`500` (default 50).                                                                                                                                                                     |
| `--include <regex>` / `--exclude <regex>`                                                   | crawl: pathname filters; repeat the flag for several patterns.                                                                                                                                                    |
| `--sitemap <mode>`                                                                          | crawl: `skip` (default), `include` (also queue the site's sitemap URLs) or `only` (seed + sitemap URLs, no link following). See [crawl](https://webforai.dev/platform/api-reference#post-v1crawl).                |
| `--respect-robots` / `--no-respect-robots`                                                  | crawl honors robots.txt by default; batch only with `--respect-robots`.                                                                                                                                           |
| `--llms-txt`                                                                                | crawl: also write [`llms.txt`](https://llmstxt.org) (site title, summary and a `- [title](url): description` link per page, grouped by top-level path) and `llms-full.txt` (every page's Markdown, concatenated). |
| `-f, --file <path>`                                                                         | batch: read URLs from a file, one per line (`#` comments allowed); up to 100 URLs per job.                                                                                                                        |
| `--json`                                                                                    | Print `{ jobId, status, credits, pages: [{ url, file, title, … }], failures: [{ url, code, message }], llmsTxt?, llmsFullTxt? }` instead of the file list.                                                        |
| `--timeout <seconds>`                                                                       | Stop waiting after this long (default 1800). The job keeps running; results stay available for 7 days.                                                                                                            |
| `--engine`, `--region`, `--extractor`, `--frontmatter`, `--api-key`, `--platform-url`, `-d` | As above.                                                                                                                                                                                                         |

stdout lists the written files, one per line (or the `--json` envelope); the progress and a summary — including every failed page with its error code — go to stderr. Failed pages are not billed and do not fail the command; it exits `1` when the job itself failed or no page was converted. `rate_limited` responses are retried after `Retry-After`; `too_many_jobs` (3 concurrent jobs on the free tier, 20 with a subscription) is reported instead.

## JSON output

`--json` prints a single envelope object:

```json
{
  "source": "https://example.com",
  "loader": "platform",
  "url": "https://example.com/",
  "engine": "browser",
  "region": "auto",
  "markdown": "# …",
  "metadata": { "title": "…" },
  "credits": 2,
    "output": "article.md"
}
```

`engine`, `region` and `credits` appear only when the platform loader ran (`screenshotUrl` also with `--screenshot`, which adds 1 credit); `output` only when `-o` also wrote a file. A `warning` appears when the result is probably degraded — most commonly a client-rendered page fetched without JavaScript; rerun with `--engine auto` or `-l playwright` when you see one.

## For AI agents

The CLI ships an [Agent Skill](https://github.com/vercel-labs/skills) that teaches coding agents when and how to use it:

```plain
# print the SKILL.md
npx webforai@latest skill
 
# install it via Vercel's skills CLI (agent picker, symlinks, lockfile)
npx webforai@latest skill --install
 
# or install straight from the repository
npx skills add inaridiy/webforai
```

`skill --install` forwards `--global`, `--agent <agents>`, `-y`, `--copy` and `--all` to `npx skills add`; `--dir <path>` writes the file directly without the skills CLI.

## Interactive mode

`npx webforai -i` walks through source, loader (including the platform), mode and output path with prompts.
