Skip to content

CLI

npx webforai converts a URL or a local HTML file to clean Markdown. It is non-interactive by default: the Markdown goes to stdout and every log goes to stderr, so the command composes with pipes, scripts and AI agents. The guided wizard is still there behind -i.

# Markdown to stdout
npx webforai@latest https://example.com/article
 
# write to a file instead
npx webforai@latest https://example.com/article -o article.md
 
# convert a local HTML file
npx webforai@latest ./page.html
 
# machine-readable output
npx webforai@latest https://example.com --json | jq -r .markdown

Flags

flagmeaning
-o, --output <path>Write Markdown to a file instead of stdout.
-l, --loader <loader>fetch (default for URLs), playwright, or platform. Local files are detected automatically.
-m, --mode <mode>default or ai — ai strips links, tables and images down to plain text (local conversion only).
--extractor <name>Main-content extraction preset: auto (default: site adapters, then the learned kiwame classifier), kiwame, takumi (the heuristic, Readability-style extractor), minimal, or none for the whole page. The platform loader accepts all but kiwame.
--frontmatterPrepend YAML front matter built from the page metadata (title, author, canonical URL, …).
--jsonPrint a JSON envelope instead of raw Markdown (below).
--engine <engine>Platform engine: auto (renders JavaScript only when the page needs it), fetch, browser, proxy-fetch, proxy-browser (proxy engines need a subscription). Implies -l platform.
--region <region>Platform proxy egress region: auto or jp (jp runs on the proxy tier, so it needs a subscription). Implies -l platform.
--screenshotPlatform browser engines only: also capture a screenshot (expiring URL).
--respect-robotsPlatform loader only: honor the target site's robots.txt rules (off by default). A disallowed URL fails with robots_disallowed and is not billed. See robots.txt.
--api-key <key>Platform API key; defaults to $WEBFORAI_API_KEY.
--platform-url <url>Platform base URL; defaults to $WEBFORAI_PLATFORM_URL or the hosted instance.
-i, --interactiveGuided prompt flow.
-d, --debugVerbose logs on stderr.

Exit codes: 0 success, 1 runtime failure, 2 usage error. API failures include the platform error code and retryAfter (seconds, when supplied) on stderr; invalid_api_key, payment_required and spend_cap_reached name the platform dashboard URL (your --platform-url when set), as does the usage error for a missing API key. Explicit local loaders cannot be combined with --engine or --region; this exits with code 2 before making a request. Interactive mode applies the same flag validation and environment defaults.

Loaders

  • fetch — plain HTTP fetch. Fast, no JavaScript execution. Follows <meta http-equiv="refresh"> redirect stubs (an HTTP 200 "Redirecting…" page) like ordinary redirects, up to 3 hops.
  • playwright — renders the page in local headless Chromium first. Install the browser once with npx playwright install chromium.
  • platform — the hosted webforai platform fetches and converts server-side. Useful for JavaScript-heavy pages without a local browser, bot-protected sites (proxy engines), region-pinned fetching, and screenshots.
export WEBFORAI_API_KEY=wfa_...   # from https://platform.webforai.dev/dashboard
npx webforai@latest https://spa.example.com --engine auto
npx webforai@latest https://blocked.example.com --engine proxy-fetch --region jp   # subscription

Batch and crawl

Two subcommands run asynchronous platform jobs end to end: submit, wait (a live progress line on a terminal, quiet when piped), page through the results and write one Markdown file per page into a directory.

export WEBFORAI_API_KEY=wfa_...
 
# crawl a docs site (same-origin), plus llms.txt and llms-full.txt
npx webforai@latest crawl https://docs.example.com -o docs --limit 100 --include '^/docs' --llms-txt
 
# convert a list of URLs (arguments, or one per line with --file; "-" reads stdin)
npx webforai@latest batch https://a.example/post https://b.example/news -o pages
cat urls.txt | npx webforai@latest batch --file - -o pages --json
flagmeaning
-o, --output <dir>Required. Files mirror URL paths: /docs/intro → docs/intro.md, / → index.md; batch prefixes the host (a.example/post.md). A query string adds a short hash; collisions get _1, _2.
--max-depth <n>crawl: link depth from the seed, 0–5 (default 2).
--limit <n>crawl: maximum pages, 1–500 (default 50).
--include <regex> / --exclude <regex>crawl: pathname filters; repeat the flag for several patterns.
--sitemap <mode>crawl: skip (default), include (also queue the site's sitemap URLs) or only (seed + sitemap URLs, no link following). See crawl.
--respect-robots / --no-respect-robotscrawl honors robots.txt by default; batch only with --respect-robots.
--llms-txtcrawl: also write llms.txt (site title, summary and a - [title](url): description link per page, grouped by top-level path) and llms-full.txt (every page's Markdown, concatenated).
-f, --file <path>batch: read URLs from a file, one per line (# comments allowed); up to 100 URLs per job.
--jsonPrint { jobId, status, credits, pages: [{ url, file, title, … }], failures: [{ url, code, message }], llmsTxt?, llmsFullTxt? } instead of the file list.
--timeout <seconds>Stop waiting after this long (default 1800). The job keeps running; results stay available for 7 days.
--engine, --region, --extractor, --frontmatter, --api-key, --platform-url, -dAs above.

stdout lists the written files, one per line (or the --json envelope); the progress and a summary — including every failed page with its error code — go to stderr. Failed pages are not billed and do not fail the command; it exits 1 when the job itself failed or no page was converted. rate_limited responses are retried after Retry-After; too_many_jobs (3 concurrent jobs on the free tier, 20 with a subscription) is reported instead.

JSON output

--json prints a single envelope object:

{
  "source": "https://example.com",
  "loader": "platform",
  "url": "https://example.com/",
  "engine": "browser",
  "region": "auto",
  "markdown": "# …",
  "metadata": { "title": "…" },
  "credits": 2,
    "output": "article.md"
}

engine, region and credits appear only when the platform loader ran (screenshotUrl also with --screenshot, which adds 1 credit); output only when -o also wrote a file. A warning appears when the result is probably degraded — most commonly a client-rendered page fetched without JavaScript; rerun with --engine auto or -l playwright when you see one.

For AI agents

The CLI ships an Agent Skill that teaches coding agents when and how to use it:

# print the SKILL.md
npx webforai@latest skill
 
# install it via Vercel's skills CLI (agent picker, symlinks, lockfile)
npx webforai@latest skill --install
 
# or install straight from the repository
npx skills add inaridiy/webforai

skill --install forwards --global, --agent <agents>, -y, --copy and --all to npx skills add; --dir <path> writes the file directly without the skills CLI.

Interactive mode

npx webforai -i walks through source, loader (including the platform), mode and output path with prompts.