Skip to content

Websites

Fetch public page HTML or visible text and discover published website email addresses.

Send a public HTTP or HTTPS URL to one of these endpoints:

Endpoint Initial credit cost Result
GET /v1/websites/content?url=… 2 Final URL, page title and HTML
GET /v1/websites/text?url=… 2 Final URL, page title and visible text
GET /v1/websites/emails?url=… 5 Published email addresses and their source URLs

Use your usual API key. All return { data, meta: { creditCost } }. JavaScript rendering, internal retries and additional discovery pages are included in the request’s price. Each request fetches fresh content.

Page text is extracted after any JavaScript rendering. Scripts, styles, hidden elements, embedded frames and other non-visible content are removed. Block elements become line breaks, paragraphs are separated by a blank line, and runs of whitespace collapse to one space. Text is capped at 2 MiB. A page with no visible text returns a target failure and is refunded.

Email discovery starts with the submitted page, then follows relevant links on the same website, prioritizing contact, support, about and legal pages in several languages. It attempts at most eight pages and stops at the first page containing email addresses. It returns up to 100 distinct addresses from that page. It does not check whether the addresses can receive mail.

Read stopReason before interpreting an empty result:

Value Meaning
found A page contained addresses.
exhausted No addresses appeared on the selected pages. Other website pages may still contain addresses.
page_limit Eight pages were attempted and more selected pages remained.
time_limit Some pages were inspected, but the time budget ended.
page_errors Some selected pages could not be inspected.

pagesChecked counts pages successfully inspected. truncated indicates that a page contained more than 100 addresses; it does not describe whether the whole website was searched.

Successful empty and partial searches retain the charge. Confirmed missing pages also retain the charge. Server failures and hard timeouts are refunded through the normal credit settlement process.

Website processing has a 35-second budget. Pages larger than 2 MiB of decoded HTML are rejected. Only public destinations on standard web ports are accepted, including redirects and requests made by page scripts. Private networks, localhost, cloud metadata addresses, URL credentials and non-web schemes are blocked.

Returned HTML and text are untrusted content from the website. Sanitize or isolate it before displaying it in your application. Website response examples in this reference are illustrative.