Integration Guide

How to Fetch Any Web Page Reliably From Code

'Blocked from a website, what are my options?' gets 50 upvotes because a plain fetch() failing silently is a common, frustrating problem. Here's how to build a fetch pipeline that either gives you the real page or tells you clearly why not.

Written by Alex P.

  • web scraping
  • fetch pipeline
  • page validation
  • data reliability

“Blocked from website, what are my options?” is a 50-upvote r/webscraping thread, and it points at a problem bigger than any one site: a plain fetch() in your own code doesn’t fail loudly when it gets the wrong page. It returns 200 OK with a cookie-consent wall, a loading skeleton with no content yet, or a page that redirected somewhere else entirely — and your pipeline happily parses garbage because nothing told it the page wasn’t real.

The fix isn’t a smarter request. It’s a fetch call that validates what it got before handing it back, and that lets you set an explicit ceiling on what a request is allowed to cost before you decide it isn’t worth finishing.


Ask for the page, and say what “correct” means

const API_KEY = process.env.FETCHLAYER_API_KEY;

async function fetchPage(url, expectText) {
  const res = await fetch('https://api.fetchlayer.dev/web-unblocker/fetch', {
    method: 'POST',
    headers: { Authorization: `Bearer ${API_KEY}`, 'Content-Type': 'application/json' },
    body: JSON.stringify({
      url,
      output: 'markdown',   // smaller than html, and the best shape for feeding to a model
      expectText,           // if this text never shows up, the request is reported unsuccessful — not returned as a "200" with the wrong page
    }),
  });
  return res.json();
}

const result = await fetchPage('https://example.com/pricing', 'Pricing');
console.log(result.status, result.textLength, result.truncated);

expectText is the difference between “got a response” and “got the page you actually wanted.” Point it at something that’s only present on the real page — a heading, a price, a product name — and a request that lands on a paywall or a bot-check page reports itself as unsuccessful instead of quietly handing you the wrong content to parse downstream.


Setting a cost ceiling instead of guessing

Some pages are simple; some need more effort to read correctly. Rather than making that decision per-request, maxEffort sets a ceiling — basic, standard, advanced, or maximum — and the cheapest approach that actually works is always used regardless of the ceiling you set. Setting it low doesn’t make an easy page slower; it caps how much a hard page is allowed to cost you before giving up:

async function fetchWithBudget(url, expectText, maxEffort = 'standard') {
  const res = await fetch('https://api.fetchlayer.dev/web-unblocker/fetch', {
    method: 'POST',
    headers: { Authorization: `Bearer ${API_KEY}`, 'Content-Type': 'application/json' },
    body: JSON.stringify({ url, expectText, maxEffort, output: 'markdown' }),
  });

  if (res.status === 401) throw new Error('Bad API key');
  const data = await res.json();

  // The response tells you which tier actually served the page, and what it
  // cost — a request that lands above `basic` says so, rather than you
  // having to infer it from response time.
  console.log(`Served at ${data.usage.step}, cost ${data.usage.billableUnits} unit(s)`);
  return data;
}

A page that needs more than your ceiling allows doesn’t come back with a wrong answer — it comes back unsuccessful, with notes explaining that an effort ceiling was hit. That’s a clean signal to retry at a higher ceiling for pages you’ve confirmed are worth the extra cost, rather than silently accepting a partial result everywhere.


Reading the cost honestly

pagesFetched (the same number as usage.billableUnits) isn’t a page count — it’s a cost figure, at three units to the credit. A simple page costs 1 unit; a harder one costs more, up to 9 for the hardest. This matters for budgeting a large crawl: two requests that both return 200 can cost very differently, and the field to watch for that is usage.step, not the HTTP status:

function estimateMonthlyCost(recentResults) {
  const totalUnits = recentResults.reduce((sum, r) => sum + r.pagesFetched, 0);
  const avgUnitsPerPage = totalUnits / recentResults.length;
  return { avgUnitsPerPage, avgCreditsPerPage: avgUnitsPerPage / 3 };
}

A 503 is never billed — so a failed request in a retry loop doesn’t compound cost the way a loop of “successful but wrong” responses would.


Validating before you trust the content

Beyond expectText, three fields are worth checking on every response before a pipeline treats a page as usable:

function isUsable(result) {
  if (result.status >= 400) return false;
  if (result.textLength === 0) return false;   // rendered nothing readable
  if (result.truncated) return false;          // cut short at maxBytes — treat as incomplete, not just smaller
  return true;
}

finalUrl is worth logging alongside the result, too — if it doesn’t match the URL you requested, the page redirected somewhere else, which for some sites (a product page redirecting to an out-of-stock or region-locked notice) is itself the signal you were checking for, not an error.


A resilient batch pipeline

async function fetchBatch(pages) {
  const results = [];
  for (const { url, expectText } of pages) {
    try {
      const data = await fetchWithBudget(url, expectText, 'standard');
      results.push({ url, ok: isUsable(data), content: data.content, step: data.usage?.step });
    } catch (e) {
      results.push({ url, ok: false, error: String(e) });
    }
  }
  return results;
}

const pages = [
  { url: 'https://example.com/pricing', expectText: 'Pricing' },
  { url: 'https://example.com/product/123', expectText: 'Add to cart' },
];

const results = await fetchBatch(pages);
console.log(`${results.filter((r) => r.ok).length}/${results.length} usable`);

Run this sequentially rather than in parallel unless you’ve confirmed your plan’s concurrency limits — a validated response you can trust is worth more than a fast batch of responses you have to re-check.


Practical notes

  • waitForSelector is a stronger guarantee than expectText for pages that load content in stages. If content only appears after a specific element exists, wait for that element specifically rather than a fixed delay — a fixed waitMs either wastes time on a fast page or isn’t enough on a slow one.
  • assets: 'minimal' is worth trying first for anything you’re only reading text from. It skips images and video entirely; only step up to standard or full if a page’s content genuinely depends on something those assets load.
  • session keeps a multi-step flow coherent. Pass the same arbitrary label across a sequence of requests against one site (search, then a result page, then a detail page) so they’re read as one continuous visit rather than three unrelated ones.

Next Steps