Skip to product information

Extract Website Markdown Chunks with Apify & HTTP Request

Extract Website Markdown Chunks with Apify & HTTP Request

 (200+Reviews)
Regular price £42.99
Regular price £42.99 Sale price
SAVE Sold out
⬇
Instant Digital Download
∞
Unlimited Downloads
★
Lifetime Access in Your Account
🔥
128+ Sold
Popular with n8n builders
âš¡
23 people viewing
High interest right now
✅
9 added today
Fast-moving digital product
Extract Website Markdown Chunks with Apify & HTTP Request

Extract Website Markdown Chunks with Apify & HTTP Request

Regular price £42.99
Regular price £42.99 Sale price
SAVE Sold out

Extract public website passages into clean Markdown chunks using Apify & n8n

This n8n workflow runs Mako’s Web Content Crawler on Apify to fetch website content, validate coverage, and return Markdown chunks ready for account research or retrieval workflows—each chunk includes its source URL, heading path, content hash, and crawl timestamp.

What this workflow does

  • Manual execution with explicit start URLs: when you run it manually, it defines 1–10 start URLs and crawl limits.
  • Triggers an Apify Actor run: it starts the actor agency-shift/web-content-crawler using the Apify REST API via n8n HTTP Request.
  • Polls until completion (or deadline/failure): the workflow polls the run status until it finishes, a polling deadline is reached, or a status request fails. It then attempts to abort the run on the stop path, with an additional 180-second Actor timeout applied.
  • Checks RUN_SUMMARY for coverage & truncation: it fetches the RUN_SUMMARY from Apify Key-Value Store and stops if the run failed, coverage is incomplete, or output is truncated/empty.
  • Returns one output item per chunk: it retrieves dataset items, validates each page’s Markdown/chunk schema, and outputs chunk content plus:
    • sourceUrl/sectionUrl, heading path
    • content hash and crawl time
    • an oversized flag when applicable

Use cases

  • Preparing public website passages for account research and company knowledge bases
  • Building a retrieval workflow (RAG indexing) from web pages split into Markdown chunks
  • QAing crawl coverage before connecting downstream processing

Technical details

  • Apify REST API via n8n HTTP Request (Actor: agency-shift/web-content-crawler)
  • Auth: set Authorization: Bearer <YOUR_APIFY_TOKEN> using an n8n HTTP Header Auth credential
  • Nodes used: Manual Trigger, if, code, wait, sticky note, HTTP Request
View full details