Extract Website Markdown Chunks with Apify & HTTP Request
Extract Website Markdown Chunks with Apify & HTTP Request
Regular price
£42.99
Regular price
£42.99
Sale price
Unit price
/
per
⬇
Instant Digital Download
∞
Unlimited Downloads
★
Lifetime Access in Your Account
Couldn't load pickup availability
🔥
128+ Sold
Popular with n8n builders
âš¡
23 people viewing
High interest right now
✅
9 added today
Fast-moving digital product
Extract Website Markdown Chunks with Apify & HTTP Request
Regular price
£42.99
Regular price
£42.99
Sale price
Unit price
/
per
Extract public website passages into clean Markdown chunks using Apify & n8n
This n8n workflow runs Mako’s Web Content Crawler on Apify to fetch website content, validate coverage, and return Markdown chunks ready for account research or retrieval workflows—each chunk includes its source URL, heading path, content hash, and crawl timestamp.
What this workflow does
- Manual execution with explicit start URLs: when you run it manually, it defines 1–10 start URLs and crawl limits.
- Triggers an Apify Actor run: it starts the actor agency-shift/web-content-crawler using the Apify REST API via n8n HTTP Request.
- Polls until completion (or deadline/failure): the workflow polls the run status until it finishes, a polling deadline is reached, or a status request fails. It then attempts to abort the run on the stop path, with an additional 180-second Actor timeout applied.
- Checks RUN_SUMMARY for coverage & truncation: it fetches the RUN_SUMMARY from Apify Key-Value Store and stops if the run failed, coverage is incomplete, or output is truncated/empty.
-
Returns one output item per chunk: it retrieves dataset items, validates each page’s Markdown/chunk schema, and outputs chunk content plus:
- sourceUrl/sectionUrl, heading path
- content hash and crawl time
- an oversized flag when applicable
Use cases
- Preparing public website passages for account research and company knowledge bases
- Building a retrieval workflow (RAG indexing) from web pages split into Markdown chunks
- QAing crawl coverage before connecting downstream processing
Technical details
- Apify REST API via n8n HTTP Request (Actor: agency-shift/web-content-crawler)
-
Auth: set
Authorization: Bearer <YOUR_APIFY_TOKEN>using an n8n HTTP Header Auth credential - Nodes used: Manual Trigger, if, code, wait, sticky note, HTTP Request
