{"product_id":"n8n-webhook-apify-openai-extract-structured-data-json","title":"n8n Webhook: Apify + OpenAI Extract Structured Data JSON","description":"\u003ch3\u003eExtract structured JSON from any webpage—on demand via an n8n Webhook\u003c\/h3\u003e\n\u003cp\u003eThis n8n workflow exposes a webhook that takes a target webpage (a direct \u003ccode\u003etarget_url\u003c\/code\u003e or a \u003ccode\u003esubject\u003c\/code\u003e + \u003ccode\u003ewebsite_domain\u003c\/code\u003e), scrapes it with \u003cb\u003eApify\u003c\/b\u003e, and uses \u003cb\u003eOpenAI\u003c\/b\u003e to return \u003cb\u003estructured JSON\u003c\/b\u003e for exactly the fields you request.\u003c\/p\u003e\n\n\u003ch3\u003eWhat this workflow does\u003c\/h3\u003e\n\u003cul\u003e\n  \u003cli\u003e\n\u003cb\u003eAccepts webhook POST requests\u003c\/b\u003e with either \u003ccode\u003etarget_url\u003c\/code\u003e or \u003ccode\u003esubject\u003c\/code\u003e + \u003ccode\u003ewebsite_domain\u003c\/code\u003e, plus a required \u003ccode\u003etarget_data\u003c\/code\u003e array describing the fields to extract.\u003c\/li\u003e\n  \u003cli\u003e\n\u003cb\u003eValidates input\u003c\/b\u003e: if \u003ccode\u003etarget_data\u003c\/code\u003e is missing or empty, the workflow returns \u003cb\u003eHTTP 400\u003c\/b\u003e.\u003c\/li\u003e\n  \u003cli\u003e\n\u003cb\u003eResolves the page URL\u003c\/b\u003e:\n    \u003cul\u003e\n      \u003cli\u003eIf \u003ccode\u003etarget_url\u003c\/code\u003e is provided, it uses it directly.\u003c\/li\u003e\n      \u003cli\u003eOtherwise, it runs an \u003cb\u003eApify Google Search Scraper\u003c\/b\u003e using a \u003ccode\u003esite:\u003c\/code\u003e query and takes the first organic result.\u003c\/li\u003e\n    \u003c\/ul\u003e\n  \u003c\/li\u003e\n  \u003cli\u003e\n\u003cb\u003eHandles missing matches\u003c\/b\u003e: if no URL is found, it returns \u003cb\u003eHTTP 404\u003c\/b\u003e.\u003c\/li\u003e\n  \u003cli\u003e\n\u003cb\u003eScrapes and normalizes content\u003c\/b\u003e using \u003cb\u003eApify Website Content Crawler\u003c\/b\u003e (optionally with provided cookies), including Markdown normalization and content length.\u003c\/li\u003e\n  \u003cli\u003e\n\u003cb\u003ePre-checks content length\u003c\/b\u003e: if the scraped content is too short, it returns \u003ccode\u003esuccess:false\u003c\/code\u003e with \u003cb\u003eHTTP 200\u003c\/b\u003e.\u003c\/li\u003e\n  \u003cli\u003e\n\u003cb\u003eExtracts with OpenAI\u003c\/b\u003e: sends the page content and the requested field definitions to OpenAI using a \u003cb\u003edynamic schema\u003c\/b\u003e, returning structured JSON and using \u003cb\u003e\"NA\"\u003c\/b\u003e for missing values.\u003c\/li\u003e\n\u003c\/ul\u003e\n\n\u003ch3\u003eUse cases\u003c\/h3\u003e\n\u003cul\u003e\n  \u003cli\u003eTurn landing pages, docs, or product pages into \u003cb\u003estructured data\u003c\/b\u003e for a CRM or database.\u003c\/li\u003e\n  \u003cli\u003eBuild an API-like endpoint for \u003cb\u003econtent extraction\u003c\/b\u003e that calls OpenAI with field definitions.\u003c\/li\u003e\n  \u003cli\u003eAutomate lead enrichment by extracting attributes from websites found via \u003cb\u003eApify\u003c\/b\u003e search.\u003c\/li\u003e\n\u003c\/ul\u003e\n\n\u003ch3\u003eTechnical details\u003c\/h3\u003e\n\u003cul\u003e\n  \u003cli\u003e\n\u003cb\u003eNodes \/ tools used\u003c\/b\u003e: \u003ccode\u003ewebhook\u003c\/code\u003e, \u003ccode\u003eif\u003c\/code\u003e, \u003ccode\u003eset\u003c\/code\u003e, \u003ccode\u003ehttp request\u003c\/code\u003e, \u003ccode\u003en8n-nodes-langchainagent\u003c\/code\u003e (plus supporting workflow logic and notes).\u003c\/li\u003e\n  \u003cli\u003e\n\u003cb\u003eApify integrations\u003c\/b\u003e: Header-authenticated \u003cb\u003eApify HTTP Request\u003c\/b\u003e steps for \u003cb\u003eGoogle Search Scraper\u003c\/b\u003e and \u003cb\u003eWebsite Content Crawler\u003c\/b\u003e (Markdown normalization).\u003c\/li\u003e\n  \u003cli\u003e\n\u003cb\u003eOpenAI\u003c\/b\u003e: OpenAI chat model configured via n8n credential to produce JSON with a dynamic schema.\u003c\/li\u003e\n  \u003cli\u003e\n\u003cb\u003eAuthentication\u003c\/b\u003e: set up an n8n Header Auth credential to send \u003ccode\u003eAuthorization: Bearer \u0026lt;APIFY_TOKEN\u0026gt;\u003c\/code\u003e, then select it on both Apify HTTP Request steps.\u003c\/li\u003e\n\u003c\/ul\u003e\n\u003cp\u003eConfigure your webhook path when calling the workflow to trigger extraction and receive the returned structured JSON.\u003c\/p\u003e","brand":"N8N Commerce","offers":[{"title":"Default Title","offer_id":45785673203891,"sku":"N8N-17910","price":6.99,"currency_code":"GBP","in_stock":true}],"thumbnail_url":"\/\/cdn.shopify.com\/s\/files\/1\/0749\/6279\/6723\/files\/5z5Q-hOe6546x_XQfqgjk_5b4kUVFf.png?v=1786180213","url":"https:\/\/buyflowscripts.com\/products\/n8n-webhook-apify-openai-extract-structured-data-json","provider":"N8N Commerce","version":"1.0","type":"link"}