{"product_id":"n8n-chat-workflow-route-prompts-to-free-llms-nvidia-ollama","title":"n8n Chat Workflow: Route Prompts to Free LLMs (NVIDIA, Ollama)","description":"\u003ch3\u003eRoute n8n Chat prompts to free LLMs—then return the reply\u003c\/h3\u003e\n\u003cp\u003eThis n8n workflow automatically takes messages from \u003cstrong\u003en8n Chat\u003c\/strong\u003e, sends the prompt to multiple \u003cstrong\u003efree LLM options\u003c\/strong\u003e (NVIDIA Integrate, Agnes AI, and Ollama), and returns responses back to the chat—plus an optional \u003cstrong\u003eOllama Cloud LangChain agent\u003c\/strong\u003e run with memory.\u003c\/p\u003e\n\n\u003ch3\u003eWhat this workflow does\u003c\/h3\u003e\n\u003cul\u003e\n  \u003cli\u003e\n\u003cstrong\u003eReceives a chat message\u003c\/strong\u003e (and optional file uploads) via the \u003cstrong\u003en8n Chat trigger\u003c\/strong\u003e.\u003c\/li\u003e\n  \u003cli\u003e\n\u003cstrong\u003eCalls NVIDIA Integrate\u003c\/strong\u003e using the \u003ccode\u003e\/v1\/chat\/completions\u003c\/code\u003e endpoint, then returns the first completion to the chat.\u003c\/li\u003e\n  \u003cli\u003e\n\u003cstrong\u003eCalls Agnes AI\u003c\/strong\u003e using the same \u003ccode\u003e\/v1\/chat\/completions\u003c\/code\u003e pattern, then returns the first completion to the chat.\u003c\/li\u003e\n  \u003cli\u003e\n\u003cstrong\u003eCalls a local Ollama model\u003c\/strong\u003e via the \u003ccode\u003e\/api\/chat\u003c\/code\u003e endpoint and returns the model’s message content to the chat.\u003c\/li\u003e\n  \u003cli\u003e\n\u003cstrong\u003eRuns a LangChain agent with Ollama Cloud\u003c\/strong\u003e: uses an Ollama Cloud chat model plus \u003cstrong\u003ebuffer-window memory\u003c\/strong\u003e, then returns the agent output to the chat.\u003c\/li\u003e\n\u003c\/ul\u003e\n\n\u003ch3\u003eUse cases\u003c\/h3\u003e\n\u003cul\u003e\n  \u003cli\u003e\n\u003cstrong\u003eMulti-model chat testing\u003c\/strong\u003e: compare responses from NVIDIA Integrate, Agnes AI, and your local Ollama setup.\u003c\/li\u003e\n  \u003cli\u003e\n\u003cstrong\u003eCost-control and flexibility\u003c\/strong\u003e: route prompts to local inference (Ollama) or Ollama Cloud while keeping the chat experience consistent in n8n.\u003c\/li\u003e\n  \u003cli\u003e\n\u003cstrong\u003eAgentic assistance\u003c\/strong\u003e: use the Ollama Cloud LangChain agent with memory for iterative, context-aware replies.\u003c\/li\u003e\n  \u003cli\u003e\n\u003cstrong\u003eSaaS operators\u003c\/strong\u003e: offer users a single n8n Chat entry point while experimenting with free LLM backends.\u003c\/li\u003e\n\u003c\/ul\u003e\n\n\u003ch3\u003eTechnical details\u003c\/h3\u003e\n\u003cul\u003e\n  \u003cli\u003eTrigger\/Chat: \u003cstrong\u003en8n-nodes-langchainchat trigger\u003c\/strong\u003e\n\u003c\/li\u003e\n  \u003cli\u003eHTTP calls: \u003cstrong\u003eHTTP Request\u003c\/strong\u003e steps for \u003cstrong\u003eNVIDIA Integrate\u003c\/strong\u003e and \u003cstrong\u003eAgnes AI\u003c\/strong\u003e (\u003ccode\u003e\/v1\/chat\/completions\u003c\/code\u003e)\u003c\/li\u003e\n  \u003cli\u003eLocal inference: \u003cstrong\u003eOllama chat\u003c\/strong\u003e via \u003ccode\u003e\/api\/chat\u003c\/code\u003e (optionally update base URL like \u003ccode\u003ehttp:\/\/host.docker.internal:11434\u003c\/code\u003e)\u003c\/li\u003e\n  \u003cli\u003eAgent mode: \u003cstrong\u003en8nn8n-nodes-langchainagent\u003c\/strong\u003e using \u003cstrong\u003eOllama Cloud chat\u003c\/strong\u003e with \u003cstrong\u003ebuffer-window memory\u003c\/strong\u003e\n\u003c\/li\u003e\n  \u003cli\u003eIncluded nodes\/components: sticky note, LangChain chat\/agent nodes (e.g., \u003cstrong\u003en8nn8n-nodes-langchainchat\u003c\/strong\u003e, \u003cstrong\u003en8nn8n-nodes-langchainlm chat ollama\u003c\/strong\u003e)\u003c\/li\u003e\n\u003c\/ul\u003e\n\n\u003cp\u003e\u003cstrong\u003eSetup notes:\u003c\/strong\u003e add an NVIDIA API bearer token and an Agnes AI bearer token, install Ollama and pull the referenced model (e.g., \u003ccode\u003egemma4:latest\u003c\/code\u003e), then configure \u003cstrong\u003eOllama Cloud (Ollama API) credentials\u003c\/strong\u003e and ensure \u003ccode\u003egemma4:31b\u003c\/code\u003e is available.\u003c\/p\u003e","brand":"N8N Commerce","offers":[{"title":"Default Title","offer_id":45830479151283,"sku":"N8N-18203","price":33.99,"currency_code":"GBP","in_stock":true}],"thumbnail_url":"\/\/cdn.shopify.com\/s\/files\/1\/0749\/6279\/6723\/files\/mFNY_m79Z0-3NT6SxsRKl_ZD8nmruq.png?v=1786699204","url":"https:\/\/buyflowscripts.com\/products\/n8n-chat-workflow-route-prompts-to-free-llms-nvidia-ollama","provider":"N8N Commerce","version":"1.0","type":"link"}