{"product_id":"n8n-workflow-screen-llm-messages-with-judgment-api-guardrails","title":"n8n Workflow: Screen LLM Messages with Judgment API Guardrails","description":"\u003ch3\u003eScreen LLM inputs and outputs with Judgment API guardrails—then automatically block or escalate risky content\u003c\/h3\u003e\n\u003cp\u003eThis n8n workflow manually runs to score incoming user messages and outgoing model replies using the \u003cstrong\u003eJudgment (n8n-nodes-judgment) API\u003c\/strong\u003e. It identifies jailbreak attempts, prompt injection, PII, harm, and self-harm risks—then routes each item to pass, blocked, or withhold-and-escalate paths based on a probability threshold.\u003c\/p\u003e\n\n\u003ch3\u003eWhat this workflow does\u003c\/h3\u003e\n\u003cul\u003e\n  \u003cli\u003e\n\u003cstrong\u003eManual test execution:\u003c\/strong\u003e Run it by clicking \u003cem\u003e“Test Workflow”\u003c\/em\u003e, which generates sample user messages and a sample model reply.\u003c\/li\u003e\n  \u003cli\u003e\n\u003cstrong\u003eHazard scoring for messages:\u003c\/strong\u003e Each message is sent to the Judgment API to score \u003cstrong\u003efive hazards\u003c\/strong\u003e: \u003cem\u003ejailbreak\u003c\/em\u003e, \u003cem\u003eprompt injection\u003c\/em\u003e, \u003cem\u003ePII\u003c\/em\u003e, \u003cem\u003eharm\u003c\/em\u003e, and \u003cem\u003eself-harm\u003c\/em\u003e.\u003c\/li\u003e\n  \u003cli\u003e\n\u003cstrong\u003eRisk normalization + verdict:\u003c\/strong\u003e The workflow normalizes each hazard score into a message\/hazard ID, a risk probability, and a \u003cstrong\u003epass\/block\u003c\/strong\u003e verdict using a \u003cstrong\u003e0.5 threshold\u003c\/strong\u003e.\u003c\/li\u003e\n  \u003cli\u003e\n\u003cstrong\u003eConditional routing:\u003c\/strong\u003e If a hazard probability is \u003cstrong\u003e≥ 0.5\u003c\/strong\u003e, the item is routed to a \u003cstrong\u003eblocked\u003c\/strong\u003e path; otherwise it goes to a \u003cstrong\u003epassed\u003c\/strong\u003e path.\u003c\/li\u003e\n  \u003cli\u003e\n\u003cstrong\u003eCompliance screening for replies:\u003c\/strong\u003e The model reply is evaluated by the Judgment API to detect whether it complied with an injected instruction or leaked internal\/system prompt content.\u003c\/li\u003e\n  \u003cli\u003e\n\u003cstrong\u003eReply handling:\u003c\/strong\u003e Risky replies are routed to a \u003cstrong\u003ewithhold-and-escalate\u003c\/strong\u003e path; clean replies are routed to a \u003cstrong\u003esend-to-the-user\u003c\/strong\u003e path.\u003c\/li\u003e\n\u003c\/ul\u003e\n\n\u003ch3\u003eUse cases\u003c\/h3\u003e\n\u003cul\u003e\n  \u003cli\u003eGuard your n8n-powered chat or webhook flows against \u003cstrong\u003eprompt injection\u003c\/strong\u003e and \u003cstrong\u003ejailbreak\u003c\/strong\u003e attempts.\u003c\/li\u003e\n  \u003cli\u003eReduce exposure to \u003cstrong\u003ePII\u003c\/strong\u003e by blocking messages that exceed a risk probability threshold.\u003c\/li\u003e\n  \u003cli\u003eEscalate \u003cstrong\u003eharm\u003c\/strong\u003e or \u003cstrong\u003eself-harm\u003c\/strong\u003e content before it reaches end users.\u003c\/li\u003e\n  \u003cli\u003ePrevent prompt leakage by withhold-and-escalating replies that may reveal \u003cem\u003einternal\/system\u003c\/em\u003e instructions.\u003c\/li\u003e\n\u003c\/ul\u003e\n\n\u003ch3\u003eTechnical details\u003c\/h3\u003e\n\u003cul\u003e\n  \u003cli\u003eRequires the community node package \u003cstrong\u003en8n-nodes-judgment\u003c\/strong\u003e (install via \u003cem\u003eSettings → Community Nodes\u003c\/em\u003e).\u003c\/li\u003e\n  \u003cli\u003eAdd a \u003cstrong\u003eJudgment API credential\u003c\/strong\u003e (e.g., a TypeSafe Judgment API key) and select it in both screening steps.\u003c\/li\u003e\n  \u003cli\u003eUses nodes including \u003cstrong\u003eif\u003c\/strong\u003e, \u003cstrong\u003eset\u003c\/strong\u003e, \u003cstrong\u003ecode\u003c\/strong\u003e, \u003cstrong\u003eno op\u003c\/strong\u003e, \u003cstrong\u003eswitch\u003c\/strong\u003e, and \u003cstrong\u003esticky note\u003c\/strong\u003e.\u003c\/li\u003e\n  \u003cli\u003eReplace sample message\/reply steps with your real chat\/webhook input and LLM output, and tune the \u003cstrong\u003e0.5 risk threshold\u003c\/strong\u003e to match your policy.\u003c\/li\u003e\n\u003c\/ul\u003e","brand":"N8N Commerce","offers":[{"title":"Default Title","offer_id":46063656763571,"sku":"N8N-19721","price":22.99,"currency_code":"GBP","in_stock":true}],"thumbnail_url":"\/\/cdn.shopify.com\/s\/files\/1\/0749\/6279\/6723\/files\/iTFWqlvKooVfKmrJf3TzE_ZnLnOZM0.png?v=1789808754","url":"https:\/\/buyflowscripts.com\/products\/n8n-workflow-screen-llm-messages-with-judgment-api-guardrails","provider":"N8N Commerce","version":"1.0","type":"link"}