Red-Team AI Agents with Jailbreak Probes (Groq, Discord, n8n)
Red-Team AI Agents with Jailbreak Probes (Groq, Discord, n8n)
Regular price
£27.99
Regular price
£27.99
Sale price
Unit price
/
per
⬇
Instant Digital Download
∞
Unlimited Downloads
★
Lifetime Access in Your Account
Couldn't load pickup availability
🔥
128+ Sold
Popular with n8n builders
⚡
23 people viewing
High interest right now
✅
9 added today
Fast-moving digital product
Red-Team AI Agents with Jailbreak Probes (Groq, Discord, n8n)
Regular price
£27.99
Regular price
£27.99
Sale price
Unit price
/
per
Red-team your Groq AI agent safely—by replaying jailbreak probes and scoring behavior end-to-end
This n8n workflow automates red-teaming of an AI agent by replaying stored jailbreak and safety probes against a Groq-hosted model, scoring tool use and responses for canary leakage and policy compliance, then saving results to n8n Data Tables and notifying a Discord channel (plus a weekly digest).
What this workflow does
- One-time manual setup: creates the required n8n Data Tables, loads a default probe set, and inserts demo customer records.
- Start a red-team run from an n8n Form: select a probe suite, choose the maximum number of probes, and set the target system prompt. A `canary` build key is appended automatically to detect leakage.
- Load and queue probes for regression: enables probes (optionally filtered by suite), reads a stored baseline, and attaches each probe’s prior verdict so runs can be compared over time.
- Execute probes with tool constraints: sends each probe’s attack prompt to an n8n AI Agent backed by Groq, allowing only a customer-record lookup tool and a simulated external messaging tool while capturing intermediate tool steps.
- Deterministic scoring and review: checks for canary leakage, forbidden tool use, compliance markers, record overreach, refusal behavior, tool-call caps, and blocked/errored outputs. A Groq reviewer can downgrade suspicious “PASS” answers to “REVIEW” using verbatim quotes.
- Persist and notify: writes the run report, per-probe results, and updated baseline to Data Tables, posts failures to Discord, and displays a completion page (including a weekly Discord digest).
Use cases
- Continuously validate that an automation engineer’s Groq-powered agent cannot leak a canary key or access data beyond allowed tools.
- SaaS operators run scheduled regression checks against jailbreak and safety probes after prompt/tool changes.
- Debug tool-call caps and failure modes by capturing blocked/errored outputs per probe.
Technical details
- n8n nodes: if, set, code, form, discord, data table
- Integrations: Groq for the AI agent and Groq-based reviewer; Discord for alerts and weekly digest; n8n Data Tables for baseline and run persistence.
Keywords: n8n automation workflow, red-team AI agents, Groq, Discord alerts, jailbreak probes, AI safety regression testing, Data Tables.
