Compare GPT-4 vs Claude vs Gemini with LMUnit in n8n
Compare GPT-4 vs Claude vs Gemini with LMUnit in n8n
Regular price
£71.99
Regular price
£71.99
Sale price
Unit price
/
per
⬇
Instant Digital Download
∞
Unlimited Downloads
★
Lifetime Access in Your Account
Couldn't load pickup availability
🔥
128+ Sold
Popular with n8n builders
⚡
23 people viewing
High interest right now
✅
9 added today
Fast-moving digital product
Compare GPT-4 vs Claude vs Gemini with LMUnit in n8n
Regular price
£71.99
Regular price
£71.99
Sale price
Unit price
/
per
Compare GPT-4, Claude, and Gemini side-by-side—then score response quality automatically in n8n using LMUnit
If you’ve ever wondered which LLM delivered the clearest, most concise answer, this n8n workflow streamlines the process. It collects responses from multiple models, evaluates them with Contextual AI’s LMUnit, and outputs consistent 1–5 quality scores with an easy model-wise summary.
What this workflow does
This workflow automates LLM response quality evaluation—a task that’s often manual, inconsistent, and hard to scale.
- Collects model outputs via a chat trigger node using the same input prompt for a fair comparison.
- Runs evaluations with Contextual AI’s LMUnit using predefined quality criteria.
- Applies natural language unit testing focused on clarity and conciseness:
- Clarity: “Is the response clear and easy to understand?”
- Conciseness: “Is the response concise and free from redundancy?”
- Generates LMUnit scores (1–5) for each test and aggregates results into a structured summary.
Use cases
- Prompt iteration: test a prompt across OpenAI GPT-4.1, Claude 4.5 Sonnet, and Gemini 2.5 Flash, then quantify improvements in clarity/conciseness.
- LLM vendor benchmarking: compare multiple models on the same task and track performance over time.
- SaaS quality monitoring: spot when specific models drift toward longer, more redundant outputs.
Technical details
- Tech stack / nodes used: set, code, wait, merge, sticky note, and n8nn8n-nodes-langchainchat.
- Evaluation engine: Contextual AI’s LMUnit with a 1–5 scoring scale.
- Output: model-wise performance plus overall averages, formatted into a structured summary.
