Evaluation-First Development
n8n-skills uses an Evaluation-Driven Development (EDD) approach: evaluations are written before the skill, not after. This ensures every skill solves real, measurable problems.Write your evaluation scenarios before writing a single line of SKILL.md. Evaluations define what “done” looks like.
Evaluation File Format
Each evaluation is a JSON file inevaluations/[skill-name]/. The filename follows the pattern eval-NNN-kebab-case-description.json.
Real examples
Here are two evaluation files from theexpression-syntax skill that illustrate good evaluation design:
- Basic usage (eval-001)
- Critical gotcha (eval-002)
- Common mistake (eval-003)
How Many Evaluations?
Every skill needs a minimum of 3 evaluations. The existing skills in the project follow this coverage pattern:
Aim for at least:
- Basic usage — the most common trigger query
- Common mistake — a specific error the skill should help fix
- Advanced scenario — a more complex or edge-case query
Running Evaluations Manually
There is no automated evaluation runner yet. Test each scenario by hand:1
Start Claude Code
Launch Claude Code with the skill loaded from your
skills/ directory.2
Ask the evaluation query
Copy the
query field from the evaluation JSON exactly as written and send it.3
Check expected behaviors
Go through each item in
expected_behavior and verify whether the response satisfies it. Be specific — vague confirmation does not count.4
Document results
Mark each behavior as PASS or FAIL. A scenario only passes when every expected behavior is present.
5
Iterate if needed
If any behaviors fail, update SKILL.md to address the gap and re-run. Repeat until 100% of scenarios pass.
An automated evaluation framework is planned for a future release. Until then, manual testing against evaluation JSON files is the standard process.
What Makes a Good Evaluation?
- Good
- Bad
Characteristics of effective evaluations:
- Specific, measurable expected behaviors (not “gives a good answer”)
- Based on real user queries that have actually been seen
- Covers both common and edge cases
- Includes a
baseline_without_skillthat shows what a generic response would miss - Each
expected_behavioritem is independently verifiable
Test Quality Criteria
Before considering a skill complete, confirm all of the following:- All evaluations pass (every
expected_behavioritem verified) - Skill activates correctly on the trigger query
- Content in the response is accurate
- All code examples in the response actually work
- Baseline comparison confirms meaningful improvement over no-skill response
MCP Tool Testing
Before writing any skill content, test the relevant MCP tools and record real responses. This ensures the skill content is grounded in actual tool behavior. Document findings indocs/MCP_TESTING_LOG.md:
- Finding 1
- Finding 2