Skip to main content

Evaluation-First Development

n8n-skills uses an Evaluation-Driven Development (EDD) approach: evaluations are written before the skill, not after. This ensures every skill solves real, measurable problems.
Write your evaluation scenarios before writing a single line of SKILL.md. Evaluations define what “done” looks like.
The full cycle for a new skill:
Why this works: It ensures skills solve real problems and can be tested objectively, rather than optimizing for content that sounds good but doesn’t measurably improve AI behavior.

Evaluation File Format

Each evaluation is a JSON file in evaluations/[skill-name]/. The filename follows the pattern eval-NNN-kebab-case-description.json.

Real examples

Here are two evaluation files from the expression-syntax skill that illustrate good evaluation design:

How Many Evaluations?

Every skill needs a minimum of 3 evaluations. The existing skills in the project follow this coverage pattern: Aim for at least:
  1. Basic usage — the most common trigger query
  2. Common mistake — a specific error the skill should help fix
  3. Advanced scenario — a more complex or edge-case query

Running Evaluations Manually

There is no automated evaluation runner yet. Test each scenario by hand:
1

Start Claude Code

Launch Claude Code with the skill loaded from your skills/ directory.
2

Ask the evaluation query

Copy the query field from the evaluation JSON exactly as written and send it.
3

Check expected behaviors

Go through each item in expected_behavior and verify whether the response satisfies it. Be specific — vague confirmation does not count.
4

Document results

Mark each behavior as PASS or FAIL. A scenario only passes when every expected behavior is present.
5

Iterate if needed

If any behaviors fail, update SKILL.md to address the gap and re-run. Repeat until 100% of scenarios pass.
An automated evaluation framework is planned for a future release. Until then, manual testing against evaluation JSON files is the standard process.

What Makes a Good Evaluation?

Characteristics of effective evaluations:
  • Specific, measurable expected behaviors (not “gives a good answer”)
  • Based on real user queries that have actually been seen
  • Covers both common and edge cases
  • Includes a baseline_without_skill that shows what a generic response would miss
  • Each expected_behavior item is independently verifiable
Example of a specific, measurable behavior:

Test Quality Criteria

Before considering a skill complete, confirm all of the following:
  • All evaluations pass (every expected_behavior item verified)
  • Skill activates correctly on the trigger query
  • Content in the response is accurate
  • All code examples in the response actually work
  • Baseline comparison confirms meaningful improvement over no-skill response

MCP Tool Testing

Before writing any skill content, test the relevant MCP tools and record real responses. This ensures the skill content is grounded in actual tool behavior. Document findings in docs/MCP_TESTING_LOG.md:
Response:
Key Insights:
  • Finding 1
  • Finding 2