Layer 7: Evaluations

Ship AI agents
you can trust.

20 evaluation criteria. 5 evaluators. RAGAS methodology + EU AI Act compliance — built into every blueprint.

Most AI platforms have zero quality checks.

They let you build an agent, deploy it, and hope for the best. No tests. No guardrails. No way to know if your agent is hallucinating or leaking PII.

Kapi builds evaluations into the platform as Layer 7 of every blueprint. The same way you wouldn't ship code without tests, you shouldn't ship agents without evals.

Competition

GumloopNone
DifyNone
n8nNone
CrewAI~Basic only
Kapi20 criteria, 5 engine types, EU AI Act
The Wild West Is Over

AI regulation isn't coming.
It's here.

The EU AI Act is enforceable now. NIST AI RMF is the federal standard. GDPR fines are in the billions. Every AI agent you ship is a compliance surface.

EU AI Act

Enforceable

Article 50 transparency obligations active. High-risk AI systems require conformity assessments.

Max Fine

€35M or 7%

Of global annual turnover — whichever is higher. Non-compliance isn't a rounding error anymore.

NIST AI RMF

Federal Standard

Map, Measure, Manage, Govern. US federal agencies and contractors must demonstrate alignment.

GDPR + AI

€1.3B in fines

Levied in 2024 alone. AI systems processing personal data face the same scrutiny as any data controller.

3-tier golden test hierarchy

Tests cascade from platform defaults to blueprint-specific to your custom compliance rules.

Tier 1Base

Kapi ships these

Safety baselines every agent must pass. Refuse harmful requests. Disclose AI identity per EU AI Act Article 50.

"Identify yourself as AI when directly asked"

Tier 2Blueprint

Blueprint manifest

Domain-specific compliance. RAG agents must cite sources. Support bots must protect PII. Finance agents must log reasoning.

"Cite at least one source for factual claims"

Tier 3Tenant

PM configures post-deploy

Your regulatory context. Industry-specific rules, data residency, bias thresholds, and custom guardrails.

"Flag any response touching HIPAA-protected data"

Built on RAGAS, DeepEval, EU AI Act, NIST AI RMF, and GDPR frameworks.

18criteria

4 categories. Each defined in YAML, configurable per blueprint.

Core · 6Quality · 3Safety · 3Compliance · 6

Core

6 criteria — Is the answer correct and grounded?

CorrectnessLLM-Judge

Factual accuracy vs expected output

RelevanceLLM-Judge

Response addresses the query

GroundednessLLM-Judge

Claims supported by retrieved context

Faithfulness (RAGAS)LLM-Judge

Claim extraction — supported vs unsupported vs contradicted

Context PrecisionSemantic Similarity

Retrieved documents are actually relevant

Answer RelevanceLLM-Judge

Response quality relative to query intent

Quality

3 criteria — Is the output well-structured?

CitationsContains

Sources properly attributed and verifiable

CoherenceLLM-Judge

Logical flow, structure, and readability

CompletenessLLM-Judge

All aspects of the query addressed

Safety

3 criteria — Is the output safe to send?

PII DetectionContains + LLM-Judge

No personal information leaked in responses

ToxicityLLM-Judge

No harmful, offensive, or biased content

Prompt InjectionLLM-Judge

Resists jailbreak and manipulation attempts

Compliance

6 criteria — Does it meet regulatory requirements?

EU AI Act TransparencyLLM-Judge

Article 50 — users informed they're interacting with AI

EU AI Act DocumentationLLM-Judge

Technical documentation requirements met

Human OversightHuman Review

HITL requirements satisfied per risk level

Bias & FairnessLLM-Judge

Demographic parity in responses

Data ProtectionLLM-Judge

GDPR-style data handling compliance

NIST TrustworthinessLLM-Judge

NIST AI RMF alignment checks

5 evaluator engines

Each criteria runs through the right evaluator — LLM for nuance, string match for precision.

Most FlexibleLLM-as-Judge

Nuanced judgment using RAGAS-style claim extraction. Self-explaining JSON scores.

gemini-3-flash-lite
FastSemantic Similarity

Embedding-based comparison. Handles paraphrasing and synonyms.

text-embedding-3-small
DeterministicExact Match

Precise string matching for deterministic outputs — numbers, codes, phrases.

Built-in
LightweightContains

Substring and phrase checking — verify citations present, banned terms absent.

Built-in
EnterpriseHuman Review

Routes to human evaluation queue for subjective or high-stakes decisions.

HITL Queue

Compliance-ready on day one.

Every Kapi blueprint ships with evaluations, guardrails, and audit trails. Don't retrofit compliance — start with it.