Glossary

Agent Eval

Agent eval is repeatable measurement of run quality on a known set of questions. Without it, a prompt or model change is only a feeling.

Also called ارزیابی عامل

Eval should score the same shape the product shows: the claim is right, evidence sticks, a rejected policy stays empty. A “writes well” score is not enough.

Every pipeline change should rerun on the same golden set so a regression is visible.

Let’s make your processes agentic

Where do we start?
Agent Eval · DIDRAH