Glossary
Agent Eval
Agent eval is repeatable measurement of run quality on a known set of questions. Without it, a prompt or model change is only a feeling.
Also called ارزیابی عامل
Eval should score the same shape the product shows: the claim is right, evidence sticks, a rejected policy stays empty. A “writes well” score is not enough.
Every pipeline change should rerun on the same golden set so a regression is visible.
Let’s make your processes agentic
Where do we start?