Upload or define test cases with expected answers. The model's outputs are compared to ground truth and scored by an LLM judge.
Custom Dataset Evaluation
Model: Not configured · Judge: Same as active
70%
Model output score must reach this threshold to pass each case
API not configured. Go to Settings
Test Cases (3 of 3 valid)
Each case needs a question and expected answer
1
What is the capital of France?
Expected: Paris
2
What is the chemical formula for water?
Expected: H2O
3
Who wrote the play Romeo and Juliet?
Expected: William Shakespeare