AI Eval Platform

Upload or define test cases with expected answers. The model's outputs are compared to ground truth and scored by an LLM judge.

Custom Dataset Evaluation

Model: Not configured · Judge: Same as active

70%

Model output score must reach this threshold to pass each case

API not configured. Go to Settings

Test Cases (3 of 3 valid)

Each case needs a question and expected answer

1

What is the capital of France?

Expected: Paris

2

What is the chemical formula for water?

Expected: H2O

3

Who wrote the play Romeo and Juliet?

Expected: William Shakespeare