AI Eval Platform

Compare two model outputs side-by-side. An LLM judge picks the winner with dimensional scoring and detailed reasoning.

A/B Pairwise Comparison

Judge: Same as active · A: Not selected · B: Not selected

No models configured. Add models in Settings

No models configured. Add models in Settings

Add at least 2 models for A/B comparison. Go to Settings