LLM Evaluation
Human reviewers assess AI-generated responses against your own rubrics and requirements.
- Factuality and accuracy review
- Instruction following
- Safety review, where defined by your guidelines
- Rubric-based scoring
- Relevance to the prompt or task
- Hallucination detection
- Pairwise comparison of responses
- AI-generated code response review