btay.io/wiki

AI evaluation resources

Official guides and project documentation for testing prompts, agents, and RAG answers.

Updated 2 min ago

Sources reviewed

Choose the failure you need to detect before choosing a metric or a tool.

SourcePublisherUse it for
AI RMFNISTRisk and evaluation scope
Agent evalsAnthropicTasks, graders, trials
RagasRagas projectRAG and answer metrics
PromptfooPromptfooRepeatable comparisons

Read the project's data-handling and provider settings before sending private examples. Calibrate automated graders against human judgments on consequential cases.

Related pages