AI evaluation resources
Official guides and project documentation for testing prompts, agents, and RAG answers.
Updated 2 min ago
Sources reviewed
Choose the failure you need to detect before choosing a metric or a tool.
| Source | Publisher | Use it for |
|---|---|---|
| AI RMF | NIST | Risk and evaluation scope |
| Agent evals | Anthropic | Tasks, graders, trials |
| Ragas | Ragas project | RAG and answer metrics |
| Promptfoo | Promptfoo | Repeatable comparisons |
Read the project's data-handling and provider settings before sending private examples. Calibrate automated graders against human judgments on consequential cases.
Related pages
- AI-assisted software engineering in 2025What DORA, Stack Overflow, and METR measured about AI adoption, trust, and productivity.
- RAG for transit planningIFP's proposal to make past transit project records searchable, citable, and reusable.
- An AI-assisted wikiLet an agent organize sources and propose linked notes while you retain editorial control.
- Building an AI-first bank cultureJPMorgan's 2025 account of self-service AI adoption alongside targeted workflow redesign.
- Who Are You Trying to Impress with Your Deadlines?A cheatsheet on why hard deadlines hurt your team, your product, and your customers — and what to do instead.