# AI tool evaluation worksheet

Prepared by SOTA — https://sota.cc/guides/choosing-an-ai-tool/

This is a blank evaluation plan, not a completed benchmark.

- Task and intended user:
- Correct result and explicit failure result:
- Project, exact package versions or commit:
- Runtime and deployment environment:
- Provider, model identifier and settings:
- Data handling and permitted tool actions:
- Source documentation and date checked:

| Case | Expected result | Observed result | Validation error | Factual error | Retries | Latency | Usage/cost |
| --- | --- | --- | --- | --- | --- | --- | --- |
| Normal input | | | | | | | |
| Missing information | | | | | | | |
| Contradictory information | | | | | | | |
| Unavailable tool/provider | | | | | | | |
| Cancellation | | | | | | | |
| Similar inputs requiring different answers | | | | | | | |

Record repeated trials and the method used to measure them. Separate documented
features from observed behavior. Do not turn GitHub stars or schema validity into
a quality score. State what remains untested before making a recommendation.
