Start with the business decisions and workflows the AI system is expected to support.
The goal is to identify the parts of the product or process where model performance actually matters. That might include high-volume workflows, tasks that are especially important to customers, new capabilities you are considering launching, or areas where mistakes would have meaningful consequences.
It can help to group candidate tasks into categories such as:
- Core workflows the system handles every day
- High-value or high-risk use cases
- Tasks that require more reasoning or judgment
- Workflows where the model already appears to struggle
- New capabilities you want to test before deployment
You do not need to evaluate every possible task equally. The eval set should concentrate on the workflows that matter most to the business and the decisions you want the evaluation to inform.