Try it
Add the skill to a bot, then ask your Chief of Staff:
“Use the Model Evaluation Specialist skill on this: [describe the job, or paste your notes].”
Model evaluation for modern AI extends beyond accuracy metrics. LLMs and generative models require multi-dimensional assessment: factuality, reasoning, instruction following, safety, and human preference alignment. This skill covers evaluation frameworks, benchmark suites, custom metrics, LLM-as-judge patterns, and reproducible evaluation pipelines.
What it covers
- LLM Benchmark Landscape
- Evaluation Frameworks
- LLM-as-Judge Patterns
- Prompt
- Response A
- Response B
- Custom Metrics Design
- Human Evaluation Protocols