MURAGE
How it worksFeaturesSolutionsUse casesComparePricingDocs
DownloadDownload
MURAGE

An AI Chief of Staff and a team that does the work. Free and open source, from Ferrox Labs.

Product

  • Features
  • How it works
  • Engines and apps
  • Privacy and security
  • Pricing
  • Murage Cloud
  • Download
  • Changelog
  • FAQ

Solutions

  • Solopreneurs
  • Founders and startups
  • Agencies
  • Ecommerce
  • Small business
  • Creators and educators
  • Developers
  • SaaS
  • Nonprofits
  • Consultants and freelancers
  • Customer success teams
  • Legal and professional services
  • Small teams
  • Sales teams
  • Operations managers
  • Industries
  • Use cases
  • Build your team

Compare

  • vs Claude Code
  • vs Codex (ChatGPT Work)
  • vs Claude (Cowork)
  • vs Paperclip
  • vs OpenClaw
  • vs Hermes Agent
  • vs Manus
  • vs Genspark
  • vs n8n
  • vs Zapier
  • vs Grok Bot
  • vs Viktor
  • vs Lindy
  • vs Tasklet

Resources

  • Docs
  • Team library
  • Skill library
  • AI Chief of Staff
  • How we test
  • Limits
  • Switching to Murage
  • Open source
  • Blog
  • Release notes

Company

  • About
  • Partners
  • Enterprise
  • Contact
  • Privacy
  • Terms
  • Cookies
© 2026 Ferrox Labs. Murage is open source under AGPL-3.0.
AI and machine learning skills

Model Evaluation Specialist

Advanced model evaluation covering LLM benchmarks, evaluation frameworks (lm-evaluate-harness, HELM, RAGAS), leaderboard interpretation, custom metrics design, human evaluation protocols, automated LLM-as-judge patterns, and evaluation pipeline architecture for both traditional ML and generative AI systems.

Download Murage

Free and open source. macOS, Windows and Ubuntu.

Try it

Add the skill to a bot, then ask your Chief of Staff:

“Use the Model Evaluation Specialist skill on this: [describe the job, or paste your notes].”

Model evaluation for modern AI extends beyond accuracy metrics. LLMs and generative models require multi-dimensional assessment: factuality, reasoning, instruction following, safety, and human preference alignment. This skill covers evaluation frameworks, benchmark suites, custom metrics, LLM-as-judge patterns, and reproducible evaluation pipelines.

What it covers

  • LLM Benchmark Landscape
  • Evaluation Frameworks
  • LLM-as-Judge Patterns
  • Prompt
  • Response A
  • Response B
  • Custom Metrics Design
  • Human Evaluation Protocols

Skill Guard: clean

Scanned and clear. A bot can use it as soon as you add it.

Subject
AI and machine learning
Level
intermediate
License
Apache-2.0
ai-mltestingguide

How to use it.

  1. 1

    Download Murage

    Free for Mac, Windows and Ubuntu.

  2. 2

    Add the skill

    Open Settings → Skills and switch on Model Evaluation Specialist for any bot. Or ask your Chief of Staff to pick skills for a job.

  3. 3

    Give it a job

    The bot reads the skill when the job calls for it, and works the way it lays out.

Put this skill to work.

  • ResearchAsk a real question. Get one page with sources.
  • Skill GuardEvery skill is scanned before a bot can use it.

Questions.

What is the Model Evaluation Specialist skill?+

Advanced model evaluation covering LLM benchmarks, evaluation frameworks (lm-evaluate-harness, HELM, RAGAS), leaderboard interpretation, custom metrics design, human evaluation protocols, automated LLM-as-judge patterns, and evaluation pipeline architecture for both traditional ML and generative AI systems.

How much does it cost?+

Nothing. It ships with Murage, which is free and open source. Your bots run on the AI plan you already have.

Can I change it?+

Yes. Open it in Settings → Skills to read, edit or duplicate it. Skill Guard scans it again after every edit.

More AI and machine learning skills.

  • AI Readiness EvaluationOrganizational AI readiness assessment evaluating data quality, team capabilities, infrastructure and governance.
  • Feature EngineerFeature engineering covering numerical features (scaling, binning, log transforms), categorical encoding (one-hot, target and ordinal).
  • ML Ops EngineerMachine learning operations covering MLflow experiment tracking and model registry, model versioning and reproducibility.
  • ML PipelineML pipeline design covering feature engineering, model training workflows, hyperparameter tuning, cross-validation, experiment tracking (MLflow and W&B).
  • MLOps EngineerMLOps and deployment covering model serving (TorchServe, TF Serving, Triton), containerized inference and A/B testing deployment.
  • Model EvaluatorML model assessment covering classification metrics (precision, recall, F1, AUC-ROC), regression metrics (MAE, RMSE, R2) and confusion matrix analysis.

All AI and machine learning skills

Give your first job to Murage.

Download the free app, connect the AI you already pay for, and tell your Chief of Staff what needs doing. Plan on about ten minutes from install to a working team.

Free · Download Murage

No account needed · Runs on the AI plan you already pay for