Complete AI TrainingYourJobSkills for your job

Catalog

agent-evaluation skills

2 skills. Every one is readable here; downloads and live use need a subscription.

agent-evaluation-reporting

Use when summarizing agent evaluations where autonomous, assisted, failed, timed-out, or invalid outcomes must remain distinct and comparable.

run-deep-swe

Run reproducible DeepSWE coding-agent benchmark evaluations through OpenRouter and mini-swe-agent.