Complete AI TrainingYourJobSkills for your job

Skills / agent-evaluation

agent-evaluation-reporting

Use when summarizing agent evaluations where autonomous, assisted, failed, timed-out, or invalid outcomes must remain distinct and comparable.

newData & AnalyticsAI & Automationagent-evaluationbenchmarkingmetricsreliabilityreporting

Agent Evaluation Reporting

Overview

Turn raw agent evaluation runs into a decision-ready report without hiding failures or overstating capability. Keep outcome populations, denominators, latency populations, and experiment conditions explicit so readers can reproduce every headline number.

When to Use This Skill

  • Use when reporting benchmark, regression, pilot, or production evaluation runs for an AI agent.
  • Use when autonomous and human-assisted completions appear in the same result set.
  • Use when failures, timeouts, infrastructure-invalid runs, retries, or partial results affect the denominator.
  • Use when comparing two agents, prompts, harnesses, or releases and deciding whether the comparison is valid.

How It Works

Step 1: Freeze the comparison contract

Record the task set and sampling, model and provider, prompt or policy version, tool and harness versions, evaluator rubric, timeout and retry policy, token or cost budget, environment, and human-intervention policy. Assign the configuration a stable label or digest.

If a material condition differs between runs, mark the comparison as non-equivalent. Report a directional observation only; do not claim that the changed agent caused the difference.

Step 2: Build a mutually exclusive outcome ledger

Classify every scheduled attempt exactly once:

OutcomeMeaning
`autonomous_succ

Subscribers only

The full skill, its 1 bundled files and every download is included with every paid Complete AI plan.

Details

Sourceself
License
Risk labelnone ("critical" means the skill may run commands or touch files — read before use)
FilesSKILL.md
Added2026-08-18

Related skills

run-deep-swe

Run reproducible DeepSWE coding-agent benchmark evaluations through OpenRouter and mini-swe-agent.