About this project
skill-up is a command-line tool for evaluating Agent Skills, agents and their workspaces. It executes declarative evaluation cases against an Agent Engine, grades the response and any workspace changes, and produces reports either locally or in CI.
Three evaluation modes are described: Skill evaluation, which measures a Skill's behavior across cases and can compare runs with and without the Skill; Agent evaluation, which runs the same cases without installing a Skill to assess an engine or custom agent on its own; and Workspace evaluation, which checks how an agent works with files and repositories using case fixtures or an existing local workspace.
Configuration is declarative YAML: an eval.yaml plus cases/*.yaml files define the environment, engine, model and cases. The tool provisions a local, Docker or OpenSandbox runtime per case, can install configured real or mocked MCP servers into supported engines, and can upload context.repo_fixture and context.files into the workspace. Runs can be skill-optional, and an existing local directory can be targeted with --workspace (requiring environment.type: none, one case at a time, and benchmark mode off).
Built-in Agent Engines include Qoder CLI, Claude Code and Codex, with user-defined agents supported through engine.custom over local transport. Judging supports rule_based, script and agent_judge strategies. Reports include Anthropic-compatible grading.json and benchmark.json, benchmark.md, plus result.json, JUnit XML and HTML output. An import command migrates Anthropic-style evals.json into the YAML layout, and --auto can auto-detect it.
A companion Skill called skill-upper is shipped in the repository and is intended to close an eval-to-evolution loop: it inspects failures, repairs or expands the eval suite, and reruns skill-up through conversation. A DeepSeek Harness plugin bundle adds durable observation capture, approval-gated regression cases, isolated runs and status comparison. A GitHub Action at the repository root runs evals on pull requests across engines; it is a Docker container action requiring a Linux runner and a model credential stored as a secret, and it pins an immutable image digest.
A user-level config supplies default OpenTelemetry environment variables and per-environment runtime kwargs, with a discovery chain from embedded defaults through user and project files to an explicit --config path. Secrets are expected to be referenced via ${ENV_VAR} rather than literals. The CLI exposes run, validate, list-cases, report, import and debug subcommands. The project is written in Go (>=1.25) and licensed under Apache 2.0.
Comments
0 Rating appears after 10 ratings
Sign in to join the discussion.