Plugin quality dashboard

Skills that make delivery disciplined.

Everything the SKRAFT plugin ships — skills, agents, workers, review lenses — with its context profile, its evaluation coverage, and the evidence from controlled skill-versus-baseline runs.

These filters apply to catalogue, quality, efficiency and model comparison.

Usage order

Product preflight, then engineering

Optional standalone product workflows precede the single engineering entrypoint. Brownfield and other roots stay independent.

Source-derived catalogue

Agents, workers, lenses and skills

Loading catalogue…

Controlled comparison

Quality and activation

How to read quality

Each model is compared with and without the skill. Models appear together only when the same judge scored them; “highest score” applies to this skill and judge, not to every task.

Baseline
Score without the skill.
Skilled
Score from the same model with the skill available.
Lift
Skilled minus baseline. Positive means the skill helped on this run set.
Activation rate
How often the skill loaded when expected.
Unexpected
Loads on prompts where the skill should have stayed inactive. Zero is the target.
Verdict
Statistical conclusion across paired trials, not a single success or failure.

Runtime profile

Efficiency

How to read efficiency

Compare each model with its own baseline. Lower resource use is useful only when quality remains acceptable; efficiency alone does not identify the best model.

Duration
Median elapsed time per trial.
Tokens
Median model input and output volume per trial.
Turns
Median number of agent reasoning cycles.
Tool calls
Median number of external actions taken by the agent.
Delta
Change from baseline. A negative percentage means the skilled run used less.

One arm per agent model

Model comparison

How to compare models

One row is one agent model. Compare rows only inside the same table: they used the same skill and were scored by the same judge. “Highest score” is local evidence, never a universal model ranking.

Agent model
Model that performed the task.
Judge
Model that scored the outputs against the rubric.
Highest score
Largest skilled score in this table; inspect lift and efficiency before deciding.