Describe the job.
Write the skill’s purpose, prerequisites, and expected behavior. Keep the trigger narrow enough to leave unrelated work alone.
The local workbench / Build. Compare. Refine.
Write what good looks like. Test whether your skill gets you there. Build focused expertise with evidence you can inspect.
v1.0.63 · signed & notarized · all platforms, checksums & install notes →
Local workspace. Your configured model provider handles model requests.
Make the improvement observable
Start with a real task. Describe the expected result, compare configurations, and inspect failures before you publish. Shorter instructions can be a better outcome.
Write the skill’s purpose, prerequisites, and expected behavior. Keep the trigger narrow enough to leave unrelated work alone.
Hold the task steady. Compare skill versions or configurations, then inspect what succeeded, failed, or became slower.
Use results to improve or retire instructions. Recheck after meaningful model, harness, or tool changes.
Inside the workbench
Your skill, test cases, and results belong together. Studio helps you work through them without turning every comparison into a custom script.
Define expected behavior in your own words. Inspect the judge’s reasoning and check that the verdict matches the outcome you care about.
Compare prompts and skill configurations side by side. Use the same task and starting conditions so the difference is meaningful.
Use supported providers and your configured models. Re-evaluate when the model, tools, or harness changes; results do not transfer automatically.
Author and manage skills on your machine. Cloud model calls send evaluation inputs to your selected provider; publishing sends content to its destination.
Model ≠ agent
A model operates inside a harness with its own instructions, tools, and permissions. A skill that helps one setup may add overhead to another.
Read the evaluation guide ↗Start with one task
Use the desktop app or launch the same workspace from your terminal. Connect the model provider you want to evaluate with.
First time? Run npx vskill@latest init to scaffold your test suite. Read the Studio guide ↗