The local workbench / Build. Compare. Refine.

Skill Studio

Write what good looks like. Test whether your skill gets you there. Build focused expertise with evidence you can inspect.

Download for MacDownload for WindowsDownload for Linux

v1.0.63 · signed & notarized · all platforms, checksums & install notes →

Or open it from your terminal

Local workspace. Your configured model provider handles model requests.

Make the improvement observable

From a useful idea
to a tested skill.

Start with a real task. Describe the expected result, compare configurations, and inspect failures before you publish. Shorter instructions can be a better outcome.

01 / DEFINE

Describe the job.

Write the skill’s purpose, prerequisites, and expected behavior. Keep the trigger narrow enough to leave unrelated work alone.

02 / COMPARE

Change one thing.

Hold the task steady. Compare skill versions or configurations, then inspect what succeeded, failed, or became slower.

03 / REFINE

Keep the evidence.

Use results to improve or retire instructions. Recheck after meaningful model, harness, or tool changes.

Inside the workbench

Designed for iteration.

Your skill, test cases, and results belong together. Studio helps you work through them without turning every comparison into a custom script.

Plain-English Evals

Define expected behavior in your own words. Inspect the judge’s reasoning and check that the verdict matches the outcome you care about.

A/B Compare

Compare prompts and skill configurations side by side. Use the same task and starting conditions so the difference is meaningful.

Your Model

Use supported providers and your configured models. Re-evaluate when the model, tools, or harness changes; results do not transfer automatically.

Local Workspace

Author and manage skills on your machine. Cloud model calls send evaluation inputs to your selected provider; publishing sends content to its destination.

Model ≠ agent

Test the setup
you actually use.

A model operates inside a harness with its own instructions, tools, and permissions. A skill that helps one setup may add overhead to another.

Read the evaluation guide ↗
Hold the task steady
Same request, same starting files, same pass conditions.
Record the configuration
Skill version, model, provider, harness, tools, and effort where available.
Interpret within scope
A result describes the tested setup. It is not a universal model ranking or a safety certification.
Recorded with vskill 1.0.132:41 · 1080p

Skill Studio in 2:41

Watch the authoring, evaluation, publishing, and installation workflow. The recording shows an earlier release.

Chapters

Start with one task

Get Started

Use the desktop app or launch the same workspace from your terminal. Connect the model provider you want to evaluate with.

Run in your project directory

First time? Run npx vskill@latest init to scaffold your test suite. Read the Studio guide ↗

Get the desktop app ↗Native app · v1.0.63