Project
Personal Eval Studio
A local workspace for writing repeatable AI evaluations and reviewing evidence.

What it is
Personal Eval Studio is a local workspace for making AI evaluation more repeatable. It brings prompt pairs, test cases, saved revisions and evidence review into a guided interface alongside the existing evaluation harness.
What it does
Start from a template or a blank evaluation, write cases and save revisions without overwriting earlier versions. Definitions can be imported and exported. Separate deterministic example recipes show how runs and named baselines can be compared.
Current status
The authoring Studio is in development. Its example runs use fixed answers; they do not test the prompts you write. Live model connections, generation and judging are still to be built. There is no public hosted service.