{"type":"project","slug":"personal-eval-studio","title":"Personal Eval Studio","status":"building","tagline":"A local workspace for writing repeatable AI evaluations and reviewing evidence.","summary":"A local studio for authoring versioned prompt pairs and test cases, alongside deterministic example runs and baseline evidence review.","canonicalUrl":"https://jonathantipper.com/projects/personal-eval-studio","markdownUrl":"https://jonathantipper.com/projects/personal-eval-studio.md","jsonUrl":"https://jonathantipper.com/projects/personal-eval-studio.json","launchUrl":null,"launchLabel":null,"launchNote":null,"started":null,"tags":["ai","evaluation","local-tools"],"links":[],"image":{"src":"https://jonathantipper.com/images/projects/personal-eval-studio-preview.webp","alt":"Personal Eval Studio comparing baseline and candidate prompts in a built-in example","mode":"screenshot"},"content":"## What it is\n\nPersonal Eval Studio is a local workspace for making AI evaluation more repeatable. It brings prompt pairs, test cases, saved revisions and evidence review into a guided interface alongside the existing evaluation harness.\n\n## What it does\n\nStart from a template or a blank evaluation, write cases and save revisions without overwriting earlier versions. Definitions can be imported and exported. Separate deterministic example recipes show how runs and named baselines can be compared.\n\n## Current status\n\nThe authoring Studio is in development. Its example runs use fixed answers; they do not test the prompts you write. Live model connections, generation and judging are still to be built. There is no public hosted service."}