I tried using Jev to make my AI slop a little less sloppy. Here’s what I did.
I dusted off an old proof of concept I’d built a few months ago. When I was building it, I checked it two ways: had AI check its own work, then used it myself to see what broke.
This time, I wanted to try something else. Jev takes information and judges it against a set of choices you provide. Could we use it to check whether the code followed the rules of the app?
Here’s one thing it caught.
A user could select suggested changes to their plan and click Apply. The suggestions disappeared. Looks done, right?
Except when a save failed, the suggestions disappeared anyway.
I can miss that testing an app myself. I click the button, the interface responds, and I keep going. Meanwhile, the saved plan tells a different story.
Jev flagged the inconsistency. We tested what happened when a save failed, reproduced the bug, and Codex fixed it. Now, changes that fail to save stay visible.
So how did we get there?
I didn’t have a specification sitting around. Codex and I put together a skill, a set of reusable instructions, to discover the app’s rules from its prompts, UI copy, and code. It saved a reusable rulebook and ran a first review across the codebase.
The promises could be as basic as: if saving a change fails, the app shouldn’t act like it succeeded. Where Codex inferred a rule, it marked it as provisional.
Codex gathered the relevant code and turned those rules into focused questions for Jev. Jev returned judgments, and Codex investigated what they pointed to.
We cleaned up a couple of bugs. That’s enough to make me want to try it again.
Would Codex have found the same things on its own? I don’t know. I’d also like to measure whether this saves a boatload of Codex tokens. I suspect it could, but I haven’t measured that.
Next, I want to try this on a few more apps, especially a more complex production app, and see what it catches.
If you want to try it on yours, here are the exact skills I used. Give it a try and let me know how it goes!
Download the skills
Download both review skills (ZIP)
The bundle includes the two review skills Codex and I put together. You can also download them individually:
- App rules (ZIP): discover the app’s rules, save a reusable rulebook, and run the first review across the codebase.
- Pre-commit review (ZIP): check later changes against that saved rulebook.
Both depend on the TypeSafe AI skill, created and maintained by TypeSafe. Get their original skill and installation instructions on GitHub.
Install TypeSafe’s skill from their repository, then unzip this bundle and ask Codex to install the two review skill folders. Then ask it to use typesafe-app-rules to inspect your app. The included README covers setup, including the TypeSafe API key needed to run Jev evaluations.