How I used TypeSafe’s new System One model (Jev) to save GitHub Actions minutes and Claude Code tokens on a Sentry Automation
Filtering Sentry noise with Jev before it reaches Claude Code and GitHub Actions.
That title is quite a mouthful! Hopefully, I’m still able to organise ideas in a clear and coherent manner to unpack it. It’s been almost a month since I submitted my Master’s thesis, and I haven’t done any significant writing (with minimal AI assistance) since then. Let’s begin with the Sentry automation setup.
Sentry automation
Sentry is an error tracking and application monitoring platform that helps developers find and fix bugs the moment they happen in an application while it’s in production. It captures the complete error context such as the exact stack traces, affected user counts and environment details such as the operating system, device type and browser on which the error occurred. For example, Sentry has helped us detect that our app blanked entirely on browsers with blocked site data. This was an interesting error to catch and fix and it highlights one of the gaps that Sentry addresses: I might not have caught the error while testing in my own browser which has no site data blocked. Additionally, we did not have to rely on the user reporting the crash, which they never did, to become aware of the issue.
Initially, after an issue was reported in the Sentry dashboard, I used to fix it in VS Code or paste the stack traces into Claude Code for fixing. However, I’ve since set up an automation that initiates fixing of issues by Claude as soon as they are filed in Sentry. Once Claude fixes the issue, it opens a PR into a staging branch in GitHub. I then enter the loop to review the PR, merge and do QA. The automation means I spend less time checking Sentry’s dashboard and pasting issue traces into Claude Code.
Here’s how the automation fits together. A new issue in Sentry triggers a webhook to be sent to a Cloudflare worker relay, with the details of the issue. The Cloudflare worker filters out issues that shouldn’t trigger the automation and sends a repository dispatch event to GitHub for issues that pass the filter checks. The event starts a GitHub Actions workflow that runs Claude Code. Claude investigates the bug and uses its judgement to decide whether or not the bug is something it should fix. If it can’t, it stops. If it can, it does and opens a PR. The setup used to use Opus 5 but I’ve since switched to Sonnet 5.5 since its release.
The problem
The setup described above is not resource free. The automation relies heavily on GitHub Actions and GitHub caps the amount of monthly actions minutes on the free tier at 2000. That sounds like a lot of minutes but given how much code I ship, thanks to coding agents, I exhausted them mid month at one point before applying optimisations. Most of the minutes were consumed by my CI/CD setup that also uses GitHub Actions. I will probably follow up in a future post sharing my thoughts on the elevated importance of automated testing in the agentic coding era.
The setup also uses Claude to fix bugs and that eats into my Claude Pro usage limits. Every issue that Claude fixes or decides it won’t fix, consumes tokens. This meant that valuable minutes and tokens were spent on issues that Claude would ultimately classify as noise. In this context, “noise” means issues that Claude consistently determines are not actionable application bugs. For instance after a new deploy, a tab that was loaded before it asked for a JavaScript chunk the new deploy no longer served. The app recovers on its own by reloading so Claude rules it as not a defect in code.
My first course of action to address such noise was to add a pre-filter to the Cloudflare Worker that relays Sentry’s webhooks. Before dispatching, it reads the issues and skips the cases Claude always stopped on. This cut a significant number of issues in my evaluations, around 68 percent. However, the prefilter relies on hardcoded rules and exact matches. Noise can still escape if it doesn’t contain the exact tags the prefilter is meant to catch. An ML classifier was still necessary.
Enter Jev
Jev is a System One model released by TypeSafe around three weeks ago. A System One model is an AI model designed to make fast, structured decisions for software to consume directly, rather than generating natural-language text or prose for humans to read like LLMs such as Claude, Gemini or GPT. As outlined in TypeSafe’s docs, Jev can choose an option from a set, score based on a rubric or answer yes/no questions and it provides the confidence in its answer as a percentage. Even though it does not generate prose or perform complex reasoning like LLMs, Jev caught the attention of the AI community because of how cheap and fast it is. As per the model’s release blog, it is 193.6x faster and 444.6x cheaper than frontier LLMs (on TypeSafe’s internal evaluations).
Jev seemed particularly well suited for the task of classifying Sentry issues as noise or not. Running issues from the prefilter through Jev before they got to Claude could lead to significant cost cutting. I implemented Jev into the automation pipeline and on initial evaluation with historical issues, it showed a lot of promise. I’m still monitoring the effectiveness in production and will share findings soon.
Thanks for reading. And check out pachena.co, Africa’s platform for workplace transparency, data and insights to know more about what we are building.