One of the most useful things AI can leave behind is a process that no longer needs a conversation to run. You explain what you want, work through the exceptions, and turn the result into software. Tomorrow, you press a button. The work starts from the same rules.
That is what interests me about AI building determinism. A model can help design and implement a repeatable system. It can also supply judgment inside that system, at the points where a fixed rule would be too brittle. Both uses become more valuable when we are precise about where the uncertainty lives.
Look closely at an agent harness, a new decision model like Jev, or the small news application I recently built with an AI coding assistant for my own Mac. Underneath the model calls, you find familiar software: classifiers, branches, validation, stored state, and limits. Those pieces have been doing useful work all along.
The ordinary software around the model
The March 2026 Claude Code source leak brought attention to the machinery surrounding a coding model. Anthropic confirmed that a release had accidentally included internal source code. The public documentation gives us a firmer basis for a specific architectural observation: some of that surrounding machinery uses ordinary pattern matching.
Claude Code’s hook reference documents exact-name and regular-expression matchers for deciding which events trigger a hook. A tool-name pattern can select the calls a check applies to. Code then runs the check at a defined point in the lifecycle. That is a small, explicit decision with a result a developer can inspect.
A regex is a perfectly reasonable way to recognize a known shape. It becomes fragile when we ask it to interpret everything a person could mean. Conversely, asking a language model whether an exact tool name matches a fixed list adds uncertainty to a question code can settle directly.
The engineering work is choosing the boundary. A model may interpret an unfamiliar request. The harness still needs to know which tool is available, whether execution is permitted, how to handle failure, and when to stop. Pattern matching is one component of that control, not a complete security boundary.
Jev makes a smaller answer interesting
TypeSafe introduced Jev on September 15, 2026 as its first public System One model. Its attraction is easy to understand: many steps in an application need a decision that software can consume immediately. Generating an explanation at every step can be unnecessary work.
Jev takes state and typed questions. Its documented primitives cover choosing an option, scoring against a rubric, and estimating whether a statement is true. The results include probabilities; Choice and Score also include confidence. A developer can use those values in a branch, a ranking, or a review queue.
TypeSafe reports large speed and cost improvements on its own workflow evaluations. Its launch post also explains that the biggest reported gains are likely toward the high end of real-world results. Those are vendor measurements, not benchmarks we have reproduced. The practical appeal remains: a fast decision can be useful at many points where a longer model call would interrupt the application.
There is a subtlety here. In its FAQ, TypeSafe distinguishes identical results for identical input from similar decisions when the meaning stays similar. It says Jev is designed for the latter: consistency. That is a better description than treating the model as a guarantee of strict determinism.
And a valid answer can still be wrong. If the allowed outputs are “relevant” and “irrelevant,” always returning one of those values solves a formatting problem. It does not prove the relevance judgment. TypeSafe’s own limitations page describes problems with exact arithmetic, adversarial content, and sensitivity to the order of choices. Those are useful boundaries to design around.
The idea is already spreading
We can point to concrete implementation work without guessing at Jev’s market share. LangChain’s integration exposes a classifier and demonstrates using it to route requests between models and check proposed tool calls. In an October 1 write-up, LangChain reports that moving its classifier to Jev made classification almost 50 times faster in that implementation. That is a result for their classifier, not a promise that every agent becomes 50 times faster.
Other projects are exploring the interface from different directions. Nokia Applied Research’s AnyJev adapts open language models to return typed decisions and probabilities. Laya offers an open decision engine and a Jev-compatible API. Compatibility and inspiration do not establish that they reproduce Jev’s architecture, training, or performance.
To me, the interesting sign is that teams are making the decision itself a distinct component. It can have its own latency budget, evaluation set, and replacement strategy. You can improve that component without asking it to own the entire application.
The classifier never left
Consider a support inbox. A classifier estimates whether a message concerns billing, a technical problem, or a sales question. Application logic decides where to send it. Another check may decide whether the message needs a person immediately. Permissions determine who can see it.
The label is one input into a larger process. Whether it comes from a traditional supervised classifier, a generative model, or a typed decision model, somebody still has to define the available destinations and what each destination does.
The same applies to a decision tree. Here we mean the branching logic of an application; a trained decision-tree model is one possible implementation, not a requirement. If the source is missing, do not claim to have read it. If a task has already completed, do not repeat its side effect. If a judgment is uncertain, take a defined review path.
AI broadens the kinds of inputs that can reach these branches. A condition can depend on the meaning of a paragraph instead of an exact keyword. The responsibility for the branch remains with the application.
That separation also makes improvement easier to explain. Was the wrong category selected? Was the category correct but its destination wrong? Did a retry repeat an action? These are different failures, and they need different fixes.
A repeatable news process on a Mac
My starting point was personal: I want to search these things and get a consolidated report back daily. I had subjects I wanted to follow, and I wanted the results in one place, with links I could open when something deserved a closer look.
So I worked with an AI coding assistant to build a small news app on my Mac. In this story, “we” means me, Trevor Ewert, and the AI helping me build it. It is a personal tool, not a KWIP product. We translated that recurring request into saved projects, topic terms, and a history of reports. The current app runs when I start a search; the daily report is my reading routine and goal, rather than an automatic delivery feature.
At the beginning of a run, my saved terms are retained. The selected model can suggest related searches, and the implementation accepts at most two of those suggestions. If that expansion fails, the app continues with the original terms and shows a warning. The model can broaden the search within that limit; it cannot silently replace the whole assignment.
- Choose topicsTopics I choose
- Expand & retrieveAt most two added searches
- Compose & validateCheck source references
- Keep the runReport and evidence
The app queries Google News RSS directly, makes small sequential requests, and deduplicates the gathered results. The selected model then composes findings from the supplied headlines and snippets. The app does not retrieve and read the full articles, so the report has that evidence limit.
Before accepting the report, code filters its source IDs against the sources actually retrieved. Findings without a surviving source reference, a title, or a summary are discarded. If no valid findings remain, the app reports an error. This establishes that the references exist. It does not establish that every claim is supported by the referenced snippet; that still needs careful reading.
Past runs preserve the original terms, provider and model information, evidence, and output locally. I can inspect what happened. A new run may retrieve different news or produce different wording, so saving a run is different from guaranteeing an identical rerun.
Jev is not part of this app. The connection is architectural: the model has a bounded contribution, and the application owns the sequence, limits, and failure behavior. The same design question applies whichever model supplies the judgment.
AI can help build the repeatability
The second role for AI was in building the software itself. I supplied the goal and worked through the behavior with the assistant: which topics to keep, what a report should contain, and what should happen when a search or model call fails. The AI helped turn those decisions into the application. Our conversation produced a process I could use again.
Now I do not have to reconstruct the entire research request in a fresh conversation each time. I keep my topics, choose when to run, and inspect the resulting report. Some of the work still needs inference, but the surrounding instructions have become application behavior.
This is a useful standard for an AI-assisted build: what repeatable capability remains when the building session ends? A saved workflow, a validator, or a small application can make the next attempt cheaper to start and easier to examine.
Make the promises separately
Three promises are easy to confuse. A stable interface means the answer has the expected shape. A repeatable process means the same rules govern execution and failure. A correct result means the answer is supported by the facts. Each needs its own evidence.
For a decision model, test ambiguous examples and wording changes. For the workflow, test missing data, failed calls, cancellation, and retries. For a report, check whether its claims follow from its sources. Record the version you evaluated: TypeSafe’s model documentation explicitly notes that a moving alias can change the model behind a request.
Keep exact arithmetic, permissions, and execution limits in code. Give a model the questions that benefit from interpretation. Keep an outcome for uncertainty instead of forcing every input into a confident action.
What this could mean for guardrails
The same separation matters when a decision protects an action. A fast classifier could help check whether a proposed tool call fits the user’s request, whether retrieved material is trying to redirect the task, or whether a result needs human review. If those checks become cheaper and faster, I expect more applications to place them at several points in a workflow, including immediately before an action has an external effect.
But a confidence score cannot grant permission. The application still needs to enforce which tools are available, which destinations are allowed, and which actions require approval. A model’s judgment can add a reason to stop or ask for review; it should not silently override a hard restriction. TypeSafe’s documented adversarial-input limitations make that distinction especially relevant.
Multimodal inputs make the opportunity more interesting. Imagine an assistant comparing a spoken request with the destination shown in a screenshot before sending a file. A guardrail could look for a mismatch between what the person asked for and what the interface is about to do. A check that sees the image could also notice a warning that never appeared in a text transcript.
That is a direction for development, not a capability I am attributing to Jev today. Its current model specification lists text input, without native image, audio, or video support. A system using it for those cases would first need another component to extract text or structured observations. That extraction can lose information before the decision model sees it.
Even with a future native multimodal decision model, perception would remain fallible. A screenshot can be ambiguous, a transcript can omit a word, and instructions embedded in an image should not acquire authority merely because a model can read them. I would want evaluations that measure both unsafe actions allowed and legitimate actions blocked, across each input type and cases where the inputs disagree. Uncertainty needs a defined path to pause, decline, or ask a person.
For my news app, the useful boundary is simple: keep my topics, gather evidence, compose a report, and make the references inspectable. For a system that can act on the world, the boundaries need to reach further. The common idea is the same: use intelligence where interpretation helps, and give the surrounding software clear rules for what happens next.
That is the version of “AI builds determinism” I want to build toward: intelligence helping us create software whose behavior we can describe, inspect, and repeat. The next run should inherit what we already worked out.