System One Model

Why "Jev-as-a-Judge" is a Game Changer for AI-Driven Automation

A "System One Model" known as "Jev" (1) is currently drawing significant attention, leading many to suggest that full-scale AI-driven automation is about to begin. In this article, taking the classification of bank customer complaints as our theme, we explore how applying "Jev" can make automation successful.

 

1.Measuring Prediction Accuracy

The dataset prepared for this evaluation consists of US bank customer complaints (2). Using "Jev," we tackled a classification task to predict which financial product each individual complaint pertains to. We built a prediction application as shown below and compared the accuracy across five AI models, including "Jev."

The task requires selecting one out of the following six financial products. A total of 50 samples were classified.

List of financial products to choose from

As shown below, "Jev" achieved an accuracy of 90%, standing toe-to-toe with the other models, while completely outperforming them in computational speed. The architecture—which returns only the prediction result without generating natural language text—is clearly proving effective here. Truly impressive.

Comparison with other AI models

 

2. Identifying Target Samples for Automation Using Prediction Confidence

A major feature of "Jev" is its ability to simultaneously calculate a prediction confidence score. For predictions made with high certainty, the confidence score approaches 1. Therefore, we use this confidence metric to select only the samples about which "Jev" is genuinely confident as targets for automation. In this test, 1,000 samples were evaluated. We set only the samples with a confidence score of 1 as candidates for automation. The qualifying sample count was 788, yielding an accuracy of 95.1%. While this accuracy is already solid, mapping legacy category names that "Jev" did not select to their updated counterparts (as product names were updated over the course of the long-term data collection period) brings the accuracy up to 98.2%. In other words, the error rate for samples targeted for automation drops below 2%, making them ideal candidates for automated processing. A summary of the results is shown below

Analysis results for samples with confidence = 1

 

3.Implementation Workflow for AI-Driven Automation

We found that "Jev" can effectively identify which samples to automate. In this experiment, samples with a confidence score of 1 accounted for roughly 80% of the total, with an error rate of about 2%. By designing an operation that routes the remaining 20% to human review (Human-in-the-loop), an operational productivity boost of approximately 5x compared to pre-automation levels can be expected. Naturally, results depend on the incoming data, and several points require evaluation—such as whether a 2% error rate is acceptable. Nevertheless, as the precision of confidence scoring is expected to improve in the future, right now is the best time to consider automation powered by "Jev." The automation workflow is as follows

Automation workflow using "Jev"

 

What did you think? Entrusting decision-making tasks to Jev in place of humans is referred to as "Jev-as-a-Judge." As the accuracy of "Jev-as-a-Judge" continues to advance, business process automation driven by AI may take substantial leaps forward. There is plenty of reason for optimism.

 

Here at Toshi Stats, we plan to continue exploring various applications of "Jev-as-a-Judge." Stay tuned!

 

You can enjoy our video news “ToshiStats AI Weekly Review” from this link.

1) Introducing System One Models & Jev, August 15 2026, Diogo Almeida, founder, TypeSafe
2) Consumer Financial Protection Bureau

Copyright © 2026 ToshiStats Co., Ltd. All right reserved.

Notice: This is for educational purpose only. ToshiStats Co., Ltd. and I do not accept any responsibility or liability for loss or damage occasioned to any person or property through using materials, instructions, methods, algorithms or ideas contained herein, or acting or refraining from acting as a result of such use. ToshiStats Co., Ltd. and I expressly disclaim all implied warranties, including merchantability or fitness for any particular purpose. There will be no duty on ToshiStats Co., Ltd. and me to correct any errors or defects in the report, the codes and the software.