AI & agents

Why raising your confidence threshold does not make it safer

The mistake

You wire a model into your support queue. Every answer comes back with a number beside it, and anything at or above 0.75 is handled without a human.

Forty tickets went through this morning. Ten of the answers were wrong, and nobody read them. The fix looks obvious, because the number looks like a safety dial. Turn it up.

The machine

Simulator · confidence threshold

Every count runs on the tested reducer, over a dataset written so that the calibrated model's confidence is exactly right. No model was called.

Overconfident model, acting alone at or above 75%. 40 tickets handled alone, 10 of them wrong with nobody looking, 0 sent to a human.

Drive it

The queue above is the one from this morning. The model has already answered all forty tickets, and it is acting on every answer it rates at 75% or over.

  • Press “Raise”, and keep pressing. The counters sit under the queue. Watch the one that says nobody looked.
  • Keep going until the button will not move any further. Then press “Show the mistakes” at the top of the simulator. Those are the ten it got wrong, by subject, and who read them.
  • Now switch the model to Calibrated and work the dial again, down with “Lower” and back up to 90%. Same forty tickets, same forty answers, same ten mistakes. Only the number written beside each answer is different.

The mechanism

A classifier gives you two things: an answer, and a number attached to it. Only the answer is about the ticket. The number is about the model.

Calibration is the property that makes that number usable. A model is calibrated when it is right as often as it says it is. Take every answer it labelled 0.8 and count how many were correct. You should get 80%. Do the same for 0.6 and you should get 60%.

The chart at the bottom of the simulator is that count. Each row is one confidence level the model used: what it said on the left, how often it was really right on the right. The model you arrived on has a single row. It says 95%, and it delivers 75%. Your threshold reads the left number. Your customers live with the right one.

That is why the dial could not save you. >= 0.9 means “act when the model claims 0.9”. It is a test on the claim, so it is only ever as good as the claim, and a model that claims 0.95 on everything passes every test you can write. No setting excludes the ten wrong answers, because the model never marked them.

The ten stay ten, whatever you do. The threshold does not change a single answer, it only decides who reads it. Switch to the calibrated model and the same ten tickets are still wrong. What moves is where they sit. At 90% nine of the ten fall below the bar, where a human reads them. At the top of the dial, where you just left it, all ten do.

So the number to trust is not the one in the conditional. It is the one you measured yourself, on your own data, against outcomes you already knew. And you measure it again when the model or the prompt changes.

In your code

The classification API in laravel/ai returns one typed answer per question, so the threshold is an ordinary conditional. Two of them, because certainty lives at both ends of the range.

use Laravel\Ai\Classification;
use Laravel\Ai\Classification\Boolean;

$answer = Classification::of($ticket->body)
    ->question('spam', new Boolean('Is this message spam?'))
    ->classify()['spam'];

if ($answer->isTrue(threshold: 0.9)) {
    $ticket->markSpam();   // sure enough that it is spam
} elseif (! $answer->isTrue(threshold: 0.1)) {
    $ticket->route();      // sure enough that it is not
} else {
    $ticket->escalate();   // anything between the two, a human reads it
}

A boolean answer carries probability, the probability that the answer is yes. That is not a confidence, and reading it as one is expensive: 0.02 is a confident no, not an uncertain anything. Escalate everything under 0.9 and you send the model’s surest answers to a human along with its least sure ones.

isTrue() takes a default threshold of 0.5, which means “act whenever the model leans yes”. That is the habit this page is about, shipped as a default.

Classification::fake() and assertClassified() test all three branches with no network call. What they cannot tell you is whether 0.9 is the right number. Only your own labelled outcomes can.

The fine print

The forty tickets are written by hand. Nothing here was measured. They exist so that “calibrated” has an exact meaning the test suite can assert: the calibrated model’s rows match its real accuracy to the percentage point, by construction. Nothing on this page is a claim about how any real model behaves.

The three models make identical predictions. Only the confidence they state differs. Real models are not so tidy, and a miscalibrated one is usually less accurate as well. They are held equal here so the page is about calibration alone.

The dial stops at 95%, and a real one would go to 99%. The ceiling is there to be hit, but nothing rests on where it sits: a model that states 0.99 on everything defeats a 0.99 threshold in the same way. There is always a number the model can state.

The third model in the picker is the opposite failure. It understates itself, so nothing clears a high bar and a human reads all forty, including the thirty the model already had right. That one costs payroll instead of customers, which is why nobody reports it.

Left out: how calibration is achieved or repaired, which is its own subject; calibration across many classes, which is harder than the two-way case here; and the cost of being wrong, which is rarely the same in both directions, while the simulator runs one threshold for both.

Further reading

  • TypeSafe: Introducing System One Models and Jev is the launch post, and the source of the calibration claim this page answers. Read it for what is being promised, and for how much of the decision it leaves to you.
  • laravel/ai PR #1010 is the diff that added the classification API, including the Boolean and Choice question types, the answer objects and the fake. Read it before any write up of it, including this one.
  • Wikipedia: calibration of probabilities covers the reliability diagram and the standard scoring rules, if you want the formal version of the chart in the simulator.

Spotted a problem, or have a way to make this clearer? Suggest an improvement.