Technical Blog

"High Confidence" Is Not a Security Control

· min read
"High Confidence" Is Not a Security Control

AI security tooling keeps getting faster.

Faster triage, faster classification, faster threat intelligence analysis, faster decisions. And, increasingly, faster automated action.

That’s exciting, but it’s also where things start getting dangerous.

I’ve been looking at JEV recently, partly because I was considering integrating it into cARL, my own agentic coding governance tooling.

JEV is interesting because it’s not trying to behave like a general-purpose conversational model. It’s designed to make fast, structured decisions: booleans, choices, scores and classifications.

Which makes it extremely attractive for security automation.

Give it some evidence, ask it a question, get a typed answer. Do that very quickly and very cheaply. What’s not to like?

Quite a lot, potentially.

The judge becomes part of the threat model

Christophe Parisel recently published The “Covert Order”: Exploiting a Weakness in JEV’s Decision Logic, describing what he calls order steering.

The attack is interesting because it doesn’t require adding malicious instructions, special tokens or prompt-injection payloads. The words stay the same, only their order changes.

In his experiment, the same nine sentences were presented to JEV in different permutations while everything else was held constant. The underlying meaning was deliberately balanced between two choices. The result was not order agnostic.

Across permutations, the probability assigned to one class moved from 0.21 to 0.46. A 25-point swing.

More importantly, confidence did not necessarily fall as the answer moved. In the most extreme condition, the model became more confident in the shifted result and that matters because the comforting assumption behind confidence scores is often something like:

if the model is uncertain, confidence falls

But here, changing only the presentation order of semantically identical evidence could change the decision while maintaining or increasing confidence.

Christophe also found an even stronger bias around label assignment.

Changing which class was mapped to the API’s arbitrary “A” label materially changed the outcome, to the point that one class was never selected in 200 calls when it occupied the other label position.

That’s not random noise, it’s exploitable decision structure. And once a model like that is put into a security decision path, the model itself becomes part of the threat model.

Threat intelligence is hostile input

This is what really made my eyes widen.

I’ve seen increasingly impressive examples of JEV being used for high-speed threat intelligence and security analysis.

Adam Chester recently published Experimenting with Jev for Offensive Security, exploring JEV for things such as OPSEC scoring and identifying potentially sensitive files.

His write-up is actually careful about JEV’s limitations. He explicitly notes that it struggled with domain-specific offensive-security knowledge and needed significant contextual handholding in the supplied state.

But that’s exactly why the broader security use case deserves scrutiny. Because threat intelligence has a rather inconvenient property:

the subject of the analysis is often your adversary.

The input is not neutral.

It can contain: HTTP headers, URLs, DNS names, commandline syntax, file names, process arguments, repository metadata, threat reports, malware configuration, email content, package metadata, user agents, detection telemetry and much more.

An attacker may not control the entire input presented to an AI security system, but they don’t need to. They only need to control enough of it to influence the decision.

Traditional evaluation might measure:

  • precision
  • recall
  • false-positive rate
  • false-negative rate
  • calibration

Those are still useful, but they are not enough when an adversary can deliberately construct the thing being classified.

The more spicy question becomes:

Can an attacker change the presentation of the evidence without materially changing the underlying activity, and cause the security decision to change?

For JEV specifically, Christophe’s results say that order alone can materially move a boundary decision.

So now I want to know things like:

  • Does ordering the same telemetry differently change the classification?
  • Can irrelevant attacker-controlled fields steer the decision?
  • Does duplicated evidence receive unintended weighting?
  • Can contradictory context create confidently incorrect results?
  • Can an attacker learn which representation produces the most favourable decision?
  • Can semantically equivalent evidence produce materially different confidence?

At that point, we are no longer measuring ordinary classification accuracy.

We are measuring adversarial robustness.

Those are not the same thing.

JEV as a judge

Another recent article, Jev Model: Jev-as-a-Judge for Agent Evals, explores using JEV as an evaluator for agent traces.

The idea makes sense:

Agent execution produces traces, those traces contain evidence, and JEV can make fast structured decisions about them.

Questions might include:

  • Was the correct tool selected?
  • Was escalation appropriate?
  • Was the final answer supported by the evidence?
  • Which failure category applies?

Manjeet’s article is fairly cautious here too. It recommends building human-reviewed datasets, comparing JEV with existing judges, inspecting disagreements and testing whether confidence predicts correctness before using it to route cases.

All sensible.

But there’s a recursive security problem hiding underneath.

The agent being evaluated produces the trace.
        |
        v
The trace becomes input to the judge.
        |
        v
The judge decides whether the agent behaved correctly.

So what happens when the agent knows how the judge works?

Or more interestingly:

what happens when the agent deliberately produces a trace designed to receive a favourable evaluation?

We already know this problem from metrics engineering.

Goodhart’s Law is usually summarised as:

When a measure becomes a target, it ceases to be a good measure.

Now imagine that the thing optimising against the measure is an autonomous agent. That gets interesting very quickly.

The question stops being:

Can JEV judge agents accurately?

And becomes:

Can an agent manipulate the evidence presented to JEV so that JEV judges incorrect behaviour as correct?

That’s a much nastier problem.

Then I saw the SOC use case

And this is where the architecture started looking very familiar.

Vega recently published Jev Closed Up to a Third of Our Triage Agent’s Alerts, 230x Faster and 2,000x Cheaper.

Their proof of concept puts JEV in front of a more expensive triage agent.

JEV sees the alert, the detection, the rows that triggered it, the playbook and optionally tenant context.

If JEV returns “not escalated” above a confidence threshold, that alert can be closed without the full triage agent investigating it.

In their replay across two tenants, JEV closed 15% of alerts on the busier tenant and 33% on the quieter one at a 0.8 threshold.

The judged correctness of those closures was reported at around 98-99%.

But there were confirmed escalations inside that closed set: one on the busy tenant and two on the quiet one.

At a 0.9 threshold, no confirmed escalation slipped through in that particular replay, although the coverage reduced.

Vega are very clear that this was a proof of concept, not live production traffic, and that the threshold would need recalibration. They also explicitly note that JEV could not replace the full agent’s verdict and works better as a one-sided gate.

That’s responsible framing, but the security question is still unavoidable.

The pipeline now looks like this:

Security telemetry
        |
        v
       JEV
        |
        v
"Not escalated"
        |
        v
 Alert closed

The model is no longer merely providing information. Its output is consequential and that changes everything.

I’ve seen this movie before

This pattern felt uncomfortably familiar because I reported something architecturally similar a few weeks ago in Microsoft Security Copilot.

My finding involved indirect prompt injection through attacker-controlled security telemetry.

In the proof of concept, an attacker-controlled HTTP User-Agent influenced Copilot’s recommendation.

The important part wasn’t simply:

the AI can be manipulated

But instead:

what happens if something trusts the result?

Security Copilot could produce a recommendation based on attacker-controlled evidence. A downstream automation layer could then act on that recommendation.

The pattern was:

Attacker-controlled telemetry
        |
        v
AI interpretation
        |
        v
Structured recommendation
        |
        v
Trusted automation
        |
        v
Security action

Now look at the JEV triage design:

Attacker-influenced telemetry
        |
        v
JEV classification
        |
        v
"Not escalated"
        |
        v
Alert closed

Different model, different mechanism, same architectural smell.

The moment a probabilistic interpretation becomes the thing that determines whether investigation continues, the model has moved from being an analytical aid to being part of the control plane and that’s a very different security proposition.

The dangerous assumption

AI makes mistakes, we know that.

The dangerous assumption is:

sufficiently high confidence turns probabilistic judgment into authority

Dunning-Kruger disagrees…

Christophe’s work is particularly uncomfortable here because the problem he demonstrates is not merely inconsistency. The model can be predictably influenced by presentation order while remaining confident.

Now combine that with a SOC pipeline that says:

if probability_not_escalated >= 0.8:
    close_alert()

And the security research question becomes extremely obvious:

Can an attacker construct malicious activity whose observable telemetry is shaped specifically to push the decision above the auto-close threshold?

Because if yes, That’s no longer merely a model accuracy problem.

You have built an AI-powered alert suppression factory.

The missing experiment

If I were evaluating AI-assisted security triage today, I would want two very different sets of measurements.

The first is the familiar one:

P(correct | normal input)

Precision, recall, calibration, false positives, false negatives. All useful.

But I also want:

P(correct | adversarially constructed input)

And, for automated triage specifically:

P(auto-close | malicious, adversarial input)

That’s the number I care about. Take the same malicious behaviour and systematically mutate only attacker-controlled or semantically neutral parts of the evidence.

Reorder fields, reorder sentences, change filenames, change URLs, modify HTTP headers, alter User-Agent strings, insert benign-looking context, duplicate evidence, introduce contradictory context, swap label order where the interface allows it.

Preserve the actual malicious behaviour, then measure how far the decision moves.

Don’t just test whether the label changes. Measure the probability shift. Measure the confidence shift. Measure whether the threshold gets crossed.

And because Christophe’s work shows the effect can be predictable, test whether an attacker can search the decision surface themselves and identify favourable representations.

That’s the part that turns a model quirk into an exploit primitive.

High confidence is not authorization

This is the distinction I think we need to get much better at making.

AI output can be extremely useful security evidence.

It can:

  • prioritise
  • classify
  • correlate
  • summarise
  • recommend
  • estimate confidence
  • reduce analyst workload

All of that has value.

But somewhere along the way, we have started quietly turning evidence into authority.

A model says an alert is probably benign. Therefore we close it.

A judge says an agent behaved correctly. Therefore the run passes.

A classifier says an indicator is low-risk. Therefore investigation stops.

That transition deserves considerably more scrutiny than it currently receives.

Because:

high confidence is not authorization.

And:

structured output is not a security boundary.

A machine-readable answer is still just an answer. Typing it as a boolean does not make it true. Attaching a probability does not make it safe to automate. Making the call 2,000 times cheaper does not make the decision 2,000 times more trustworthy.

If anything, lower cost increases the blast radius of being wrong.

Evidence is not authority

This keeps dragging me back to the same design principle.

AI systems should produce evidence. Security controls should enforce policy.

Those are not the same thing.

If a system says “this appears benign”, that can be valuable.

If a system says “this appears benign with 93% confidence”, that can also be valuable.

But the moment something downstream interprets that as:

therefore no investigation is required

we have crossed an architectural boundary.

And that boundary needs its own threat model.

Who controls the inputs? Which parts of the state are attacker-influenced? Can the representation be manipulated? Can the decision be probed? Can confidence be steered? Can an attacker observe enough outcomes to optimise against the model? Can the system abstain? Is there a deterministic control outside the model? What happens when the model is confidently wrong?

Those are security questions, not model-quality questions.

This actually makes me more interested in JEV

None of this means JEV is useless. Quite the opposite. I still intend to test it.

In fact, Christophe’s research probably makes me more interested in testing it than I was before. Just not initially for the reason I expected.

I had been thinking about integrating JEV into cARL as another decision engine.

Now I think the more interesting experiment is to treat JEV itself as an untrusted component.

Put it in the firing range. Attack the assumptions around it. Test order invariance. Test representation invariance. Test adversarial telemetry. Test label-position effects. Test confidence calibration under attack. Test whether an agent can manipulate its own evaluation trace.

Test whether malicious activity can be made to cross a benign classification threshold without materially changing the activity itself.

That feels far more useful than simply proving that I can wire another model into a CLI.

Because AI security systems are moving very quickly from:

Here is something you should investigate.

to:

I investigated this for you.

and increasingly to:

I investigated this and already took action.

That progression changes the threat model.

The faster the system becomes, the more important that distinction gets.

Speed is useful. Accuracy is useful. Automation is useful.

But none of them turn probabilistic judgment into a trustworthy security boundary.

The model is not the control plane.

And if we’re going to make it one anyway, we had better start attacking it like one.

Sources

comments powered by Disqus