Writing
The quiet shift to right-sized models in security pipelines
published: 2026-10-01 · status: canonical · expanded from the original post
Over the next year, I expect more security teams to reconsider defaulting to frontier LLMs for high-volume classification work. The economics are starting to argue against it. The reflex for the past couple of years has been to throw the biggest, most capable model at any problem that looks like language understanding. That made sense when the only viable option was a general-purpose model with a convenient API. But classification tasks—especially in security—are often narrow, repeatable, and defined by well-understood failure modes. When you are processing millions of events, the difference between a few cents per call and a few tenths of a cent per call stops being a rounding error.
Vincenzo Iozzo's test of Jev on identity resolution is a concrete signal. He ran a small calibrated model against five real organizations' data and got a 0.948 F1 at $0.62 per 1,000 accounts. That beat Claude Haiku 4.5 and Sonnet 5 on accuracy while costing roughly a tenth to a fortieth of their price, with far lower latency. Iozzo documented the setup and results in his article "Test-driving Jev on a security task: identity resolution." The numbers are specific enough to be useful: this is not a synthetic benchmark, and the task is one that security teams actually face.
That's not a marginal gain; it's a different operating point for bounded, repeatable tasks.
Identity resolution is a bounded task: given a set of records, determine which ones refer to the same entity. It requires some judgment about fuzzy matches, but the input space is constrained and the correct output is often checkable against ground truth. The same pattern applies to adjacent problems like fraud scoring and entity matching. If a small model that has been calibrated on your own data can outperform a general-purpose frontier model at a fraction of the cost, the default "just prompt GPT-5" reflex starts to look expensive and imprecise by comparison. That does not mean frontier models are useless. They still matter for open-ended investigation, summarization, and tasks where the input distribution is unpredictable. But for high-volume, structured classification, they are often overkill.
The open question is whether security teams can operationalize these specialized models without building an ML team. Running a small calibrated model is not the same as calling an API. You need to collect labeled data, fine-tune or otherwise adapt the model, evaluate it against your own distribution, and maintain it as that distribution drifts. Some of this is getting easier with tooling, but it is not yet a turnkey process for most security engineers. The signal to watch over the next year is whether teams without significant in-house ML expertise can adopt this approach, or whether the tooling still needs to mature first. If the tooling catches up, the cost advantage will be too large to ignore.
My guess is that we will see a quiet shift toward right-sized models in production security pipelines long before the market consensus catches up. Security teams rarely announce cost optimizations; they just change what they run. The market narrative will keep focusing on bigger models and new capabilities, but the teams actually processing millions of events will be paying attention to the unit economics. The shift will not be dramatic, but it will be real.
Originally covered at vincenzoiozzo.com ↗