Webneuron
Artificial Intelligence

What it costs to keep an AI feature running after launch

Inference is the line item everyone models. Drift, review, evaluation and incident load are the ones that decide whether the feature survives its second year.

December 2, 20256 min readBy Webneuron Engineering Team

Business cases for AI features are usually built on a single number: cost per call. It is the easiest figure to obtain and the least representative of what the feature will actually cost you. The bill that arrives in year two is composed almost entirely of things the business case had no row for.

The costs that show up later

Drift is the obvious one and the most misunderstood. Models do not decay on a schedule; they decay when the world moves — a new product line, a redesigned form, a customer segment that behaves differently from the one you trained on. Detection requires monitoring somebody had to build. Correction requires retraining somebody has to fund.

Then there is human review. Almost every AI feature touching a consequential decision keeps a person in the loop, and that person is a permanent operating cost. If the business case assumed review would taper to zero, it assumed the model would become perfect. It will not, and the tapering rarely survives contact with the first serious error.

Evaluation is the cost that catches engineering teams off guard. Every prompt change, model upgrade and vendor deprecation requires re-establishing that quality has not regressed. Without an automated harness that means a person spot-checking outputs and forming an impression, and impressions do not catch a two percent regression — they catch it three months later, in a complaint.

A more honest cost model

  • Inference and retraining compute, projected at realistic volume rather than pilot volume.
  • Human review time, held flat unless you have evidence rather than hope that it will fall.
  • Evaluation and regression testing, including the engineering time to maintain the harness itself.
  • Monitoring, alerting, and the on-call load the feature adds to a rotation that already exists.
  • Vendor change absorption. Model deprecations and pricing revisions arrive on the vendor calendar, not yours.
  • The exit cost of switching providers, which is mostly prompt and evaluation rework rather than integration work.

Why this is worth doing before you ship

None of this argues against building. It argues for building fewer things, better justified. A feature that clears a realistic cost bar will survive its first budget review. One justified on inference pricing alone tends to be switched off quietly in a cost-cutting cycle, taking the team’s credibility with it and making the next proposal harder to fund.

The organisations getting durable value from AI are not the ones running the most experiments. They are the ones that priced the second year before they shipped the first.

Let's build the system your business will run on next.

Tell us where it hurts. We'll bring the architects, engineers, and delivery model to fix it — and scale it.