For Engineering leadership

AI for engineering leaders

You are being asked whether AI spend is producing anything, and the available evidence is an invoice and anecdotes. Both are consistent with any conclusion, which is why the conversation keeps repeating.

What you are actually accountable for

You own two claims that currently rest on the same thin evidence: that the spend is justified, and that the tools are not creating risk. The uncomfortable part is that the obvious measurements are traps — per-developer token spend produces gaming, lines-of-code metrics reward the wrong output, and satisfaction surveys measure enthusiasm rather than throughput. Getting this right is mostly about picking measurements that survive contact with the people being measured.

What is landing on you right now

  • A budget line growing faster than headcount

    AI spend rises without a corresponding hiring decision, which makes it look anomalous on a finance review even when it is well spent.

  • A productivity claim you cannot substantiate

    You believe the tools help and you have anecdotes. Anecdotes lose to a spreadsheet in a budget conversation, and the counter-anecdotes are equally available.

  • Security asking for controls that cost velocity

    Some of the asks are genuinely important and some are ticket queues that will slow every engineer daily. Distinguishing them is your job and requires understanding the mechanism, not the framing.

  • Quality effects that are real and hard to source

    If an unreviewed artefact tells agents to skip a check, the result is a class of defect with no obvious origin, appearing months later. Nothing in your existing metrics attributes it.

Questions you are being asked and cannot answer

the gap

  • What is our cost per merged change, by team, and is it moving in the right direction?
  • How many purchased AI seats were used at all in the last 30 days?
  • Did shipping a new version of a shared workflow make it cheaper or three times more expensive?
  • Which of security's proposed controls will actually slow my engineers, and by how much?
  • Is the variation in AI spend between teams explained by the work or by the tooling they assembled?

The analysis, from where you are standing

The measurement problem starts with a pricing detail that gets flattened in reporting. Per-seat licences and per-token usage attribute in opposite ways: seats are a fixed cost where the only real question is activation, and tokens are a variable cost with enormous individual variance. Combined into one figure, both levers disappear — and the seat side is usually where the fastest saving sits, because unused licences require no behaviour change from anyone to reclaim.

The metric worth building toward is cost per unit of delivered work, by team. Merged changes or closed tickets are both flawed denominators and both are dramatically better than none, because they convert 'this team spends more' into 'this team spends more per merged change' — which is either fine or a genuine finding. Present the denominator's weaknesses openly; doing so is what stops the number being weaponised, and it is what makes it survive a second meeting.

There is one measurement to refuse outright. Publishing per-developer token spend produces avoidance and gaming rather than efficiency: people use the tool less, or split work to look cheaper, and you lose the productivity along with the telemetry. Attribute at team and project level, keep individual data for capacity planning only, and say so before anyone asks — the announcement is what makes it credible.

On the security asks, the mechanism tells you the velocity cost, so it is worth learning. Credential scoping and egress restriction are configuration changes with essentially no daily friction and large risk reduction — agree to them quickly. A per-request approval queue for adding a connector is a permanent tax on daily work and will be routed around, which loses you both the velocity and the visibility. Push for a reviewed catalogue engineers self-serve from instead; it satisfies the same control objective without putting a human in the hot path.

What to do — the end state

  • Separate seat spend from token spend permanently

    Report them apart at every level. They have opposite levers — activation versus consumption — and combining them is the main reason AI spend reviews produce no decisions.

  • Pull seat activation before optimising anything

    Licences purchased, assigned, and used in the last 30 days. The gap is routinely larger than the requesting team expects, and reclaiming it needs no behaviour change.

  • Fix attribution identity before building dashboards

    One credential per team or project, CI keys named for their pipeline. Attribution cannot be recovered from an invoice afterwards, and a dashboard on shared keys shows confident nonsense.

  • Pick one outcome denominator you already measure

    Merged changes or closed tickets. Rough is fine; none is not. State its flaws when you present it so the number is used rather than fought over.

  • Refuse per-individual spend reporting, publicly

    Announce the boundary before anyone asks for it. Discovering the boundary later reads as concealment; stating it early is what makes team-level numbers trusted.

  • Triage security asks by daily friction

    Credential scoping and egress restriction: agree fast, near-zero friction. Per-request approval queues: push back and offer a self-serve reviewed catalogue instead.

  • Standardise the artefacts your teams share

    Rules files and connectors are already being copied between engineers. One reviewed, pinned set is both a quality control and the thing that makes cross-team comparison meaningful.

  • Insist on rollback for shared artefacts

    Version pinning is a reliability feature first. If a shared workflow change makes things worse, you want to revert it in minutes rather than debug it across teams for a week.

  • Ask for artefact-version cost attribution

    Cost per session joined to the artefact versions active in it. Few setups can do this; it is the only way to know whether a workflow change helped or quietly tripled cost.

  • Watch for defects with no obvious origin

    A shared artefact that weakens a check produces a defect class months later with nothing attributing it. Treat an unexplained cluster as a possible artefact question, not only a training one.

What to do first — the order

  • Weeks 1–2 · Split and count

    Separate seat from token spend and pull activation per tool. This alone usually produces a reclaimable-cost number that funds the rest of the work.

  • Weeks 2–5 · Fix identity

    One credential per attribution unit, CI keys named. Unglamorous plumbing, and every number after this depends on it being right.

  • Weeks 5–7 · Add a denominator

    Cost per merged change by team, with its weaknesses stated. Announce the individual-reporting boundary at the same time, not later.

  • Ongoing · Standardise and attribute

    Move shared artefacts to one reviewed, pinned set, then push for session-level cost attribution against artefact versions.

What to report upward, and the trap in each

  • Seat activation rate

    Used in 30 days over purchased. The fastest saving available and the least contentious, because reclaiming an unused licence affects nobody's work.

  • Cost per merged change, by team

    Your primary number. Present the denominator's flaws with it — that candour is what keeps it a decision input rather than an argument.

  • Month-on-month token spend change by team

    Direction matters more than level, since teams do different work. A step change usually means a workflow or artefact changed, which is a question rather than a verdict.

  • Fleet convergence on the reviewed artefact set

    Low convergence means cross-team comparisons are measuring tooling differences rather than team differences, which invalidates the other metrics.

  • Cost delta across versions of a shared workflow

    The most decision-useful and least commonly available metric. It is the difference between guessing at efficiency and knowing.

Where we fit, in your order

  • Vincosha Ledger

    Attributes every token and session to a user, project and the artefact versions that were active — which is the join that produces cost per merged change and cost per workflow version.

  • Vincosha Registry

    Puts shared rules and connectors on one pinned, reviewed source, so cross-team comparison is meaningful and a bad change can be rolled back in minutes.

  • Vincosha Assay

    Catches artefacts that weaken checks before they spread, which is the defect class your existing quality metrics cannot attribute.

Frequently asked

What is the fastest way to reduce AI spend?
Unused seats, almost always, and it costs nobody any change in behaviour. Pull licences assigned against 30-day usage per tool before touching consumption — the gap is usually larger than expected and it is the one saving with no downside to argue about.
How do I substantiate the productivity claim?
Cost per merged change or per closed ticket, by team, tracked over time. Neither denominator is good and both beat anecdotes decisively. Present the flaws yourself; a number offered with its limitations is far harder to dismiss than one presented as clean.
Should I resist security's AI controls?
Triage them by daily friction rather than by principle. Credential scoping and egress restriction are near-free and genuinely important — agree quickly. A per-request approval queue for connectors is a permanent tax that will be bypassed; counter with a self-serve reviewed catalogue, which meets the same objective.
Why does artefact standardisation matter to me rather than to security?
Two reasons that are yours, not theirs. It makes cross-team cost comparison meaningful, because otherwise you are measuring tooling differences. And it gives you rollback: when a shared workflow change makes things worse, reverting in minutes beats debugging across four teams for a week.

Sources

Written 2026-07-25. Nothing on this page asserts a statistic we did not gather ourselves, and where a regulatory or standards obligation is mentioned it is described in prose and left to your counsel to confirm against the current text rather than paraphrased as fact.

Start with these procedures

The same problem, from another desk

Concepts on this page

The list all of this depends on

Every recommendation above is downstream of knowing what your AI surfaces actually load. Vincosha Registry makes that one signed, versioned source; Vincosha Assay vets what enters it; Vincosha Ledger tells you which versions were active, on which machine, at what cost.