All AI news
AI-Native EngineeringAug 11, 2026 · 2 min read

Two releases in a week aimed at agent token costs, which tells you where the pain is

Gemini 3.6 Flash and NVIDIA's Nemotron 3.5 Lightning both landed with always-on agents and their running costs as the headline.

Source: Oracle

  • AI economics
  • AI agents
  • Inference cost
  • Model selection
$/token

Within a week, Gemini 3.6 Flash arrived with enterprise agent token costs as its stated focus, and NVIDIA's Nemotron 3.5 Lightning became available on Oracle's cloud pitched at always-on agents. When two releases target the same problem in the same week, the problem is the story.

The constraint moved

For two years the question was whether a model could do the task. That question is mostly settled for the work most companies want automated. The live question is what it costs to run continuously, because an agent that watches a queue all day makes calls a chatbot never did.

What I'd do about it

Instrument cost per outcome before you scale anything, not after. The number that matters is not cost per token or per call, it is cost per invoice matched, per ticket resolved, per order routed. Teams that skip this discover the economics on an invoice, usually one quarter after the pilot everyone was pleased with.

It is also worth designing for model substitution from the start. Two releases in a week is the tempo now, and a system welded to one provider cannot take advantage of it.

More from AI in the news
AI Strategy · Aug 16, 2026

The watermark is Anthropic's half. The labeling duty is yours.

AI Strategy · Aug 16, 2026

The EU writes the labeling rules in September, and deployers should be in the room

Operating Philosophy · Aug 16, 2026

Loud churn and measured churn are rarely the same number