Current assessments
The token economics of many AI agents do not add up
What is happening. Agentic systems consume many times as many tokens as a chat assistant on a single task: they plan, check and retry. In many applications, the inference bill grows faster than the benefit.
What it means for Germany. Falling prices per token do not solve the problem if consumption per task rises faster. For well-defined routine tasks, smaller locally run models become attractive: predictable costs, data that stays in-house, and less dependence on any single vendor.
What to do now. For every use case, measure the cost per completed task, not the price per token. And test early which tasks a small local model handles just as well.