When it comes to AI, the sticker price is often a red herring. We’ve been conditioned to think cheaper means better value, but in the world of AI services, that logic can lead you astray. Personally, I think this is one of the most misunderstood aspects of AI adoption today. The real cost isn’t just about the price per token—it’s about how effectively the model completes the task. What makes this particularly fascinating is that a model with higher token costs might actually be the more economical choice in the long run.
Take Databricks’ recent internal benchmark, for example. They tested various AI models on real engineering tasks, and the results were eye-opening. Open-weight models like Z.ai’s GLM 5.2 held their own against frontier models like Anthropic’s Opus 4.8. What’s more, GLM 5.2 cost just $1.28 per task compared to Opus’s $1.94, despite Opus being the more expensive model per token. This raises a deeper question: Are we focusing on the wrong metrics when evaluating AI costs?
From my perspective, the per-token pricing model is a relic of early AI economics. It’s like judging a car’s value by the price of its individual parts instead of its overall performance. What many people don’t realize is that cheaper tokens often come with hidden costs—like higher token consumption or lower task completion rates. Databricks’ findings underscore this: Anthropic’s Sonnet 5, while 1.7x cheaper per token than Opus 4.8, ended up costing more per task because it was less efficient. If you take a step back and think about it, this isn’t just about pricing—it’s about productivity.
But here’s where it gets even more interesting: the tooling around these models matters just as much as the models themselves. Databricks’ Matei Zaharia pointed out that the ‘harness’—the software that interfaces with the model—can dramatically impact cost and performance. A detail that I find especially interesting is how a simple harness like Pi’s achieved the same success rate as more complex ones from LLM vendors, but at half the cost. This suggests that the real innovation in AI might not be in the models themselves, but in how we deploy and optimize them.
What this really suggests is that we’re still in the early days of understanding AI economics. Academics have already noted that in about a third of model comparisons, the cheaper model ends up costing more. For instance, Gemini 3 Flash is 80% cheaper than GPT-5.4 per token, but its actual cost across tasks is 38% higher. This isn’t just a quirk—it’s a pattern that challenges our assumptions about value in AI.
In my opinion, the future of AI pricing will shift from per-token to per-task metrics. Tools like Databricks’ Omnigent, which acts as a wrapper for multiple coding agents, are a step in this direction. They allow businesses to focus on outcomes rather than inputs, which is how AI should be evaluated in the first place.
If you’re adopting AI, here’s my takeaway: Don’t be seduced by low token prices. Instead, focus on task completion rates, efficiency, and the tooling ecosystem. Cheap can indeed be expensive, and in the world of AI, value is measured in results, not tokens.