Watching the Race in Real Time
Something has been bothering me about how people talk about AI progress.
The conversation usually goes in one direction: benchmarks. A new model ships, someone posts a chart showing it scored 87.4 on some evaluation suite, and the feeds light up with either celebration or scepticism. What almost nobody does is ask the follow-up question: at what cost, and can it actually do anything useful on its own?
That gap is what I built Stack Futures to close.
The Three Numbers That Actually Matter
When you’re evaluating a frontier AI model, there are really three things you need to know.
The first is raw intelligence. How well does it reason? How does it perform on hard problems against every other model in the world? This is what Chatbot Arena captures: real humans, real conversations, a live ELO ladder that moves when models improve or fall behind. It’s the closest thing we have to an honest signal.
The second is agentic capability. Can the model actually do things without hand-holding? Can it use tools, write code that runs, navigate a codebase, complete a multi-step task end-to-end? This is where the gap between impressive demo and production reality lives. A model can be very good at answering questions and completely useless as an autonomous agent. The industry has finally started measuring this properly (SWE-bench, Tau2-bench) but it rarely gets reported alongside the headline benchmarks.
The third is cost. A model that costs $180 per million output tokens is a different product to one that costs $3. Capability matters. But capability per dollar is what most people are actually buying.
Stack Futures combines all three into a single composite score updated every six hours. Not to flatten nuance, but to give you something you can actually track over time.
Why Track It at All?
Because the rate of change is the story.
We tend to notice individual moments: a new model ships, a benchmark record falls, a company announces a funding round. What’s harder to see is the slope. How much better is the frontier than it was six months ago? How much cheaper? How much more capable of doing real work autonomously?
Those three curves, combined, are the most important economic signal of our time. They determine when AI stops being a productivity tool and starts being something that replaces categories of work entirely. They determine when the cost of intelligence drops below the cost of human judgement for a given task class. They determine which companies are accelerating and which are stalling.
Most people don’t have a clear view of that slope. They see the headline events but not the accumulation.
The Acceleration Is Real
When you look at the data running back, a few things become clear.
The intelligence ceiling has moved fast. Models that would have been state-of-the-art eighteen months ago are now mid-table. The frontier has pulled ahead so quickly that last year’s best is today’s commodity option.
Agentic capability has moved even faster. Twelve months ago, an AI completing a real software engineering task end-to-end, without a human in the loop, was a research demo. Today it’s a product feature. The SWE-bench numbers tell this story clearly: what scored 30% eighteen months ago now scores 90%+.
Cost has moved fastest of all. GPT-4 cost $30 per million input tokens when it launched in 2023. Comparable-quality models today are under $1. That’s a 97% drop in three years, and the race is still running.
Put those three curves together and what you’re watching is something genuinely unprecedented: the price of intelligence collapsing, while the capability of that intelligence to act independently accelerates. That’s not a product cycle. That’s a structural shift.
What Stack Futures Is For
I wanted somewhere to watch this happen clearly.
Not just a model leaderboard. Not just a pricing database. Something that holds all three signals together and shows you how they move relative to each other over time. When a new model ships, what’s actually changed? Is it smarter? Can it do more? Does it cost less? All three, or just one?
That’s the question Stack Futures is built to answer.
The composite index, the SFX-10, tracks the top ten models by that combined score. It moves when models improve, when pricing shifts, when new entrants appear and old ones fall behind. It’s designed to be a clean signal in a noisy market.
We’re in the middle of something that moves faster than it looks. Stack Futures is my attempt to watch it properly.