Useful work per dollar — what delivery and analytics should change first
If agents become a unit of work, stop counting runs and start defining useful outcomes — the same discipline that makes a Monday dashboard trustworthy.
The question
If agentic systems are becoming the unit of work, what should a delivery or analytics lead change first — the model, or the definition of useful?
The default pitch says buy agents and measure how often they run. I have already watched that movie in analytics: more extracts, more tiles, more meetings that reconcile numbers instead of deciding what to do. Agents will replay it faster, because they can generate plausible work without improving the call.
What the signal is really saying
How to manage AI investments in the agentic era (OpenAI Blog) is easy to skim as another enterprise AI post. The useful idea underneath is sharper: measure useful work per dollar (or per hour of scarce human attention), not activity.
That lands because the scarce resource is not capability anymore. Judgment, ownership, and rollback are.
What people get wrong
"Add agents" is being sold the way "add a dashboard" used to be sold — as if the tool is the operating model. Teams then track tokens, runs, tickets closed, and demo velocity. None of those answers the only question that matters in delivery or portfolio work: did a decision get better?
Counter-case: raw activity metrics are not useless. They catch cost blow-ups and stuck loops. Treat them as health checks, not proof that the investment worked.
Change the definition of done first
Before you change the model, change what "done" means for an agentic loop:
- What decision or artifact improved? A scoped analysis, a release risk call, a refreshed metric with an owner — not "the agent ran."
- Who owns the failure when it is wrong? If the answer is "the platform," you do not have an operating model.
- What would falsify the hype next week? An eval slice, a golden set, a Monday exception list — something that can fail a release.
- How do you disable the loop without a war room?
Without those, you have expensive autocomplete with a project plan.
This is the same instinct as a credit portfolio review. I do not want thirty charts. I want a short decision set: is emerging risk accelerating, where is the engine drifting from the model, what changed that we did not expect. Everything else links out.
The analytics parallel (and why it is not a metaphor)
When two teams disagree on "delinquency," a dashboard becomes a debate club. When two teams disagree on what an agent was allowed to decide, an automated workflow becomes the same debate with a worse audit trail.
So the first artifacts I want before the exciting demo are boring on purpose:
- A one-page scope: what the agent may draft vs what a human must sign
- An eval slice: ten real cases where wrong is costly
- A kill switch: how we stop the loop cleanly
Delivery leads should refuse agent workflows that cannot name a rollback. Analytics leads should refuse agent outputs that cannot name a definition.
How this shows up in what I am building
On Orbit the same pressure shows up in miniature. Related articles refresh daily. A draft can be generated from that cache. The useful unit is not "a draft existed." It is a publish decision with an owner: Approve in Slack, commit the Write, redeploy — or Skip, and leave the Blog alone.
I am also writing the same rule into portfolio analytics: Landing → Pulse → Drivers → Queue only earns its slot when the queue names who to contact next. An agent that fills tiles without that path is just a faster empty report.
Mistakes I refuse to repeat
- Optimizing for demo wow before the metric has a definition doc
- Letting "useful work per dollar" become another vanity KPI with no owner
- Automating the meeting before the exception list is stable
- Shipping a loop I cannot disable cleanly
What I am not doing yet
Chasing every agent framework. Replacing human review on anything that touches risk language or production change. Treating a vendor narrative as a roadmap.
Takeaway
If agents are becoming a unit of work, do not start with the agent. Start with the unit of useful work: the decision, the owner, the eval, and the rollback. That is how you keep Monday honest — whether the artifact is a Power BI queue or an automated draft waiting for Approve.