Most AI projects start with a stack. Useful ones start with a number.
Someone picks a model, connects a few tools, and shows a fluent draft in a meeting. The room nods. Two weeks later the same team cannot say whether the system returned an hour a day, cleared a queue, or simply created a new place to babysit unfinished work. The model did not fail. Nobody defined what “worked” meant.
Version one needs a success number. Not a vision deck. Not a platform roadmap. A number you will check on a calendar date you already wrote down.
Why stacks come first (and fail second)
Buying tools feels like progress. There is a catalog, a pricing page, and a demo that looks like the job. Measuring the job requires an uncomfortable conversation: which repetitive work actually matters, who owns the outcome, and what you will stop doing if the system ships.
So teams reverse the order. They assemble capability, then hunt for a use case large enough to justify the spend. That hunt rarely ends with a tight first release. It ends with a pilot that touches too many systems and cannot prove anything in two weeks.
What a useful success number looks like
A good metric is narrow, already painful, and checkable without a new analytics project.
Examples that work: hours per week returned on one named workflow; tickets cleared to a definition of done; cycle time from request to first reviewable draft on real cases.
Examples that do not: “adoption,” “AI maturity,” “number of agents deployed,” or any metric that only improves if you redefine the job midstream.
If two people on the team would argue about whether the number moved, pick a different number.
Scope the job before you scope the software
Write four sentences before you open a vendor page:
- The job. One repetitive workflow that already costs time or money every week.
- Done. What a finished instance looks like in language a skeptic accepts.
- The number. The metric and the date you will check it (two weeks is enough for a first signal).
- The boundary. What version one will not touch, other teams, payments, production writes without review, the second workflow someone wants to add “while we are in there.”
Those four sentences are the brief. Models and tools are implementation details that serve them. If the brief will not fit on one page, it is not version one.
Keep version two from eating version one
The failure mode is not “we chose the wrong model.” It is “we kept adding scope until the success number became unverifiable.”
Park new ideas in a dated list. Do not stretch the current build to absorb them. When the two-week check arrives, either the number moved or it did not. If it moved, you earned the right to widen. If it did not, fix the handoff or the definition of done before you buy more capability.
The same rule applies whether you are commissioning a done-for-you agent, subscribing to a productized automation, or building custom software because an automation is not enough. The constraint chooses the path. The success number chooses whether you shipped.
How we use this at Viking Labs
We start client work the same way we run our own products: one bounded workflow, a clear first release, and a result we can measure. ResearcherFlow is the production case, live billing, permissions, APIs, and a research agent that drafts but does not write into the workspace until a human approves. That discipline is the point of running software ourselves. Advice that only works in a slide deck is not useful.
A two-week test
Pick the painful job. Write the four sentences. Build the smallest path that can produce a reviewable result against that definition of done. Run real cases, not polished demos. On the check date, read the number out loud.
If you cannot name the number today, you are not behind on AI. You are early on scoping. That is the cheaper place to be stuck.
Bottom line
The impressive part of an AI system is rarely the release notes. It is a finished job with an owner, a boundary, and a success number somebody actually checks.
That is what we build. If you have one workflow ready for that treatment, map the first build: tell us the workflow
Jon Marrs · About Viking Labs
Leave a Reply