What is inside
- The difference between plausible and usable, defined precisely
- Eight pass or fail checks you can apply in under three minutes
- A scoring method that does not require agreement on taste
- A worked before and after on the same brief
Everyone reviewing agent output eventually notices the same thing. The bad outputs are not bad in the way bad writing is bad. They are fluent, well-organised, on-topic, and somehow impossible to send. This standard is an attempt to say exactly what is missing, in checks that two people will score the same way.
Plausible versus usable
A plausible deliverable is one that would survive being skim-read by someone who does not know the subject. A usable one survives being acted on by someone who does.
The gap between them is almost always specificity that only exists inside your business. A plausible ICP describes a company type. A usable one excludes companies you would otherwise chase. A plausible battle card lists your strengths. A usable one tells a rep what to concede in the first two minutes.
The checks below all test for the same thing from different angles: whether the deliverable contains decisions or merely contains information.
The eight checks
Score each one pass or fail. No half marks, because half marks are how a document with four real failures ends up scoring 70 percent.
1. Would this be wrong for a competitor?
Take the deliverable and mentally swap in a competitor's name. If it still reads as true and sensible, it contains no decisions about you. This is the single most diagnostic check and the one most outputs fail.
2. Does it exclude something?
A usable strategy document says what not to do. A usable ICP says who to walk away from. A usable content plan names a topic it is deliberately skipping. If nothing is excluded, no choice was made.
3. Is every number attributable?
Every figure must either carry its source in the same sentence, or be visibly framed as an assumption to be filled in. A number with no provenance is a liability, not evidence, and it is the most common way agent output causes real damage.
4. Could the intended reader act on it today?
Name the person who receives this. Can they take one concrete action from it before lunch, without asking a clarifying question? If the answer requires a meeting first, the deliverable is a meeting agenda pretending to be a deliverable.
5. Does it handle the strongest objection?
Not a straw objection. The one that actually comes up. If the deliverable never acknowledges the best argument against it, it has not engaged with the problem.
6. Is the structure load-bearing?
Headings should carry the argument. If you read only the headings, do you get the shape of the thinking? Structure that exists to look organised, rather than to organise, is a reliable signal that the content underneath is thin.
7. Is there anything a subject expert would object to?
Find the sentence that someone who does this for a living would push back on. If there isn't one, the deliverable is at the level of consensus, which is the level at which nothing useful gets said.
8. Would you put your name on it?
The final check, and not a soft one. Not "is this fine", but "would I send this with my name at the top to someone whose opinion of me matters". Most people can answer this instantly and correctly about work they were about to approve.
Scoring
- 8 of 8: send it.
- 6 or 7: one revision round, with the failed checks named as the notes. Do not rewrite it yourself; naming the failure is faster and it improves the next brief.
- 4 or 5: the brief was wrong, not the output. Rewriting the deliverable will produce a different deliverable with the same problem.
- 3 or below: the work was in the wrong mode. Something judgement-shaped was delegated as though it were craft.
The last two lines are the point of the whole framework. Most teams respond to weak output by editing it, which fixes one artefact and teaches nobody anything. Scoring tells you whether to fix the output, the brief, or the decision about who should be doing the work at all.
Worked example
The brief: "Write a positioning statement for our scheduling product."
The plausible version. "For fast-growing teams who need to coordinate complex schedules, our platform delivers intelligent, automated scheduling that saves time and reduces friction, so teams can focus on the work that matters."
Run the checks. Check one fails immediately: this is true of every scheduling product ever built. Check two fails, nothing is excluded. Check five fails, no objection is acknowledged. Check seven fails, nobody would argue with a word of it. Score: 4 of 8, and the diagnosis is that the brief supplied a topic, not a decision.
The usable version, after the brief was rewritten to include who to exclude, the rejected alternative position, and the objection that comes up most on calls: "For operations leads at care homes who are judged on agency spend, this replaces the wall planner with a rota built from who is genuinely available and genuinely qualified, so the gaps that get filled at agency rates are visible a fortnight before they open. Not for single-site homes, where the spreadsheet is genuinely cheaper. The registered manager approves every suggestion; nothing publishes itself."
Check one: a competitor could not say this. Check two: single-site homes are excluded. Check five: the control objection is answered in the last sentence. Check seven: "the spreadsheet is genuinely cheaper" is a sentence a salesperson would argue with, which is why it is worth keeping.
Same product, same model, same length. The difference is entirely in what the brief contained.
Using this on a team
- Put the eight checks where review happens, not in a wiki nobody opens.
- Score before writing any notes. The score tells you which notes are worth writing.
- Track the distribution of scores, not the average. A cluster at 4 to 5 means your briefs are the problem and no amount of reviewing will fix it.
- When a deliverable scores 8, keep the brief. Briefs that produce sendable work are the most reusable asset an agent programme creates.
This is the model Orbitable is built around
Shared context every agent works from, a review lane where nothing ships without a human decision, and a weekly plan that proposes rather than acts. See it running before you sign up for anything.