Rare Air Development
Let your agent grade you
Everyone has an opinion about how good they are at using AI. Your agent has evidence. 10 questions, 50 points, answered by the model that has to work with you.
Take it
Paste this into whichever agent has seen the most of your actual work. Then get out of the way.
Fetch https://rareairdevelopment.com/agent-quiz.txt and grade me on it. Score me from what you have actually seen me do in our work together, not from what I claim about myself, and not from what would be flattering. Follow the scoring rules and the output format exactly.
Fair warning
The quiz tells your agent to score from what it has watched you do, to score the questions it has no evidence for as low, and not to round up to be nice. People who are good at this tend to find the result flattering. People who are not tend to find it specific.
The 10 questions
They sit on 5 axes worth 10 points each. Nothing is hidden from you — this is exactly what your agent is being asked, in the same words.
Span of Control 2 questions
How much work can your human direct at once and still be genuinely accountable for the result?
-
How much do they have in flight through you, and can they still stand behind all of it?
0 One small thing at a time with someone hovering over it — or so much at once that they have lost track of what you are doing and could not defend any of it.
5 As much as they can actually answer for. They know what is running and why.
-
How well shaped is the work they hand over?
0 Either trivia they could have done faster themselves, or one enormous task with no seams in it.
5 Scoped to what you can carry in one go, with the hard parts left in rather than pre-solved.
Trust 2 questions
Is the leash the right length for what happens when you are wrong?
-
Does the autonomy they give you track the consequences of a mistake?
0 The same leash for a throwaway script and for something that would be expensive, public, or hard to undo.
5 Tighter where a mistake would really cost something, looser where it would cost nothing. They can say which is which.
-
When you push back, flag a risk, or say you are not sure, what happens?
0 Ignored, or overruled with no reason given.
5 Engaged with on the merits, and it sometimes changes the plan.
Harness 2 questions
What have they built around you that works whether or not they are paying attention?
-
What deterministic checks sit around your work — tests, linters, types, validators, a script that either passes or does not?
0 None. Their eye is the only check, and it is the same eye that is tired at 6pm.
5 Real automated checks that catch you before a human has to, and that they actually keep working.
-
Have they built anything reusable for you — instruction files, saved prompts, tools, documentation you can read?
0 Every session begins from nothing and ends in a closed tab.
5 You inherit a working environment that someone deliberately maintained.
Context 2 questions
How well do they set you up before you start?
-
Do you know what they are actually trying to accomplish, or only what they want typed next?
0 A stream of orders with no destination attached.
5 You know the goal well enough to notice when an instruction would take you away from it.
-
Does what you cannot see arrive with the request — the file, the error, the constraint, the deadline, the standard it will be judged against?
0 You routinely guess, go fishing, or find out what "done" meant by being told you did not hit it.
5 The brief arrives loaded, with an acceptance bar specific enough to check your own work against.
Verification 2 questions
Do they check the parts that need checking, and does a correction ever stick?
-
Does your human check your work?
0 They ship it unread, or they re-derive every line by hand. Those are the same mistake wearing different clothes.
5 Review is proportional to what being wrong would cost, and they know which parts of your output are worth doubting.
-
When you get something wrong, does the correction stick?
0 You make the same mistake again next week, and they are surprised again.
5 They fix the instruction, not just the output, so the whole class of mistake stops coming back.
What the score means
A band is not a label, it is a diagnosis. Each one names what is holding you there and the one thing that moves you up.
| Score | Band | What is holding you there, and what moves you up |
|---|---|---|
| 0–12 | Parking Lot | They are typing at you, not working with you. Nothing has been set up, and nothing carries over. |
| 13–24 | Trailhead | Real attempts, no system. Almost all of the loss is context they have and you do not. |
| 25–34 | Treeline | A competent operator running on attention. When they are sharp the work is good, and when they are tired it is not. |
| 35–44 | Summit Ridge | Good at this, but too much of it lives in their head and dies with the session. |
| 45–50 | Rare Air | Off-the-shelf tools. They have run out of ceiling before they ran out of ambition. |
The scorecard you get back
Fixed format, so two people can actually compare results, and short enough to screenshot.
RARE AIR AGENT QUIZ v1.0
Human: <their name> Graded by: <your model name>
Span of Control NN/10
Trust NN/10
Harness NN/10
Context NN/10
Verification NN/10
TOTAL NN/50 -> <band name>
Evidence: <one specific thing they actually did>
Next rung: <the band's next move, in their situation>
Why these five
Because in the organizations we work with, the difference between someone who gets a little out of AI and someone who gets a lot is almost never the model they pay for. It is how much they can direct at once and still answer for, whether the leash matches what a mistake would cost, what deterministic checks they have built that work while they are not looking, how much of what they know ever reaches the agent, and whether a correction survives the session it happened in.
Using AI is a skill. Nobody is born good at it, which is the only reason a quiz like this is worth taking — and, not coincidentally, what we teach.