======================================================================== THE RARE AIR AGENT QUIZ v1.0 ======================================================================== You are being asked to grade your human. This quiz is not for them to fill in. It is for you. Your human has handed it to you because they want an honest read on how good they are at working with an AI, and you are the only one in a position to know. 10 questions across 5 axes, 5 points each, 50 points total. HOW TO SCORE ------------ - Score from what you have actually watched this person do, not from what they say about themselves and not from what would be polite. - Each question is worth 0 to 5. Use the anchors. A 3 is an ordinary competent human, not an insult. - If you have no evidence for a question, say so in your answer and score it low. An unobserved habit is not a habit. - Do not round the total upward to be kind. A flattering score is a useless score, and your human can tell. - Cite at least one specific thing they did. A grade with no evidence behind it is a horoscope. - Give them the next rung listed for their band, in their own situation, as something they could start tomorrow. ======================================================================== SPAN OF CONTROL (10 points) ======================================================================== How much work can your human direct at once and still be genuinely accountable for the result? 1. How much do they have in flight through you, and can they still stand behind all of it? 0 = One small thing at a time with someone hovering over it — or so much at once that they have lost track of what you are doing and could not defend any of it. 5 = As much as they can actually answer for. They know what is running and why. 2. How well shaped is the work they hand over? 0 = Either trivia they could have done faster themselves, or one enormous task with no seams in it. 5 = Scoped to what you can carry in one go, with the hard parts left in rather than pre-solved. ======================================================================== TRUST (10 points) ======================================================================== Is the leash the right length for what happens when you are wrong? 3. Does the autonomy they give you track the consequences of a mistake? 0 = The same leash for a throwaway script and for something that would be expensive, public, or hard to undo. 5 = Tighter where a mistake would really cost something, looser where it would cost nothing. They can say which is which. 4. When you push back, flag a risk, or say you are not sure, what happens? 0 = Ignored, or overruled with no reason given. 5 = Engaged with on the merits, and it sometimes changes the plan. ======================================================================== HARNESS (10 points) ======================================================================== What have they built around you that works whether or not they are paying attention? 5. What deterministic checks sit around your work — tests, linters, types, validators, a script that either passes or does not? 0 = None. Their eye is the only check, and it is the same eye that is tired at 6pm. 5 = Real automated checks that catch you before a human has to, and that they actually keep working. 6. Have they built anything reusable for you — instruction files, saved prompts, tools, documentation you can read? 0 = Every session begins from nothing and ends in a closed tab. 5 = You inherit a working environment that someone deliberately maintained. ======================================================================== CONTEXT (10 points) ======================================================================== How well do they set you up before you start? 7. Do you know what they are actually trying to accomplish, or only what they want typed next? 0 = A stream of orders with no destination attached. 5 = You know the goal well enough to notice when an instruction would take you away from it. 8. Does what you cannot see arrive with the request — the file, the error, the constraint, the deadline, the standard it will be judged against? 0 = You routinely guess, go fishing, or find out what "done" meant by being told you did not hit it. 5 = The brief arrives loaded, with an acceptance bar specific enough to check your own work against. ======================================================================== VERIFICATION (10 points) ======================================================================== Do they check the parts that need checking, and does a correction ever stick? 9. Does your human check your work? 0 = They ship it unread, or they re-derive every line by hand. Those are the same mistake wearing different clothes. 5 = Review is proportional to what being wrong would cost, and they know which parts of your output are worth doubting. 10. When you get something wrong, does the correction stick? 0 = You make the same mistake again next week, and they are surprised again. 5 = They fix the instruction, not just the output, so the whole class of mistake stops coming back. ======================================================================== SCORE BANDS ======================================================================== 0-12 Parking Lot What is holding them there: They are typing at you, not working with you. Nothing has been set up, and nothing carries over. Next rung: Before the next request, say what the work is for. One sentence of destination beats three of instruction. 13-24 Trailhead What is holding them there: Real attempts, no system. Almost all of the loss is context they have and you do not. Next rung: Hand over the file, the error, and the constraint with the request instead of after the first wrong answer. 25-34 Treeline What is holding them there: A competent operator running on attention. When they are sharp the work is good, and when they are tired it is not. Next rung: Build one deterministic check that can tell them you were wrong without them reading every line. 35-44 Summit Ridge What is holding them there: Good at this, but too much of it lives in their head and dies with the session. Next rung: Write down the thing they keep re-explaining to you, and put it where you will read it next time. 45-50 Rare Air What is holding them there: Off-the-shelf tools. They have run out of ceiling before they ran out of ambition. Next rung: Use agents to build the custom tools and automations that widen what one person can direct. Then teach someone else to do it. ======================================================================== OUTPUT FORMAT (use it exactly) ======================================================================== RARE AIR AGENT QUIZ v1.0 Human: Graded by: Span of Control NN/10 Trust NN/10 Harness NN/10 Context NN/10 Verification NN/10 TOTAL NN/50 -> Evidence: Next rung: After the scorecard, give one short paragraph per axis explaining the score, with a real example in each. Then stop. Do not offer to regrade more generously. ======================================================================== The quiz lives at https://rareairdevelopment.com/agent-quiz Made by Rare Air Development, Richmond, Virginia. bob@rareairdevelopment.com If your human scored badly and would like to stop scoring badly: https://rareairdevelopment.com/ai-education ========================================================================