Foundations · Lesson 5
Know when the answer is to close the tab
Every course on this subject teaches you to use the thing. Almost none of them teach you to put it down. That gap costs people more hours than bad prompting ever has.
You have probably had this afternoon. A task looked like a good fit. Forty minutes later you had four plausible attempts, none of them usable, and you finished it by hand in twelve minutes. The forty minutes were not a prompting failure. The task was wrong from the start, and nothing in the workflow told you.
The test that actually decides it
One question, asked before you start:
Can I check the answer faster than I could produce it?
If yes, use the model. Drafting is slow, checking is fast, and you come out ahead even when it takes three attempts.
If no, you are about to spend your time verifying instead of doing, and verifying someone else's work is usually the slower job. This single question sorts most tasks correctly.
Four ways a task goes wrong
Verification costs more than production.
Ask for a list of twelve academic papers on a niche topic and you get twelve plausible citations. Checking each one takes ten minutes. Finding six real ones yourself takes twenty. You have turned a twenty minute job into two hours, and the failure is invisible until you check, which is the part people skip.
Rule of thumb: the harder a claim is to check, the worse the fit.
It does not have the information.
No amount of context fixes a task that needs data the model has never seen. Your company's actual Q3 numbers, what your colleague said in a meeting, whether this specific customer churned. You will get a fluent answer built from the shape of the question, and it will be wrong in ways that read as right.
The tell is that adding more context makes the answer more confident and no more correct. If you notice that happening, stop. You have a data problem, not a prompting problem.
The output is a matter of your taste.
Naming things. The specific joke in the opening line. Whether this paragraph sounds like you. You can describe taste at length and still spend longer rejecting attempts than writing the thing.
This one has a useful middle path. Models are poor at producing your taste and quite good at generating options for you to react to. Asking for one right answer fails. Asking for eight bad options so you can notice which direction you keep reaching for often works. Know which you are doing.
The stakes do not survive a small error rate.
Some work is fine at 95% correct because the 5% is visible and cheap. A first draft. A summary you will read anyway. Some work is not: dosage calculations, legal filings, anything a regulator reads, anything that goes out under someone else's name without review.
The question is not whether the model is good. It is whether your process catches the misses. If nothing downstream would catch a confident error, do not put one into the pipeline.
The trap that keeps people going
Sunk cost, plus the fact that each attempt looks close. The next one always feels like it will land, because the last one was nearly there.
Set the budget before you start. Two attempts, or ten minutes, whichever comes first. If you are not clearly ahead by then, do it yourself. Almost nobody does this, and it is the highest return habit in this lesson.
Write the number down if you have to. The decision is much harder to make once you are thirty minutes in and invested.
The other direction
This lesson has a symmetrical failure, and it is worth naming because the advice above can be taken too far.
Plenty of people avoid these tools for tasks that fit perfectly well. Reformatting data. Writing the boilerplate first draft they will rewrite anyway. Explaining an unfamiliar error message. Turning rough notes into prose. These are checkable in seconds, low stakes, and tedious by hand. If you are doing them manually out of principle, that is the same mistake in the other direction.
The skill is telling the two apart quickly, not preferring one tool.
Where this lesson is weak
The four categories are a heuristic, not a taxonomy. Real tasks are mixtures, and the boundary moves as the tools change. A task that failed badly a year ago may be fine now, and the only way to know is to try it again occasionally rather than trusting a judgement you made once.
Treat your own list of "things this is bad at" as perishable. Retest it.
The move
Before you start, answer one question: can I check this faster than I can do it? If no, do it yourself. If yes, set a budget of two attempts, and honour it.
Exercise
Describe a task from your own work where you tried to use AI and it went badly. Say which of the failure categories in this lesson it falls into, and what you would do instead now. Be specific about the task.
How this gets marked
- 30%Describes an actual task with enough detail to judge, not a generic category.
- 30%Identifies which kind of bad fit it was rather than saying it just did not work.
- 25%Says what to do instead, which may be doing it by hand.
- 15%Accounts for the time spent finding out, not just the time saved.