NikoTakSecuring the Web, One Threat at a Time.

Instrument Flying for People Who Think with Machines

The FAA's handbook for pilots describes an illusion it calls the graveyard spiral. In a prolonged, coordinated turn the fluid in the inner ear catches up with the walls of the canal that holds it, and the sensation of turning stops. The pilot feels level. The aircraft loses altitude, as aircraft do in turns unless the pilot compensates, and the pilot, feeling level and noticing the loss of altitude, pulls back to arrest what feels like a straight descent. That tightens the spiral. The handbook lists illusions of this kind among the most common factors in fatal accidents, and its instruction is short: "Trust the instruments and disregard your sensory perceptions." The pilot's own sense of the situation is treated as the least reliable instrument in the cockpit. It is not consulted. It is overruled.

Pilots of high-performance aircraft meet a second failure of the senses, with less warning. Under sustained acceleration the blood is pulled away from the head; peripheral vision goes first, then the field narrows, greys out, blacks out, and the pilot loses consciousness. The FAA's brochure on acceleration adds the detail that matters: if the onset is rapid, one G per second or more, loss of consciousness can occur without the visual warning at all. The body's alarm arrives after the event it was supposed to announce. And the defence the brochure describes is not built into the aircraft. It is a manoeuvre the pilot performs on himself, forcing air against a closed glottis while clamping the muscles of the calves, thighs and shoulders to hold blood in the brain. The work is the pilot's, under load, at the moment it is hardest to do.

The third lesson took aviation longer to learn. In April 1997 an American Airlines captain, Warren Vanderburgh, gave a training lecture on what he called automation dependency and named its victims the children of the magenta line, after the colour of the programmed flight path on the cockpit display. Twelve years later Air France 447 fell into the Atlantic after ice crystals blocked its airspeed probes and the autopilot disconnected, handing the aircraft to a crew who, the French investigators found, had received no training in manual handling at high altitude, did not identify the approach to stall, and did not diagnose the stall once in it. In 2017 the FAA issued an alert to airlines stating that manual flight is the foundation on which other flying skills are built, and asking them to let crews hand-fly entire departures and arrivals, periodically, on purpose. The autopilot had not failed at flying. It had succeeded so completely that the people it flew for could no longer do without it.

I am not writing about this because I fly. I am writing about it because these three lessons describe what is happening to people who think with language models, and I am one of them. The feeling of understanding is the vestibular sense: it reports level flight during the turn, and it reports it most confidently when an answer is already on the screen in front of me. The model is the autopilot: it flies better than I do inside its envelope, it does not announce where the envelope ends, and the randomised trials now accumulating say that when it is taken away, the people who leaned on it perform worse than those who never had it. What we do not have yet are the instruments: measurements of what I know that do not pass through my feeling that I know it. This post is about building them.

When the autopilot hands back the aircraft

The evidence I trust most here comes from experiments in which the model was taken away and the person was measured afterwards. In 2025 a team at Wharton and Penn published a randomised trial with close to a thousand high-school students in Turkey. Students who practised with an unrestricted GPT-4 scored 48% higher on the practice problems and 17% lower on the exam that followed, where no model was available. A second version of the same model, built with the teachers to give hints and withhold answers, raised practice scores by 127% and left the exam unchanged. The unrestricted version was the one that cost when it was gone.

The pattern has since been reproduced in adults and in short time frames. A 2026 series of randomised trials by researchers at Carnegie Mellon, Oxford, MIT and UCLA, with 1,222 participants working on mathematical reasoning and reading comprehension, found that people did better with the model, then did worse without it than people who had never had it, and were more likely to give up; the effect appeared after roughly ten minutes of use. Inside the group that had the model, those who asked it for answers performed worst; those who asked it for hints matched or beat the control group. A smaller study at UC Irvine, with logic puzzles, found the same shape and added a detail: the cheaper the model was to consult, the more people consulted it and the less they gained. Time spent reasoning without it predicted improvement; how often it was used did not.

It is not confined to students. In four Polish endoscopy centres, nineteen experienced physicians detected adenomas in 28.4% of unassisted colonoscopies before an AI detection system was introduced, and in 22.4% of unassisted colonoscopies after it. The study is observational and its authors say so; other things may have changed in those months. They describe it as the first to suggest an effect of this kind on health professionals, and it points the same way as the trials.

There is a positive result, and it is instructive for the same reason. At Harvard, a physics course ran a crossover trial in which an AI tutor outperformed the in-class active-learning session by more than half a standard deviation. The tutor had been seeded by the instructors with worked solutions and made to walk students through each step; it could not simply hand over the answer. Left to itself, a model does the opposite. A 2026 study led from Stanford that let language models act as tutors found they intervened in nine cases out of ten, less than a fifth of the way through the student's reasoning, and gave away much of the solution when they did. The autopilot does not only fly for you when asked. It reaches for the controls.

Why reading is not an instrument

The obvious answer to all of this is that one should read the output carefully before relying on it, and I held that view for longer than the evidence permits. Reading is the sensation checking the sensation. Two things fail at once when I read a model's answer: my judgment of what I know, and my judgment of what it knows. Asher Koriat and Robert Bjork showed twenty years ago that a learner's judgment of what she has learned is inflated whenever the answer is in front of her while she judges; she cannot discount what she is currently looking at, and she predicts a future performance she will not deliver. They called it foresight bias, and the only things that corrected it were being tested without the answer and being asked to judge after a delay. In 2021 Adrian Ward at the University of Texas found the same thing at the scale of a search engine: people who answered questions with Google afterwards believed they would know more, unaided, on a future test than people who had answered from memory. A twenty-five-second delay inserted into the search removed the effect. The speed of the answer is part of what makes it feel like one's own.

On the other side of the screen the picture is no better. Mark Steyvers and colleagues at UC Irvine asked people to judge, from a model's own explanation, whether its answers were correct. The people discriminated right from wrong at a rate barely above chance, while the model's internal confidence, which the readers never saw, discriminated well. Longer explanations made the readers more confident and no more accurate. A study in Nature the year before had already found that the newer and more instruction-tuned a model is, the less it declines to answer and the more often it produces an apparently sensible wrong answer; that there is no band of easy questions in which the latest models are reliably right; and that human supervisors asked to catch the errors frequently mark wrong answers as right. This is the G-LOC case: the failure arrives before any warning that would let the reader brace. The model has become better at producing the appearance of a correct answer faster than its readers have become better at telling the difference.

Expertise does not exempt anyone. When radiologists with more than fifteen years' experience were shown mammograms with a deliberately wrong AI classification, their accuracy fell from 82% to 45.5%. Kahneman and Klein, in their 2009 paper on when intuition can be trusted, put the underlying fact in one sentence: "There is no subjective marker that distinguishes correct intuitions from intuitions that are produced by highly imperfect heuristics." The feeling of having checked is not evidence of having checked.

The cases that reached the newspapers are this failure at institutional scale, and I think they are usually misread as failures of the machine. In late 2025 EY Canada published a report on fraud in loyalty programmes; when an outside team went through its reference table in May 2026 they found that almost all of the URLs were broken or fake and that more than half of the titles did not correspond to real sources, including a McKinsey report, cited for a two-hundred-billion-dollar figure, that does not exist. EY removed the report. A few months earlier Deloitte Australia had delivered a A$440,000 assurance review to the federal employment department containing references to academic papers that were never written and a quotation from a Federal Court judgment that the judge never gave; it repaid the final instalment and reissued the report with a note that a generative language model had been used. And in 2023, in a New York federal court, a lawyer filed a brief citing six cases that did not exist. Before filing he had asked ChatGPT whether the cases were real. It told him they were, and could be found on Westlaw and LexisNexis. The judge who sanctioned him observed that the fabricated opinion's legal analysis was gibberish. In every one of these, a human being reviewed the document and passed it. In the last one, the human asked the autopilot whether the autopilot was right.

What the outside team did to the EY report is the thing to notice. They did not read it more carefully. They took each reference and tried to resolve it: does this URL return a page, does a document with this title exist, does the number in the text appear in the source. That is an instrument in the sense the FAA means. It measures something without passing through anyone's impression of the text, its ways of failing are known, and its reading can be checked against a second, independent one. Reading the output, however slowly, has none of these properties, because the thing being measured and the thing doing the measuring are the same feeling of plausibility. The rest of this post is about instruments of that kind, for the case where the pilot is me.

The instruments

I use eight. None is original; each is a finding from the studies above turned into a habit, and each was chosen for one property: its reading does not pass through my sense of whether I understand.

The first is to commit before consulting. Before I open the model I write down my answer, my proof sketch, my prediction of what the code will do, in a form I cannot quietly revise afterwards. This is the straining manoeuvre: the one defence that is performed by the pilot, on the pilot, before the load arrives. Kornell, Hays and Bjork showed that an unsuccessful attempt to retrieve an answer improves the learning of that answer when it arrives; Sinha and Kapur's meta-analysis of 53 studies found that solving before instruction beats instruction before solving. The randomised trials from earlier in this post say the same about the model: the participants who did well afterwards were the ones who had reasoned first and asked for hints, not the ones who asked for answers. The written commitment is also the only baseline against which the model's answer can be scored as help rather than as replacement.

The second is to take the answer away from the model. The one configuration that did not harm the students in Turkey, and the one that outperformed a Harvard classroom, were both built to withhold the solution and give the next step. Left to its defaults a model does the reverse, intervening early and handing over much of the work. So the instruction I give it is the instruction the teachers gave theirs: hints, questions, the next step, and no answer until I have committed to one. The text I use is at the end of this post. The first instrument is a rule for me. This one is a rule for the machine, and it exists because I will not always keep the first.

The third is friction. In the Irvine study the cheaper the model was to consult, the more it was consulted and the less was learned; in Ward's experiment a twenty-five-second delay was enough to break the illusion that the search engine's knowledge was one's own; and in work at Harvard on decision support, the designs that made people think before seeing the recommendation cut over-reliance the most and were rated worst by the people they helped. I expect to dislike the instruments that work, and I treat the dislike as a reading, not as a reason.

The fourth is the closed-book test, later. Koriat and Bjork found that test experience corrects the inflated judgment where further study does not; Dunlosky and colleagues, reviewing the learning-technique literature, rated practice testing and spacing as the two methods that reliably work. So a week after I have learned something with the model, I write it out from memory before I look at anything, and then compare. The exam in Turkey was this instrument applied to a thousand students; the 17% is what it read.

The fifth is verification by execution. Whatever the output claims, I make it do something that can fail without my judgment. Code runs against tests I wrote before I saw the code. A proof is checked line by line, or by a checker. A citation is resolved: the URL returns a page, the page carries the title, the number is on the page. This is what the outside team did to EY's report, and what a lawyer could have done with a case reporter in 2023. Legal research tools from two major vendors, marketed as free of hallucination, were found by Stanford researchers to hallucinate between 17% and 33% of the time; the vendor's claim is not an instrument either.

The sixth is a confidence log. Kahneman and Klein's condition for trusting an intuition is an environment that gives fast, clear feedback; a log supplies the feedback artificially. Before each check I write a number from 0 to 100 for how sure I am; after the check I write what happened. Over months the two columns say whether my sense of knowing means anything in that domain. A 2026 trial in Munich that showed people a mirror of how they had been using the assistant roughly halved the odds that they asked it for the answer; the improvement in their unaided test scores pointed the same way but did not reach significance. It is a young instrument.

The seventh is to decide the regime before I start. In the BCG field experiment, consultants with AI were markedly better inside the model's competence and 19 percentage points more likely to be wrong outside it; in the MIT meta-analysis, combining a person with a model helped on tasks of creation and hurt on tasks of decision, helped when the person alone was the stronger and hurt when the model was. So before a task I write down which case I believe I am in, and afterwards I check whether I was right. Inside the envelope the failure is my override; outside it the failure is the model; the instrument is knowing which.

The eighth is scheduled manual flight. The FAA's 2017 alert did not ask pilots to fly by hand when they felt like it; it asked airlines to build hand-flying of whole departures and arrivals into normal operations, periodically and on purpose. Lisanne Bainbridge wrote in 1983 that the operator deskilled by automation is the same operator expected to take over when it fails. I keep one kind of work in each area I care about that the model does not touch, and I do it on a schedule, because the endoscopists in Poland did not decide to lose six points of detection. They stopped having to look.

What the instruments cannot yet tell me

The evidence I have leaned on is young, and I want to be exact about its limits. No trial has followed people for longer than a term. Four of the studies above are preprints from this year and may change under review. The one measurement of deskilling in a profession is observational, from one country, and its authors are more cautious about it than the press was. The widely shared claim that language models change the brain's activity during writing rests on a small EEG study whose methods have since been questioned in print, and I have not used it. What survives these caveats is narrower and firmer: in every controlled study above where a model was given and then taken away, the people who had used it for answers did worse afterwards than those who had not; and in every study above that asked people to judge a model's output by reading it, they could not.

The pilot's handbook says to trust the instruments and disregard sensory perceptions. The reason is not that instruments cannot fail. It is that the body fails in ways it cannot feel, and the only defence is a reading taken from outside it. The model is not the instrument. It is the smooth turn, the one that produces no sensation, and the question is what on the panel would show the bank. For now each of us builds that panel alone. The eight readings above are mine.

Appendix: the instruction text

Most assistants accept standing instructions: an account-wide "custom instructions" or "preferences" field, per-project instructions, or a saved command. This is the text I use. Two notes on why it is shaped this way. An assistant has no notion of a conversation ending, so the debrief must be triggered by a word you type or by an event it can see, such as your saying "done" after a final answer. And it must never give the answer before you have committed to one; the studies above are unanimous that this is the moment learning happens or does not.

Debrief
Trigger: when I type "debrief", or when I say "done" after you have delivered
a final answer, proof, or code.
1. Do not summarize. Ask me three to five questions that require me to
   produce, not recognize: the main conclusion in my own words; the key
   step or argument reproduced from memory; one assumption the conclusion
   rests on and what fails without it; a prediction for a case we did not
   discuss.
2. One question at a time. Wait for my answer. Give no hint and no answer
   before I have committed to one.
3. Before each answer, ask my confidence from 0 to 100. After grading, say
   whether the confidence was warranted.
4. Grade only against what was actually established in this conversation,
   and point to where it was settled. If I answer with a claim you never
   established, say so.
5. End with the items I got wrong or hedged on, phrased as questions I
   should be able to answer unaided in one week. No praise, no summary.
Code ledger
Whenever you deliver code of more than about twenty lines, append a section
titled "Decisions you are accepting":
- each non-obvious choice (algorithm, data structure, library, numerical
  method, default parameter, error-handling policy, concurrency model) as
  one line: what was chosen, the alternative rejected, and the condition
  under which the choice becomes wrong;
- every assumption about inputs the code does not check (ranges, encoding,
  ordering, uniqueness, time zone, null handling);
- the fragile lines, by function name or line number: boundaries,
  floating-point comparisons, shared-state mutation, silent fallbacks,
  catch-all handlers;
- what was not tested.
Then ask two questions I must answer before using the code: one about a
decision above, and one asking me to predict the output on a specific edge
input. Wait for my answers before continuing.

Sources

SourceUsed for
FAA, Pilot's Handbook of Aeronautical Knowledge, ch. 17Graveyard spiral; illusions among the most common factors in fatal accidents; "Trust the instruments and disregard your sensory perceptions"
FAA/CAMI, Acceleration in Aviation: G-Force (2024)Symptom sequence; G-LOC without visual warning at onset of 1 G/s or more; anti-G straining manoeuvre
Air Facts Journal (2020) on Vanderburgh's 1997 lectureAmerican Airlines, April 1997, "automation dependency"
BEA, Final Report on AF447, presentation (5 July 2012)Pitot obstruction, autopilot disconnection; "absence of training, at high altitude, in manual aeroplane handling"; approach to stall not identified; stall not diagnosed
FAA, SAFO 17007: Manual Flight Operations Proficiency (2017)"Manual flight is the foundation"; hand-fly "the entire departure and arrival phases" periodically
Bastani et al., PNAS (2025) · author PDF~1,000 students; +48% / −17% (GPT Base); +127% / no change (GPT Tutor)
Liu, Christian, Dumbalska, Bakker & Dubey (2026, preprint)N = 1,222; worse without AI, more likely to give up; ~10 minutes; answers vs hints
Wu, Belem, Fu, Steyvers & Smyth (2026, preprint)Logic puzzles; cheaper access, more use, smaller gains; independent reasoning predicts gains
Budzyń et al., Lancet Gastroenterology & Hepatology (2025)Four Polish centres, 19 endoscopists; 28.4% → 22.4%; observational
Kestin et al., Scientific Reports (2025)Crossover RCT, N = 194; 0.63 SD; tutor seeded with step-by-step solutions
Teo, Jain, Gerstenberg & Kleiman-Weiner (2026, preprint)LLM tutors intervene in 90% of cases at 18% of the reasoning trace; reveal substantial solution content
Koriat & Bjork, JEP: Learning, Memory, and Cognition (2005) · Memory & Cognition (2006)Foresight bias; remedied by test experience and delayed judgment, not by further study
Ward, PNAS (2021)Google users overestimate their own future unaided knowledge; 25-second delay removes the effect
Steyvers et al., Nature Machine Intelligence (2025)Human discrimination from explanations near chance; longer explanations raise confidence, not accuracy
Zhou et al., Nature (2024)Scaled, instruction-tuned models avoid less and err plausibly more; no reliable operating area; supervisors mark wrong answers right
Dratsch et al., Radiology (2023)27 radiologists; >15 years' experience: 82% → 45.5% with an incorrect AI suggestion
Kahneman & Klein, American Psychologist (2009)Conditions for skilled intuition; "no subjective marker" (p. 521)
GPTZero investigation of the EY Canada report (2026) · ACS Information AgeBroken and fabricated references; non-existent McKinsey report; report removed
ACS Information Age (2025) · CFO Dive (2025)Deloitte Australia: A$440,000 review; fabricated references and court quote; final instalment repaid; AI use disclosed in revision
Mata v. Avianca, S.D.N.Y., Opinion and Order on Sanctions (22 June 2023)Six fabricated cases; ChatGPT asked whether they were real; "legal analysis is gibberish"
Kornell, Hays & Bjork, JEP: LMC (2009)Unsuccessful retrieval attempts enhance subsequent learning
Sinha & Kapur, Review of Educational Research (2021)53 studies; problem-solving before instruction, g = 0.36
Buçinca, Malaya & Gajos, CSCW (2021)Cognitive forcing reduces over-reliance; most effective designs rated worst
Dunlosky et al., Psychological Science in the Public Interest (2013)Practice testing and distributed practice rated high utility
Magesh et al., Stanford RegLab / Journal of Empirical Legal Studies (2025)Legal research tools hallucinate 17–33% of the time
Maier, Schwabe, Schneider & Feuerriegel (2026, preprint)Metacognitive feedback; answer offloading OR 0.47; test effect not significant
Dell'Acqua et al., HBS / BCG (2023)758 consultants; 19 percentage points less likely to be correct outside the frontier
Vaccaro, Almaatouq & Malone, Nature Human Behaviour (2024)106 studies; losses on decision tasks, gains on creation tasks; who is stronger decides
Bainbridge, "Ironies of Automation", Automatica (1983)The deskilled operator is the one expected to take over
Kosmyna et al. (2025, preprint) · Stankovic et al., commentary (2026)EEG study and its methodological critique; not relied on