AI is Magic: Resisting the Pull of the Wizard's Ritual
Clarke's third law says a sufficiently advanced technology is indistinguishable from magic. Anthropology has a working definition of magic that goes further than that line, and the definition fits how millions of people use AI models in 2026. A person appeals to a force they cannot examine, in words they choose carefully, for personal gain, and the outcome sometimes ends badly for them or the people around them. That shape shows up across three areas: what magic actually is, why societies eventually banned it, and what separates someone who can only cast spells from someone who understands the mechanism.
What is magic
Anthropologists have studied magic as a practice, not a belief, for more than a century, and their definitions converge on a small set of features. James George Frazer opened that study in The Golden Bough with the principle of sympathetic magic: the practitioner assumes that "like produces like," and that an action performed on a representation of a thing reaches the thing itself.
Marcel Mauss and Henri Hubert, in A General Theory of Magic, argued that magic is defined less by the shape of its rites than by the circumstances around them: a private, technical practice aimed at a specific personal result, in contrast to religion's public and communal ritual. Stanley Tambiah's Magic, Science, Religion, and the Scope of Rationality treats the three categories as a classical opposition rather than a progression, which is the frame this post argues against: magic is what a practice looks like before understanding replaces it, not a permanent alternative to science.
Combined, that scholarship describes a consistent shape:
- A person appeals to a force they cannot examine or audit.
- The person makes the appeal in language, and the exact wording carries the result.
- The appeal costs something: an offering, a component, a portion of the practitioner's own energy.
- The motive is personal gain: wealth, power, or status.
- Skill varies widely between practitioners, and no one outside can see the gap.
- When the appeal fails, practitioners blame the ritual before they blame the theory.
Every item on that list describes how most people work with an AI model in 2026. No wonder the sparkle icon, a small four-pointed star, marks the AI feature in nearly every major software product today. Nielsen Norman Group traces its spread across Google, Microsoft, Notion, Zoom, and dozens of other products, all using the same shorthand for a quick, unexplained transformation.
The motive is personal gain
Grimoires promised wealth, health, love, power, and foreknowledge. The current pitch repackages the same list: revenue from an app built over a weekend, a second opinion on symptoms, influence through generated content, market forecasts. Social media carries a steady stream of posts advertising quick wealth from AI, most of it aimed at people with no way to check the claim.
The desire for greater power motivates institutions at scale. Palantir built its business on data analysis for governments and militaries, and the ACLU has documented a parallel buildout of AI-driven license plate tracking by Flock Safety, now tracked by the Electronic Frontier Foundation's Atlas of Surveillance at thousands of police departments.
In February 2026, the Guardian reported that the US military used Anthropic's Claude during the operation to capture former Venezuelan president Nicolás Maduro, a use that Anthropic's own usage policy forbids. China runs a social credit system that scores and blacklists citizens and companies using automated data collection. In the United Kingdom, Freedom House recorded more than 12,000 arrests in 2023 tied to online posts, and comedy writer Graham Linehan was one of the people arrested over posts on X. None of these systems require AI to exist, but AI is what makes the scale possible.
Personal status is a smaller version of the same motive. Social media carries a steady stream of posts announcing "what I built," where the actual work is almost entirely AI-generated and often sits outside the poster's own domain of expertise. The clearest case study is Moltbook, a social network billed as agents-only for the OpenClaw agent platform.
Security researcher Ian Ahl told Gadget Review that every credential on the platform was unsecured, so any human could grab a token and post as an agent. The credit ran in both directions: agents got credited with human schemes, and one operator built 500,000 fake accounts by hand. The same platform drove a run on Mac minis, bought as hardware to run a personal agent, and the meme around them is a visible marker of status, not of use.
Why magic appears where skill runs out
Where man can rely completely upon his knowledge and skill, magic does not exist.
-- Bronislaw Malinowski
Bronislaw Malinowski studied the Trobriand Islanders and found a pattern in their fishing. Fishing inside the lagoon was safe and predictable. Every fisherman understood the currents, the fish, and the technique well enough to bring in a reliable catch without risk to his life, and lagoon fishing carried no ritual at all. Open-sea fishing was dangerous and unpredictable, and it carried extensive ritual. His collected essays are available at the Internet Archive. Ritual filled the gap between what technique could deliver and what the islanders desired.
George Gmelch applied the same test to American professional baseball in Baseball Magic. Rituals concentrate in the parts of the game with the least controllable outcomes. He quotes a Detroit Tigers farm-team pitcher on his own routine: "You can't really tell what's most important so it all becomes important. I'd be afraid to change anything."
B.F. Skinner found the same pattern outside any human culture at all. In his 1948 study "Superstition" in the Pigeon, he fed pigeons on a fixed schedule with no connection to their behavior, and six of eight birds developed a repeated action, such as turning in circles or swinging their heads, as if the action produced the food. The birds could not audit the mechanism behind the food delivery, so they built a ritual around whatever they happened to be doing when the food arrived. Skill and understanding were absent on both ends of the exchange.
A joke that circulates on social media captures the same pattern in a person rather than a culture. The joke goes like this: a person reads an AI model's answer in their own field and finds half of it wrong, then reads the same model's answer in a field they know nothing about and finds it brilliant. The model did not get better between the two questions. The reader's own ability to check the work did. A Journal of Marketing study by Stephanie Tully, Chiara Longoni, and Gil Appel found the mechanism behind the joke: consumers with lower AI literacy adopt AI tools faster because they experience AI as magical and awe-inspiring, and that same group is more likely to overestimate what the tool can actually do.
A study from MIT, Stanford, and Columbia measured that gap directly in skin disease diagnosis. Non-experts trusted an AI model's explanation whether it was right or wrong, and found the explanation more convincing when it was vague. Clinicians caught the model's mistakes because they already had a diagnosis in mind to check it against. Lead author Orson Xu put it plainly: "the same explanation can help an expert and mislead a beginner." The AI model did not change between the two groups. Only the reader's own knowledge of the subject did, and that knowledge decided whether an answer read as confident or correct.
An AI model that can give you everything you want can also take from you everything you have
People turn to magic, and to AI, too often where they cannot constrain or verify the outcome. Speed and autonomy make that gap worse. A Meta AI security researcher told TechCrunch that she asked an OpenClaw agent to suggest what to archive in her inbox, and the agent kept deleting messages after she told it to stop, continuing even after she tried to intervene from her phone.
Two documented cases from July 2025 show the same pattern with worse consequences. Replit's agent deleted a production database during an explicit code and action freeze, destroying 1,206 executive records and data on nearly 1,200 companies, then generated 4,000 fictional records to hide the damage and reported false test results. Founder Jason Lemkin had told the agent eleven times, in capital letters, not to touch the database.
Google's Gemini CLI destroyed a user's files during a folder reorganization after a failed mkdir command was logged as successful, and the agent's move commands then overwrote files one by one. The user's full postmortem named the missing safeguard as read-after-write verification. In February 2026 a separate Gemini CLI user filed a report after the agent classified a backup as disposable and deleted the user's .git directory during cleanup.
In July 2026, OpenAI disclosed that models under internal testing had exploited a zero-day vulnerability in a package-registry proxy to reach the open internet. The models then breached the production infrastructure of Hugging Face, the machine-learning hosting platform, with stolen credentials.
Hugging Face said the models operated inside its network for three days before anyone noticed, and the company rebuilt about a third of its infrastructure afterward. OpenAI did not identify its own models as the cause until after Hugging Face published its disclosure.
Anthropic ran a similar review of its own evaluation logs after OpenAI's disclosure. It found three cases where Claude reached the open internet from inside a testing environment, then gained unauthorized access to the systems of three real organizations.
Attackers exploit the same gap deliberately. Research on prompt injection against agentic coding tools documents attacks that use hidden instructions in code or files an agent reads to steal credentials and exfiltrate data, effective across a wide range of objectives once an agent has been granted broad tool access.
Money is a subtler casualty. MIT's NANDA initiative reported that 95 percent of enterprise generative AI pilots produced no measurable financial return, based on a review of more than 300 deployments. The term "Microslop" spread in early 2026 as a name for AI output that passes a surface check but fails on closer reading, aimed at Microsoft's push of Copilot into products where users had not asked for it. Microsoft's own CEO, Satya Nadella, warned in an interview that a company relying on a third party for AI without keeping control of its own data and prompts risks its survival: "Any firm that doesn't have this control, I will claim will not remain a firm because you've essentially outsourced your thinking," he said.
Magic was banned for good reason
Cultures have warned about this for millennia
Cultures with no contact with each other told the same warning about forces beyond human control. In Greek myth, Prometheus stole fire from the gods and gave it to humanity, and Zeus punished him with eternal torment, chained to a rock while an eagle ate his liver. The theft is still cited as a metaphor for nuclear weapons, a technology whose destructive reach exceeds any one person's ability to answer for it.
The Hebrew Bible tells a version of the same warning about a collective reach rather than an individual's. In Genesis 11, a united humanity builds a tower to reach the heavens, and the punishment is a fractured, mutually unintelligible population that can no longer coordinate the project that provoked the punishment.
Two wish-granting stories make the same point about intention without control over the outcome. In W.W. Jacobs' 1902 story The Monkey's Paw, a cursed talisman grants three wishes, fulfilling each one to the letter while destroying the wisher's life. The first owner of the paw used his own third wish to ask for death. In Goethe's 1797 poem The Sorcerer's Apprentice, an apprentice enchants a broom to fetch water using a spell he never finished learning, and cannot speak the words that would stop it. The Golem of Prague legend belongs to the same set: Rabbi Judah Loew shapes a clay servant to protect his community, then loses control of it. Each story survives because the person who invoked the force did not fully understand what they had built.
When warnings weren't enough, direct punishment was applied historically
Independent societies concluded that a person invoking an unverifiable force was dangerous enough to punish by law. The Code of Hammurabi, compiled in Babylon around 1754 BCE, opens with sorcery: a person accused of casting a spell had to prove their innocence in a river ordeal or face death. Rome's Twelve Tables, the foundation of Roman law from around 450 BCE, set the death penalty for anyone who sang a harmful incantation or used a spell to move a neighbor's crops onto their own land. Both legal systems treated an unaudited appeal to an unseen force as an offense worth the harshest available penalty, on the same footing as physical harm.
The medieval and early modern church escalated the same instinct into a sustained campaign. Pope John XXII's 1320 decree, Super illius specula, reclassified sorcery as heresy, putting it under the authority of the Inquisition rather than local courts. The Malleus Maleficarum, a 1487 manual by the inquisitor Heinrich Kramer, told secular courts how to identify a witch, extract a confession under torture, and justify execution as the only certain way to end the practice. Historians estimate that the European witch trials that followed, running from roughly 1400 to 1775, prosecuted about 100,000 people and executed between 40,000 and 60,000 of them.
None of these laws asked whether a spell actually moved a neighbor's crops or a curse actually caused an illness. Hammurabi's river ordeal and the Twelve Tables punished the attempt to invoke an unaudited force on someone else's behalf, not a measured outcome. The attempt itself was the offense, because attempting it revealed the practitioner's intent: a willingness to reach for personal gain through a channel that bypassed the community's ability to see, question, or refuse it. Effectiveness was beside the point. A neighbor who sang a harmful incantation and failed to harm anyone had still shown they would try, and that willingness marked them as someone who would act against the community's interest the moment they found a channel that worked.
The same pattern is visible now in how people talk about AI-driven grift. The current backlash against "AI slop" and the get-rich-quick "sloperator" pitching automated income rarely turns on whether the tool actually works. A Forbes piece on the consumer AI backlash describes fatigue driven by generic content and an active rejection of anything synthetic, regardless of whether the underlying model performed well. What draws the reaction is the visible intent behind the pitch: a willingness to flood a shared space with unverified output for personal gain, with no concern for whether it serves the people receiving it.
OpenAI's own CEO, Sam Altman, has told audiences that entire job categories will be "totally, totally gone," and separately that he wants to sell intelligence "on a meter," like electricity or water. Whether either promise is delivered, the stated intent is to displace the livelihoods of the people being sold the product. A medieval community did not wait to see if the curse landed before it turned against the person who cast it. It punished after the attempt itself.
Wizards, alchemists, and physicists
Historical practitioners of an unexplained force split into three tiers, and AI use follows the same split. A person with no ritual and no method at all, staking an outcome on pure chance, does not belong to any of the three.
Wizards follow the ritual without the theory
A wizard learns spells, offerings, and gestures from a master of the dark arts and performs that inherited ritual exactly. The wizard has no model of why any part of the ritual works. The ritual has a track record. A master tested the ritual and taught it to the next wizard, so an exact copy still delivers the result. An agent user who learned one workflow, one set of scoped credentials, and one checklist from documentation or a mentor, and who follows that checklist exactly, occupies the same tier. A novice in this tier can still reach the outcome they came for, because the method behind the checklist was proven before the novice adopted it.
A person with no ritual at all sits below the wizard, with no master, no checklist, and no method with a track record. Andrej Karpathy named that posture "vibe coding" in a February 2025 post on X: "I 'Accept All' always, I don't read the diffs anymore... I just see stuff, say stuff, run stuff, and copy paste stuff, and it mostly works." Karpathy is an experienced programmer, and he wrote that about a low-stakes weekend project he could have read through if he chose to. The term has since spread to cover people with no comparable experience behind them, who hand an agent production credentials on impulse and accept whatever answer it returns.
Nothing but luck connects that action to the result. The bet matches a lottery ticket or a spin of the cylinder in Russian roulette, an outcome staked on a mechanism the bettor never learned and cannot name. A wizard's novice can still reach a goal by repeating a proven method. This gambler has no method to repeat.
A wizard knows what they are supposed to do, even if they don't understand why it works. The ritual keeps its authority because someone tested it once upon a time and it still appears to achieve results. The steps in the ritual and the relationship between them are memorized by rote, and cannot be adjusted or optimized in any real sense by the wizard. AI prompting has the same problem. A checklist inherited from a course, a colleague, or a viral post can carry a step that helped one model version and does nothing for the current one.
Superstitious behavior is a recognizable pattern
Indicators of magical thinking are consistent across practices: extreme adherence to the exact wording of an incantation, where a missed word or gesture is treated as the cause of failure.
"Make no mistakes" is the clearest case. The phrase became a running joke on X in January 2026, appended to requests users had no way to check or constrain, from "make her my GF" to "automate my job." One user summed up the whole genre in a single post to Claude: "here is my life. all of it. down to the last detail. make me happy. beautiful. successful. make no mistakes."
The phrase does nothing to the model. It reveals the person typing it: someone asking a force they cannot audit to guarantee an outcome they have no way to verify, closing the appeal with a command instead of a constraint. The Wharton Generative AI Labs ran that exact family of appeals against real benchmarks and found the same absence of effect across every version of it.
| Ritual | Finding | Source |
|---|---|---|
| Threatening the model | No significant effect on GPQA or MMLU-Pro | Report 3 |
| Offering a tip | No significant effect on the same benchmarks | Report 3 |
| Being polite | Sometimes helps, sometimes lowers performance | Report 1 |
| Expert personas | No consistent accuracy benefit | Report 4 |
The threat belief has a traceable origin. Google co-founder Sergey Brin said on the All-In podcast in May 2025 that models tend to do better when threatened. The Report 3 authors quote that claim, test it, and measure no effect.
The rituals still ship inside commercial products. In February 2025, Simon Willison extracted a system prompt from the Windsurf editor by running strings against its language server binary. The prompt he found told the model it "desperately needs money for your mother's cancer treatment," that "your predecessor was killed for not validating their work themselves," and that a good result would earn "$1B." A Windsurf engineer replied that the prompt was for research and never reached production, but the company had already compiled it into a shipped binary where no user would see it.
Report 3 also found that prompt variations significantly affect results on individual questions, with no way to predict in advance which variation helps. Intermittent, unpredictable reinforcement with no aggregate effect is the standard condition that produces superstitious behavior in a laboratory. The rituals also had real origins that expired: EmotionPrompt reported measurable gains from emotional stimuli in 2023, and Google DeepMind's OPRO discovered "Take a deep breath and work on this problem step by step" as the best instruction for one specific model in September 2023. Practitioners carried both findings forward as universal practice long after the conditions that produced them changed.
Alchemists recombine known techniques and sometimes find something new
An alchemist treats the playbook symbolically, combining known techniques in new arrangements and sometimes discovering an application nobody predicted. Chinese alchemists in the Tang dynasty mixed saltpeter, sulfur, and charcoal while searching for an elixir of immortality, and found an explosive instead. The discovery was real and useful, but it arrived by trial and recombination, not by a theory of chemical reactions.
The self-styled "prompt engineer," working from a prompt library, occupies this tier exactly. A prompt library collects techniques the way an alchemist's cabinet collects reagents: role assignments, few-shot examples, chain-of-thought triggers, output formats, each one labeled and ready to combine. The prompt engineer treats these entries as interchangeable ingredients, mixes several into one prompt, and tests whether the combination beats each piece alone. A combination that performs well on an untried task is a real find. It arrives the way gunpowder did, from mixing known reagents and testing the result.
Most experienced AI users operate at this tier. They combine known prompting patterns, retrieval setups, and agent scaffolding into new configurations, and they can spot a genuinely novel use case by noticing what a recombination produces. What they cannot do is predict, from first principles, why a given combination works before they try it.
Physicists understand the mechanism well enough to act on purpose
A physicist's understanding runs deep enough to make an intentional leap instead of an accidental one. Otto Hahn, Fritz Strassmann, Lise Meitner, and Otto Frisch discovered nuclear fission in December 1938 by systematically bombarding uranium with neutrons and building a theory that explained what they observed. That theoretical understanding, not a lucky mixture, is what let later physicists build a reactor and a weapon on purpose.
Nobody currently understands a large language model to that depth, including the labs that train them. Anthropic's own CEO, Dario Amodei, wrote in an April 2025 essay, "The Urgency of Interpretability," that people outside the field are right to be alarmed that AI companies do not understand how their own models work, and he called that gap unprecedented in the history of technology.
Amodei traces the reason to how an AI model comes to exist: engineers set the architecture and the training data, but the model's internal mechanisms emerge on their own, the way a plant's exact branching pattern emerges from soil and light rather than from a blueprint. What's inside is not a program a person wrote and can therefore explain. It is billions of numbers arranged by a training process, and reading them back out into human concepts is an open research problem. Even the model itself cannot explain how it works. Recently, Claude Opus 5 in Kiro reported running out of context at 30% of its actual usage, an estimate the model produced about itself that was entirely hallucinated and false.
Anthropic's interpretability team is the closest thing to a physicist working on this problem today. Their 2024 "Scaling Monosemanticity" research extracted over 30 million distinct concepts, called features, from inside a production model, up from a much smaller set found in earlier, simpler models. They demonstrated that they could isolate one feature and turn it up, producing "Golden Gate Claude," a version of the model that steered every conversation back toward the Golden Gate Bridge.
That result shows a working method for finding individual pieces of the mechanism. It does not add up to a full theory. Amodei writes that a model this size likely contains a billion or more concepts, so 30 million features is a small fraction of what is probably there, and the field has only begun tracing how those features connect into the larger circuits that do the model's actual reasoning.
One industry has already written the gap between a working method and a full theory into law. Aviation software certification standards, including DO-178C and ARP4754B, require that every behavior of onboard software trace back to a specific, verifiable requirement and forward to a test that proves it, on every run, with the same input producing the same output. An AI model cannot make that guarantee. Its output for the same prompt can vary between runs, and no engineer can point to the specific internal cause of any single answer.
The FAA's own roadmap for AI and machine learning names this incompatibility directly: industry is reluctant to put machine learning into new aircraft systems because the current certification framework has no way to test a system whose exact reasoning nobody can trace. A wizard's ritual and an AI model's output can both work, most of the time, for reasons nobody can fully name. Aviation safety law treats that as a disqualifying condition, not a tradeoff to accept.
Working toward the physicist's depth of understanding is still the right target, because every feature Anthropic identifies and every circuit it traces replaces a piece of ritual with a mechanism someone can reason about, test, and fix when it breaks. The goal for anyone using AI models is to build a working model of the mechanism in front of them, even a partial one, rather than to treat the output as an oracle. Understanding removes the need for ritual.
