先审后信
Critical Thinking — A General Protocol for Evaluating Any Claim
多数论证不是错在逻辑,而是错在前提。批判性思维不是一种气质而是一套要跑的流程——因为「我方偏误」与智商几乎不相关,聪明只会给你更多弹药去维护既有立场。
#1. Executive Summary
The central thesis of this chapter is simple to state and hard to live by: you can reliably separate good arguments from bad ones and strong evidence from weak, but only by running an explicit procedure on the claim in front of you — because your unaided intuition is systematically and predictably corrupted by your prior beliefs, and no amount of raw intelligence fixes this by itself. Critical thinking is not a personality trait or a feeling of skepticism; it is a disposition plus a procedure, and the procedure has to be practiced deliberately, not merely read about.
Three conclusions organize everything that follows. First, most real reasoning fails at the level of premises, not logic. People rarely commit formal logical errors; they accept false or unexamined starting assumptions. The high-yield skill is therefore reconstructing an argument’s hidden premises and grading its evidence — not memorizing a list of Latin fallacy names. Second, rational thinking is only weakly correlated with intelligence. Keith Stanovich’s decades of work show that “myside bias” — evaluating evidence in a manner biased toward your own prior opinion — is essentially uncorrelated with IQ. Smart people believe stupid things not because they cannot reason but because they deploy their reasoning to defend conclusions they already hold. Third, the discipline can be taught, but the average effect is modest and depends heavily on how you teach it. Abrami et al.’s 2015 meta-analysis found a weighted random-effects mean effect size of g+ = 0.30 (p < .001) across 341 effect sizes — real, but between “small” and “moderate” — with the largest gains coming from dialogue, real-world problems, and mentoring combined, and from intensive argument mapping.
In the AI era, all of this becomes more urgent. When text, citations, images, and video can be machine-fabricated at near-zero cost, the bottleneck skill for intellectual independence is the ability to audit a source rather than trust it. The chapter closes with an operable middle path between credulity (“if it’s published, it’s true”) and corrosive cynicism (“everything is biased, so nothing can be known”), and a five-question protocol you can run today.
#2. Why This Topic Matters
Every other skill in this curriculum is exercised on claims. Decision-making (Chapter 1) requires you to judge which forecasts and options are credible. Probabilistic and Bayesian thinking (Chapters 3–4) require you to assign likelihoods to evidence and priors to hypotheses — but garbage evidence produces garbage posteriors. Second-order thinking (Chapter 6) asks “and then what?” about claims whose first-order truth you must first establish. Critical thinking is the meta-skill: it is the general protocol for evaluating any claim, and it is domain-neutral in a way almost nothing else in this curriculum is.
The urgency is now structural. Before 2023, the overwhelming majority of text and images a person encountered were made by other humans, however biased or mistaken. That assumption is dead. Large language models generate fluent, confident, syntactically perfect prose — including fabricated citations — at essentially zero marginal cost. A 2024 Stanford study (“Large Legal Fictions,” Dahl, Magesh, Suzgun & Ho, Journal of Legal Analysis) found general-purpose models hallucinated on legal queries between 58% and 88% of the time; the authors called legal hallucinations “pervasive and disturbing.” When the ratio of generated-to-human content explodes, the naïve default of “trust what looks authoritative” becomes actively dangerous. The ability to audit — to ask who made this, how do they know, and what would change my mind — becomes the difference between an independent mind and one that is quietly steered.
#3. Foundations
#Core vocabulary
A claim (or proposition) is a statement that is either true or false. An argument is a set of statements in which some (the premises) are offered as support for another (the conclusion). This is the atomic unit of critical thinking: not the isolated fact, but the inferential structure connecting facts to conclusions.
The most consequential distinction in the whole field is between validity and soundness:
- An argument is valid if the conclusion follows from the premises — if the premises were true, the conclusion would have to be true. Validity is about form, not content.
- An argument is sound if it is valid and its premises are actually true.
You can have a perfectly valid argument that is worthless because a premise is false (“All birds can fly; penguins are birds; therefore penguins can fly” — valid form, false first premise). For inductive arguments — which is most real reasoning — we speak instead of strength: how much the premises raise the probability of the conclusion.
There are three broad modes of inference: deduction (from general premises to a conclusion that must follow, if valid), induction (from observed cases to a probable generalization), and abduction (inference to the best explanation — reasoning from an observation to the hypothesis that would best account for it). Deduction preserves truth; induction and abduction, which dominate everyday and scientific reasoning, only make conclusions more or less probable — which is why “strength,” not “validity,” is the operative concept for most claims you will meet.
A fallacy is a defect in an argument — a pattern of reasoning that fails to support its conclusion. This is the crucial contrast with Chapter 2: a cognitive bias is a property of minds (a systematic tendency of the reasoner), whereas a fallacy is a property of arguments (a defect in the product). The two interlock — confirmation bias in a mind produces cherry-picked arguments — but they are analyzed differently.
An evidence hierarchy ranks types of evidence by how well they control for the ways we fool ourselves. Motivated reasoning is the tendency to reach conclusions we want to reach by selectively weighting evidence. Intellectual humility is the recognition — and the operational habit — that any given belief of yours may be wrong.
#Historical development
The tension at the heart of critical thinking is ancient. The Sophists of 5th-century BCE Athens taught rhetoric — how to make the weaker argument appear stronger — while Socrates and later Aristotle sought methods to distinguish sound reasoning from mere persuasion. Aristotle’s Prior Analytics formalized the syllogism, giving us the first systematic theory of valid inference, and his Sophistical Refutations catalogued fallacies.
The modern lineage runs through John Stuart Mill’s A System of Logic (1843) on inductive inference and causation, and John Dewey’s How We Think (1910), which reframed the goal as “reflective thinking” — the active, careful consideration of beliefs in light of the grounds that support them. The mid-20th century produced the informal-logic movement: Stephen Toulmin’s The Uses of Argument (1958) offered a model of real-world argument beyond the syllogism — decomposing arguments into claim, data (grounds), warrant, backing, qualifier, and rebuttal; Charles Hamblin’s Fallacies (1970) revived serious study of fallacies; and Johnson & Blair built informal logic into a teaching discipline in the 1970s–80s. Douglas Walton later recast fallacies not as always-errors but as defeasible argumentation schemes with attached critical questions — appeals to authority, for instance, are often legitimate, and the job is to ask the right probing questions.
Running in parallel, the cognitive-science turn of the 1970s–90s — Kahneman and Tversky’s heuristics-and-biases program, and Stanovich’s research on individual differences in rationality — supplied the descriptive counterpart: not how we ought to reason, but how we actually do, and why we fail. The 2020s have added a fourth front: information science and AI studies, grappling with an environment where the provenance of content can no longer be assumed.
#4. Current Scientific Understanding
#Robust findings (well-established)
Reasoning is reliably biased by prior belief. This is one of the most replicated findings in cognitive science. Charles Taber and Milton Lodge’s “Motivated Skepticism in the Evaluation of Political Beliefs” (2006, American Journal of Political Science) showed people uncritically accept confirming arguments while vigorously counter-arguing disconfirming ones — and that this disconfirmation bias was strongest among the most politically sophisticated. Ziva Kunda’s “The Case for Motivated Reasoning” (1990, Psychological Bulletin) established the mechanism: we reason to conclusions we want, constrained only by the need to construct a plausible justification.
Myside bias is largely independent of intelligence. Stanovich, West & Toplak’s review “Myside Bias, Rational Thinking, and Intelligence” (2013, Current Directions in Psychological Science) reports that “the magnitude of the myside bias shows very little relation to intelligence.” As Macpherson & Stanovich (2007) summarize the experimental results, “cognitive ability displayed near zero correlations with myside bias as measured in two different paradigms.” This makes myside bias an outlier bias: most cognitive biases correlate weakly and negatively with cognitive ability (i.e., the smarter show slightly less bias), but this one barely correlates at all. The implication is stark — high IQ does not protect you here, and may make it worse by giving you more ammunition to defend your side. (This is well-established for the specific myside/one-sided reasoning paradigms; note that a broader claim that intelligence never helps against any bias would be false — most biases do show small negative correlations with ability.)
Thinking dispositions predict rationality beyond intelligence. Stanovich’s Comprehensive Assessment of Rational Thinking (CART; Stanovich, West & Toplak, The Rationality Quotient, MIT Press, 2016) demonstrated that dispositions like actively open-minded thinking (AOT) and need for cognition predict rational-thinking performance independently of IQ. Most CART subtests correlate with intelligence only in the 0.25–0.50 range; overconfidence (measured by the Knowledge Calibration subtest) shows about a 0.38 correlation with intelligence — “a substantial amount of dissociation,” in Stanovich’s words. Rationality and intelligence are related but genuinely distinct.
Untrained people mostly cannot name the manipulation. Absent training, most people fail to identify the specific fallacy or evaluate source credibility well. Wineburg & McGrew (2019, Teachers College Record), sampling 10 PhD historians, 10 professional fact-checkers, and 25 Stanford undergraduates, found even the historians and the elite undergraduates were routinely fooled by superficial website cues — official-looking logos, .org domains — that professional fact-checkers ignored. The fact-checkers “read laterally” (leaving the site to check it against other sources) and “arrived at more warranted conclusions in a fraction of the time.”
Training works, modestly, and depends on pedagogy. Abrami et al. (2015) found a mean effect of g+ = 0.30 across 341 effect sizes . The dominant story, however, is heterogeneity (the collection was statistically heterogeneous, p < .001) — meaning how you teach matters more than whether you teach (see §7 for the moderator breakdown).
#Live debates (unsettled)
- General skills vs. domain knowledge. Do we teach critical thinking as a transferable skill (Robert Ennis) or is it inseparable from domain knowledge (John McPeck, Daniel Willingham)?
- Does debiasing transfer? Lab gains are easy; real-world transfer to new contexts weeks later is hard and often absent.
- Do fallacy lists help or backfire? Teaching a catalogue of fallacies risks producing people who wield “that’s a strawman!” as a conversation-ending weapon (the “fallacy fallacy”) rather than reasoners who understand why an inference fails.
- “Lazy” vs. “biased.” Pennycook & Rand (“Lazy, not biased,” 2019, Cognition) argue susceptibility to fake news is driven more by a failure to reason than by motivated reasoning — analytic thinkers discern true from false better regardless of partisan alignment. This directly challenges Dan Kahan’s “identity-protective cognition” / motivated System 2 reasoning account. The debate is active and unresolved; replications (e.g., a Hungarian representative sample) have tended to support the “classical reasoning”/lazy account, but the question is not closed.
#5. Interdisciplinary Perspectives
Four disciplines converge on critical thinking, and their frictions are instructive.
| Discipline | Core question | Contribution | Characteristic blind spot |
|---|---|---|---|
| Philosophy (informal logic, argumentation theory) | What ought good reasoning look like? | Normative standards: argument structure, validity/soundness, Toulmin’s model, Walton’s schemes | Idealizes a rational agent that humans cannot be unaided |
| Cognitive science / psychology | How do we actually reason and fail? | Descriptive account: motivated reasoning, myside bias, dual-process theory, dispositions | Documents failure vividly but is weaker on fixes that transfer |
| Education science | Can it be taught and transferred? | Intervention evidence, effect sizes, moderators | Modest effects; measurement and transfer problems |
| Information / AI studies | How do we evaluate this new environment? | Lateral reading, SIFT, provenance, hallucination auditing | Fast-moving, thin longitudinal evidence |
The sharpest conflict is between the philosopher’s ideal and the psychologist’s finding. Philosophy describes the norms a perfectly rational agent would follow; psychology shows those norms are unreachable by unaided humans, especially on identity-relevant topics. The resolution the education and information sciences offer is structural: don’t try to become the ideal reasoner in your head — externalize the process. Map the argument on paper; read laterally in another browser tab; run a checklist. The philosopher supplies the standard, the psychologist supplies the warning, and the applied fields supply the scaffolding. A second friction: popular culture treats critical thinking as obviously and easily teachable, while the education literature’s honest answer is “yes, somewhat, if you do it deliberately and intensively — and it may not transfer.”
#6. Mental Models
1. Argument reconstruction + steelmanning. Before evaluating any argument, restate it in your own words in its strongest form — including any unstated premise (an enthymeme is an argument with a suppressed premise). If someone says “She went to a good school, so she’ll be a great hire,” the hidden premise is “graduates of good schools make great hires” — which is exactly the debatable part. Steelmanning — engaging the best version of an opponent’s case rather than a caricature (the opposite of strawmanning) — is both an epistemic and an ethical discipline. Works when: you genuinely want truth. Fails when: the original argument is so vague there is nothing to reconstruct, or when steelmanning becomes an excuse to invent a better argument nobody actually made.
2. The evidence hierarchy. Rank evidence by how well it controls for self-deception: anecdote → case series → observational study → randomized controlled trial (RCT) → systematic review/meta-analysis. Works when: comparing evidence within a question. Fails when: treated as an automatic ranking that ignores context — a well-designed observational study can beat a small, biased RCT, and the hierarchy says nothing about whether the studies were honestly conducted.
3. The “who benefits?” source audit. For any claim, ask: who is making it, what is their expertise, what do they gain if you believe it, and what would they lose if it were false? This is the cui bono test, and it connects directly to conflicts of interest.
4. Bayesian likelihood ratio as argument strength (connects to Chapter 4). A piece of evidence is strong exactly insofar as it is much more likely under one hypothesis than under its rival — that is the likelihood ratio. “This treatment worked for me” is weak evidence because it is nearly as likely whether or not the treatment works (placebo, regression to the mean, natural recovery). Reframing “how strong is this evidence?” as “how much more likely is this observation if the claim is true than if it’s false?” is the single most powerful upgrade to everyday argument evaluation.
5. The “steelman + probability + update” triple. (a) State the best version of the claim and its best counter; (b) assign a rough probability; (c) specify what evidence would move that probability. If nothing could move it, you are not reasoning — you are rationalizing.
6. The Dunning-Kruger self-check. Regardless of the effect’s contested statistical status (see §9), the practical prompt is sound: on topics where you feel most confident and least informed, deliberately seek the strongest disconfirming source.
7. The five-question claim test. Presented in full in §10.
#7. Common Misconceptions
“Critical thinking means being negative / skeptical about everything.” No. It equally requires knowing when to accept a claim. A critical thinker who cannot ever be convinced is as broken as one who believes everything. The goal is calibration, not doubt.
“All claims are equally biased, so truth is just opinion” (the relativism trap). That everyone has biases does not make all claims equally supported. A systematic review and a Facebook meme are both “biased” in the trivial sense of having a perspective, but they are not equally reliable. Symmetric-sounding skepticism that flattens all sources to the same level is itself a manipulation technique (see §9, manufactured doubt).
“Logical fallacies are the main problem in real reasoning.” They are not. As emphasized throughout, most real arguments fail at the premises — false or unexamined assumptions — not at the logical step. Fallacy-spotting is a minor skill; premise-auditing is the major one.
“Smarter people think more critically.” Only weakly. Intelligence and rational disposition are distinct (Stanovich). Intelligent people are often better at constructing elaborate defenses of what they already believe.
“If it’s published in a peer-reviewed journal, it’s true.” Peer review is a filter, not a guarantee. The Wakefield MMR paper (§9) was peer-reviewed and published in The Lancet. Publication is necessary but nowhere near sufficient; p-hacking, publication bias, and undisclosed conflicts survive review routinely.
“Critical thinking is the same as skepticism.” Skepticism is one half. The other, harder half is proportioning belief to evidence — which sometimes means firmly accepting a well-supported claim despite social pressure to doubt it.
Why do intelligent people fall for these? Because each misconception feels like intellectual sophistication. Reflexive doubt feels smarter than belief; “everything’s biased” feels worldly; fallacy-spotting feels rigorous. The felt sense of insight is not a reliable guide — which is the whole lesson.
#The fallacy toolbox as patterns of manipulation
A short field guide. The point is not to memorize names but to recognize where each is deployed and to feel the manipulation:
| Fallacy | The move | Where it lives |
|---|---|---|
| Ad hominem | Attack the person, not the argument | Politics, comment sections |
| Straw man | Refute a distorted, weaker version of the claim | Debates, op-eds |
| False dilemma | “Either X or Y” when other options exist | Advertising, political rhetoric (“us or chaos”) |
| Slippery slope | Assert one step inevitably leads to a catastrophe | Policy fights |
| Appeal to authority | Cite status rather than relevant, demonstrated expertise | Marketing (“9/10 dentists”) |
| Appeal to emotion | Substitute fear/outrage/pity for evidence | The attention economy’s core lever |
| Post hoc ergo propter hoc | “After, therefore because of” — correlation as causation | Health claims, superstition |
| Texas sharpshooter | Draw the target after seeing where the bullets landed — cherry-pick a cluster | p-hacking, “predictions” that fit only in hindsight |
Note that appeal to authority is not always fallacious: deferring to a genuine expert within their domain of demonstrated competence is often the rational thing to do. It becomes a fallacy when the authority is irrelevant to the claim, when experts disagree, or when status substitutes for evidence. This is exactly Walton’s point: the test is a set of critical questions (“Is this person an expert in this field? What do other experts say? Is there a conflict of interest?”), not a blanket ban.
#8. Real-World Applications
The principle is constant across domains: reconstruct the claim, grade the evidence, audit the source, state what would change your mind, and price the cost of being wrong. The direction of application:
- News and social media: Practice lateral reading — leave the page and check the source in other tabs before believing or sharing. Notice the emotional hook; strong affect is the attention economy’s lever (Caulfield’s SIFT “habit”: if you feel a strong reaction, stop).
- Health claims and marketing: Ask for the evidence tier. “Clinically proven” usually means a tiny study or none. Apply cui bono to the seller.
- Political argument: Assume your own myside bias is operating most strongly where you feel most certain. Steelman the other side before critiquing.
- Workplace meetings: Separate the claim from the confident person making it. Ask “what evidence would change our decision?” before committing.
- Science in the press: A press release is not the study; the study is not the field. Correlation reported as causation is the default failure mode.
- Evaluating AI outputs: Treat the model as a source to be audited, not an oracle — verify every specific, checkable claim (names, numbers, citations) independently. This connects to Chapter 5’s automation bias: humans over-trust algorithmic output.
- Personal belief management: Keep a decision/belief journal (Chapter 1). Periodically ask which of your confident beliefs you have never seriously tried to disconfirm.
#9. Case Studies
#(a) FAILURE: Wakefield’s MMR-autism paper (1998)
On 28 February 1998, The Lancet published a paper by Andrew Wakefield and colleagues describing 12 children and suggesting a link between the MMR vaccine, bowel disease, and regressive autism. The study was a triple failure.
The study itself was a case series of n=12 — no control group, no comparison, at the very bottom of the evidence hierarchy. Yet it triggered a global panic. A first-pass evidence evaluation would have flagged this immediately: twelve hand-selected children cannot establish causation about a vaccine given to millions. Worse, investigative journalist Brian Deer later established that the data were falsified (contrary to the paper, five of the twelve children had documented developmental problems before vaccination, and diagnoses were altered), that patients were recruited through anti-MMR campaigners, and that Wakefield had an enormous undisclosed conflict of interest. Per Deer’s Sunday Times and BMJ investigation, Wakefield was personally paid £435,643 in fees plus £3,910 in expenses — routed through solicitor Richard Barr of the firm Dawbarns at roughly £150/hour from 1996 — by lawyers building a lawsuit against vaccine manufacturers.
Media amplification: a press conference converted a 12-child case series into a worldwide health story. Motivated publics: frightened parents seeking an explanation for autism found one.
The reckoning: the UK General Medical Council, after its longest-ever fitness-to-practise hearing, found Wakefield guilty of serious professional misconduct (announcing its findings on 28 January 2010; the striking-off from the medical register followed on 24 May 2010). The Lancet fully retracted the paper on 2 February 2010. In January 2011, the BMJ (Deer) labelled it an “elaborate fraud.” A proper evaluation at each step — evidence tier, conflict-of-interest check, source audit — would have caught it. The MMR scare contributed to falling vaccination rates and the UK’s loss of measles-elimination status in 2018.
#(b) FAILURE: Mata v. Avianca (S.D.N.Y. 2023)
Attorney Steven Schwartz, of Levidow, Levidow & Oberman, used ChatGPT to research a brief opposing a motion to dismiss in a personal-injury case against the airline Avianca (docket 22-cv-1461, before Judge P. Kevin Castel). ChatGPT produced six entirely fabricated cases — including Varghese v. China Southern Airlines — complete with fake citations, fake internal quotations, and fake reasoning attributed to real judges. When opposing counsel could not find the cases, Schwartz asked ChatGPT whether they were real; it falsely confirmed they were, even claiming they could be found on Westlaw and LexisNexis.
On 22 June 2023, Judge Castel sanctioned Schwartz and his co-counsel Peter LoDuca and their firm $5,000 (joint and several) under Rule 11, and required them to notify each real judge falsely named as author of a fabricated opinion. Castel described one fabricated analysis as “gibberish.” The court stressed that the sanction stemmed less from the initial AI error than from the attorneys’ failure to verify and their doubling down after being challenged.
The lesson is exact: the ordinary discipline of checking a citation against the actual reporter — ten minutes of work — would have caught every fabrication. The AI era does not require a new epistemology; it requires that we actually apply the old one to a source that produces fluent, confident, verifiable-looking falsehoods. The case became the canonical AI-hallucination precedent, cited in bar advisories worldwide — and it was no isolated event. According to Paris-based researcher Damien Charlotin’s AI Hallucination Cases database (HEC Paris), 1,598 court cases involving AI-fabricated citations or content had been identified worldwide as of 9 June 2026 — up from roughly 200 a year earlier; Bloomberg Law reports that about 90% of those decisions were written in 2025, with Charlotin noting the pace reached “two cases per day or three cases per day” by spring 2025.
#(c) SUCCESS (institutional): Evidence-based medicine
The evidence-based-medicine (EBM) movement — the Cochrane Collaboration (founded 1993), the systematic review, and the GRADE framework for rating evidence quality — is critical thinking institutionalized. Its core move was to replace “the eminent professor says so” (authority) with auditable, reproducible evidence syntheses: explicit search strategies, pre-registered protocols, formal risk-of-bias assessment, and transparent grading of how much confidence a body of evidence warrants. This is the evidence hierarchy operationalized at civilizational scale, and it has demonstrably improved medicine.
But it has honest limits, and applying the chapter’s own standards to it is instructive. John Ioannidis — famous for “Why Most Published Research Findings Are False” (2005) — argued in “Evidence-Based Medicine Has Been Hijacked: a report to David Sackett” (2016, Journal of Clinical Epidemiology) that EBM has been captured by commercial interests: industry runs and funds the influential trials, asks the convenient questions with surrogate outcomes, and even sponsors meta-analyses that reliably reach favorable conclusions. He warned that even Cochrane reviews “may cause harm by giving credibility to biased studies of vested interests through otherwise respected systematic reviews.” The deeper lesson: a hierarchy is only as trustworthy as the integrity of the studies feeding it. “Systematic review” at the top of the pyramid is not a magic word; a review of biased trials launders bias into apparent consensus. Critical thinking about EBM means using the hierarchy and auditing the inputs.
#(d) SUCCESS (individual): Richard Muller’s climate reversal
Richard A. Muller, a physicist at UC Berkeley, was a prominent, publicly skeptical voice on climate science, arguing that flaws in existing temperature records cast doubt on the warming record itself. Rather than merely asserting this, he did the disciplined thing: he co-founded the Berkeley Earth Surface Temperature project (2010) to reanalyze the data independently — funded in part by $150,000 from the Charles G. Koch Charitable Foundation (about one-quarter of the roughly $600,000 project cost), a foundation that had supported efforts opposing mainstream climate science.
The data did not cooperate with his priors. In a New York Times op-ed on 30 July 2012, “The Conversion of a Climate-Change Skeptic,” Muller wrote: “Call me a converted skeptic.” He concluded that global warming was real, that prior estimates of its rate were correct, and — going further than his own earlier position — that “humans are almost entirely the cause.”
Dissect this with the steelman-and-update protocol: Muller (1) took the skeptical hypothesis seriously enough to test it rigorously (steelman); (2) specified in advance what the data would have to show; (3) updated publicly when the evidence went against his prior and against the funding interest. This is intellectual humility operating exactly as it should — belief following evidence rather than identity. (A fair caveat, applying the chapter’s own standard: Berkeley Earth released its findings online before full peer-review publication, which drew methodological criticism — a reminder that even exemplary updating is not immune to process critique.)
#(e) PERSONAL SCALE: the group-chat argument
Consider a realistic exchange. Someone posts: “My uncle took vitamin C megadoses and never got sick all winter. Doctors won’t tell you this because Big Pharma can’t profit from cheap vitamins. Anyone who trusts ‘the studies’ is just a sheep — those studies are all funded by the drug companies anyway.”
Run the protocol. Reconstruct the argument: Conclusion — vitamin C megadoses prevent illness. Hidden premises — my uncle’s experience generalizes; absence of profit motive explains absence of medical endorsement; funding bias invalidates all contrary studies. Spot the fallacies: (1) anecdote / hasty generalization (n=1 uncle, no control for the counterfactual winter); (2) ad hominem / genetic fallacy (“sheep,” and dismissing studies by their funding rather than their content); (3) false dilemma (either trust the uncle or be a mindless sheep). The motivated reasoning: the conspiracy premise conveniently makes the belief unfalsifiable — any contrary evidence is re-cast as proof of the cover-up. Evaluate the evidence: one anecdote sits at the bottom of the hierarchy; the claim of universal funding bias is empirically false (many nutrition trials are publicly funded) and, even where conflicts exist, they must be shown case by case, not assumed. What would change a mind: a large RCT or systematic review. In fact, the Cochrane review (Hemilä & Chalker, 2013, CD000980) found that regular vitamin C (≥0.2 g/day) did not reduce cold incidence in the general population (pooled relative risk 0.97, 95% CI 0.94–1.00) but did shorten cold duration modestly — by about 8% in adults and 14% in children. So the honest verdict is “small reduction in duration, not prevention” — nowhere near the uncle’s claim. The productive response is not “that’s three fallacies” (the fallacy-fallacy trap) but “here’s the strongest version of your concern about incentives — and here’s why one anecdote can’t settle it, and what evidence could.”
#10. Practical Framework — The Claim Evaluation Protocol (Five Questions)
Q1 — What is the claim, exactly? Reconstruct it as an argument: identify the conclusion and the premises, and name the implicit premise. Half the work of critical thinking is done here, because most bad arguments hide their weakest assumption.
Q2 — What is the evidence, and how strong is it? Place the evidence on the hierarchy (anecdote → … → systematic review). Then audit the source: Who is making this claim? What is their expertise and track record? Who benefits if I believe it? What would they lose if it were false?
Q3 — What would the opposite look like? Steelman the best counterargument. Then ask the Bayesian question: What specific evidence would change my mind? If the honest answer is “nothing,” stop — you have found a belief you hold for non-evidential reasons.
Q4 — What is your confidence, and are you overprecise? Give an actual number (Chapter 3: humans are systematically overprecise). “I’m about 70% confident” beats “I’m sure.” Then interrogate the number: is it high because the evidence is strong, or because the conclusion is comfortable?
Q5 — What is the cost of being wrong in each direction? Errors are rarely symmetric. Being wrong about a restaurant recommendation is cheap; being wrong about a medical decision is not. Decide the stopping rule: investigate further, provisionally accept, or suspend judgment — matched to the stakes. Under low stakes, act on decent evidence; under high stakes and asymmetric costs, demand more.
#Printable one-page checklist
THE CLAIM EVALUATION PROTOCOL
[ ] Q1. CLAIM: I can restate it in one sentence. Hidden premise: __________
[ ] Q2. EVIDENCE: Tier = anecdote / case / observational / RCT / review
SOURCE: Who? Expertise? Who benefits? What would they lose if false?
[ ] Q3. OPPOSITE: Best counterargument = __________
What evidence would change my mind? __________
[ ] Q4. CONFIDENCE: ____% . Is it high for evidential or emotional reasons?
[ ] Q5. COST of error: symmetric? asymmetric which way?
DECISION: investigate / accept / suspend
RED FLAGS: strong emotion • urgency to share • "everyone knows" •
unfalsifiable • attacks the person • single anecdote • no named source
#Two-week exercise
Each day, pick one public claim from your feed. Run the Five Questions in writing (5–10 minutes). Record your confidence before investigating and after. At the end of two weeks, review: Where were you overconfident? Which claims did you get right, and did your process (not just your conclusion) hold up? This trains calibration and makes the protocol automatic.
#LLM debate-practice exercise
Take a claim you hold. Ask an LLM to argue the strongest case against it, then the strongest case for it. Now evaluate both outputs with the same protocol — checking every specific factual claim, name, and citation independently (remember Mata v. Avianca). You get three benefits at once: exposure to a steelmanned opposing case, practice auditing machine output, and a check on your own myside bias. Treat the model as a sparring partner whose every punch you verify — never as a referee.
#11. Can It Be Taught? — The Education Evidence
This section deserves its own treatment because the popular claim (“critical thinking is just teachable”) and the research reality diverge sharply.
The anchor is Abrami, Bernard, Borokhovski, Waddington, Wade & Persson (2015), “Strategies for Teaching Students to Think Critically,” Review of Educational Research 85(2):275–314 — a meta-analysis of 341 effect sizes from experimental and quasi-experimental studies using standardized critical-thinking measures. The headline: a weighted random-effects mean effect size of g+ = 0.30 (p < .001), with the collection statistically heterogeneous (p < .001). By Cohen’s conventions (0.2 small, 0.5 moderate, 0.8 large), that is real but modest — and the heterogeneity is the real story: pedagogy drives the variation.
Two moderator findings matter most.
The three-strategy combination. Abrami et al. classified instruction by whether it used dialogue (discussion), authentic/anchored instruction (real-world problems), and mentoring. Individually each helped modestly (when the technique favored the experimental group: dialogue g+ ≈ 0.32, authentic instruction ≈ 0.34, mentoring ≈ 0.39). But mentoring alone was not reliably effective — it “may serve in a catalytic capacity … but is not especially successful if pursued in isolation.” The striking result was the combination of all three (Authentic + Dialogue + Mentoring), which produced g+ = 0.57 (k = 19), significantly larger than dialogue-plus-authentic without mentoring. The practical takeaway: critical thinking is best learned by arguing about real problems under the guidance of a more experienced thinker — not by passive exposure to fallacy lists.
Ennis’s taxonomy. Robert Ennis distinguished four ways to organize instruction: General (CT as a stand-alone course, no subject content), Infusion (CT as an explicit objective within a content course), Immersion (thought-provoking content, but CT not made explicit), and Mixed (a general CT track running alongside content). A crucial correction to a widely-repeated statistic: in the 2015 paper, the four approaches did not differ significantly (Q-between p = .25); Mixed was numerically largest at only g+ = 0.38. The often-cited “Mixed = 0.94” figure comes from the earlier Abrami et al. (2008) “Stage 1” meta-analysis (RER 78(4):1102–1134), a larger and less methodologically restricted dataset in which explicit approaches (infusion g+ ≈ 0.54, mixed g+ ≈ 0.94) beat implicit immersion (g+ ≈ 0.09). The two studies should not be conflated — but both point the same direction: making critical thinking an explicit goal beats hoping students absorb it implicitly.
Argument mapping is the most striking single intervention. Tim van Gelder’s work (summarized in The Palgrave Handbook of Critical Thinking in Higher Education, 2015) reports that intensive computer-aided argument-mapping courses produce gains of roughly 0.8 SD — more than double the typical CT course — with a meta-analysis (Álvarez-Ortiz, 2007) finding ~0.68 SD for courses using some argument mapping and ~0.78 SD for high-intensity mapping practice. Van Gelder estimates a semester of intensive mapping can yield gains comparable to what would otherwise take an entire undergraduate education. The mechanism is deliberate practice with feedback (Ericsson): mapping forces the hidden premises into the open and gives students something concrete to be corrected on.
The debate underneath all this is Willingham vs. Ennis. Willingham (2007, American Educator, “Critical Thinking: Why Is It So Hard to Teach?”) argues critical thinking “is not a skill” that transfers freely; it depends on domain knowledge and practice, and the failure of transfer is why decades of general CT courses disappointed. Ennis and van Gelder hold that there are genuine general skills worth teaching directly. The evidence supports a synthesis: explicit, effortful, feedback-rich practice on argument evaluation helps — but it helps most when anchored in real content you actually know something about.
#12. Criticisms and Limitations
Applying the chapter’s standards to itself:
- The transfer problem. This is the deepest limitation. Training gains often fail to generalize beyond the trained context. Willingham argues bluntly that critical thinking depends heavily on domain knowledge; if so, a general “critical thinking course” may teach less than we hope.
- Domain dependence. You cannot think critically about a field you know nothing about; you lack the background to know which premises are plausible. This tempers the whole “general meta-skill” framing of the chapter.
- The fallacy vocabulary as a weapon. The “fallacy fallacy” (argument from fallacy) is the formal error of concluding that because an argument contains a fallacy, its conclusion must be false — when in fact a badly-argued claim can still be true. Online, fallacy names are routinely used to dismiss rather than understand — a form of intellectual signaling that shuts down inquiry.
- The Dunning-Kruger caution. The famous claim that the least competent most overrate themselves is contested on statistical grounds. Nuhfer et al. (2016, 2017) and Gignac & Zajenkowski (2020, Intelligence, “The Dunning-Kruger effect is (mostly) a statistical artefact”) argue the classic quartile-plot pattern can be reproduced from random noise plus the better-than-average effect and regression to the mean. The debate continues — Dunkel et al. (2023) replicated Gignac & Zajenkowski’s method but still found a weak, statistically significant effect, and Hiller (2023) contests their recoding choices. Bottom line: treat “Dunning-Kruger” as a useful self-check heuristic, not a robustly established magnitude.
- Replication caveats in the debiasing literature. Many social-psychology findings this field draws on suffered in the replication crisis; some debiasing effects are fragile. Present all effect sizes as provisional.
- Cultural and linguistic relativity. The norms of “good argument” taught here are largely those of the Western analytic tradition. Argumentation styles and standards of evidence vary across cultures; the framework is powerful but not culturally neutral.
- The uncomfortable selection effect. Critical-thinking instruction may work mainly on people who already value it. The people most in need of it may be least disposed to adopt it — and since myside bias is uncorrelated with intelligence, we cannot assume the “smart” will self-correct.
No perspective here is offered as absolute. The honest position is that critical thinking is trainable but hard, general but knowledge-dependent, powerful but abusable.
#The skeptical middle: failure modes of over-skepticism
Critical thinking has two failure modes, not one. Under-skepticism (credulity) is obvious. But over-skepticism is equally corrosive and less recognized:
- Nihilistic relativism — “everything is biased, so nothing can be known.” This confuses the presence of some uncertainty with the impossibility of any knowledge, and it flattens the real, large differences in reliability between sources.
- Weaponized skepticism / manufactured doubt. Oreskes & Conway’s Merchants of Doubt (2010) documents how the tobacco industry, and later fossil-fuel interests, deliberately manufactured doubt to delay action — “Doubt is our product,” as one tobacco executive wrote. The strategy is not to win the argument but to “keep the controversy alive” past the point of scientific consensus. Reflexive skepticism is exactly the vulnerability this exploits: a mind that treats all claims as equally doubtful can be paralyzed on demand.
- Analysis paralysis — endless investigation as a substitute for decision.
The cure is a stopping rule (Q5): match the depth of scrutiny to the stakes and the asymmetry of error costs, then act. The ability to reach a warranted conclusion and commit to it is as much a part of critical thinking as the ability to doubt.
#13. Future Directions
AI as threat and tool. The same technology that fabricates citations can help verify them. AI-assisted search, cross-checking, and adversarial LLM dialogue (§10) are emerging verification aids — used carefully, with every output audited.
Provenance over detection. Because detecting fakes after the fact is a losing arms race — deepfake generators improve as fast as detectors , and human deepfake-detection ability is near chance and poorly calibrated (Köbis et al., 2021, iScience: people “cannot detect deepfakes but think they can”; Diel et al.‘s 2024 meta-analysis of 56 papers found manipulated video especially hard to spot and training gains only modest) — the more promising direction is provenance: cryptographically signed content credentials via the C2PA standard (Coalition for Content Provenance and Authenticity, founded by Adobe, Arm, Intel, Microsoft, and Truepic), embedding a tamper-evident history of how a piece of media was made and edited. Adoption is accelerating (Google joined the steering committee in 2024; Cloudflare and camera makers such as Leica, Sony, Nikon are implementing it; OpenAI’s Sora inserts C2PA data). But provenance is fragile: metadata can be stripped, adoption is incomplete, and missing credentials do not prove content is fake. Watermarking (e.g., Google’s SynthID) faces similar robustness limits. The realistic near-term posture is: provenance where available, lateral reading always, and default skepticism toward unsourced media in high-stakes contexts.
Educational reform. The evidence points toward explicit, intensive, dialogue-and-mentoring-rich argument-evaluation training embedded in content domains — not standalone “critical thinking” units. Whether institutions (journals, platforms, courts) can redesign incentives to enforce better epistemic standards at scale — pre-registration, provenance requirements, verification duties like the one Mata imposed on lawyers — is the central open question.
#14. Recommended Resources
Beginner
- Carl Sagan, The Demon-Haunted World (1995). The “baloney detection kit” — the most inspiring introduction to skeptical thinking as a positive, wonder-preserving discipline. Read for the ethos: skepticism married to openness.
- Anthony Weston, A Rulebook for Arguments. A short, practical manual on argument structure. Read to internalize the premises/conclusion mechanics.
- Browne & Keeley, Asking the Right Questions. A workbook for the habit of interrogating claims; excellent for Q1–Q2 of the protocol.
- Steven Novella, The Skeptic’s Guide to the Universe (2018). A modern, media-literate baloney-detection kit from a working neurologist.
Intermediate
- Bergstrom & West, Calling Bullshit (2020). The best current book on evaluating data-driven and statistical claims in the digital age; directly targets p-hacking, misleading graphs, and viral nonsense.
- Oreskes & Conway, Merchants of Doubt (2010). The definitive account of manufactured doubt — essential for the §12 warning about weaponized skepticism.
- Sloman & Fernbach, The Knowledge Illusion (2017). Why we know far less than we think, and how understanding is distributed across communities — the antidote to overconfidence.
- Kahneman, Thinking, Fast and Slow (2011). The canonical dual-process account — but read with the explicit caveat that several priming and social-psych findings it cites have failed to replicate; treat the framework as illuminating, not gospel.
Advanced
- Stanovich, West & Toplak, The Rationality Quotient (2016). The empirical case that rationality is measurable and distinct from intelligence.
- Landmark papers: Abrami et al. (2015, Review of Educational Research) on what teaching works; Stanovich & West (2008, Journal of Personality and Social Psychology / Thinking & Reasoning) on dispositions vs. intelligence; Kahneman & Klein (2009, American Psychologist) on when intuition can be trusted; Pennycook & Rand (2019, Cognition) on reasoning and misinformation; Wineburg & McGrew (2019, Teachers College Record) on lateral reading; and the Mata v. Avianca opinion itself.
- Influential researchers to follow: Stanovich, West, Toplak, Kahneman, Pennycook, Rand, Abrami, Willingham, Caulfield, Wineburg, Novella.
- Courses: the free SIFT / Check, Please! materials (Caulfield) for lateral reading; introductory informal-logic courses for argument structure.
#15. Self-Check
Attempt each from memory before checking anything:
- Reconstruct a claim you recently believed as a formal argument. What were its implicit premises, and which one is actually doing the load-bearing work?
- Why is “appeal to authority” sometimes valid and sometimes a fallacy? What specifically distinguishes the two cases?
- State, from memory, the strongest finding and the strongest limitation of the Abrami et al. (2015) meta-analysis.
- What is the difference between being open-minded and being empty-minded? Why is the asymmetry important?
- Explain why myside bias is called an “outlier bias,” and what its (near-zero) relationship to intelligence implies for the phrase “smart people believe stupid things.”
- Run the Five Questions on one claim in your feed right now, in writing.
- Using the Wakefield case, explain why “published in a peer-reviewed journal” is necessary but not sufficient evidence.
- When should you stop investigating and act? State your stopping rule in terms of stakes and error costs.
Synthesis (not an answer key): If you could reconstruct the hidden premise in Q1, you have the core skill — most arguments live or die on an unstated assumption. Q2 turns on the difference between citing an authority within their domain of demonstrated expertise (legitimate) and citing status or credentials irrelevant to the claim (fallacious) — Walton’s critical questions are the test. Q3’s strongest finding is that training works (g+ = 0.30, and up to ~0.57 when dialogue, authentic problems, and mentoring combine); its strongest limitation is the heterogeneity of results and the transfer problem — gains often don’t generalize. Q4: the open mind considers new evidence and updates; the empty mind holds no position and mistakes indecision for virtue — open-mindedness is actively seeking disconfirmation, not the absence of belief. Q5: myside bias barely correlates with IQ, so intelligence is no shield — which is precisely why the fix must be a procedure you run, not a smartness you possess. If your recall was thin on Q5 or Q7, reread §4 and §9; those are the conceptual and case anchors of the whole chapter.
#Knowledge Card
## Knowledge Card — Critical Thinking
- Core terms:
- Claim: a statement that is true or false.
- Argument: premises offered in support of a conclusion.
- Validity: conclusion follows from premises (form); Soundness: valid + true premises.
- Fallacy: a defect in an argument (contrast: bias, a defect in a mind).
- Evidence hierarchy: anecdote → case series → observational → RCT → systematic review.
- Enthymeme: an argument with an unstated (often load-bearing) premise.
- Myside bias: evaluating evidence in favor of one's prior view; near-zero correlation with IQ.
- Intellectual humility: the operational habit of "I could be wrong."
- Core mental models:
- Reconstruct + steelman before you evaluate; audit the premises, not just the logic.
- Grade evidence by likelihood ratio (Bayesian strength) and by "who benefits?"
- Run the Five Questions; set a stopping rule matched to the cost of error.
- Connections to prior chapters:
- Ch.2 (Biases): fallacies are to arguments what biases are to minds.
- Ch.3 (Probability): confidence must be numeric and is usually overprecise.
- Ch.4 (Bayes): "what would change my mind?" = specifying the likelihood ratio.
- Ch.5 (Behavioral Econ): automation bias = over-trusting AI outputs.
- Ch.1 (Decisions): judge process, not outcome; use journals and checklists.
- Recommended next chapter: Formal Logic & Epistemology (Tier Six) — where the "good enough" argument structure and evidence norms used here are formalized and philosophically grounded.
- One habit to keep: Before believing or sharing anything, ask "what would change my mind?" — and if the answer is "nothing," treat that as a red flag, not a strength.