Evolutionary Pedagogy begins with an unsettling fact observed in the world's best classrooms: students who can recite formulas and score high marks routinely fail the simplest conceptual understanding tests. The author — a school principal with twenty years of frontline experience — argues that this illusion of "understanding" is neither a failure of teachers nor laziness on the part of students. It is a systemic flaw rooted in a faulty premise: the belief that "taught clearly + practiced thoroughly = truly learned."
This book offers a unified alternative explanation. Drawing on Donald T. Campbell's "Blind Variation and Selective Retention" (BVSR), it reframes learning as a miniature evolutionary process — not the transfer of knowledge from the teacher's head into the student's, but Variation → Selection → Retention: ideas, attempts, and errors are continuously generated (Variation), filtered by feedback and application (Selection), and the survivors crystallize into durable, transferable understanding that can be retrieved and deployed (Retention).
This algorithm is then traced across three nested scales — the micro-layer of the individual brain, the meso-layer of classroom culture, the macro-layer of the school ecosystem — and the cross-layer loops that connect them. The unified framework built across this journey is named the Constructivist Learning Ecology (CLE). Along the way, the book develops a set of actionable conceptual tools: Cognitive Niche, Punctuated Equilibrium, Productive Failure, Retrieval Practice, Wait Time, and the teacher's role reframed as "designer of selective pressure."
Written for practitioners, not specialists, this book deliberately replaces academic jargon with plain language, and abstract terminology with vivid cases and next-day-ready tools. Its ambition is to find for education what Darwin found for biology — a unifying principle. The question it answers is not "what to teach," but "how does learning actually happen?"
Keywords: Evolutionary Pedagogy · Constructivist Learning Ecology (CLE) · Variation–Selection–Retention · BVSR · Cognitive Niche · Punctuated Equilibrium · Productive Failure · Micro / Meso / Macro Layers · Selective Pressure
This book begins with a fact that keeps me up at night.
After all these years running schools, I have seen this scene too many times: an excellent teacher delivers a crystal-clear lesson, the board work is immaculate, the example problems are drilled inside out, and the students nod in unison — "Got it." But come the exam, or when the context shifts, those kids who "got it" are exposed. They memorized the principle but never truly understood it. Between "clearly taught" and "truly grasped" lies a crack we have all been pretending not to see.
I have no standing to criticize from the shore. I myself was once that tireless instructor — convinced that if I just explained clearly enough and drilled thoroughly enough, knowledge would naturally flow into students' minds. It took me a long time to see the truth: when a student is merely a passive "receiver," cognition cannot grow.
Something else troubled me even more. For twenty years as a principal, every instructional reform I launched began with the same passion and ended with the same regret. The classroom did change — but before long, everything gradually slid back to where it started. I have also seen more than one school where the administration laid down detailed systems to "encourage innovation": monthly awards for innovative teachers, innovation bonuses — yet as long as the core metric remained standardized test rankings, once the wind passed, teachers eventually returned to their familiar old paths. Not because they didn't want to change. Because two opposing forces were pulling at them, and the deeper one won.
I couldn't figure it out: why does every sincere effort drift back to square one?
The turning point came from an unlikely place — Darwin. When I tried looking at the classroom through the evolutionary lens of "Variation → Selection → Retention," many things that had baffled me suddenly made sense: why errors are valuable, why performance stagnates, why some classrooms buzz with life while others feel dead, why reforms keep circling back to where they began. I realized that learning, the classroom, the school — at their core, they are all ecosystems that evolve on their own. You cannot "assemble" a tree. You can only give it the right soil, light, and water, and let it grow. Education is the same. This book is the result of systematically working through that insight.
Let me be honest: I believe Evolutionary Pedagogy describes something real, but it is by no means a definitive answer, and certainly not yet another educational panacea. Any theory that cannot honestly state its own boundaries is not worth trusting. It is simply a new pair of glasses — I believe it will help you see things that were once blurry, but it will also have its own blind spots. I am writing it for every teacher, principal, and parent who has been stung by that "crack" — so that together we can see it a little more clearly.
One more thing I need to be upfront about. This book was not written before AI arrived — it was written after GPT-4 scored 130 out of 150 on China's college entrance math exam. That is not a technical footnote. When an artificial system can achieve, in a closed domain, what you spent twelve years training humans to do, many things in education that once seemed solid — the definition of "fundamentals," the boundary between "understanding" and "rote memorization," even the question of "what is worth teaching" — all need to be re-examined. This book was written standing at that watershed.
Wang Sai
This book has one central thread and three layers. The central thread is a single algorithm: Variation → Selection → Retention. The three layers are the three scales across which it unfolds — the individual, the classroom, and the school.
If you are a classroom teacher — I suggest starting with Chapter 4 (Variation) and Chapter 7 (Punctuated Equilibrium). These two chapters are closest to your daily classroom reality, using scenes you will recognize. After reading them, flip to Chapter 15 for the first-week launch plan — not ten principles, but three things you can try starting Monday. Chapters 9 and 10 can wait until you have experimented for a week or two. The theory in Chapters 2 and 3 can be saved for summer break — they explain why these things work, but they won't stop you from trying them first.
If you are a school leader — read Chapter 1 first to build a sense of the problem, then jump directly to Chapter 11 (the Macro Layer), Chapter 12 (Three-Layer Cycle), and Chapter 13 (Why Reforms Go in Circles). These three chapters address the systemic challenges you care about most. After those, circle back to Chapters 4–7 for the core mechanisms — you will understand exactly what is jamming those mechanisms at the institutional level.
If you are a parent or researcher — Parents can read Chapter 1, then jump to Chapter 7 ("Stagnation is actually a good sign" — this may be the most anxiety-relieving chapter in the whole book), followed by Chapter 6 and the closing section of Chapter 15. Researchers should read Chapters 2–3 to build the framework, paying special attention to Chapter 8 (Cognitive Niche), Chapter 14 (Boundary Statements), and Appendix D (Gaokao Phase-Space Analysis).
If you are a parent, you do not need to read cover to cover — the theoretical arguments in Chapters 2–3 can be skipped entirely. This streamlined route is enough to help you understand what your child is actually experiencing, and you can use it tonight: Chapter 1 (know that "understood ≠ can do" is not your fault) → end of Chapter 4 parent guide → Chapter 6 ("Don't ask 'did you understand?' — have them explain it to you") → Chapter 7 ("Stagnation is the best data" — this chapter is an anti-anxiety medicine in itself) → Chapter 15 closing. After five chapters, you will have a new pair of "learning glasses" — when you watch your child learn, you won't feel anxiety; you'll feel curiosity.
You don't need to memorize every concept. What this book really wants to give you is a pair of "glasses" — as you read, think of real examples from your own classroom, your child, your school. The moment you start seeing familiar scenes through the lens of Variation, Selection, and Retention, this book has done its job.
It was the fall of 1991. In a Harvard University physics lecture hall, Eric Mazur stood at the podium, staring at the test papers in his hands, brow furrowed.
He was teaching introductory physics — the brightest young minds at Harvard, SAT math near perfect, physics entrance scores among the nation's top. For years, he had taught the traditional way: derive formulas in class, assign problem sets for homework, test calculation skills on exams. Student performance was solid, and feedback was positive. Then one day, a colleague urged him to attend a workshop on "conceptual understanding."
The workshop organizers handed him a test that looked "ridiculously simple." It was called the Force Concept Inventory — the FCI for short.¹ The questions required no calculation whatsoever. They only tested the most basic conceptual understanding of physics. For example:
A typical FCI question:
You throw a ball straight up into the air. After the ball leaves your hand — ignoring air resistance — how many forces act on the ball in mid-flight?
A) Gravity + an upward force
B) Gravity only
C) Gravity + a forward force
D) No force at all
E) None of the above
(Correct answer: B — gravity only. More than half of students chose A.)
Mazur handed this test to his own students. The results shocked him.
More than half of Harvard physics students firmly believed that after the ball leaves the hand, it is still being pushed by an "upward force." These were not students who had never studied mechanics. They had just spent a semester solving complex projectile motion problems with calculus — precise flight times, landing positions, maximum heights. But after all that computation, their conceptual understanding of "force" was indistinguishable from that of someone who had never taken a physics course.
Mazur later wrote a sentence in his book that still keeps countless educators up at night. He said:
"Reciters" — the word is devastatingly precise. These students had memorized the problem-solving procedures, memorized the formulas, memorized the techniques for scoring points on exams. But they had never truly understood the concept of "force." They "got" every step of the derivations — they nodded when the professor wrote them on the board. They "mastered" every calculation problem — plug in the numbers, spit out the answer, score the points. But when Mazur stripped away the math and asked only about concepts — those "got it" and "mastered it" collapsed like a sandcastle.
This is not a story about students who are behind but progressing. This is a story about the best university in the world, the brightest students, the most dedicated teacher — and real understanding still didn't happen.
But Mazur didn't stop at shock. He did something most professors would never do: he changed the way he had taught for twenty years. He flipped his classroom from "I lecture, you listen" to "I ask a question — you think on your own — then discuss with the person next to you — then vote — I see where you're stuck, and then I explain." This became what thousands of teachers worldwide now use: Peer Instruction.¹b
His transformation did not come from theory — it came from that sleepless moment: discovering that after twenty years of teaching, your students had never truly learned. That moment, sooner or later, hits everyone who takes education seriously.
I
Mazur's story is not an isolated case. It is an extreme version of the most pervasive, most invisible, and most underestimated phenomenon in education.
This phenomenon is called the illusion of "I got it" — the teacher believes she explained it clearly, the student believes he understood, but when it comes time to actually use it, the student's mind is empty. It consists of three nested illusions:
Illusion One: Being able to repeat it = understanding.
The teacher asks in class: "Does everyone understand?" Students nod. The teacher asks: "Who can explain it back?" A student raises a hand and does a decent job. The teacher, satisfied, moves on to the next section. But "being able to repeat" and "truly understanding" have no causal relationship. The student simply reproduced, from short-term memory, what the teacher just said. That is not understanding — that is an echo.
Illusion Two: Being able to solve problems = mastery.
After learning a formula, students do ten practice problems and get eight right. The teacher says: "Mastered." But "plugging into a formula to solve problems" is precisely what Mazur's students excelled at — and it is the most dangerous kind of success. When the same concept appears in a different context, students don't recognize it. Because what they memorized was "how to plug in the numbers," not "why this concept explains this phenomenon."
Illusion Three: Being able to pass the exam = having learned.
High scores on the test, everyone is happy. But what the exam tests is "answering a limited set of questions using recently learned knowledge in a highly structured context" — that is not real-world learning. Real-world learning means facing an open-ended problem, selecting from existing knowledge, combining, trying, failing, and trying again. Exam scores, precisely, only measure the former — the part that is best at faking it.
In 1972, cognitive psychologists Craik and Lockhart proposed a framework that still has explanatory power today: the Levels of Processing theory.² They distinguished between "shallow processing" — attending only to the surface form and sound of words ("I've heard the word the teacher just said") — and "deep processing" — attending to meaning and relationships ("How does this concept connect to what I learned earlier?"). Listening to a lecture mostly activates shallow processing. Understanding requires deep processing. The classroom "I got it" is often nothing more than the familiarity of shallow processing.
An observation of my own:
Over the past twenty years, I have witnessed this scene countless times: the teacher finishes explaining a problem, the entire class nods. The teacher asks, "Any questions?" No one raises a hand. Ten minutes later, the teacher tweaks one condition and asks the same question again — more than half the students are stuck. The disconnect between classroom "nod rate" and "transfer rate" is the most universal and most overlooked teaching signal I have ever seen. Nodding does not mean understanding. Nodding only means "I was paying attention just now."
The illusion of "I got it" is so pervasive not because teachers are irresponsible or students are lazy — but because the traditional "lecture → practice → exam" pipeline is structurally designed to bypass any real test of understanding. Understanding needs to be made visible, exposed, challenged. Every link in the traditional pipeline gives students a chance to disguise "recitation" as "comprehension."
II
If you are a Chinese educator reading this, you may be thinking: that is a Harvard problem. Chinese basic education is different. We push harder, drill more intensely, test more frequently — our standard for "understood" is higher.
You are half right. Chinese basic education does have unique strengths in the systematic mastery of foundational knowledge. But the "Reciter" problem is not unique to the West.
PISA data consistently shows Chinese students leading the world in reading, math, and science.³ At the same time, a substantial body of research points to an uncomfortable contrast: Chinese students' performance in "creative problem-solving," "cross-disciplinary transfer," and "self-directed inquiry" does not positively correlate with their PISA rankings.⁴ In short — the scores won, but understanding lost.
"High scores, low ability" is a crude label. It is unfair — it ignores the excellence of countless Chinese students — and it is imprecise, because the issue is not "ability" but the mismatch between evaluation systems and genuine learning. But its very existence is a signal: our teaching system is far too efficient at producing Reciters.
A student who can solve fifty similar problems using a single method has proven the volume of his practice, not the depth of his conceptual understanding. A teacher who can deliver a lesson with impeccable logic and clarity has not proven that students have constructed that logic for themselves. None of this is a question of "right" or "wrong" — it is a question of "enough" versus "not enough."
What is truly unsettling is that we cannot even bring ourselves to ask the question. Which principal dares to ask her teachers — "Are you sure your students really understand?" Which teacher dares to ask her students — "You didn't just memorize that, right?" The question itself makes us flinch. We vaguely know the answer is not good, so we don't ask. But the data is not waiting.
In 1991, Eric Mazur discovered the "Reciter" at Harvard University.
Thirty-three years later, the problem has not disappeared — it has simply hidden itself inside the nodding heads mouthing "I got it" in every classroom.
In an ordinary physics class, the teacher is lecturing on buoyancy. The explanation is clear, the board work is neat, and the example problems are worked through repeatedly. Five minutes before the bell, the teacher draws a simple diagram on the board — a block of wood floating on water — and asks: "The buoyant force on this block equals what?" The whole class answers in unison: "The weight of the displaced water!" The teacher nods, satisfied.
Then he changes the question. He replaces the beaker with half a cup of water and asks: "On the Moon, what is the buoyant force on this block?" The room falls silent for three seconds. Of the students who had just answered in perfect chorus, not a single hand goes up. Not because they don't know Archimedes' principle — they had just recited it flawlessly. They had simply never considered: does Archimedes' principle still hold at one-sixth of Earth's gravity? They memorized the principle but never understood the conditions under which it holds.
This is not an exception. This is the silent, unspoken "I don't know" buried beneath the chorus of right answers in every Chinese classroom that is "taught clearly, drilled thoroughly, and tested well." Mazur is not fighting alone — he is just the first person who wrote it into a paper.
▎ Why are these phenomena so universal?
If the crack between "taught clearly" and "truly understood" only existed in one particular classroom or one particular teacher's class, we could chalk it up to individual problems. But Mazur's story happened at Harvard, with the best students in the world. The illusion of "I got it" appears in any classroom with lectures, problem sets, and exams. The debate about "high scores, low ability" has persisted for two decades and shows no sign of fading away.
This tells us — the crack is not an isolated phenomenon. It is systemic.
The problem lies in our educational assumptions. We assume that "clear teaching + thorough practice = deep learning." But this assumption is shattered in the face of FCI data. We need to know why. And traditional educational theories — behaviorism, cognitivism, even classical constructivism — honestly, they cannot explain this contradiction.
Next chapter, we take them apart: why traditional explanations fall short.
¹ Hestenes, D., Wells, M., & Swackhamer, G. (1992). Force Concept Inventory. The Physics Teacher, 30(3), 141-158. The FCI later became one of the most widely used conceptual diagnostic tools in physics education research, administered at hundreds of universities worldwide.
¹b Mazur, E. (1997). Peer Instruction: A User's Manual. Prentice Hall. Spurred by the FCI data, Mazur flipped his Harvard physics classroom and developed Peer Instruction. Its core principle — having students think independently before discussing with peers — has since been validated by extensive empirical research.
² Craik, F. I. M., & Lockhart, R. S. (1972). Levels of processing: A framework for memory research. Journal of Verbal Learning and Verbal Behavior, 11(6), 671-684. While the theory has been tested and revised over the decades, its core distinction — shallow vs. deep processing — remains the most concise framework for understanding why "listening to a lecture" does not equal "learning."
³ OECD. (2019). PISA 2018 Results. The China consortium (Beijing-Shanghai-Jiangsu-Zhejiang) ranked first in reading, mathematics, and science.
⁴ See the body of work by Zhao Yong, especially Who's Afraid of the Big Bad Dragon? (Jossey-Bass, 2014), for a systematic analysis of the strengths and limitations of Chinese education. See also OECD reports on PISA creative problem-solving assessments.
Chapter 1 left us with a question: why does "taught clearly + practiced thoroughly" not equal "learned deeply"? Have traditional educational theories — behaviorism, cognitivism, constructivism — provided an answer?
The answer: each has explained part of the phenomenon, but none has touched the underlying algorithm of learning.
In this chapter, we take them apart. For each theory, we first give it three minutes of respect — what it actually explains, where its contribution lies. Then we ask: can it explain the "Reciters" in Mazur's classroom?
I
Behaviorism is probably the oldest and most stubborn shadow in education. Every teacher who has ever assigned "copy this problem three times" is a behaviorist — whether they have ever heard Skinner's name or not.
What did it contribute?
Behaviorism grasped the most critical point: behavior can be shaped by the environment.¹ Skinner's operant conditioning proved: when a behavior is rewarded, its frequency increases; when punished, it decreases. This is astoundingly efficient for training basic skills — multiplication tables, phonetic decoding, Chinese character writing, keyboard typing — any foundational skill that requires high-frequency repetition and standardized output falls squarely within behaviorism's domain.
Reciting multiplication tables — recite it eighty times if eight doesn't work, until it becomes a conditioned reflex. The "muscle memory" of Chinese character writing — stroke order and structure solidified through repetition. Any skill that requires high-frequency repetition and standardized output is adequately covered by the behaviorist framework.
But what can't it explain?
Let's return to the core question from Chapter 1: Mazur's students could solve complex calculus problems — a massive behavioral success — but their minds went blank when facing the question "what forces act on a ball in mid-air?" Why?
Because within the behaviorist framework, "understanding" is not a concept that can be defined or measured. Behaviorism cares only about inputs (stimuli) and outputs (responses). Whatever happens inside the student's mind — whether they truly understand the concept of "force" — it does not care, and it does not need to care. A student who can solve calculus problems correctly has, from the behaviorist's perspective, "already learned." Whether they grasped the physics concepts — that question has no place in behaviorism.
This is behaviorism's fundamental limitation: it explains how "training" changes behavior, but not how "understanding" changes thinking.
Many of the most educationally significant phenomena in the classroom — a student suddenly "gets it," another reorganizes a concept in her own words, a third discovers a connection between two seemingly unrelated ideas — none of these can be described as "stimulus → response." They happen inside students' minds, and behaviorism is precisely a theory that refuses to look inside the mind.
A student memorizes all the mechanics formulas but fails the FCI's basic conceptual questions — behaviorism cannot explain this crack. Because formula recitation and conceptual understanding both produce the behavioral output of "getting the answer right," yet their cognitive mechanisms are entirely different. Behaviorism simply lacks the language to describe "cognitive mechanisms."
So, if we use only the behaviorist framework, the "Reciters" in Mazur's classroom are not even a problem. Their exam scores are fine — behaviorally, they meet the standard. But our intuition tells us something is wrong. Whether a student "understands" matters more than whether they "answered correctly." But behaviorism cannot tell us why.
II
If behaviorism refuses to look inside the mind, then looking inside the mind — that is cognitivism's starting point.
What did it contribute?
The "cognitive revolution" of the 1950s and 1960s put the brain back at the center of research. Cognitive psychology proposed a computer metaphor: the brain is an information processor — information enters, is encoded and stored, and is retrieved when needed.²
This framework yielded two enormously important findings. The first is the capacity limit of working memory — Miller's (1956) famous "magical number 7±2" tells us the human brain can process only about seven chunks of information at any one time.³ Sweller built on this to develop Cognitive Load Theory, revealing a fundamental constraint on instructional design: you cannot dump all the information on students at once, because their working memory will crash.⁴
The second contribution is the distinction of levels of processing — already cited from Craik & Lockhart (1972) in Chapter 1 — shallow processing (repetition, rehearsal) and deep processing (semantic, relational) produce radically different memory outcomes. This explains why "going over it ten times" does not equal "learning it."
Why do students "zone out" after forty minutes of lecture? — Working memory is saturated. Why does "understanding" last longer than "rote memorization"? — Deep processing produces more durable memory traces than shallow processing. Why can't a single class session cram in too much content? — Cognitive load exceeds the learner's processing capacity. These findings are now basic common sense in instructional design.
But what can't it explain?
Cognitivism's computer metaphor carries a hidden premise: knowledge can be "transmitted." Information is encoded in the teacher's mind, "transmitted" via language into the student's mind, decoded, stored, retrieved. The problem is — if knowledge really could be transmitted like data, Mazur's students should not exhibit the "can solve it, don't understand it" phenomenon. The definition of "force" was transmitted — in fact, more than once — yet the concept of "force" never took hold in their minds.
Why?
Because knowledge is not data, and understanding is not copying. When a teacher describes a concept in words, what the student hears is not "the concept itself" but "their own interpretation of that language." The teacher's "force" and the student's "force" may be entirely different things — the teacher is thinking Newton's Second Law F=ma, the student is thinking "you need to push something to make it move," a piece of everyday experience — and yet they are both using the same word. In the transmission model, this counts as "successful transmission." In real learning, nothing has happened at all.
Why does the same teacher, using the same textbook, explaining the same problem, produce completely different things that different students "hear"? The information-processing model assumes "the input is fixed" — but what students actually take in is not the content the teacher transmitted. It is something the student constructed themselves. Cognitivism describes the processing that happens after this "construction," but it does not explain how the "construction" itself occurs.
Cognitivism looked "inside the mind," but what it saw was a factory already up and running — information flows in, is encoded, stored, and retrieved. But how was the factory built? From what raw materials, through what processes, under what conditions did it "grow"? Cognitivism does not answer.
Behaviorism looks only at behavior. Cognitivism sees cognition. But both share a default premise: knowledge already exists, and the job of teaching is to "deliver it inside" — the only difference being that behaviorism delivers through training, cognitivism through optimized encoding. Yet Mazur's story tells us: delivered it was. Received, it was not.
III
Constructivism offered a fundamentally different answer: knowledge is not "delivered in" — it is "grown" by the student.
What did it contribute?
Constructivism says: there is no such thing as "knowledge transmission." The learner is not a blank slate — they enter the classroom with existing cognitive structures and use those structures to interpret new experiences. When new experience aligns with old structures, it is smoothly assimilated (Piaget's assimilation); when new experience conflicts with old structures, the old structures must be adjusted or even rebuilt — this is accommodation.⁵ The essence of learning is not receiving information; it is actively constructing meaning.
Vygotsky approached from a different direction: learning is not an isolated individual act but happens within social interaction.⁶ He identified that each student has a Zone of Proximal Development — the region between what they can do independently and what they can do with help. Teachers and peers play a "scaffolding" role within this zone: not providing answers, but providing support that lets the student reach further than they could alone.
These two directions — Piaget's individual construction and Vygotsky's social construction — together form what education today calls "constructivism." Its core claims are now widely accepted: students are active, knowledge is constructed, context matters.
Why, after a hundred explanations, do students still "not get it"? — Because "getting it" is not information reception; it is meaning construction. The same information produces different meanings in different students' cognitive structures. Why is peer discussion often more effective than listening to a lecture? — Because social interaction forces students' understanding to be exposed, tested, and revised. The spread of these insights through the education world represents enormous progress.
But what can't it explain?
Constructivism offers a profound philosophical stance — "knowledge is constructed" — but it fails to give precise answers on two critical questions.
Question One: How does construction actually happen?
Piaget says "assimilation and accommodation," but this is more a naming of the phenomenon than a description of the mechanism. When a student faces a new concept that confuses them, what actually happens inside their brain? What determines whether a failed "assimilation" leads to "accommodation"? What factors tip the conflict between old and new cognitive structures toward deeper understanding rather than toward giving up and going numb? The textbooks say "construction is active" — but why do so many students in class clearly appear to be "listening" yet not engaged in "constructing"?
This is the question I most want to pose to classical constructivism. It tells us knowledge is "grown," but it does not tell us what the underlying code for "growing" is. How a tree grows from a seed into a full tree is clear — biological principles trace it down to the molecular level. How learning goes from "not understanding" to "understanding" remains, in classical constructivism, a black box.
Question Two: What kind of construction is effective?
Not all construction leads to better understanding. A student may have "constructed" an extremely stubborn misconception — for example, "force is the cause of motion" (Aristotelian mechanics). This misconception explains a great deal of everyday experience (you push, it moves), so it is well "constructed" within the student's cognitive framework. And it is remarkably stubborn — even after studying Newton's laws, many students still hold Aristotelian understandings at the conceptual level. This problem has been repeatedly verified in FCI data.⁷
Classical constructivism acknowledges the existence of prior conceptions but offers no operational framework for "making effective construction happen." It says "teachers should provide a rich environment" — which is entirely correct but far too vague. Rich to what degree? What kind of environment triggers effective construction? What kind of environment, paradoxically, reinforces misconceptions?
Classical constructivism explains that "knowledge is constructed" but not "what the algorithm of construction is." It tells us students are not empty vessels — but how exactly do their minds operate? When a student in Mazur's class passes the exam through "recitation" — does that count as a kind of "construction"? From a constructivist standpoint, in some sense, yes — the student built a cognitive structure of "solving problems = scoring points." But that is clearly not the kind of construction we want. Where is the difference? Classical constructivism lacks the precise language to distinguish "effective construction" from "ineffective construction."
IV
Now let's place the three theories side by side and see what each saw and what each missed.
Three Theories of Education: A Comparison
Do you notice something?
The three theories are almost completely aligned on the row labeled "Attitude toward error": in traditional frameworks, error is something to be eliminated.
Behaviorism sees error as behavior to be corrected — reinforce the correct response, extinguish the incorrect one. Cognitivism sees error as a glitch in processing — optimize the encoding strategy, reduce the "mistakes." Constructivism is slightly gentler — seeing error as a "mismatch" in cognitive structures — but its ultimate goal is still "matching": replace the wrong structure with the right one.
The three theories' biggest shared blind spot:
Error, in traditional frameworks, has only negative value.
But Mazur's "Reciters" failed to build genuine understanding precisely because they never made errors. They kept solving correctly, reciting correctly, scoring correctly — and it was this very "correctness" that masked the absence of understanding.
Let's return to the stinging fact from Chapter 1: Mazur's students were not struggling learners. They were the most diligent learners and the most successful test-takers. If error is "something to be eliminated," they should have been the most successful students — because they barely ever made exam errors. But it was precisely because they were never exposed, challenged, or forced to "get it wrong" at the conceptual level that they never truly constructed an understanding of "force."
This paradox — without errors, there is no real learning — cannot be explained by any of the three traditional theories. Because their shared premise is: the correct goal is already known, and learning is simply moving the learner from "incorrect" to "correct." In this framework, error is a roadblock to cross. But in genuine cognitive development, error is the road itself.
▎ The Dichotomy Trap: "Skill" and "Knowledge" Are a Spectrum, Not a Split
So far we have critiqued each theory's individual blind spots. But they share a deeper assumption: that educational goals can be divided into "lower-order skills" and "higher-order thinking." Bloom's taxonomy, Marzano's new taxonomy, Piaget's "concrete operations → formal operations" — a century of educational taxonomies has revolved around this binary.
And then AI arrived. GPT-4 solves Gaokao math — not because it "understands" mathematics, but because its hundreds of billions of parameters cover this closed space. When computing power is no longer the constraint, the wall between "understanding" and "memorization" collapses. What we used to treat as differences of "cognitive kind" — skill versus knowledge, training versus construction — are, under sufficient computation, simply different parameter settings for the same BVSR algorithm.
The human brain was never the problem. Pedagogy has simply been mistaking the computational limits of the human brain for the natural boundaries of cognition. AI tore that boundary down. As for the measurement framework to replace that old dichotomy — the Three Rulers — we will find a more suitable place for it when we reach the chapter on Retention (Chapter 6).
V
Behaviorism explains "training" but cannot explain "understanding." Cognitivism reveals the mechanism of "processing" but cannot answer how "construction" happens. Constructivism acknowledges the active nature of knowledge but lacks a precise description of the "algorithm of construction."
Each of the three theories holds a piece of the truth. But they all sidestep the same fundamental question:
If such an algorithm exists, it must satisfy three conditions:
First, it must be able to explain the difference between the "Reciter" and the "Understander" — not because the former is lazy or the latter smarter, but because the two kinds of learning follow different speeds and qualities of the same fundamental process.
Second, it must give "error" a positive role — an explanatory framework that not only tolerates error but makes error its core driver.
Third, it must be cross-level — it must explain both the micro-level changes inside one student's brain, and how a class's culture takes shape, and why a school's reform succeeds or fails.
Three puzzles, one algorithm — just running at different scales.
One framework has done exactly this. It comes from an unexpected source — not education, not psychology, but biology.
It has already explained 3.8 billion years of life's evolution. It has explained the growth of scientific knowledge. It has explained cultural evolution. Yet almost no one has systematically applied it to education.
Before we go further, I want to make one thing clear. What this book aims to build is a new lens for understanding learning — Evolutionary Pedagogy. The theoretical framework that supports it, which I call CLE (Constructivist Learning Ecology), belongs to the constructivist tradition. It inherits the deepest insights of Piaget, Vygotsky, Dewey, and others: knowledge is not passively received; it is actively constructed by the learner. This position does not change.
But what CLE sets out to do is what classical constructivism left unfinished: to provide "construction" with an operational evolutionary algorithm. Classical constructivism made clear that "knowledge is constructed" — but never explained "how construction operates." CLE uses the three verbs Variation → Selection → Retention to fill this missing mechanism layer. In this sense, CLE is both a deepening of constructivism and a transcendence of it — not a negation, but a completion.
And here, an unexpected source of corroboration appears — AI.
How do large language models like GPT-4 learn? The training loop is strikingly simple: predict → compare → correct → predict again. If you have studied Piaget, these four steps do not need translation: prediction is assimilation, error is cognitive conflict, correction is accommodation, convergence is equilibration. This is not just an analogy — it points to the same underlying logic, independently realized on two completely different material substrates. To be precise: the gradient descent of machine learning and Piaget's "cognitive conflict → accommodation" differ in the mechanism of "how to correct," but at the higher structural level, they share the same fundamental shape: predict → encounter obstacle → revise → converge. This structural isomorphism deserves serious attention.
A silicon-based system, with no preset "understanding," no prior knowledge framework — through nothing but a constructivist training loop, learned language, reasoning, and mathematics. What does this tell us?
Construction is not one school of educational philosophy. Construction is the foundational law of intelligent systems.
AI did not "believe" in constructivism — it simply proved, with mathematics: if you want to learn anything, the only path is construction.
This is the greatest gift AI has given education: constructivism is no longer just a psychological hypothesis Piaget proposed a century ago while watching children play with blocks on the shores of Lake Geneva — it has received independent engineering verification. A system with a physical architecture entirely different from the human brain took the same path and arrived at the same destination. What this tells us is not "AI learns like humans" — it is that the algorithm of human learning is not uniquely human. It is the mechanism that any system capable of "Variation → Selection → Retention" will naturally converge upon.
If the algorithm of learning is not unique to humans, then it must have a more fundamental name. In the next chapter, we go find it.
▎ Next chapter, we turn our attention to this framework
We are not "borrowing a biological metaphor" — as in, "learning is like evolution," an analogy. We are saying: learning and evolution are the same thing — the same algorithm, running on different material substrates (DNA and neural networks).
Next chapter: In Search of a Unifying Principle — From BVSR to Evolutionary Pedagogy
¹ Skinner, B. F. (1953). Science and Human Behavior. Macmillan. Skinner's operant conditioning is the theoretical foundation of all behaviorist applications in education, and its influence persists to this day — from classroom token economy systems to the instant-feedback design of digital learning platforms, the fingerprint of behaviorism is everywhere.
² Atkinson, R. C., & Shiffrin, R. M. (1968). Human memory: A proposed system and its control processes. Psychology of Learning and Motivation, 2, 89-195. The classic multi-store model — sensory memory, short-term memory, long-term memory — in its three-tier architecture.
³ Miller, G. A. (1956). The magical number seven, plus or minus two: Some limits on our capacity for processing information. Psychological Review, 63(2), 81-97. The foundational study on working memory capacity.
⁴ Sweller, J. (1988). Cognitive load during problem solving: Effects on learning. Cognitive Science, 12(2), 257-285. The original paper on Cognitive Load Theory. Sweller pointed out: if instructional design ignores the capacity limits of working memory, even the most "content-rich" instruction can backfire.
⁵ Piaget, J. (1952). The Origins of Intelligence in Children. International Universities Press. Assimilation and accommodation are the core mechanisms of Piaget's theory of cognitive development. Interestingly, Piaget himself was a "genetic epistemologist" — he was less concerned with "how to teach" than with "how knowledge emerges from nothing." This perspective shares a deep affinity with the evolutionary approach of this book.
⁶ Vygotsky, L. S. (1978). Mind in Society: The Development of Higher Psychological Processes. Harvard University Press. The classic texts on the "Zone of Proximal Development" and the concept of "scaffolding."
⁷ Hestenes, D., Wells, M., & Swackhamer, G. (1992). Force Concept Inventory. The Physics Teacher, 30(3), 141-158. FCI data show that even after formal physics instruction, large numbers of students still hold the Aristotelian misconception that "force is the cause of motion." This indicates that "constructed" misconceptions are extremely stubborn — they cannot be corrected by simply "explaining it again."
▎ If you want to skip this chapter — jump to Chapter 4 (Variation). This chapter explains "why learning is evolution inside the human brain." It is an argument, not an instruction manual. Come back to it once the case studies in the later chapters have persuaded you.
If learning is evolution inside the human brain, then the answer should not be sought within education alone — or at least not only within education.
At the end of the previous chapter, we identified a gap: behaviorism, cognitivism, constructivism — each theory holds a piece of the truth, but none has answered the fundamental question: does learning have a unified underlying algorithm?
A familiar scene: After the final exam, two students in the same class both score 72. But Student A got there by memorizing templates; Student B got there by understanding the derivational relationships between formulas. The test paper cannot tell them apart. But when the new semester begins and new content is introduced, Student A stalls in the first week. Student B does not. The underlying algorithm running in these two students' brains is completely different. So what exactly is this "underlying algorithm"?
The time has come to go find it.
I
In 1859, Darwin did something remarkably "simple."
More than a century later, evolutionary biologists condensed his insight into three verbs — and with them, explained 3.8 billion years of life's history.
No complex equations. No theoretical framework requiring a decade to master. Three verbs: Variation. Selection. Retention.
A process driven by these three verbs — with no purpose, no direction, no designer. Yet from the first primitive cell, it produced the blue whale, the human being, and the brain capable of contemplating its own existence.
Pause and let the weight of that sink in: a process with absolutely no "plan" has produced the most complex structures the human mind can imagine.
The evolutionary biologist Dobzhansky wrote a line that has been quoted to death: "Nothing in biology makes sense except in the light of evolution."¹ Applied to learning, the sentence holds equally true.
A simple thought experiment:
If what runs in the classroom and what has run on Earth for 3.8 billion years are the same algorithm —
then our "teaching" and "learning" are simply a special form of evolution.
Not a process "like" evolution — it is evolution.
Once this lens is in place, many things suddenly find new anchor points.
II
For more than a century after Darwin, most people assumed "natural selection" was a biological affair.
That changed in 1974, when psychologist Donald T. Campbell published a paper with a boldly titled name: Evolutionary Epistemology.²
Campbell's core argument comes down to a single sentence: every process of knowledge growth — biological evolution, scientific discovery, individual learning, cultural innovation — runs the same algorithm.
He called this algorithm BVSR: Blind Variation and Selective Retention.
Notice the word "Blind." Campbell deliberately added it because he wanted to emphasize: at the moment of generation, variation does not know which direction is correct. If variation already "knew" the answer, it would not be variation — it would be transportation.
What does Campbell's BVSR imply? In one sentence: "learning" is not a uniquely human privilege.
It means: when a student racks her brain trying a clumsy approach, a microorganism undergoes a random mutation, a scientist proposes a new hypothesis — structurally, these three events are the same thing. Their differences lie not in "whether thinking is involved" (microorganisms do not think) but in the material substrate on which the algorithm is implemented: DNA, neural networks, or scientific communities.
This framework brings about a fundamental shift in perspective:
Not "learning is like evolution" —
Learning is evolution inside the human brain.
Learning is not downloading. It is evolving.
After Campbell's paper appeared, Popper immediately picked up the thread. Popper's "conjecture and refutation" — the proposal of scientific hypotheses (Variation) and their being overturned by experiment (Selection) — is BVSR in the specific domain of science.³ Popper even went so far as to say: from the single-celled organism to Einstein, the pattern of knowledge growth is the same.
The same? Einstein and a bacterium?
Yes. Popper's logic: bacteria in a sugar solution — some swim left, some swim right — that is Variation. Whichever side has the sugar, those that swim that way survive — that is Selection. Einstein proposing relativity — that too is a pile of "conjectures" (Variation) being filtered by the "selective pressure" of experimental data. The difference is, the bacterium uses flagella and Einstein used thought — but the underlying algorithm is the same.
But Campbell himself acknowledged: BVSR is a universal framework, not an education-specific toolkit. It tells you "learning is evolution," but it does not tell you, when teaching a specific student, what to do next.
That is exactly what this book aims to do.
III
Before entering CLE, we need to honestly confront a core tension.
Campbell's BVSR deliberately added the prefix "Blind" to emphasize: when variation is generated, it does not know which direction is correct. If variation already "knew" the answer, it would not be true variation — it would be directed search.
But "variation" in education — the various attempts a student makes when facing a new problem — is never entirely blind. A student does not generate answers randomly from scratch. Her attempts are always constrained by prior knowledge, teacher guidance, and problem framing. This is "guided variation."
The question arises: if variation in CLE is "guided," can it still be called "variation" in the BVSR sense?
The answer is yes — but "blind" needs to be re-understood.
Campbell's "blind" does not mean variation is completely unconstrained.
DNA mutations are also constrained by chemical laws — not all mutations are equally probable. Campbell's "blind" means: the variation mechanism itself does not foresee the selection outcome. After variation is generated, the environment decides what survives. The variation mechanism does not happen "for the sake of" some outcome — it simply happens.
In this sense, educational "guided variation" is still blind. A teacher can design a problem that steers students' thinking in a certain direction — but the teacher cannot predict which specific attempt will succeed. Within that direction, students are still probing blindly. Problem design narrows the search space of variation (from "guess everything" to "guess what's relevant"), but it does not eliminate the blindness of variation.
So CLE's modification of BVSR is not turning "blind variation" into "sighted variation," but rather:
Why are these modifications necessary? Because education is not natural evolution — we have forty minutes per class. Natural evolution can afford millions of years. A classroom cannot. Guided variation lets the BVSR algorithm run on an actionable timescale, and this is the key to CLE's upgrade from a "descriptive framework" to a "designable framework."
These two modifications are, in fact, the fundamental purpose of this book: not to describe how learning "happens to" occur, but to help teachers design the conditions under which learning happens more efficiently.
IV
BVSR is a master key. The problem, precisely, is that it is too masterful.
It explains what is common to all knowledge-growth processes, but it does not answer three core educational questions:
First, how does an individual's "cognitive variation" arise? When a student faces a new problem, why doesn't she "guess randomly" but instead produces specific types of attempts? What kind of instructional design causes students to generate more valuable variations?
Second, how does a classroom's "selective pressure" form? A teacher's remark, a classmate's glance, a score on a test paper — how do these feedback signals shape students' cognitive structures in a real instructional environment?
Third, how does a school's "ecological environment" evolve? Why do so many educational reforms — with sound principles, sufficient evidence, and proper execution — ultimately fail? What happens inside classrooms that a department head cannot dictate? Why are principals' reforms always dragged back to square one by "slow variables"?
These three questions, BVSR does not answer. They require a set of theoretical tools purpose-built for educational settings.
At the end of the previous chapter, the example of GPT-4 showed us: construction is the foundational law of any intelligent system. But a "foundational law" needs a formal name and structure — and CLE is that name.
This is what I set out to do in this book — on the universal framework of BVSR, re-examine the three core dimensions of education. I call this theoretical framework CLE (Constructivist Learning Ecology) — and it is the kernel of the "Evolutionary Pedagogy" perspective.
Three new terms are about to appear — Cognitive Niche, Punctuated Equilibrium, Three-Layer Nesting. For now, build a general impression; you do not need to memorize them all here. Each subsequent chapter will unfold them one by one, grounded in specific educational settings.
These three lenses are not parallel perspectives; they unfold along three different dimensions: Cognitive Niche defines the space in which variation occurs (within the individual mind), Punctuated Equilibrium describes the temporal rhythm of the process, and Three-Layer Nesting reveals that the algorithm operates simultaneously at multiple scales. In real learning scenarios, they happen all at once:
A student (micro-layer) encounters an open-ended problem and generates various attempts (Variation). The teacher's timely feedback and peer discussion form the classroom's "selective pressure" (meso-layer), filtering out more effective approaches. If the school's evaluation system is "only the final exam score matters" (macro-layer), then the teacher dare not give students much time for trial and error — the meso-layer's selective pressure is distorted by the macro-layer's assessment regime. In the end, at the micro-layer, the student can only select "test-taking strategies" rather than genuine understanding.
This is why reforms that change only the classroom but not the evaluation system always fail. You change one layer, and the other two layers pull you back.
These three lenses will run through the entire book. The upcoming chapters will unfold along the central axis of BVSR — Variation (Chapter 4), Selection (Chapter 5), Retention (Chapter 6) — one by one, then re-examine everything through the frameworks of Punctuated Equilibrium (Chapter 7) and the Three Layers (Chapters 9–11), and finally return to boundaries and operational principles (Chapters 13–15).
Placing BVSR and CLE side by side, the differences become clear:
| Dimension | BVSR (Campbell) | CLE (This Book) |
|---|---|---|
| Scope | All knowledge-growth processes | Specifically for educational settings |
| What is Variation? | Blind, random attempts | Active probing within a cognitive niche — modulated by problem design and error-tolerant culture |
| Who performs Selection? | Physical environment / natural selection | Teacher-designed selective pressure (feedback, commentary, peer interaction) |
| Vehicle of Retention | DNA / genes | Neural plasticity / cognitive structures |
| Time scale | Generational transmission | Minutes to semesters (three layers, three speeds) |
| Can it explain stagnation? | No | Yes — Punctuated Equilibrium: plateau periods are systemic reorganization |
| Can it guide instructional design? | No (universal framework) | Yes — what kind of selective pressure each variation needs, what kind of intervention each layer requires |
CLE is not a negation of BVSR; it is BVSR's educational implementation. BVSR tells you that "Variation → Selection → Retention" exists. CLE tells you: in a real classroom, how this algorithm actually runs — what accelerates it, what blocks it, and how to design for it simultaneously across all three layers.
V
We now have our theoretical anchor. We have a map of the three layers. One clarification is needed: the proposition that "learning is evolution" is offered here as a working hypothesis, not an established theorem. BVSR provides a unifying framework; CLE adapts it to educational settings. But the ultimate test lies not in theoretical deduction — it lies in the classroom. What follows in this book is a sustained effort to test this hypothesis using the cases and evidence in every chapter.
Now it is time to return to the fundamentals.
The "V" of Variation, the "S" of Selection, the "R" of Retention — three verbs. Part Two of the book takes them apart, one at a time.
We start with the simplest one — Variation.
It is not error.
It is the raw material of learning.
▎ Next chapter, we begin with Variation
If learning is evolution, then Variation is the first step — without variation, there is nothing for selection to act upon. But what does "Variation" mean in the classroom? How is it permitted, provoked, and designed? Why does a classroom that forbids mistakes strangle the core engine of learning?
Next chapter: Variation — The Source of Cognitive Diversity
¹ Dobzhansky, T. (1973). Nothing in biology makes sense except in the light of evolution. The American Biology Teacher, 35(3), 125-129. This sentence is widely cited, but few notice the year Dobzhansky published it: 1973 — just one year before Campbell's evolutionary epistemology. It was the eve of the "evolutionary worldview" expanding from biology into cognitive science.
² Campbell, D. T. (1974). Evolutionary epistemology. In P. A. Schilpp (Ed.), The Philosophy of Karl R. Popper (pp. 413-463). Open Court. The classic formulation of BVSR (Blind Variation and Selective Retention). Why did Campbell emphasize "blind"? Because if variation already "knows" the answer at the moment of generation — then it is not true variation but directed search. All genuine innovation comes from "blind" variation. The educational implication of this distinction is profound: when we assign students "exploratory tasks" that in reality have only one standard answer waiting for them, that is not variation — it is a treasure hunt with a map.
³ Popper, K. R. (1972). Objective Knowledge: An Evolutionary Approach. Clarendon Press. Popper's "conjecture and refutation" methodology is an independent expression of BVSR within the philosophy of science. His famous claim that "from the amoeba to Einstein, the growth of knowledge is always the same" is often misread as "Einstein is the same as an amoeba" — Popper's real meaning is: the underlying algorithm is the same, but the complexity of its implementation differs by orders of magnitude.
The first three chapters laid out the problem. From this chapter onward, we enter the core mechanism—Variation → Selection → Retention—examining each link one at a time.
The first link to unpack is "Variation."
To most educators, the word "variation" probably seems unrelated to learning. Variation is a biological concept—genetic mutation, population divergence, species evolution. In the world of education, we talk about "standards," "norms," and "correct answers." Variation sounds like evidence of educational failure.
But if we are willing to shift our perspective, a new chapter in the history of education opens.
Have you ever seen a student like this? Faced with a new problem, he first scribbles on a piece of scrap paper for a while, writes down an answer, crosses it out, writes another—and what he finally hands in is wrong. But when you look at his scratch paper, you realize: he's not incapable. He is trying.
This "trying" is the starting point of learning. The first step of biological evolution is variation—errors that occur during gene replication, producing offspring that differ from their parents. Without variation, natural selection would have no raw material. The logic of learning is identical: a student who only repeats known answers will never see his cognition evolve. The impulse to make mistakes is the instinct to learn.
Kapur calls this "productive failure"—letting students try to solve a difficult problem before they have been taught how. They will, of course, fail. But it is precisely during those twenty minutes of struggle that their minds generate abundant "cognitive variation": Could this approach work? Is that method viable? Oh, so that path leads nowhere… When the teacher explains later, their brains are already "primed" with failure signals—no longer blank receivers, but detectives searching for answers with questions in hand.
You might think "make mistakes → fail → learn" is common sense. Yet the vast majority of classrooms around the world are designed precisely to eliminate "making mistakes" as something to be eradicated. This is not merely a flaw in instructional design; it is a fundamental misunderstanding of the nature of learning. A classroom that forbids mistakes cuts off the very first step of learning.
Theoretical Framework · 30%
Within the CLE framework, "cognitive variation" is simply defined as: the various spontaneous attempts a student generates when confronting a new problem—guesses, intuitions, possible approaches, erroneous assumptions, and even "outrageous" ideas.
These attempts look like "mistakes," but their essence is diversity. Nature does not know in advance which evolutionary path is correct, so it lets variation occur randomly and allows the selection environment to filter. The brain works the same way—it does not know which understanding is effective, so it lets different interpretations compete, allowing feedback to select the winner.
Key Distinction — Random Variation vs. Guided Variation:
Why does the classroom need variation?
Without variation, there is no selection. If all thirty-two students in a class solve the same problem using the same method—then that classroom effectively has only "one person" learning. If that single method fails, the entire class fails. But if those thirty-one students use twenty-one different approaches—even if fifteen of them are wrong—that classroom has generated rich "cognitive diversity," and diversity is precisely the source of systemic resilience.
The Irish Potato Famine (1845–1849) is the most devastating illustration of this principle. At the time, virtually all potatoes grown in Ireland were of a single variety—the Lumper. Genetically, they were clones of the same plant. In 1845, the late blight pathogen (Phytophthora infestans) arrived—and because all potatoes were genetically identical, not a single plant had resistance. Within four years, approximately one million people starved to death and one million fled. Not because there weren't enough potatoes, but because the potatoes weren't diverse enough.
A classroom that does not allow students to make mistakes is a field of "monoculture cognitive potatoes." It looks orderly, efficient, and compliant—until the exam takes a turn.
Manu Kapur is a cognitive scientist at the National University of Singapore. In 2008, he conducted a simple experiment whose results sparked years of discussion in educational psychology circles.
Experimental Design: Two groups of students learning the same mathematical concepts (e.g., variance, standard deviation). Group A: Traditional instruction—teacher explains first, students practice afterward. Group B: Productive Failure—teacher provides no explanation, instead giving students a complex problem they had not yet learned about and letting them attempt it for 20 minutes (inevitable failure), followed by instructional guidance.
Results: On the post-test, Group B students (failure first, then instruction) significantly outperformed Group A (instruction first, then practice) on transfer tests. In other words: students who experienced "failure" performed better when asked to apply knowledge to new situations.
Why? Kapur's explanation: those twenty minutes of failure were not wasted—they were activating deep processing. Students first generated rich cognitive variation (various immature attempts), then approached the teacher's explanation with these "raw ideas"—they weren't "receiving" information, they were "testing" it: Why was the direction I tried wrong? How does the correct answer differ from my earlier guess?
Key Mechanism: Failure itself does not produce learning. The cognitive variation generated by failure is the raw material of learning. If failure is not followed by effective feedback to filter these variations (i.e., "selection"), then failure is just failure. Productive failure works because it follows variation with a carefully designed selection environment—the teacher's explanation, peer discussion, self-reflective questioning.
Kapur's experiments revealed the positive value of variation. But variation is not magic—it needs to be designed into the classroom. Here are three of the most critical sources:
Source One: Problem Design — Open-Ended vs. Closed
"What is the answer to this question?"—a closed question activates "recall." There is only one answer, leaving the student with only one door in mind. "What are some possible ways to solve this problem?"—an open-ended question activates "variation." Students begin to think: this could work here, that could work there. A well-designed problem is not one that "the whole class got right" but one that "the whole class approached in different ways."
Source Two: Wait Time — The Difference between 1 Second and 3 Seconds
Mary Budd Rowe discovered an astonishingly simple fact in 1986: teachers wait an average of only 1 second after asking a question before intervening. When this wait time is extended to 3 seconds or more, the length and complexity of student responses doubles. Three seconds—just two extra seconds, and variation has a chance to emerge from students' minds. But most teachers cannot wait. Because that 1 second of silence makes them think "the student doesn't know."
Source Three: Error-Tolerant Culture — The Teacher's First Response
When a student gives a wrong answer, the teacher's first response determines whether the entire class will "dare to make mistakes" for the rest of the year. If the teacher says "No, think again" (in an impatient tone), the variation rate in that classroom plummets. If the teacher says "Interesting, how did you arrive at that?" (in a curious tone), the variation rate rises. Not because the latter teacher is gentler, but because the latter teacher signals to the entire class: errors themselves are worth discussing.
Real Stories · 30%
In the "Cosmic Theme Creation" section of the AI literacy course, JAM needed to use an AI image tool to render his imagined cosmic scene. His initial prompts were fragmented and low in logical coherence—the AI generated extremely mediocre images, even producing distorted color blocks, failing entirely to capture the effect he wanted.
The teacher could have handed him a "standard prompt template"—Subject + Details + Style + Lighting—letting him plug and play. Within minutes, he could have produced an image that "looked good." But the teacher chose a slower approach: no correction, let him experiment on his own.
Over the remainder of the class, JAM went through a frustration phase (the images kept coming out wrong) → a turning point (because no one criticized his "mistakes," he began trying to cram knowledge about accretion disks, event horizons, and spacetime curvature into the dialog box) → a breakthrough phase (he spontaneously researched precise definitions of astronomical terms, observing how word order changes affected the images). He was no longer "inputting text"—he was manipulating physical laws.
The results exceeded expectations. He independently summarized his own set of "cosmic creation logic"—and those images that initially looked "wrong" were precisely the material from which he learned the most.
Key Insight: JAM shifted from "fearing mistakes" to "learning through mistakes." Zero-error output equals zero variation, which equals zero learning.
Seven is in first grade. Three months into school, he barely speaks in class. When the teacher asks a question, he lowers his head. During group activities, he watches from the sidelines. In English class, when the teacher asks "What's your name?", the other children can answer, but Seven just purses his lips and stays silent.
The parents were frantic—"Did I do something wrong? Should we sign him up for a public speaking class?"
But this is not "not knowing." Seven talks a lot at home and expresses himself very clearly. His silence at school is because his brain is undergoing a massive process of variation—matching the language he hears against his existing cognitive structures, trying out various possible interpretations, eliminating incorrect assumptions. This process is invisible, but it is happening.
They held back their anxiety. No extra classes, no pressure, no pushing. Every morning, they simply said, "I hope you have a happy day today."
Three months later, one afternoon, the homeroom teacher sent a voice message—Seven had done a self-introduction in front of the entire class in complete English sentences. His voice was soft, but every word was crystal clear.
Key Insight: "I don't know how" is often not "I have no ideas," but "I'm still thinking." Cognitive variation requires not only "permission to err" but sometimes also "permission to be silent"—permission for that period when no progress is visible.
Scientific Evidence · 20%
Singaporean education researcher Manu Kapur conducted an experiment that many found counterintuitive (Kapur, 2012, 2014). He had two groups of students learn variance and standard deviation. Group A took the traditional route: teacher explains first, students practice later. Group B, however, was thrown directly into a complex problem—with no prior instruction, they were guaranteed to fail. After they had struggled enough and failed, the teacher then began instruction.
The result? On tests requiring knowledge transfer to new situations, the group that failed first and then received instruction scored significantly higher than the group that received instruction first and then practiced. The effect size was d = 0.72—considered a large effect in educational research.
Failure itself, of course, does not produce learning. But failure forces out cognitive variation—students try out various approaches during the struggle; some are correct, some are wrong, but all become "raw material" for subsequent instruction. When the teacher explains, they are not receiving unfamiliar information; they are filtering what they themselves just attempted. Without variation there can be no selection—make mistakes first, then correct them, and you'll go further than a path paved smooth from the start.
Experiment Source: Rowe (1986).
Findings: The average teacher wait time is only 1 second. When wait time is extended to 3 seconds or more:
Key Insight: Extending wait time appears to be merely "waiting two more seconds," but in reality it is the teacher returning the "right to choose" from their own hands to the students—you are no longer the one who decides for them ("This answer is right"); you are the one who gives them space to generate variation on their own.
Action Guide
One day I stood outside a classroom window and listened to an entire lesson. The topic was reading comprehension. The teacher did not begin by going over the passage directly. Instead, she asked the students first: "If you had to summarize this passage in one sentence, what would you say? Don't open your book yet—go with your first impression." The classroom was silent for three seconds. Then one student raised a hand: "I think this passage is about… friendship?" The teacher neither said "right" nor "wrong." She said: "Good. 'Friendship' is one keyword. Does anyone else have a different keyword?" Another student said: "I think it's about 'growth.'" A third said: "I think it's both—friendship promotes growth." The teacher said: "Excellent. Now we have three keywords. Please read the passage with these three keywords in mind and see which one fits best."
Throughout the entire lesson, this teacher never gave the "standard answer." She simply kept asking questions, pressing students to probe further, and letting them challenge each other. In the end, the students arrived at the main idea of the passage on their own.
After class, I asked her: "Why didn't you just tell them what the main idea was? Wouldn't that be more efficient?" She said something I have remembered ever since:
"If I tell them the answer directly, they'll remember it for three days. If they figure it out themselves, they'll remember it for three years."
Behind this sentence lies a secret about selection.
The previous chapter discussed variation—students need to generate abundant cognitive attempts for learning to have raw material. But "generating" alone is not enough. If a student produces an idea and no one tells them whether it is right or wrong, where it is good or where it falls short—then that idea is like a seed dropped in the desert; it will neither sprout nor die, it will simply hang there, having nothing to do with the student's brain.
The fate of variation depends on environmental feedback.
This is the subject of the present chapter—Selection.
Theoretical Framework · 20%
Within the CLE framework, "selection" is defined very succinctly: the filtering of cognitive variations by the environment—effective understandings are retained, ineffective ones are revised, and uncertain ones are set aside for testing.
This definition carries several key implications:
First, selection is not as simple as "judging right or wrong." Right and wrong form a binary. But most understandings in the cognitive world are not "completely right" or "completely wrong"—they are typically "partly right, partly wrong," or "right in this context, wrong in another context." Good selection pressure is not simply marking checks and crosses; it is providing students with enough information for them to judge for themselves the extent to which their understanding is valid.
Second, selection is everywhere, not only in exams. A teacher's glance, a peer's "Huh? I don't quite follow what you're saying," the moment of self-rereading when you think "Something doesn't feel right here"—all of these are selection. Every minute of classroom time involves selection; it is just that most of this selection is implicit, unconscious, and inefficient. The teacher's task is to design these implicit selection processes into explicit, conscious, and efficient feedback loops.
Third, a teacher is not a pipeline for "transmitting knowledge" but an environmental engineer who "designs selection pressure." If a teacher simply tells students the correct answer, students will never learn "how to judge whether an understanding is valid." But if a teacher designs a system—in which students' own ideas are continuously tested, fed back, and revised—then students not only come to know the correct answers but also learn "how to self-test." The latter is the true goal of education.
In a single sentence: A teacher's value lies not in how much information she provides, but in how much effective selection pressure she designs.
So, in what contexts are a student's cognitive attempts filtered? I classify them into three types:
The student has just completed a cognitive attempt and immediately receives a response. A peer's comment, a teacher's follow-up question, a check mark on a pop quiz—these all complete selection within seconds to minutes.
After completing their learning, students test their understanding again after a period of time. Homework grading, unit tests, midterms—these complete selection within days to weeks.
Continuous testing through interactions with others. Peer discussions, group work, teacher-student dialogue—this selection pressure has no fixed interval; it is embedded in every day's interaction.
These three are not substitutes for one another—each has its own function and its own optimal timing. A healthy learning ecosystem requires all three feedback loops; none can be missing.
Let's begin with the first: immediate feedback.
Jack is our science teacher, recently arrived from overseas, teaching seventh grade. The students are bright, but their English foundation is just beginning, and their science vocabulary is nearly zero.
The first class—I sat in the back row and listened for forty minutes. It was an uncomfortable lesson to sit through. Jack had prepared a wealth of content and taught every point earnestly. But the students' expressions grew more and more bewildered—they couldn't keep up. Seeing their lack of response, he stopped and went back to re-explain. Re-explaining put him behind schedule. Growing anxious—there was so much content still to cover—he quickened his pace. With a faster pace, students fell even further behind. Falling further behind meant stopping to re-explain again…
The bell rang. Jack hadn't finished. His expression—I still remember it—was that of a conscientious, responsible teacher facing a situation of "No matter how I explain, you just don't get it," the frustration written all over his face.
After class, Jack and I talked for less than five minutes. I offered one suggestion: After finishing each main point, stop. Leave 10 seconds. Don't speak, don't supplement, don't explain. Just wait in silence for 10 seconds.
The next class, he did it. Explain a concept, pause for 10 seconds. During those 10 seconds, students lowered their heads to think, quietly discussed, underlined key words in their notebooks. After 10 seconds, he continued. The bell rang—he had finished, everything covered. Jack looked at me, and that expression of delight I still remember today: a teacher experiencing for the first time that "I don't need to work this hard, and students can still learn."
Why did merely 10 extra seconds make such a world of difference? In the first class, although the total output was large, the student comprehension rate was extremely low, and actual effective learning time might have been less than a dozen minutes. In the second class, although a few minutes were "wasted" on pauses, the student comprehension rate doubled, and the effective learning time far exceeded the first class. Jack's first lesson lost because he mistook "teacher output volume" for "teaching efficiency."
Those 10 seconds of silence were, in fact, a redesign of selection pressure. Without those 10 seconds, the classroom was selecting "who responds fastest"; with those 10 seconds, the classroom was selecting "who understands most deeply." Two completely different survival rules grew two completely different classrooms.
Experiment Source: Mazur (1997), Peer Instruction: A User's Manual.
In the 1990s, Harvard physics professor Eric Mazur conducted an experiment that transformed university physics teaching. His classroom flow went as follows:
Step 1: Explain a concept (about 10 minutes).
Step 2: Pose a "concept test" question—not a calculation problem, but a conceptual understanding question that tests whether students truly grasp the concept, not just whether they can apply formulas.
Step 3: Students think independently and vote on an answer.
Step 4: Students discuss in small groups—discussing their answers and reasoning with nearby classmates. This is the most critical step in the entire process.
Step 5: Vote again.
Step 6: Whole-class discussion, Mazur explains.
Results: After peer discussion, the correct rate on the second vote increased on average by about 15–30 percentage points compared to the first. Moreover, students who answered incorrectly the first time and correctly the second time demonstrated stronger mastery of that concept on subsequent exams than students who "answered correctly the first time."
Why does it work? Mazur's explanation: peer instruction creates a form of social immediate feedback. Students must not only think for themselves but also explain to peers. If their answer is wrong, peers will question them during the explanation—"Why do you think that?" "I don't think so, because…" This social feedback is more direct and more forceful than teacher correction, because it is symmetric—students will not stay silent "for fear of being criticized by the teacher."
Key Mechanism: Peer discussion is essentially exposing "one person's cognitive variation" to "another person's selection pressure." In speaking, you discover gaps; in listening, peers point out contradictions—both sides are undergoing selection and both sides are revising their understanding.
Implications for Educators: A teacher cannot provide each student with immediate, one-on-one feedback. But peers can. Designing a good "peer discussion" mechanism means letting the entire class become each other's "selection pressure"—immediate, efficient, requiring no teacher intervention.
Mazur's experiment has a detail that is often overlooked. His "concept tests" themselves were not exams—each occupied only 5–10 minutes of class time and did not count toward grades. But it is precisely this "low-stakes, immediate feedback" format that allows students to dare to expose their genuine understanding. If every vote counted toward a grade, students would choose a "safe-looking" answer rather than their "real thought"—and the selection pressure would fail.
This points to an important design principle: For immediate feedback to be effective, the precondition is "safety"—students must believe that exposing errors brings not punishment but opportunity for improvement.
The second type of feedback loop is delayed feedback.
If immediate feedback is "micro-level selection within the classroom," then delayed feedback is "macro-level selection within the assessment system." Homework, quizzes, midterms, final exams—these are all typical forms of delayed feedback.
What is the value of delayed feedback?
The advantage of immediate feedback is speed—as soon as a student completes an attempt, they know whether they are right or wrong and can adjust rapidly. But immediate feedback also has a weakness: it is too "close." A student might only have temporarily memorized the correct answer without truly reconstructing their understanding—next time the question is phrased differently, they will fall into the same trap.
Delayed feedback addresses this problem: it forces students to face the same knowledge again after "forgetting"—and often in a new context. This "re-encounter" is itself a powerful selection event: if you truly understood, you will remember; if you only memorized short-term, you will forget. And "forgetting" is a powerful selection signal—it tells you: here, you have not yet truly mastered.
But this raises a core tension: Are current exams really providing effective delayed feedback?
Look at reality. The flow of most exams is: students answer → teacher grades → assign a score → return the paper. Where is the problem in this flow?
Kulik & Kulik (1988), synthesizing dozens of studies, discovered an interesting split: for basic skills (computation, spelling, vocabulary), immediate feedback is better—correct mistakes immediately to prevent solidification, with about 15–20% improvement. But for higher-order skills (problem solving, critical thinking, conceptual understanding), the opposite holds—moderately delayed feedback is actually better.
The reason is not complex: immediate feedback "interrupts" the student's deep thinking. You've just thought of an approach that's not quite right but somewhat plausible, the teacher immediately says "Wrong"—the train of thought is broken. Delayed feedback gives students more opportunity to "discover contradictions for themselves"—after they've hit a wall and realized the path goes nowhere, feedback appears. At that moment, feedback is not "correction" but "confirmation and deepening." Good delayed feedback should allow students to have already partially self-corrected before being fed back.
So good delayed feedback should not just be "a single scoring event" but rather an iterative process.
Jack—the same Jack who taught science a decade ago, now teaching design technology—once in his class, a student drew an orthographic projection: three views of a remote control. He glanced at it: crooked lines, distorted proportions, missing details. By the rubric, about 1 to 2 points—out of 8.
He didn't say "this is wrong here, that needs to change there." He placed the drawing back in front of the student and had her evaluate it herself first. "What do you think?" The student looked it over and said, "A decent first attempt." He accepted that assessment without supplementing it. Then he asked: "What do you think could be better?"
The student pointed out several issues she saw herself. He gave her a fresh sheet: "Try again."
This process repeated five times. Each time, he didn't say "where the mistake was." Each time, he had her look for herself, compare for herself, find the problems herself. By the third attempt, her work was approaching mid-high quality. He encouraged her to try once more. The fifth attempt—line precision was near millimeter-level, proportions were accurate, the structure was complete. In an exam, close to a perfect score.
Five drawings laid side by side on the desk. From the crooked first attempt to the precise, fluid final version—she was not "taught." She "learned" on her own through five iterations. Jack transformed a single "summative assessment" (one score) into five rounds of "formative assessment"—each cycle of self-discovery → self-correction → retry was a complete Variation → Selection → Retention loop.
The third type of feedback loop is social feedback.
Immediate feedback comes from classroom interaction; delayed feedback comes from the assessment system. But there is another form of selection pressure that is neither on the scale of seconds nor weeks—it is embedded in daily teacher-student and student-student interaction, dynamic, continuous, and multidirectional. This is social feedback.
I've been a principal for nearly twenty years. A recurring scenario gave me the most direct understanding of "social feedback." My phone, a WeChat notification—a parent sent a photo of their child's workbook, covered in red crosses. The parent was frantic: "Has my child been paying attention in class at all? Why does this question type keep going wrong, over and over?" The last sentence, I remember to this day: "Is it that the teacher can't teach, or is my child just not cut out for this?"
I didn't reply immediately. That sentence was too heavy—it simultaneously struck at external attribution (the teacher is no good) and internal attribution (the child is no good), with no buffer zone in between. The parent had backed both herself and her child into a corner. And this anxiety is typically not a one-time thing—it seeps into the classroom in various forms: the child, scolded at home, comes to school the next day with a mindset of "I'm afraid of getting it wrong again." The teacher, receiving pressure from the parent, subconsciously becomes more "cautious" in the classroom—less willing to let students experiment freely.
This WeChat message revealed a form of selection pressure that is rarely seen within the classroom: parental anxiety, channeled through the child and the teacher, indirectly shapes the "safety factor" within the classroom. If the chain of social feedback is: student makes a mistake → teacher evaluates → student withdraws → parent applies pressure → teacher tightens control → student fears mistakes even more—then it becomes a self-reinforcing negative loop.
The first step to breaking this loop is not to change the parent, but for the teacher to realize: there is a layer of invisible social feedback in their classroom—coming from outside the classroom walls.
So, what kind of social feedback is effective?
Experiment Source: Palincsar & Brown (1984), "Reciprocal Teaching of Comprehension-Fostering and Comprehension-Monitoring Activities."
Palincsar and Brown conducted a reading instruction experiment at a middle school. They divided students into small groups of four to five. Each student took turns serving as "little teacher," explaining the reading material to peers. The "little teacher's" task was not "reciting the text" but doing four things:
Results: Students trained in reciprocal teaching improved reading comprehension scores by about 20–30%. Moreover, the effect showed transfer—after training ended, students still performed better on new reading materials.
Why does it work? Palincsar and Brown's explanation: "teaching others" is the strongest selection pressure. When your explanation does not make sense, peers will immediately tell you—"I didn't understand," "What you just said contradicts what came before," "Can you give an example?" This immediate, specific, repairable feedback compels you to continuously adjust your understanding until it can be "taught out."
Another key finding: students who served as "little teachers" improved more than students who merely listened to explanations. This shows that "output" promotes learning more than "input"—because output is a process of continuously receiving selection pressure.
The power of reciprocal teaching lies in this: it transforms social feedback from "passive receiving" into "active exposing." The student is not waiting for the teacher to test them—they actively stand before the platform to be tested by the entire group. At that moment, all the blind spots in their understanding are exposed—because someone will always ask the question they hadn't thought of.
This is the core advantage of social feedback: it can provide feedback that is more diverse, more immediate, and more attuned to students' cognitive levels than any single teacher can offer. No matter how skilled a teacher is, they can only judge where a student doesn't understand from the "perspective of someone who understands." But peers are different—peers themselves have just emerged from "not understanding," so they know exactly where your gaps are.
But social feedback has one critical boundary condition: it requires a "safe" environment. If the teacher's first response is criticism, students won't dare to ask questions. If students' first response is mockery, peers won't dare to expose gaps. For social feedback to function, the prerequisite is that everyone in the classroom believes that exposing one's lack of understanding is safe.
Design Principles · 20%
We have now covered the three types of feedback loops. But knowing they exist is not the same as being able to design effective selection pressure. The following five principles have been repeatedly validated in our classrooms.
If you can take only one sentence away from this chapter, I hope it is this one:
Variation has been generated, and it has been selected. Next step—how does the chosen understanding become fixed?
Chapter 6: Real Learning vs. Fake Effort
Let me describe an experiment for you. After you've done it, you may never be able to "study" the same way again.
Have a group of college students read an article about the history of science. After reading, divide them into two groups. The first group is asked to "read it again." The second group is asked to "close the book and write down everything you can recall from memory." Tested immediately after reading, Group 1 (the re-reading group) remembered more content than Group 2 (the retrieval group)—"reading it again" does have a short-term advantage.
But one week later, the same people return to the lab, no advance notice, no opportunity to review—tested directly. The results flipped: Group 1 remembered about 30% of the content; Group 2 remembered over 50%. The retrieval group outperformed the re-reading group by roughly 50%.
This is not an isolated case. This is the research published by Roediger and Karpicke in 2006. They gave this phenomenon a name: the Testing Effect—the very act of retrieving a memory changes what you "remember."
"Read it one more time" versus "close the book and try to recall once"—the difference between these two actions is merely the distance of a flipped page, yet the difference one week later is 50%.
Why does this happen?
The answer lies in the nature of "Retention."
In the preceding chapters, we walked through the process of "Variation → Selection": students generate abundant cognitive attempts, and the environment filters out the effective understandings from among them. But where do the selected understandings go? If they are not solidified, they are like writing in sand—a single wave and the traces are gone.
This is the question this chapter seeks to answer: Retention—how does selected understanding become permanent brain structure?
Let me first clear up a common misunderstanding.
Most people understand "retention" as one thing: remembering. The brain is like a hard drive—you learn a piece of knowledge and "save" it. When needed, you "retrieve" it. This metaphor does more harm than good—it makes people think that as long as they've "seen" something enough times, knowledge will automatically get saved.
The brain is not a hard drive. The brain is a jungle.
Jungle Law (how the brain remembers things) / Bamboo Law (why progress is not linear) / Cognitive Potatoes (why a class can't have only one way of learning). These are not metaphors—each corresponds to a core CLE mechanism.
Imagine a tropical rainforest. There are no ready-made paths on the ground. The first person to walk in must use a machete to hack through vines and undergrowth, arduously carving out a passage. The next day, the vines have grown back somewhat—but not as thickly as the first day. The second person following the same direction finds it slightly easier than the first. The third person, the fourth… as more people walk it, a faint trail begins to form. A few months later, this trail becomes a clearly recognizable path. Years later, it becomes a dirt road wide enough for a horse-drawn carriage.
The jungle never "stored" a road. The road was walked into existence. Each traversal was a reconstruction of the road—weeds trampled, soil compacted, obstacles cleared. Untraveled roads get swallowed back by the jungle. Frequently traveled roads grow wider and smoother.
The process of memory in the brain is exactly the same.
The first time you learn a concept, there is no "ready-made road" in the brain—the connections between neurons are extremely weak, and signal transmission is slow and prone to interruption. Every time you recall this knowledge, it is like walking the same path again; with each traversal, that neural pathway is "tamped down" a little more. Pathways not traveled get swallowed back by the jungle—that is forgetting. Frequently traveled pathways transmit signals faster and more accurately—that is proficiency.
Jungle Law 1: Animals that set out together travel the same path.
In 1949, Canadian psychologist Donald Hebb proposed a principle—"Neurons that fire together, wire together." Translated into the jungle language: if two groups of animals frequently travel the same path at the same time, that path is maintained by both groups simultaneously and grows wider and wider. When you activate two neural pathways simultaneously—say, "photosynthesis" and "energy conversion"—the connection between them strengthens. This is not a metaphor; it is a physical change: the synaptic gap narrows, and signal transmission speeds up. This is the most fundamental unit of learning. Much of what you learned in middle school later disappeared not because it was "deleted" but because that road went untraveled for too long and was swallowed by the jungle.
Jungle Law 2: Frequently traveled roads get paved with gravel.
Surrounding nerves is a biological insulator called "myelin"—like spreading a layer of crushed stone over a dirt road. The thicker the myelin, the faster and more precise the signal transmission. The difference between a novice and an expert is largely not a difference in "amount of information"—it is a difference in myelin thickness. A piano novice must consciously think "what fingering for this chord?" while a master performer "just plays" without thinking—because the master's neural pathways have been paved into highways by myelin. This process is called myelination. Why are some skills, once learned, never forgotten—like riding a bicycle? Because that road has been paved so thickly that the jungle can no longer swallow it.
Jungle Law 3: Untraveled roads disappear.
Every time you recall a piece of knowledge, you are not retrieving an original file from a hard drive—your brain "walks" that road again. This is why each recall differs slightly from the original—the positions of the weeds have changed, the previous day's rain carved a small gully, you discover a shortcut through a fork in the path. This is precisely the deeper meaning of "retention" in CLE: retention is not about leaving knowledge lying intact; it is about letting it be re-selected and re-consolidated with each activation. Untraveled pathways get swallowed by the jungle (synaptic pruning); repeatedly activated pathways are widened and strengthened (long-term potentiation).
Key Conclusion: Why is "use it or lose it" correct? Because the brain rewires itself daily according to your patterns of use. What you mainly study today is what your brain grows into today. A math teacher's jungle is crisscrossed with logical deduction paths; a musician's jungle has sound-finger movement connections as wide as highways. It's not that they "remember" more—it's that the roads they frequently travel are different.
These three jungle laws point to one fact: the essence of learning is not "receiving"—it is "effort."
Why is effortful retrieval more effective than effortless re-reading? Because the very act of retrieving a memory is an act of "selection"—as the brain searches for a memory, it activates and strengthens the relevant neural pathways. Even if retrieval ultimately fails ("I can't remember"), the effortful process itself generates neural activity that triggers structural change. Passive reception does not trigger this change—the signal is too weak, and the jungle has no incentive to blaze a new trail. Only effortful retrieval leaves real traces in the brain.
You don't "store" knowledge; you "grow" it. Knowledge is not moved from the textbook into the brain; it is a path laboriously walked into existence inside your mind. And "walking it into existence" requires time, repetition, and the right rhythm—the next three sections present the three most effective methods for walking that path.
Having understood that retention is structural remodeling, the next question is: What kind of repetition most effectively promotes neural remodeling?
Intuition tells us: the more repetitions, the better. Learn and forget? Just do it a few more times. But science gives a counterintuitive answer: the timing of repetition matters more than the number of repetitions.
In 1885, German psychologist Hermann Ebbinghaus published a study that has been cited for over a century. He quantitatively measured his own rate of forgetting and drew the first Forgetting Curve: 20 minutes after learning, 42% forgotten; 1 hour later, 56% forgotten; 1 day later, 74% forgotten; 1 week later, 77% forgotten. In other words: without review, within a single day you will forget three-quarters of what you learned.
But Ebbinghaus also discovered: with each review, the slope of the forgetting curve becomes gentler. After the first review, the one-day forgetting rate drops from 74% to about 50%; after the second, to about 30%; after the third, to about 10%. Review does not prevent forgetting—it slows the rate of forgetting. Each review slows the speed at which memory "recedes."
Cepeda and colleagues' 2006 meta-analysis published in Psychological Bulletin, synthesizing hundreds of studies, reached a clear conclusion: the same amount of study time distributed across several days yields 10–30% better memory retention than cramming into a single day. Moreover, the longer the interval, the better the long-term retention—but there is an optimal window; spacing too long is like starting from scratch.
A practical rule of thumb: to remember for one week, review every other day; to remember for one year, review every three weeks.
From the BVSR perspective, spacing is effective because it leaves room for forgetting, turning each repetition into a "difficult retrieval." The problem with massed practice is that on the second repetition, the previous session's information is still "warm," so retrieval takes no effort. Without resistance, there is no selection pressure; the brain is not truly "efforted," and thus structural remodeling is not triggered.
But you might think—isn't this just "don't cram"? We already knew that. But there is a deeper, counterintuitive point here: distributed practice is not only more effective, but it also feels worse while you're doing it. That is why the vast majority of people instinctively choose the wrong approach.
Imagine two students, both studying quadratic functions.
Student A: Friday evening, does 50 problems in one marathon session. The first 20 are done with careful thought; the middle 20, attention wanders; the last 10, on autopilot. Closes the book feeling that they've "practiced a lot"—a sense of accomplishment.
Student B: Friday evening, does 10 problems. Saturday evening, does another 10—upon opening the book, thinks, "Huh? I've kind of forgotten what we studied yesterday…" Grapples with this uncomfortable feeling, flips through yesterday's notes, re-recalls the method, completes the 10. Sunday evening, another 10—forgotten a bit more, must again exert effort to retrieve. Tuesday evening, 10. Thursday evening, 10.
Student B doesn't have the refreshing feeling of "done!" that Student A has—every time they open the book, things feel a little rusty, a little effortful, and silently they wonder "Have I not learned this well?"
Exam one week later: Student A did 50 problems, but fewer than 20 problems' worth of effect remain in their brain. Student B did 50 problems, and all 50 problems' worth of effect are there. Student A felt "good" during practice—because the "smoothness" produced by continuous problem-solving was misread by the brain as "I've learned it." Student B felt "terrible" during practice—because each time required effortful retrieval, and that feeling of "struggling" was misread by the brain as "I haven't learned well." But the truth is exactly the opposite: feeling effortful is the signal of effective learning. Feeling smooth is merely the illusion of short-term memory.
This is the most counterintuitive thing about spaced repetition: the best learning method often feels like "not having learned well."
Everyone who works out understands a principle: muscles don't grow during training; they grow during rest. You tear muscle fibers, then give them time to repair; the repaired muscle is thicker and stronger than before. If you train to exhaustion one day and then train again the next—muscles simply have no time to complete repair and only grow weaker.
This is exactly the same as learning. Massed practice is like training until you can't pull anymore in a single day—it looks like working hard, but the neural pathways have no time to complete structural remodeling. Spaced repetition is like stopping each session when you're "almost at the limit but can still do one or two more," leaving space for repair and growth for the next session.
Behind spaced repetition lies a very elegant explanation within the BVSR framework: spacing leaves room for selection pressure to operate.
Massed practice is inefficient because it dilutes selection pressure. Doing fifty similar problems in one hour, most of it is "repeating known answers"—your brain is not "selecting" at all; it is merely "playback." Spaced practice is efficient because each repetition occurs right at "the edge of almost forgetting"—you are forced to exert effort to retrieve the answer from memory. That effortful retrieval process is the true selection: what can be retrieved is strengthened; what cannot be retrieved exposes gaps.
Every spaced retrieval is a "micro-selection" process. What was forgotten must be reconstructed through new attempts; what was remembered is further consolidated through the process of recall. Spacing—or "time" itself—is the most efficient designer of selection pressure. It doesn't tell you "where you don't know"; it simply allows "what you don't know" to surface automatically.
If spaced repetition is the "temporal strategy of selection pressure," then retrieval practice is the "behavioral strategy of selection pressure."
Return to the opening Roediger & Karpicke experiment. Two groups of students read the same passage; one group re-read, the other group retrieved. The re-reading group felt more "comfortable"—"I've read it again; I remember it." The retrieval group felt more "painful"—"I can't remember that much; I feel like I haven't learned well." But one week later, the retrieval group's performance far exceeded the re-reading group's.
This is the most common cognitive trap in learning: the fluency illusion. When you read your notes over and over, the text becomes increasingly familiar—you think you've learned it. In reality, the text has only grown "familiar-looking," but the brain has not constructed a robust neural pathway. True "mastery" can only be tested in the moment you close the book and force yourself to recall.
Experiment Source: Karpicke & Roediger (2008), "The critical importance of retrieval for learning," Science, 319(5865), 966-968.
This follow-up experiment further ruled out one possibility: could the retrieval group simply have read the passage more attentively?
Design: Four groups of participants learning foreign-language vocabulary. Group 1: repeated study + repeated testing. Group 2: repeated study only, no testing. Group 3: repeated testing, but no feedback after each test (no correct answers even when wrong). Group 4: repeated study, testing plus study—but only studying the incorrectly answered words.
Results: One week later, Group 1 (study + test) performed best. But the most critical comparison was Group 3 vs. Group 4: Group 3 (testing only, no feedback) performed significantly better than Group 4 (studying only incorrect words).
This finding is startling: even when you answer incorrectly and receive no correct answer at all, retrieval practice itself is still more effective than "only studying what you got wrong."
Why? Because the very act of retrieving a memory is itself a form of "selection"—as the brain searches its memory with effort, it activates and strengthens the relevant neural pathways, even if retrieval ultimately fails. This "effortful process" generates neural activity that is more capable of triggering structural change than passive reception of information.
What does this mean for the classroom?
A simple substitution can dramatically improve learning efficiency: replace "explain it again" with "closed-book recall for five minutes."
The structure of a traditional classroom is "teacher explains → students listen → teacher explains again → students listen again." The most effective structure is "teacher explains → students do closed-book recall → gaps are exposed → targeted supplementation." Retrieval practice requires no extra resources—it only requires one simple action: closing the book.
In fact, more and more schools are already putting this into practice. Some American high school AP courses use "closed-book quizzes at the start of every class"—not to "test whether you studied," but so that "after the test, you realize what you haven't learned." These quizzes do not count toward grades; their sole function is to trigger retrieval.
Spaced repetition addresses "when to repeat"; retrieval practice addresses "how to repeat." But one question has yet to be touched: Repeat what?
Isolated knowledge points, even when retrieved multiple times, are easily forgotten. Because they lack an "ecological niche" within the cognitive ecology—they haven't established connections with other knowledge, so the brain cannot find cues for "what it relates to." When recalling a knowledge point, the more triggers there are, the easier retrieval becomes. And the number of triggers depends on how many connections that knowledge point is embedded in.
This is how knowledge mapping works.
Within the CLE framework, we repeatedly use a metaphor: cognitive niche. A concept is not merely a "knowledge point"—it occupies a position within a network composed of other concepts, experiences, contexts, and problems.
Take an example: the concept of "photosynthesis." If studied in isolation—"chlorophyll absorbs light energy, converting carbon dioxide and water into organic compounds and oxygen"—it is merely a sentence that can be recited. But if this concept's niche looks like this:
When "photosynthesis" is embedded in such a network, it is no longer "a knowledge point"—it is a node that can be accessed through countless pathways. And each pathway is a potential retrieval cue. The more connections, the stronger the retention.
Knowledge mapping is not "organizing notes." Most students' notes are linear—recording knowledge points by chapter and in chronological order. But the brain's storage is not linear; it is network-like. Transforming linear notes into a networked knowledge map is itself an active process of cognitive construction.
So, how do you create a knowledge map? A simple but effective method: concept mapping. After finishing a chapter, take out a blank sheet of paper, draw a central concept, then radiate connections outward—"What concepts is this concept related to? How are they related?"
But this is more than just a "study method." From the CLE perspective, drawing a knowledge map is itself an active process of selection:
Core Finding: Memory consolidation occurs not only during waking retrieval practice but also during sleep.
Neuroscience has confirmed: memories encoded during the day are "replayed" by the hippocampus during slow-wave sleep at night and transferred to the neocortex. This process is like copying files—from the temporary folder (hippocampus) to long-term storage (neocortex).
Key Data: Walker & Stickgold's research shows that sleep-deprived groups perform roughly 40% worse on memory tests than groups with normal sleep. Not because they didn't learn—but because the learned content wasn't successfully transferred to long-term storage.
Connection to Knowledge Mapping: One key function of sleep consolidation is integration—weaving new knowledge into existing knowledge networks. Scattered information learned during the day is reactivated during sleep, cross-associated, and new connections are established. In other words, while you sleep, your brain is "drawing concept maps."
Implication: "Staying up late to study" may be the most common form of self-deception in learning. You spent the time, but not only was the daytime learning not consolidated—the next day's efficiency also suffers from lack of sleep. A lose-lose.
Biology Gave Us the Best Answers
The previous two chapters featured real stories from Yingjian. This chapter, let's shift perspectives—nature has prepared the best case studies for us:
You hold an acorn in your hand. You look at it, touch it, even look up its species classification and growth habits online—you've "understood" what an acorn is.
But an acorn is not an oak tree.
Bury it in the soil, water it, wait for it to break through the earth and grow its first tender leaf—that is "can do it." Understanding is holding a seed in your hand; mastery is the seed having grown into a seedling. And the germination process—struggling in the dark soil to send roots downward and push through clods of earth upward—is the most critical effortful phase of learning. Without this effort, a seed forever remains just a seed.
This is the biological version of "understanding ≠ mastery." Knowledge points are passively received like seeds; comprehension actively grows like a seedling. The only difference: one requires no effort; the other must struggle through darkness.
Migratory birds fly thousands of kilometers from the north to the south each autumn to overwinter. Young greylag geese have no idea of the route during their first migration—they follow behind the older geese. They fly it once; autumn passes and spring returns.
The second autumn, they follow again. The third year, they begin leading the way.
They didn't "suddenly remember" that route—it was three years of repeated flying that grew a complete navigational map in their brains. If they flew it only once and expected to remember the route forever, they'd have gotten lost long ago during migration.
So-called "poor memory" is never "a bad brain"—it's flying only once and expecting to remember a three-thousand-kilometer route. The brain is not a photograph, stored forever after a single glance. The brain is a migratory bird's navigation system—the route is flown into existence.
Five
In Chapter 2, we noted that the "skills vs. knowledge" dichotomy in education has had its foundations dismantled by AI. Here, we rebuild it—replacing that old classification with a more precise set of measurement tools.
Any learning task can be placed on three continuous axes for measurement. We call this the Tri-Axial Cognition (TAC) model—three yardsticks.
First Yardstick: How variable is the context? Multiplication tables have an almost invariant context—you're always "doing multiplication." But "when to use Newton's First Law rather than the Second Law" varies with every problem. The more uniform the context, the greater the benefit of automation; the more variable the context, the more recognizing the shape of a problem outweighs rapid execution.
Second Yardstick: How broadly does it apply? Some strategies are like keys—they open only one lock. Others are like skeleton keys—the same framework holds from a billiard table to the Milky Way. The breadth of applicability determines how far it can take you after being learned.
Third Yardstick: Can it still turn? The cognitive system has an "elasticity dial" (0–1). Close to 1 is complete rigidity—only executing automatic programs; when the environment changes, it collapses. Close to 0 is a loose pile of sand—no strategy can be retained. The optimum is at 0.5–0.7—stable enough to hold, flexible enough to adapt.
Core Insight from the Three Yardsticks:
Automation and construction are the natural expressions of the same BVSR algorithm under different requirements for contextual breadth, scope of coverage, and elastic demand.
Uniform context, narrow coverage, no need to adapt—automate. Variable context, broad coverage, constant need to pivot—keep exploring. Both are rational. It's not "skills bad, knowledge good"—it's what the yardsticks say.
So returning to this chapter's core question: what kind of retention is good retention? Answering with the three yardsticks—maintain elasticity in domains with sufficiently variable contexts; pursue automation in domains with sufficiently stable contexts. Next, let's see how this principle plays out in real learning processes.
Six
Finally, let us address a question that no educator can avoid: Why does "understood" not equal "can do it"?
Within the CLE framework, moving from initial understanding to true automated retention requires passing through three stages. The three stages are not linearly arranged—they are nested, but with each rise in stage, the operating speed and selection pressure of the BVSR algorithm differ.
Level 1: Recall. The student can recall a concept, can recite a formula. This is the minimum standard of retention. Teachers typically think "you know it" is enough—but "knowing" has a hidden trap: it can be tested immediately after learning, yet may disappear three days later. Recall is the starting point of retention, not the endpoint.
Level 2: Transfer. The student not only "remembers" this knowledge but can also recognize it and use it in new contexts. You learned to solve two-variable linear equations but cannot do three-variable ones—this means retention has only reached Level 1. True transfer means: your brain can "see" the identical underlying structures across different contexts. This is where cognitive flexibility comes into play.
Level 3: Automation. Execution without conscious thought. When you ride a bicycle, you don't think "shift the center of gravity by how many degrees, turn the handlebars by how many radians"—these procedures have been myelinated into bodily intuition. The highest form of retention is not "I remember," but rather "I don't have to think."
BVSR Interpretation of the Three Stages:
Level 1 (Recall) BVSR cycles are "slow, coarse-grained"—you retrieve a concept, the brain makes a quick match; success means retention, failure exposes gaps. Level 2 (Transfer) BVSR cycles are "medium-speed, multi-pathway"—you must simultaneously activate multiple concepts, compare their similar structures, and select the best fit for transfer. Level 3 (Automation) BVSR cycles are "ultra-fast, subconscious"—BVSR has descended from the conscious level to the unconscious level; variation and selection occur on a millisecond timescale, and you are not even aware of "making a decision."
These three stages correspond to three different educational strategies:
The predicament of most traditional classrooms lies in this: teachers invest the most effort in the first stage, far too little in the second stage, and leave the third stage almost entirely to luck. The student "understands"—the teacher's job is done. But the road from "understood" to "can use it" must be walked by the student alone. Students who can walk it become "top performers"; students who cannot become the puzzled ones who "seem to understand but can't solve problems."
This is not the students' fault—it is that we have only designed for "understanding" and failed to complete the journey to "can use it" and "don't need to think."
Seven
In the previous section, we described "automation" as the highest form of retention—executing without thinking. But here lies an ultimate question that every educator must confront: Where does the endpoint of automation lie—in what environment was it automated?
There is a bacterium called Buchnera. Two hundred million years ago, it was swallowed into the cell of an aphid—not digested, but permanently settled. The aphid cell is a perfect home: constant temperature, ready-made nutrition, no predators. Buchnera no longer needs to cope with any surprises.
Two hundred million years. Buchnera's genome shrank from over 2,000 genes to 600. It lost motility, lost environmental sensing, lost stress responses—lost everything "unnecessary" inside this room.
This is not degeneration—at least, not degeneration in the usual sense. This is optimal adaptation. In the absolutely stable niche of an aphid cell, maintaining those "useless" genes is pure energy waste. Regression is evolution's rational choice.
Buchnera is one of the most perfectly adapted organisms there is. As long as its aphid host never dies.
But it died.
Every time you have a student spend an extra hour on standardized question types, you are also helping them atrophy some ability that is "unnecessary" within this closed space: raising their own questions. Tolerating confusion without standard answers. Finding direction amidst ambiguity.
Adaptation to the point of perfection inherently carries the cost of regression.
The question is no longer "Should we automate or not?"—the question is "In what environment are we automating? How long can this environment last?"
Ireland's potatoes provide a larger-scale version. In 1845, a potato variety called Lumper fed one-third of the nation's population. It was high-yielding and adaptable—the Irish converted all their farmland to potato fields. Under the constraints of the time, this was the optimal solution. Then a fungus—late blight—drifted over. Vast fields of potatoes rotted simultaneously overnight. Not because there weren't enough potatoes—potatoes were everywhere. It was because they were all the same kind.
Ireland's potato fields, in 1845.
Buchnera, on the day it left the aphid cell.
The gaokao student, the moment the standardized exam hall closes.
When uniformity is taken to its extreme, success itself becomes the prelude to disaster.
This principle has a specific name in evolutionary biology: Overspecialization. In a stable environment, specialization is the optimal strategy—you invest all your resources in the most efficient mode of adaptation. But the moment the environment changes, specialists are the first to go extinct. Generalists, while not optimal in any single environment—survive.
The First Commandment of Evolutionary Biology:
Never bet all your adaptive capacity on a single strategy.
What your cognition needs is not extreme automation—but reversible automation.
What is "reversible automation"? Return to the three-stage model of this chapter. Level 3—automation—is not itself the problem. The problem is after automating, can you return to Level 2 or even Level 1 and reconstruct?
GPT-4 can score 130 on the gaokao math exam. Not because it "understands mathematics"—a 6-dimensional problem space, indirectly covered within a trillion-parameter training dataset. You spent 12 years training yourself into an automaton capable of high-speed operation within a closed space. That's 120,000 cognitive iterations, countless sweat-soaked and defiant nights. GPT-4 achieved the same thing in a few weeks of training. And it's faster than you, steadier than you, and will never get nervous.
Twelve years. 120,000 generations of cognitive evolution. In exchange for extreme efficiency within a 6-dimensional space. AI did it in a few weeks.
So what should be done? The answer is not "don't practice fundamentals"—that would mean abandoning any form of automation. The answer lies in designing a multi-track retention strategy: some content should be automated (multiplication tables, spelling rules, basic operations)—these are the basic language of the human brain; without them, higher-level thinking has no material. But we must simultaneously retain enough "non-automated" cognitive pathways—pathways for raising questions, for connecting concepts, for persisting through ambiguity.
Of course, a word of fairness is due here. Students who chose Strategy A are not victims—they made the most rational choice within a given exam framework. The problem is not that they chose test-prep drilling; it is that this framework did not leave sufficient reward-punishment space for "understanding." Students scoring 145 and 141—the precision of the gaokao cannot distinguish these 4 points. But the former spent all 12 years investing nearly all cognitive resources in a low-dimensional closed space; the latter spent some of that time reading extracurricular books, doing projects, hitting walls in real problems. These two strategies differ by only 4 points on the gaokao, but the gap in adaptability and room for growth at university is enormous.
Action Guide
But retention has a curious property: it is rarely linear.
You practice a skill, make progress—then suddenly hit a wall. No matter how you practice, nothing changes.
This is not you learning poorly. This is the learning curve telling you: the next chapter is on its way.
Next: Chapter 7 "Punctuated Equilibrium: Stagnation Is the Best Data"
In the winter of 1972, two young paleontologists walked into the fossil vault of the American Museum of Natural History. They were about to do something "heretical"—overthrow Darwin.
More precisely, not overthrow Darwin's core idea—natural selection. What they were challenging was one of Darwin's implicit assumptions: that evolution is slow, continuous, and gradual. Like climbing a mountain, each step representing an advance, only at a magnitude too small for the naked eye to perceive.
Niles Eldredge and Stephen Jay Gould pored over thousands of fossil records. The pattern they discovered unsettled the entire paleontological community: species' morphological changes were anything but uniform in pace.
For the vast majority of time—over 90% of the entire geological record—a species' morphology remains virtually unchanged. Stable, conservative, showing no sign of "progress."
Then, in a geological "instant" (perhaps a few thousand years, no more than the span of a single page on the fossil scale), the species undergoes sudden and dramatic differentiation: new traits emerge, new forms are born, new species arise.
Then—another long period of stability.
Eldredge and Gould named this pattern "Punctuated Equilibrium."
Not a gently rising slope, but steps: stability → leap → new stability → new leap.
The reason this discovery shook the academic world was that it destabilized the default assumption of "gradualism." Darwin himself had said: "Natural selection cannot produce great leaps; it must proceed by slow and gradual steps." The fossil record said otherwise: no, most change occurs within extremely short time windows; the rest of the time, nothing happens.
"Nothing happens"—this is precisely the most thought-provoking part. Those long periods of stability are not evolution pausing. They are the system accumulating the potential energy for a leap. The genetic variations, environmental pressures, and niche explorations accumulated during stable periods are released at a critical threshold, producing explosive morphological change.
When I moved this theory into education, everything fell into place.
Students' learning curves and organisms' evolutionary curves look exactly the same.
If cognition is also a form of evolution—the core hypothesis of CLE—then cognitive change should also follow the pattern of punctuated equilibrium. A substantial body of research evidence supports this.
Theoretical Source: Anderson (1982), ACT-R Model (Adaptive Control of Thought-Rational).
John R. Anderson divided the acquisition of any complex skill into three non-skippable stages:
Key Insight: Stage 2 (knowledge compilation) is precisely the punctuated equilibrium of learning. The leap from declarative to procedural does not happen smoothly—it requires a "compilation period," during which nothing looks like progress, and there may even be regression. But once compilation is complete, the leap occurs.
Phenomenon Source: Research on child second-language acquisition (Krashen, 1981; Hakuta, 1974).
When children learn a second language, they almost all go through a "silent period"—lasting from a few months to as long as a year. During this stage, the child barely speaks. Parents and teachers begin to worry: "Is he not understanding?" "Is there a language ability issue?"
But neuroscience tells us: during the silent period, the brain is working intensively—the speech recognition system is being calibrated, grammatical rules are being constructed, vocabulary networks are being established. Before a child can understand aloud, they need to speak first to themselves.
Then, on some ordinary morning, the child suddenly produces a complete, grammatically mostly correct sentence. Not "improvement"—a leap.
Chinese students' "English plateau"—easy to go from 40 to 65, but stuck at 65 to 70 for an entire year—works by the same mechanism. It's not that English has become harder; it's that the brain is switching from "learning English by memory" to "using English by language intuition," and this switch cannot happen overnight.
This reminds me of Seven's story. In Chapter 4, we recounted: during Seven's first three months of first grade, he barely spoke at school. The parents were frantic—"Did I do something wrong?" But in reality, his brain was undergoing an extremely intense "silent period": repeatedly matching the English he heard against his existing cognitive structures, eliminating incorrect assumptions, building a grammatical framework.
Then, one day—the homeroom teacher sent a voice message. Seven had done a self-introduction in front of the entire class in complete English sentences. His voice was soft, every word crystal clear. Not "improvement"—a direct leap from the declarative stage to the autonomous stage. The three months of silence before were not wasted; they were the brain completing knowledge compilation in the background.
Theoretical Source: Ohlsson (2011), Deep Learning: How the Mind Overrides Experience.
Every math teacher has seen this student: a concept explained for two weeks, no matter how you explain it, they don't get it. The student is anxious too—brow furrowed, eyes glazed. Then—on a certain problem in a certain homework assignment—they suddenly say: "Oh, so that's it."
This "aha moment" is not a linear extension of logical reasoning. It is the brain completing some form of representational reorganization in the background—stitching scattered information fragments into a new framework. The old framework was no longer adequate (e.g., "applying formulas" fails when facing complex figures); the new framework had not yet been built ("analyze structure → locate conditions → reason step by step"). Thus the brain entered the chaotic period of "knowledge compilation." Every problem attempted during the chaotic period is trial-and-error variation; most attempts fail, but each failure tells the brain "this path leads nowhere."
When enough "dead ends" have been accumulated, the brain suddenly finds the passage through—and the leap occurs.
Three lines of evidence point to the same conclusion: learning is not climbing a slope; it is climbing steps. The plateau period between each step is not empty—it is the brain reorganizing cognitive structures. The more paths you attempt during a plateau, the higher the next step will be.
The theory is clear. But the reality is: nearly everyone is afraid of plateaus.
Chapter 5 told the story of an anxious parent—their child's workbook covered in red crosses, asking "Is it that the teacher can't teach, or is my child just not cut out for this?" Chapter 5 used that scenario to explain the "attribution trap": caught between "the teacher is no good" and "the child is no good," we often cannot find a third path. But behind that WeChat message lies an even deeper issue: what looks like "regression" may not be regression at all. I had her pull out all of her child's workbooks from the past month, arrange them chronologically, and look at only one type of problem—she discovered her child had gotten it wrong at the start, then right, and then recently wrong again. This was not regression. This was him going through one of the most critical stages of learning: a plateau. That parent was silent for a long time, finally replying with two words: "Got it." With a tearful emoji. But I know clearly—not every parent can find this door amidst their anxiety.
Why are we afraid of plateaus? Three reasons.
First, the assessment system measures only "burst power," not "incubation periods." Our exam system is fundamentally a "peak instantaneous output" test—within a set time, how much correct information can you recall and how many correct methods can you apply? A brain in a plateau period is reorganizing, and performance during reorganization is inevitably fluctuating and unstable. But exams don't give you the option of "currently reorganizing"; they report only a score. A plateau-period score and a decline-period score look identical, so plateaus are misdiagnosed as "regression."
Second, social comparison amplifies anxiety. "Other people's children are improving; mine is stagnating"—the background music to this sentence is: stagnation is dangerous; stagnation means being overtaken; being overtaken means being eliminated. Within this comparative framework, stagnation has no room for positive interpretation.
Third, and the deepest fear: we cannot tell the difference between "healthy stagnation" and "unhealthy stagnation." Parents and teachers aren't opposed to plateaus—they just don't know when to wait and when to push. If a child is genuinely falling behind and you tell them "this is normal, wait a bit longer," they will genuinely miss the intervention window. But if a child is in the middle of reorganization and you pile on extra classes and pressure—you've personally dismantled their leap engine right when the brain most needs consolidation.
Label: Bamboo Rhizome Phase — CLE's Renaming of the Plateau Period
Moso bamboo is extremely common in southern China. After planting a bamboo shoot, in the first year, it grows only a few centimeters. The second year, still just a few centimeters. The third year, the fourth—seemingly almost motionless.
You water it, fertilize it, care for it attentively; it remains silent. Neighbors passing by might say: "Is this bamboo stunted? Should we replace it?"
But come the rainy season of the fifth year, Moso bamboo shoots up over twenty meters within six weeks—taller than a six-story building. Overnight, the "little sprout" that had disappeared from people's sight for over four years becomes the tallest stalk in the entire grove.
Where was it during those first four years? Botanists tell us: during those four years, the bamboo's underground rhizome system—the bamboo whip—never stopped growing for a moment. It extended hundreds of square meters in all directions beneath the soil, weaving a vast and dense network. Without this "underground infrastructure," the fifth-year burst would have been simply impossible—a twenty-meter bamboo stalk, without a deep and extensive root system, would have been uprooted by the first typhoon.
The most counterintuitive fact: the apparent "stillness" of the first four years is precisely the necessary precondition for the fifth-year burst. If, in the third year, you pulled it out because you "couldn't see any change," you would never witness the rainy season of the fifth year. The bamboo's growth logic is not "grow a little each day"; it is "weave the network first, then burst forth."
The plateau periods in education are the same thing as the bamboo's first four years. The seemingly progress-free months—even semesters—that students experience may very well be their "bamboo rhizome phase": old cognitive structures are being digested and reorganized, new conceptual frameworks are silently spreading underground, just not yet breaking the surface.
If you pile on extra classes and pressure at this moment, it is like pulling up the seedling before the rainy season arrives. Bamboo does not fear growing slowly—it fears being told to grow tall before its root system is woven.
Shifting from looking at "scores" to looking at "root system status," from attending to "above ground" to attending to "what is happening underground"—this is a new yardstick CLE offers to educators.
Not all stagnation is regression. Some seemingly motionless stretches of time are the bamboo weaving its roots—only when the roots are deep does the above-ground growth race upward.
First, do not add more workbooks. First look at the scratch paper—if the types of errors in recent homework assignments are changing (last time calculation error, this time conceptual confusion, next time skipped step), it means the child is "trying different paths"; this is a healthy plateau.
Second, do not ask "Did you understand?" Change the question: "Without looking at your notes, can you explain it to me?" Only what can be explained has truly been mastered.
Third, give your child a "permit." Tell them—"The fact that you're stuck at this score now doesn't mean you're regressing. When the brain is paving roads, it doesn't show on the outside. Once the paving is done, one lesson and it clicks."
Judging whether a plateau is healthy or not isn't about looking at scores—it's about looking at the quality of variation.
This is a core insight of the CLE framework: the role of plateaus in the BVSR algorithm is that of a "variation accumulation period." The quality of variation determines the height of the next step. A plateau filled with high-quality variation is the eve of a leap. A plateau with no variation at all is genuine stagnation.
✅ Healthy Plateau — rich cognitive variation
• The student tries different problem-solving methods (even if most are wrong)
• Error types change—first time calculation error, second time conceptual misunderstanding, third time skipped step
• The student asks questions: "What if I changed this condition to… what would happen?"
• After making an error, the student can articulate "why I got it wrong"—even if not entirely accurately
• The student's subjective experience is "confusion" or "the more I study, the less I feel I know"—this is the signal of cognitive reorganization
❌ Unhealthy Plateau — cognitive variation exhausted
• Every error is in the same place, in the same way, with no change
• The student is mechanically repeating—using the same wrong method on the same type of problem
• The student no longer asks questions, no longer thinks about "why I got it wrong"
• The student actively avoids challenges, only doing problems they already know
• The subjective experience is not "confusion" but "giving up"
Diagnostic Criterion: If three or more "healthy" signals appear simultaneously in a plateau, a leap is brewing. If "unhealthy" signals dominate, intervention is needed.
This diagnostic framework matters because it transforms the "plateau" from a vague mass of anxiety into observable, judgeable data. Parents and teachers don't need to guess "is this normal?"—they only need to observe the quality of variation.
When behavior stalls, the brain may be quietly "paving roads."
Neuroscientists ran a classic experiment: have a group of people do working memory training every day for three consecutive weeks. The first two weeks, their test scores were completely flat—a classic plateau. But brain scans showed that during those fourteen days, the brain was not idle: the insulating layer around nerve fibers—the equivalent of rubber coating on electrical wires—was significantly thickening. Its function is to make signals travel faster and more steadily between neurons, like widening a country lane into a highway.
By the third week, performance suddenly jumped. And the "highway widening project" had already been largely completed during the first two weeks.
This discovery reveals a counterintuitive fact: the days when behavioral performance shows no improvement are precisely the days when the brain is spending resources paving roads.
This explains why "adding classes during a plateau" is counter-efficient. What a plateau most needs is not more knowledge input—the brain is already busy paving roads—but rather leaving time for the project to proceed. Sleep is the road crew's most important work period. The more classes you add, the less sleep there is; the roads don't get paved, and the leap naturally never arrives.
What parents and children lack most during a plateau is not "more effort" but "permission to pause."
What is this "permit"? It is a single sentence: "Stagnation is normal—you are not regressing; you are upgrading."
This sentence is not comfort—it is a direct corollary of Evolutionary Pedagogy.
Within the BVSR framework, a plateau is not a pause in BVSR but rather an active period of BVSR's variation stage. Variation must occur before selection (feedback) and retention (consolidation). Without sufficient accumulation of variation, there can be no high-quality selection and retention. Suppressing the plateau is suppressing the source of variation—which is suppressing the very possibility of cognitive evolution.
From this understanding, we can do three things.
These indicators need not be incorporated into formal exams—they only need to become the "internal language" through which teachers and parents observe students.
The core logic of the three actions above is consistent: transform the plateau from a problem to be "overcome" into a data source to be "read".
Plateaus tell us: the construction of knowledge is not synchronized.
Some knowledge constructs quickly; some knowledge requires a long "silent period."
But what exactly is being constructed during this "silent period"?
It is constructing what each learner uniquely possesses—their cognitive niche.
Next: Chapter 8 "Cognitive Niches: Knowledge Is Not Storage, It Is Adaptation"
Eldredge, N., & Gould, S. J. (1972). Punctuated equilibria: An alternative to phyletic gradualism. In T. J. M. Schopf (Ed.), Models in Paleobiology (pp. 82-115). Freeman, Cooper.
Anderson, J. R. (1982). Acquisition of cognitive skill. Psychological Review, 89(4), 369-406. (ACT-R three-stage model of cognitive skill acquisition)
Takeuchi, H., et al. (2010). Training of working memory impacts structural connectivity. Journal of Neuroscience, 30(9), 3297-3303. (Evidence of myelination changes during working memory training)
Gould, S. J. (2002). The Structure of Evolutionary Theory. Harvard University Press. (Comprehensive exposition of punctuated equilibrium theory)
The first seven chapters established CLE's core theory — the BVSR algorithm. The next five chapters unfold this same algorithm across different scales of analysis: the micro-level is BVSR inside the individual brain, the meso-level is BVSR in classroom interaction, and the macro-level is BVSR in institutional systems. The three verbs are "the unchanging algorithm"; the three levels are "the variable scales."
* This chapter uses extensive biological analogies to build intuition. Not interested in evolutionary stories? Skip directly to Section 8.3 — that's where we start talking about education.
There is a concept in biology that I consider one of the most important anchors in the entire CLE theory — the niche.
What is a niche? It is not a "physical location." A species' niche is its functional position and way of living in the ecosystem: what it eats, where it lives, whom it competes with, how it reproduces. Two birds living in the same tree do not necessarily share the same niche — one eats insects in the canopy, the other eats larvae in the trunk; their daytime activity zones are entirely different. Physically they are "together," ecologically they "live their own lives."
A niche is not a map, nor is it an address. It is the taut string between species and environment — a relational pattern, not a positional coordinate.
A niche is never static. The moment the environment moves, it must move too.
The classic example is the finches Darwin observed on the Galápagos Islands. When the dry season arrived, small soft seeds on the island were exhausted, leaving only large hard-shelled seeds. Thick-beaked finches could crack them open and survived; thin-beaked finches could not and died in large numbers. Not because the thin-beaked finches "didn't try hard enough" — it was because the environment changed, and what had been a useful tool was no longer useful. When the rainy season returned and small seeds reappeared, the proportion of thin-beaked finches slowly rose again. Same species, same place, same niche — but different environments meant entirely different selection criteria.
Another example: the peppered moth of industrial England. Before the Industrial Revolution, tree trunks in England were covered with pale lichen; light-colored moths resting on them were nearly invisible, while dark moths were easily spotted by birds. After the Industrial Revolution, coal soot blackened the trunks, and suddenly light moths became conspicuous while dark moths were hidden — within fifty years, the proportion of light moths in the Manchester area plummeted from 98% to under 5%. After the Clean Air Act of the twentieth century, trunks returned to pale, and light-colored moth populations gradually recovered. The most striking thing about this example: it wasn't the moths that changed — it was the habitat. The same genes, the same traits — an advantage in the old habitat became fatal in the new one.
These two examples reveal a crucial fact: the value of a niche depends on its environment. A "good" niche in one environment may be "bad" in another. Conversely, a "poor" niche may become "good" in a different context. The two are always bound together.
Biological Niche = a species' "way of living" in the ecosystem: what it eats, where it lives, what it avoids, how it reproduces. Not a location — a relational pattern.
Take a squirrel: its niche is "eat nuts, live in trees, dodge hawks." Its teeth can crack hard shells, its spatial memory can retrieve thousands of buried nuts, its tail keeps it balanced mid-leap — every one of these traits evolved around this specific way of living: "eat nuts, live in trees, dodge hawks." Change the way of living, and these traits lose all validity — throw a squirrel into water and all of its adaptations fail.
Cognitive Niche = each learner's "way of understanding" in the cognitive world: their habits of thought around a concept, their preferred problem-solving strategies, the way they organize knowledge.
Take the same concept of "ratio": one student is used to understanding it through mathematical formulas (a:b = c:d → cross-multiply), another through physical meaning (density = mass/volume → ratio is not a computational tool, it's a property of matter). Each has constructed a different cognitive niche — a different "way of understanding." Both niches are valid; they apply to different situations.
Cognitive Niche = a piece of knowledge's "way of living" inside a particular learner's brain
It is not a static unit of knowledge; it is a mode of being used. The same piece of knowledge "lives" differently in different students' minds.
Key Corollary — same logic as the finches and moths: Whether a student's cognitive niche is "good" or "bad" depends on the cognitive environment they face. A student who derives formulas in math class, when confronted with a physics problem requiring experimental data to model a functional relationship, suddenly finds their "formula-first" niche inadequate. Not because their understanding is wrong — but because the habitat changed. Conversely, a student accustomed to "finding the physical picture first, then writing the equation" in physics may also lose their rhythm in a pure mathematics proof. There is no absolutely "correct" way of understanding — only a way of understanding that matches — or fails to match — the current cognitive environment.
This analogy is not rhetoric. It rests on a deep theoretical foundation: the evolution of cognitive structures and biological structures share the same underlying algorithm — BVSR. Biological variations are selected and retained in new environments, forming adaptive physical structures. Cognitive variations are selected and retained in new problems, forming adaptive mental structures. Both are processes of constructing through adaptation.
The only difference is timescale: biological evolution operates across generations; cognitive evolution operates from minutes to semesters.
Traditional education has one deeply entrenched metaphor: knowledge is like food, "fed" to the brain. The teacher is the "feeder," the student the "receiver." Learning well = receiving fast and storing securely.
This metaphor is so pervasive that most people have never questioned it. Yet it is wrong.
Section 8.1's stories of finches and moths revealed half the truth about the relationship between niche and environment: environmental change drives adaptive change. In education, the same thing happens every day — when a student transitions from middle-school plane geometry (look at the diagram → find relationships → prove) to high-school functions (abstract → model → compute), their "cognitive environment" changes, and the thinking tools that used to work suddenly don't. This is not regression — their habitat changed.
But there is another half to the story: organisms also change the environment. And this half, in education, is almost entirely overlooked.
Let's return to the squirrel's niche. A squirrel eats nuts — this looks like "adapting to the environment." But in the process of burying nuts, it carries seeds into new soil, causing more nut forests to grow. The squirrel does not merely adapt to the forest — it builds more forest with its own hands. Each generation of young squirrels faces not the forest the previous generation first arrived at, but a forest modified by that generation — one with more nut trees. The squirrel's behavior changed the entire habitat's structure, and this altered structure in turn fed back to change the selection pressures the squirrel faces — because there are more nuts, competition patterns shift, and predators' pursuit paths change too.
This bidirectional loop has an extreme example in nature — coral reefs.
A coral polyp is a tiny marine organism, only a few millimeters in diameter. Living alone in seawater, it is so fragile that even a current could sweep it away. But it does one thing: it secretes calcium carbonate skeletons. One polyp's secretion is just a speck of dust, but countless generations of polyps secreting continuously in the same place — over thousands of years, these microscopic skeletons accumulate into the Great Barrier Reef, stretching two thousand kilometers, visible from space.
Once the Great Barrier Reef formed, it fundamentally transformed the ecology of that entire ocean region. What had been a seabed of nothing but currents and sand became one of the most biodiverse places on Earth — fifteen hundred species of fish, four hundred species of coral, four thousand species of mollusks making their home there. The habitat built by coral polyps in turn became the adaptive environment for thousands of other species. And the activities of these species on the reef — grazing algae, boring holes, excreting — continually reshape the reef's structure.
This is the complete picture: the environment shapes organisms, and organisms in turn shape the environment — a continuous, bidirectional interaction with no endpoint.
→ Applying this to education: A student does not passively "adapt" to the classroom — the small choices they make every day in note-taking strategy, questioning style, and discussion habits with classmates accumulate, day by day, to "build" their cognitive niche. And these constructions, in turn, determine how they can learn when they encounter new content in the next class. The learner is not "receiving an environment" — they are transforming the learning ecology in which they are situated.
In 2003, evolutionary biologist John Odling-Smee and colleagues formally proposed the concept of Niche Construction. The core claim: organisms do not merely passively adapt to their environment; they actively modify it, and these modifications in turn alter the selection pressures they themselves face. Beavers build dams to change water flow, earthworms loosen soil to change soil structure, humans farm to change the entire terrestrial ecology — niches are not given, they are built.
This claim challenges an implicit assumption in Darwinian evolutionary theory: that the environment is given and organisms can only passively adapt. No — the environment is also part of evolution, and changes in the environment are themselves altering the direction of evolution.
Educational philosopher John Dewey independently touched upon the same insight a century earlier. In Experience and Education (1938), Dewey defined "experience" as the interaction between organism and environment — not passive reception of stimuli, but active exchange of energy, information, and meaning with the environment. He argued that the most fundamental principle of education is "the continuity of experience": every experience transforms the learner, and the transformed learner then interacts with subsequent environments in a new way.
Translating Dewey into the language of CLE: learning is the process of cognitive niche construction. Every act of thinking, every problem worked, every discussion — quietly rewrites the relational pattern between the learner and knowledge.
Regrettably, Dewey is often reduced in Chinese education to the three characters "learning by doing" — as if he simply advocated getting students to use their hands. His deeper insight — that the essence of experience is the continuous co-construction between learner and cognitive environment — has almost entirely been overlooked. And it is precisely this insight that is the most direct philosophical precursor to "cognitive niche construction."
Bringing this logic into education, the conclusion is clear:
Biological Niche Construction: Squirrels bury nuts → change the forest → change the next generation's survival environment
Cognitive Niche Construction: Every "extra" thing a student does during the learning process — organizing notes, drawing knowledge maps, asking questions — remodels their cognitive environment, thereby changing the way they will understand new knowledge in the future.
Let's look at four of the most common "construction behaviors":
The common feature of these behaviors: the learner is not "receiving knowledge" — they are constructing the "way of living" through which they understand the world.
This raises a deeper question: in the traditional classroom, how much "construction space" do we leave our students?
In a typical 45-minute class, the teacher's lecture occupies 35 minutes, leaving 10 minutes for problem sets — problems that are predetermined, with standard answers, fixed procedures, and time limits. The student's only "construction activity" is memorizing these procedures and reproducing them on the next exam. This is not cognitive niche construction — it is the suppression of cognitive niche construction.
I am not criticizing teachers here — I myself was once that tirelessly instructive person. But we must confront this fact: when a student is merely a "receiver," their cognitive ecology can never grow.
A cognitive niche must be used to exist.
The previous two sections discussed the "cognitive niche" — the "way of understanding" a piece of knowledge has inside a single learner's brain. But the concept of a niche has two levels in biology: the individual niche (how a squirrel lives) and the ecosystem composed of multiple individuals (the forest itself).
It is the same in education. A student's way of understanding "ratio" is their cognitive niche. But this student sits in a classroom — with other students, with a teacher, with classroom rules, with an examination system — and all of these together constitute a larger Cognitive Ecology, that is, the "ecosystem of the cognitive world."
Clarifying the distinction between these two concepts stabilizes the terminology for the entire book:
Cognitive Niche = the individual's way of understanding
How a piece of knowledge "is used, when it is used, and what other knowledge it is connected to" inside a particular learner's brain.
Cognitive Ecology = the environment in which cognitive niches reside
Classroom atmosphere, assessment methods, peer interaction, institutional rules — the "soil, temperature, and humidity" in which learners' cognitive niches grow.
The Cognitive Niche is the "animal"; the Cognitive Ecology is the "forest."
Animals adapt to the forest, and animals also transform the forest. The forest in turn determines what the animals can become. The two are in constant interaction.
Using this framework, let's revisit the squirrel example from §8.1: the squirrel has its niche (eat nuts, live in trees, dodge hawks). The entire forest is its ecosystem — with pines, oaks, various birds, and climate variability. It is not only the squirrel's niche that influences how the squirrel lives; the state of the entire ecosystem — whether rainfall is abundant this year, whether the acorn harvest is good, whether hawks are numerous — also determines the squirrel's behavior.
Likewise, a student's cognitive niche for "ratio" does not exist in isolation. It is continuously shaped by cognitive ecology factors — the classroom atmosphere (will I be mocked if I answer wrong?), the teaching approach (does the teacher allow exploration time?), the assessment system (what kinds of questions appear on exams?). Conversely, when this student raises a distinctive question, that question enters the classroom environment and becomes part of the class's cognitive ecology — it is changing other students' cognitive niches.
This leads us to the three levels of cognitive ecology in the CLE framework. In the living world, ecology is never single-layered. Take a pond:
Level 1: The Fish's Own Niche
Grass carp eat aquatic plants, silver carp eat plankton, crucian carp eat insects at the bottom. Each layer of fish has its own way of living — different diets, different water layers for swimming, different breeding seasons. This is the "individual niche" — analogous to each student's cognitive niche.
Level 2: The Food Web Inside the Pond
Grass carp eat the water plants, changing the water's transparency; phytoplankton decrease, changing silver carp's food source. One fish's behavior alters the survival conditions of every other fish in the pond. This is not "the sum of every fish's individual way of living" — it is the sum of their interactions. Whether the pond is polluted, whether there are too many or too few fish, whether the plants are thriving — these are the states of the pond as a whole ecology. This is analogous to the classroom's "class-level cognitive ecology."
Level 3: The Watershed the Pond Belongs To
A pond is not sealed off. It is affected by upstream inflow — does the factory upstream discharge pollutants? By rainfall — is this a drought year or a flood year? By human activity — does fertilizer from surrounding farmland seep in? All these external factors determine the pond's "ceiling" — how clear the water can be, how dense the vegetation can grow. This is analogous to the school's "cognitive ecology" — institutional design, assessment methods, and management culture determine the ceiling of the classroom ecology.
The relationship among the three levels is exactly the same as in the pond
The three-level structure in the CLE framework is this pond:
Level 1: Individual Cognitive Niche (Intrapersonal)
Each person's way of understanding. Student A and Student B attended the same physics lecture; A's cognitive niche leans toward "memorize the formula," B's leans toward "understand the physical picture." What suits A's learning does not necessarily suit B. This level is the core of Chapter 9.
Level 2: Classroom Cognitive Ecology (Classroom)
The discussion norms, error culture, and knowledge-sharing patterns of the classroom community. It is not "the sum of every student's cognitive niche in the class" — it is the sum of the interactions among those individual cognitive niches.
If a class develops an atmosphere where "getting an answer wrong gets you mocked," every student's cognitive niche will converge toward the "safe" direction — speak less, risk less. If a culture develops where "errors are discussable," individual cognitive niches will stretch toward the "exploratory" direction. This level is the core of Chapter 10.
Level 3: Institutional Cognitive Ecology (Institutional)
The school's institutional design, assessment methods, and faculty culture — these determine the "ceiling" of the classroom cognitive ecology. If the school evaluates teachers solely by exam scores, teachers cannot genuinely encourage students' cognitive variation in the classroom. They want to — but the system forbids it. This level is the core of Chapter 11.
The relationships among the three levels are bidirectional:
Individual Cognitive Niche → Classroom Cognitive Ecology: A student raises a "strange" but profound question in class. This question (a cognitive variation) enters the classroom environment. It may be mocked, or it may be discussed. If discussed, it becomes part of the classroom cognitive ecology — altering the selection pressures that all classmates subsequently face.
Classroom Cognitive Ecology → Individual Cognitive Niche: When "discussing errors" becomes a class habit, students who previously dared not raise their hands begin to do so. Their individual cognitive niches change, enveloped by the classroom cognitive ecology.
Classroom Cognitive Ecology → Institutional Cognitive Ecology: As multiple classes gradually develop similar classroom cultures, the school's teaching-research activities begin to document this pattern, and eventually the administration institutionally encourages "error-tolerant teaching."
Institutional Cognitive Ecology → Classroom → Individual: School policies change, assessment methods change, classroom culture follows, and eventually every student's cognitive niche changes too.
This is a continuous bidirectional cycle — the very framework that Chapters 9 through 12 will unfold in detail.
Every teacher has encountered this: a student masters the method of completing the square in math class, then in physics class, seeing a parabola equation, has absolutely no idea that completing the square could be used. Or understands the law of conservation of mass in chemistry class, then in biology class, when discussing the products of photosynthesis, completely forgets that conservation of mass even applies.
Teachers usually call this "didn't learn it solidly enough" or "can't transfer." But CLE offers a more precise explanation: niche mismatch.
Traditional Explanation: The student didn't "truly understand" the concept, so they can't use it.
CLE Explanation: The cognitive niche the student constructed in Context A does not match the cognitive niche in Context B. The "completing the square" niche they built in math class (a formal, symbol-manipulation-focused cognitive environment) and the actual problem context they face in physics class (a concrete, physical-meaning-first cognitive environment) are entirely different niches. Their knowledge grew well in that niche — but when migrating to another niche, they can't find the entrance.
Analogy — The Woodpecker's Tongue: How long is a woodpecker's tongue? Its hyoid bone starts at the beak tip, wraps around the entire skull, passes through the right nostril, and finally anchors beneath the eye socket. The total length of the tongue is more than three times the length of the beak. This exquisite structure serves a single purpose: hooking insects out of deep holes in tree trunks. The woodpecker uses this "precision tweezer" to peck fifteen times per second on trees, with an extremely high success rate.
But put that same woodpecker on the ground and ask it to catch ants — its long tongue becomes a handicap, unable to reach into ant holes at all. It's not that it "has forgotten how to catch insects." It's that the "insect-catching method" it learned was constructed only within the specific context of "tree trunk depth." Change the context to "ground level," and the same set of tools becomes useless.
It hasn't forgotten how to forage — it simply never constructed its foraging method within the "ground foraging" niche. Transfer failure is not "the knowledge is gone" — it's niche mismatch.
Educational Corollary: If a concept is constructed in only one context, its niche is very narrow — it works well in that context but becomes "unusable" in another. The solution is not to make students "memorize the concept harder" — it is to construct the same concept repeatedly across multiple contexts, broadening its cognitive niche.
The importance of this explanation lies in the fact that it transforms "transfer failure" from a student-ability problem ("you're not smart enough") into a teaching-design problem ("we didn't give you enough niche variation").
Bransford and Schwartz (1999), in their paper "Rethinking Transfer," proposed a crucial insight: traditional transfer research designs every experiment as a "learn A → test B" model — learn in Context A, then test whether you can apply it in Context B. But genuinely complex transfer happens when a student encounters an entirely novel problem context and can "construct a cognitive niche for themselves." Most transfer tests measure not the student's "transfer ability" but whether the student encountered, in Context A, a niche similar to that of Context B.
In other words — practice without niche variation produces only "dead knowledge."
Every experienced teacher has done this: place a single piece of knowledge in different settings for students to experience it repeatedly. An English teacher teaching a new word doesn't teach it just once — it appears in reading comprehension, in listening materials, in writing models, and is used by students in classroom dialogue. The same word, four or five settings. The teacher hasn't explicitly told the students "you are learning this word," but the students "encounter" it across different settings, use it, and eventually "grow" their understanding of it.
This is not a special case from some experimental class — it is the norm that happens every day in teaching. Physics teachers teaching density don't reteach mathematical ratios; math teachers teaching similar triangles don't mention chemical solution concentrations. Yet a single concept of "ratio" appears in three subjects simultaneously — a computational tool in math, a property of matter in physics, a concentration relationship in chemistry.
In a classroom where three subjects have coordinated lesson planning, the teachers did one thing: at the beginning of each class, they spent three minutes having students discuss: "Today, all three of these subjects are actually talking about the same thing — ratio — but does their 'ratio' look exactly the same?"
Month one: students found the question strange. "Ratio is just ratio — what's the same or different about it?"
Month two: some students began to say, "Ratio in physics and ratio in math are not the same thing — physical ratio has physical meaning, while mathematical ratio is a computational tool."
Month three: one student said to the teacher, "I think all three subjects are building different platforms for the same concept — now I know, 'ratio' is not a formula; it's a way of thinking about relationships."
This student did not "learn" ratio — they had constructed the same concept repeatedly across three different cognitive niches. Each round of construction widened the concept's niche a little more. After three months, it was wide enough for them to "see" the essence of ratio.
Note: this is not the story of some "magical classroom" — it is what excellent teachers do every day. They just don't necessarily use the term "cognitive niche" to name it.
Let's return to the opening question: Is knowledge stored in the brain?
If "stored" means like a file on a hard drive — retrieved intact and unchanged — then the answer is no. Knowledge is not "deposited" into the brain; it is grown into it. Every retrieval, every use, every transfer, may transform its form. This is not a defect of knowledge — it is precisely the evidence that knowledge is alive.
Knowledge that has been repeatedly constructed across multiple contexts can be flexibly deployed in diverse environments. Knowledge constructed in only a single context can survive only in that one environment.
The task of education is not to help students install more "files" —
it is to help them build a broad space of understanding for every core concept.
In Chapter 6 we told a story — the perfect adaptation and fatal degeneration of Buchnera bacteria inside aphid cells. A stable environment paired with extreme optimization caused it to shed every "unnecessary" capability over two hundred million years. The moment the environment changed, it could only silently await death.
But the niche is never unidirectional. Evolutionary biologist Richard Lewontin pointed out as early as 1983: organisms don't just adapt to their environment — they actively modify it. This insight later developed into niche construction theory.
The North American beaver. The beaver is not satisfied with finding a place within the existing environment. It uses branches and mud to build dams, alter water flow, flood forests, and create its own wetlands — it was not originally the beaver's niche; it built it for itself.
Buchnera and the beaver tell the same evolutionary principle from two directions:
In the cognitive world, there is exactly the same bidirectional movement.
Some students sit at their desks flipping through workbooks, every problem looking just like the last — they are building a Buchnera-style cognitive niche: stable, efficient, but the moment a problem changes its face, they are utterly at a loss.
Other students in the same classroom are constructing entirely different cognitive niches: reading an extracurricular book, doing a small project, having an argument with a classmate, retrying three times after failure. They proactively pull unfamiliar concepts, different perspectives, and unfamiliar tools into their cognitive world, making the environment adapt to their construction needs, rather than merely adapting themselves to the environment's existing structure.
These two types of niches may show no difference on a single exam — in fact, the Buchnera type might even score two points higher. But the moment the environment changes — a new textbook, a new teacher, graduation to college — the gap becomes vast as heaven and earth.
Cognitive niche construction is never a one-way "adaptation to the environment" —
Are you a Buchnera, or a beaver?
This is not just a question about learning methods — it is a diagnostic tool for whether the educational system you are in is training you to adapt to old environments, or empowering you to construct new ones.
Every student's brain is a unique forest.
But this forest does not grow chaotically — it is governed by a fundamental algorithm.
How does this algorithm operate within the individual brain?
In the next chapter, we begin at the smallest scale —
Every instance of learning, every act of thinking, every "aha" moment —
the micro-evolution unfolding deep inside the brain.
Next Chapter: Chapter 9 — "The Micro-Level: Micro-Evolution Inside the Brain"
When I first read Piaget's descriptions, I experienced a strange sense of familiarity — he was not the old man who appeared in Educational Psychology 101. He was talking about evolution.
Jean Piaget (1896–1980) spent his life observing one thing: how children's knowledge grows. He didn't sit in a lab giving children tests — he observed every cognitive behavior of his own three children from infancy onward: one infant grabbed a toy, another threw a toy. In these everyday actions, he discovered a pattern.
I briefly mentioned this pattern in Chapter 3. Now, we will place it under CLE's microscope and examine it carefully.
Piaget discovered that children's cognitive development follows two simultaneous processes: assimilation and accommodation.
Assimilation — incorporating new experiences into existing cognitive structures. A two-year-old learns the word "dog," then sees a cat on the street and points, saying "dog." What is the child doing? They are placing the cat (a new experience) into their existing cognitive structure ("four-legged animal = dog"). This is not a mistake — it is cognition operating normally.
Accommodation — modifying existing cognitive structures to fit new experiences. When the mother corrects them, "That's not a dog, that's a cat," the child faces a choice: ignore the information and keep calling it "dog," or adjust their cognitive structure by creating a new subcategory called "cat" under "four-legged animals." The second option is accommodation. It is a reorganization of cognitive structure.
From the CLE perspective, these two processes correspond exactly to two forms of variation within the BVSR algorithm.
Assimilation = Variation During Stable Periods
When existing cognitive structures can accommodate new experiences, the brain produces small-range interpretive variation during assimilation. The same fact is assimilated differently by different students — they use their own existing "cognitive filters" to interpret it. This is why, faced with the same concept of "conservation of energy," one student understands it as "energy doesn't disappear," another as "energy changes form," another as "energy eventually becomes heat" — the direction of assimilation differs, and cognitive niches have already begun to diverge.
Accommodation = Variation During Burst Periods
When existing cognitive structures cannot accommodate new experiences, the brain is forced to reorganize. Piaget called this moment cognitive disequilibrium. Disequilibrium is painful, disorienting, and makes one want to flee — but this is precisely the "plateau phase" described in Chapter 7: apparent stagnation on the surface, internal reorganization underneath. The result of accommodation is a qualitative transformation of cognitive structure — a new concept is born, old concepts are integrated.
Assimilation produces quantitative variation; accommodation produces qualitative variation.
This distinction is profoundly important. It tells us two things:
First, "errors" in learning are not failures — they are the byproducts of failed assimilation. The student is not "ignorant"; they are using their existing cognitive structure to assimilate new information, but the existing structure is insufficient. Error is a signal that the cognitive structure needs upgrading — I elaborated this point in detail in Chapter 4.
Second, truly deep learning happens in disequilibrium, not equilibrium. When everything goes smoothly and the student "gets it right away," what is most likely happening is assimilation — new knowledge is compatible with the existing structure and merges in painlessly. But when the student gets stuck, confused, even angry — that is when the conditions for accommodation have just matured. This is precisely the underlying cognitive mechanism for the assertion in Chapter 7 that "plateau phases are good data."
Piaget showed a five-year-old child two identical glasses containing equal amounts of water. Then he poured the water from one glass into a taller, narrower glass. He asked the child: "Which glass has more water?"
The five-year-old would say: "The tall glass has more water." — They were assimilating: taller = more. They had not yet developed the concept of "conservation" — that the amount of water does not change with the shape of the container.
A child around age seven would correctly answer "same amount" — they had completed the cognitive structural reorganization from "single-dimension (height) judgment" to "multi-dimension (height × width) judgment." This is accommodation — a leap from intuitive perceptual operations to concrete logical operations.
This leap is not taught — it is "reorganized" by the brain itself after numerous failed assimilation attempts. This is why Piaget emphasized: cognitive development cannot be accelerated. You can teach a five-year-old to say "same amount," but they have not truly constructed the concept of conservation. Their performance "says the right thing," but their understanding remains the old one.
For educators, this is a hard law: you cannot skip accommodation. You can make a student memorize the definition of "conservation" and write it on an exam — but you cannot bypass that "micro-evolution" inside their brain. They must experience cognitive disequilibrium to construct their own concept of conservation.
This is constructivism's bottom line — and its dignity.
But there is a trap here: if we accept that "cognitive development cannot be accelerated," what is left for the teacher to do? The answer lies in the next section.
Before entering the discussion of metacognition, one key distinction needs clarification. Piaget's "cannot be accelerated" refers to qualitative leaps in cognitive structure — such as the transition from pre-operational to concrete operational — which have their own intrinsic temporal logic. Metacognitive training does not accelerate this transition, but it gives learners at the same cognitive stage a clearer perception of "what they are understanding and what they are still missing." Metacognition does not give cognition "more speed"; it gives cognition "more direction."
In 1979, developmental psychologist John Flavell, in a paper of just six pages, proposed a concept that would become one of the most powerful tools in educational psychology: metacognition.
Metacognition literally means "cognition about cognition." More colloquially: knowing what you know, knowing what you don't know, knowing what strategy to use to come to know.
From the CLE perspective, metacognition is not "another skill" — it is a special layer within the BVSR algorithm. I define it in one phrase:
Remember the three steps of BVSR? Variation generates diverse cognitive options, selection pressure filters among them, and retention stabilizes those that succeed.
But here is a hidden layer: who decides "what counts as variation"? Who sets the "criteria for selection"?
This meta-level behind them is metacognition — it is not "making a single selection," but selecting upon the selection process itself.
Specifically:
• Monitoring one's own cognitive resources: Knowing "I haven't understood this part" — this is the prerequisite for initiating variation
• Choosing strategies: Knowing "Should I read the problem first or draw a diagram first?" — this is the directionality of selection
• Evaluating understanding: Knowing "Have I truly understood, or have I merely memorized?" — this is the calibration of selection
Learning without metacognition is like a traveler without a map — variation is happening, selection is proceeding, retention is continuing, but they don't know which direction they're heading in.
This "second-layer selection" mechanism changes a fundamental question: who is doing the selecting?
Without metacognition, selection pressure comes mainly from outside — teacher feedback, exam scores, peer evaluation. Selection is "passive." With metacognition, students begin to apply internal selection pressure to their own cognitive activity — they notice that this solution path isn't working and actively switch strategies; they realize they haven't grasped this concept and actively seek a new explanation. At this point, selection is no longer "being evaluated" — it is "self-regulation."
Dunlosky et al. (2013) published a meta-analysis systematically evaluating the effectiveness of ten common study methods. The results surprised many:
These "most effective" methods share a common characteristic: they all require students to actively monitor and regulate "how they are currently learning." Retrieval practice is not just "doing test papers" — it is the student actively probing "what they know and what they don't know." Distributed practice is not just "reviewing after an interval" — it is the student planning "when to review."
And the ineffective methods — rereading, highlighting — are precisely passive behaviors that require no metacognitive engagement.
This finding tells us one thing: metacognition is not "icing on the cake" — it is a necessary condition for effective learning. Study methods without metacognitive engagement, no matter how much time students invest, have limited effect.
So the question arises: can metacognition be taught?
The answer is yes. But the method is not "lecturing about the concept of metacognition in class" — that would itself be a cognitive contradiction. Metacognition can only be learned through doing: as a student completes a task, they can ask themselves three questions —
1. Before the Task (Planning): "What does this task require of me? How is it different from tasks I've encountered before? What strategies do I have available?"
2. During the Task (Monitoring): "Is my current approach working? Where am I stuck? Do I need to switch approaches?"
3. After the Task (Evaluation): "How did the process of completing this task go? Which strategy was most effective? How should I improve the next time I face a similar task?"
These three questions appear simple, but they transform a passive learner into an active cognitive manager. The student is no longer merely "doing the problem" — they are "observing themselves doing the problem," "making choices within the process of observing themselves doing the problem."
There is a deeper transformation here: when students begin to ask themselves these questions, their cognitive niches shift from a "single layer" to a "two-layer structure." The first layer is disciplinary knowledge itself (math, physics, chemistry). The second layer is the monitoring and regulation of the first layer's knowledge. The second layer does not directly solve specific problems — but it makes the BVSR algorithm on the first layer run more efficiently: generating more directional variation, executing more precise selection, achieving more durable retention.
But metacognition is not yet the most microscopic level. Beneath metacognition, an even more fundamental mechanism is operating — self-explanation.
Imagine this: you're reading a somewhat difficult book. You reach a certain passage, pause, and say to yourself — "Wait, the author's logic here is... which means..." Or "How does this relate to that case we discussed earlier?"
You have just performed one of the most powerful actions in cognitive science: self-explanation.
In 1989, cognitive scientist Michelene Chi and colleagues published a paper that would become a classic in educational psychology. Her experimental design was simple: two groups of students studied the same physics content (text and worked examples on Newtonian mechanics). Group A was asked to "explain each step to themselves" during the learning process — "why is this step done this way," "what is the relationship between this part and the next." Group B read normally without additional requirements.
The result: Group A's comprehension and transfer ability were significantly higher than Group B's. Self-explaining students not only solved more problems but also performed better when faced with problem types they had never seen before.
Participants: University of Pennsylvania psychology undergraduates
Learning Material: Newtonian mechanics text and worked examples
Key Findings:
From the CLE perspective, the mechanism of self-explanation is remarkably clear:
Variation Phase: When a student reads new information, the brain automatically generates multiple possible ways of understanding — "What does this formula mean," "Why is this step done this way," "How does this conclusion conflict with concepts I've learned before." This is the phase where cognitive variation is produced.
Selection Phase: When the student says to themselves "because... therefore...," they are applying a kind of internal selection pressure to their own understanding — "Does my explanation make sense? Is it consistent with known facts? Does it solve the current problem?" Explanations that pass muster are retained; those that don't are abandoned or revised.
Retention Phase: Verified understandings are easier to retain than unverified information. Because self-explanation builds a causal chain — "a leads to b leads to c" — such chains are far easier for the memory system to capture and store than isolated facts.
The essence of self-explanation is the internalization of external selection pressure into a self-selection loop inside the brain.
This insight explains why "passive learning" has limited effects: when a student merely listens or reads without actively performing self-explanation, cognitive variations are indeed being produced (the brain cannot avoid producing some understanding), but these variations have not been tested by the "selection" phase. They merely float inside the brain, waiting for an external exam to test them — and exams, as a form of external selection pressure, often come too late and are too crude.
The power of self-explanation lies in this: it moves the selection pressure from "exam time" to "learning time." When students begin "testing themselves" while still reading, their learning efficiency rises dramatically, and — more importantly — they learn to independently judge "I understand" versus "I don't understand" without any external feedback. In the CLE framework, this is a milestone capability: it marks the upgrading of the individual's cognitive niche from "dependent on external selection pressure" to "possessing an internal selection pressure loop." This is the cognitive foundation of autonomous learning.
"Why is this step done this way?" — After each worked example, the teacher can ask rather than explain directly. Let students speak first.
"How does this relate to what we learned before?" — The connection between old and new knowledge is not for the teacher to declare — it is for the student to construct. Let students "connect" the two things.
Peer Reciprocal Explanation: "Tell him what you understood, then he'll tell you what he understood" — When a person is forced to "say it out loud," their self-explanation shifts from implicit to explicit, making it easier for themselves and their partner to test.
"The Third Column of the Error Log": In addition to "the wrong problem" and "the correct solution," add a column to the error log: "Why did I think of that method at the time?" — This forces the student to retrospectively examine their own variation (the erroneous line of thought) and subject it to analysis of selection criteria.
"The Step-Back Question": In classroom discussion, when a student proposes an answer, the teacher follows up — "Can you tell me how you arrived at that answer?" — This question shifts the student's attention from "Is the answer right?" to "Is my reasoning process sound?"
The common feature of these five strategies: the teacher acts as a "designer of selection pressure," not a "provider of answers" — this is precisely the core definition of the teacher's role in the CLE framework.
But while Piaget described the evolutionary mechanisms of cognitive structures within the individual, he left one question unanswered: How are cognitive variations, once they enter the social environment, collectively selected and shaped? Vygotsky's Zone of Proximal Development (ZPD) theory provides an entry point to this question. ZPD describes the level a student can reach with the help of others — in the CLE framework, this is precisely the guidance of individual cognitive variation by social selection pressure. Piaget tells you how the individual "constructs"; Vygotsky tells you how that construction process is socially "selected." Only together do they form a complete BVSR cycle.
Up to this point, we have been inside the individual — inside the microcosm of the brain — observing how the BVSR algorithm operates: how assimilation and accommodation cause the brain to produce different ideas, how metacognition imposes directionality on the "upper layer" of selection, how self-explanation provides internal selection pressure in the absence of external feedback.
But learning has never happened only inside the individual.
When a student produces a new understanding — a cognitive connection that did not exist before, a different problem-solving strategy, an "aha" moment — they do not lock it inside their head. They speak it aloud, write it on paper, raise their hand to ask, share it in a group discussion. This is the most crucial moment in learning: a cognitive variation shifts from "private" to "public."
Micro produces variation → Individual expression → Variation enters the meso environment → Others in the environment receive this variation → Produce new variation → Express → ...
This is the starting point of emergence.
The definition of emergence: when large numbers of micro-level individuals interact locally according to simple rules, at the macro level a novel, unpredictable, stable collective pattern arises.
In the classroom —
• Micro-level individuals: Every single student, running the BVSR algorithm inside their brain
• Simple rules: If you have a question, raise your hand; if you have a different opinion, say so; if you make an error, it can be discussed
• Local interactions: One student voices their understanding, another responds, a third supplements
• Collective pattern: Three months later, "the way our class discusses math" — a distinctive meso-level classroom culture — has emerged
This process is not a metaphor — it is the core mechanism by which the CLE framework transitions from the micro level (Chapter 9) to the meso level (Chapter 10).
Let me illustrate with a concrete scenario.
In a physics class, the teacher was explaining Newton's Third Law (action and reaction). Midway through, a student raised a hand and asked: "If I'm sitting on a small cart and push a bigger cart, wouldn't I not feel the reaction force?"
This question stirred a ripple through the class — the conflict between "can't feel it" and "it must exist" created cognitive disequilibrium. This student, without realizing it, had turned a private confusion into a public topic of discussion.
The teacher did not answer directly, only saying: "Good question. Everyone, sketch this scenario on a blank sheet of paper and discuss with your deskmate for two minutes."
Two minutes later, some students had drawn force analysis diagrams, some argued "the small cart is sliding, so where did the reaction force go?", some proposed "you need to consider friction." One student said: "If the carts are on a frictionless surface, then when I push the big cart, my own small cart will also slide backward — that is the reaction force."
The classroom filled with that "ohhh—" sound.
From the CLE perspective, what happened in those five minutes was a micro → meso → micro cycle:
One student's cognitive variation (the conflict between "can't feel it" and "must exist") → Expression (raised hand to ask) → Entry into the classroom environment (others received this confusion) → New variation produced in each person's brain (force analysis, friction considerations, where the reaction force goes) → Selection and retention in deskmate discussion ("this explanation makes more sense") → A new understanding emerges ("the backward slide of the small cart on a frictionless surface is the evidence of reaction force").
This "new understanding" did not come from the teacher — it came from a single ordinary question that, after undergoing "collective selection" within the classroom environment, emerged as collective construction.
This scenario reveals a profound law: classroom culture is a shared cognitive evolutionary system. Every question raised, every difference of opinion voiced, every "I've rethought this" remark — is a variation being input into the class ecology. The class ecology's discussion atmosphere, error tolerance, and the teacher's response style — these constitute the selection pressure on those variations. What eventually settles is shared understanding, default modes of thinking, and a common "way of speaking" for the entire class.
This is the subject we will explore in depth in the next chapter: the meso level — the emergence of classroom culture.
But before that, we need to remember this chapter's core conclusion —
The micro-level tells us: every learner's brain is a stage for evolution.
But when these "micro-evolutions" begin to collide with, select, and shape one another —
a new level is born.
It is not the simple sum of individuals — it is emergence.
Next Chapter: Chapter 10 — "The Meso-Level: The Emergence of Classroom Culture"
Two classrooms at opposite ends of the corridor. Same teacher, same lesson.
One class is lively beyond description — students rush to speak, unafraid of being wrong. The other is silent as an exam hall — even raising a hand requires a long hesitation.
Same textbook. Same discipline requirements. Same teacher. Same forty-five minutes.
Every teacher has seen this kind of difference. But few have asked: How does this difference "grow"?
The traditional explanation is "different classroom atmospheres" — but "atmosphere" is a black-box label; it explains the result, not the mechanism. A more precise question is: in the dozens of interactions during the first week of school, which micro-behaviors formed different feedback loops between teacher and students, eventually emerging into two entirely different collective patterns?
In the previous chapter we stepped inside the individual learner's brain and observed the "micro-evolution" happening in every student's mind. But learning does not only happen inside the mind — it happens in the spaces between people. When a student's cognitive variation is spoken aloud, it is no longer just that student's business: it enters the air of the classroom, is received, judged, transformed, rejected, or embraced by others. A new level is born.
We call this level the meso level — the emergence of classroom culture.
10.1
Emergence is a core concept in complex systems science. It describes a phenomenon where individuals, following simple rules in local interactions, produce at the macro level a novel, unpredictable, stable collective pattern. This collective pattern is not "designed" by any single individual — it "grows" spontaneously out of the interactions.
The classic example is the ant colony. No single ant ever "planned" the colony's routes or division of labor. Each ant follows only a few simple rules — "release pheromones when food is found," "follow the direction of higher pheromone concentration," "pheromones gradually evaporate" — but when thousands upon thousands of ants follow these local rules simultaneously, the colony can find the shortest path from nest to food source. This global capability does not appear in any single ant's "brain." It emerges from the interactions among the ants.
Similar phenomena appear in the migration of bird flocks, the swimming of fish schools, and the learning of neural networks. The four characteristics of emergence can be summarized as follows:
First: Local rules, global patterns. Individuals follow only simple rules; there is no "global design." But the net effect of countless local interactions is a stable collective pattern.
Second: The pattern is irreducible. You cannot understand the ant colony by analyzing the behavior of a single ant — collective behavior has its own logic, existing between individuals rather than within them.
Third: Feedback-driven. Individual behavior changes the environment, and the changed environment in turn influences the individual's next behavior. Once this loop is activated, it is self-reinforcing.
Fourth: Sensitive to initial conditions. The tiniest deviation at the start, amplified through feedback, can lead to entirely different collective patterns. This is the so-called "butterfly effect" — what determines the direction of classroom culture is often what the teacher said the very first time, not how much they said overall.
Now, replace the ant colony with a classroom. Each ant is a student. Pheromones are the signals released by the teacher's every response. Ant trails are the class's gradually forming "way of speaking," "error safety rules," and the default pattern of "who gets seen."
Emergence in the classroom is happening every day. But you won't know it's happening — because its raw materials are those things most easily overlooked: a nod, a frown, a pause of a fraction of a second, a "Wrong, sit down."
Classroom culture is not established through declarations — it is established through modeling.
Students do not become brave because you said "be brave." They become brave because of how calmly you face mistakes, slowly feeling that "saying something wrong really isn't a big deal." They do not learn to cooperate because you posted a "Cooperate" slogan. They learn because they see you carefully listening to a struggling student finish an entire sentence, slowly feeling that "everyone's voice carries weight."
An equally important corollary: Since culture emerges through modeling, changing the culture can only happen by changing what is modeled — not through meetings, not through setting rules, but through continuously changing the pattern of every micro-interaction.
Educational researcher R. Keith Sawyer, in his book Creative Teaching: Collaborative Discussion as Disciplined Improvisation, was the first to systematically introduce the concept of "emergence" into classroom teaching analysis. His core argument:
Sawyer's contribution lies in this: he revealed a core tension in teaching — the teacher needs to prepare, but preparation does not equal control; genuine learning happens in the space where the teacher "lets go" and allows student ideas to collide, compete, and synergize through interaction. This space is the space where emergence happens.
Sawyer's framework gives us an important insight: emergence is not disorder — it is structured self-organization. The teacher's role in the classroom is not to "design the culture," but to "design the conditions for emergence."
But to make this insight actionable, we need a concrete case. What kind of differences in micro-interactions can, with the same teacher, the same lesson, in two parallel classes, "emerge" into radically different classroom cultures?
10.2
Let's return to the scene that opened the chapter.
Two parallel classes, taught by the same teacher. During the first week of school, both classes are quiet — all the students are doing the same thing: observing. "How should one speak in this class? What happens if you answer wrong? What behavior will be affirmed? What behavior will be ignored?"
The answer is hidden in the teacher's response to the first wrong answer.
First day of school. The teacher asks a question. The first student to raise their hand stands up — and answers wrong.
The gaze of the entire class converges on this student, and on the teacher's face. Everyone's nervous system is doing the same thing: reading the signal — "In this class, what happens if you answer wrong?"
Version one: the teacher frowns and says, "Wrong, sit down." Just an instant. But the signal has been sent — the entire class has received it: Wrong answer = embarrassment. Wrong answer = being negated. The safe strategy is: don't raise your hand.
Version two: same question, same wrong answer. The teacher pauses, and says, "That's an interesting angle. Can you elaborate on how you arrived at that?"
The signal is completely different — what the entire class receives is: The idea was taken seriously. Even if the direction is wrong, the thinking process has value. The safe strategy is: keep thinking, keep speaking.
The next day, subtle differences begin to emerge. In Class A, someone murmurs a response under their breath when the teacher speaks. In Class B, no one does. The source of the difference may simply be that Class A's teacher at some moment said "That's a good addition," while Class B's teacher subconsciously nodded without verbalizing. One was reinforced; one was ignored.
One week later — the difference is now observable. In Class A, a culture of "blurting out answers without raising hands" has appeared. Not because the teacher permitted it. It grew naturally. In Class B, an unspoken rule of "must raise hand and wait to be called" has formed. Also not because the teacher demanded it. It also grew naturally.
By the end of the semester, the "ways of speaking" in the two classes have fully diverged — same textbook, same exam, same teacher.
| Dimension | Class A | Class B |
|---|---|---|
| Frequency of continuing to speak after a wrong answer | High | Low |
| Frequency of voluntary questioning | High | Low |
| Duration of classroom silence | Short | Long |
| Frequency of explaining reasoning in one's own words | High | Low |
| Frequency of standard-answer recitation | Low | High |
| Number of times the teacher actively invites "silent students" | Many | Few |
No one "designed" this difference — it is the accumulation and amplification of every tiny difference across dozens of lessons and hundreds of interactions, eventually emerging into two radically different collective behavioral patterns.
A positive cycle locks itself in this way: students speak boldly → teacher responds positively → more students dare to speak → the classroom becomes increasingly open → students become bolder. A negative cycle locks itself in this way: a student answers wrong and is negated → students choose silence → the teacher feels "this class doesn't like to participate" → the classroom becomes increasingly restrictive → students are even more afraid to speak. Both cycles are self-reinforcing; once you enter one of them, it is very difficult to jump out on your own.
This is how the fourth characteristic of emergence — sensitivity to initial conditions — operates in a real classroom. Those few minutes on the first day of school are not a dramatized "fate-sealing" moment, but the initial condition of emergence. The signal released by that first response becomes the starting point from which all subsequent interactions are amplified.
10.3
The difference between Class A and Class B is not caused by a single factor. If we view the classroom as a micro "selective environment" — where certain cognitive variations are encouraged and retained, and others are suppressed and eliminated — we need to identify the selection pressures that are actually operating in this environment.
So, what are these selection pressures? At first glance, the factors affecting students in a classroom seem countless — peer gaze, task difficulty, seating proximity, exam deadlines... But if we return to the sentence at the end of Chapter 9 — the crucial moment in learning is when a cognitive variation shifts from "private" to "public" — the question becomes clear.
For an idea to travel from "inside one student's head" into "the shared selective environment of the entire class," it must pass through three gates in succession:
Gate One · Can it be produced? — Does the student have time to think the idea through? (Time Pressure)
Gate Two · Dare they say it? — Is it safe to speak the idea aloud? (Evaluation Pressure)
Gate Three · Will it be heard? — Once spoken, does anyone listen? Will the student be called on? (Power Pressure)
These three gates do not overlap: producing ≠ expressing ≠ being received. They are the three necessary stages a variation must pass through before entering the classroom's selective environment — miss any one, and the variation dies before ever reaching the "selection pool." All those other factors (peers, difficulty, seating...) either operate within these three gates or are their downstream effects. So, to understand the classroom as a selective environment, first see these three gates clearly.
After a teacher asks a question, how long, on average, do they wait for a student to answer? Mary Budd Rowe, in her classic 1986 study, found that most teachers wait only 1 second. If no hand goes up within 1 second, the teacher answers the question themselves, rephrases it, or calls on a specific student.
Mary Budd Rowe, after extensive classroom observation, found:
Rowe's conclusion: A 1-second wait time is not a "neutral" number — it is itself a powerful selection mechanism. It selects for students who are "fast responders" and eliminates those who "need more thinking time." And this latter group, in nearly every classroom, is the majority.
From the CLE perspective, the length of wait time directly affects the supply of variation. A 1-second wait only allows those variations already "ready" to enter the environment; 3–5 seconds of wait gives those cognitive variations that need time to assemble, retrieve, and associate the opportunity to compete. Teachers think they are "controlling the pace of the lesson," but in reality they are invisibly determining which students' cognitive variations get the chance to be expressed.
How it operates: Teacher wait time → Depth and breadth of student responses → Which students' variations are "permitted entry" into the classroom's selective environment.
Signal Interpretation: Waiting 1 second → "You should know the answer immediately." Waiting 5 seconds → "Thinking takes time. I'll wait for you."
CLE Diagnosis: Shortening wait time is essentially reducing the scale of the space in which everyone can generate their own ideas — this is the most invisible act of variation suppression.
If time pressure determines who has the opportunity to speak, evaluation pressure determines what kind of thing they dare to say.
A student answers wrong. What is the teacher's next sentence? There are three common response patterns:
Pattern One: "Wrong, sit down." — This negates the person. The signal transmitted: the correct answer is the only measure of value. Safe strategy: only speak when you're certain you're right.
Pattern Two: "Close — think again." — This affirms the direction of effort, but retains the right of judgment over right and wrong in the teacher's hands. Safe strategy: it's okay to try, but the ultimate standard is still "right."
Pattern Three: "That's a really typical mistake — it happens to remind us to pay attention to..." — This transforms the error into a public resource. The teacher is not evaluating the student; they are using this error case to move the entire class's cognition one step forward. The signal transmitted: error is not a dead end; it's a signpost.
Three responses, three different intensities of "selection pressure." Pattern One applies the maximum selection pressure — requiring the variation to be highly close to the correct answer to survive. Pattern Two relaxes the pressure, but the teacher remains the sole "selector." Pattern Three redefines the student's erroneous variation as useful variation — it is not eliminated, but reintegrated into the learning process.
In the CLE framework, the key to evaluation pressure is not "whether the teacher corrects errors," but how errors are retained. What Pattern Three retains is not "the wrong conclusion," but "the error as an object of analysis" — this variation is dissected, discussed, and its useful parts are collectively selected by the whole class. This is the most efficient strategy for variation retention.
The third selection pressure comes from the power distribution of classroom discourse.
Who does the teacher subconsciously call on to answer questions? The high-achieving students. Those sitting in the front rows. Those who raise their hands enthusiastically. These students obtain more "speaking rights." And those students who never raise their hands gradually become "invisible."
This observation appears trivial, but in the CLE framework, it points to a neglected problem: classroom culture is not just the culture of "active students." The ignored students are also sending signals — "I don't belong in this class's discussion." When participation in a class displays a highly asymmetric "long-tail distribution" — a minority of students occupying the majority of speaking opportunities — the class's selection pressure exhibits a systematic bias: only a subset of students' cognitive variations have the opportunity to enter the "selective environment."
Describing this mechanism in behavioral terms:
When 20% of the students in a class occupy 80% of the speaking opportunities, the classroom's effective "space where everyone thinks in their own way" shrinks to 20% of the class's total cognitive resources. The result:
CLE Core Insight: The teacher's choice of whom to call on is the most easily overlooked "power lever" in the classroom. It determines the diversity of variation — and the reduction of diversity means a decline in the entire class's adaptive capacity.
Consciously calling on students who don't raise their hands, inviting their participation with smaller steps — this is not "taking care of the students who are behind but progressing"; it is restoring the diversity of the space where everyone thinks in their own way. Every "seen" silent student sends a more inclusive signal to the entire class — "There is a place for everyone here."
Gate One · Time Pressure — After asking a question, how long did you wait? (1 second vs. 5 seconds — different outcomes)
Gate Two · Evaluation Pressure — After a student answers wrong, what was your first reaction? (Negate the person vs. Affirm the effort vs. Make the error a public resource)
Gate Three · Power Pressure — Who is speaking? (How much speaking time do the top 20% of students occupy? How long since a silent student was last seen?)
These three questions can be answered by any teacher immediately after their next class. They are not theory — they are observable classroom indicators. The wider each gate opens, the more and the more diverse the cognitive variations that enter the classroom.
10.4
Having understood the mechanism of emergence and how the three selection pressures operate, a natural question surfaces: What can a single teacher do, without being able to change class size, schedule, or textbooks?
The good news: the properties of emergence dictate — starting from small changes in initial conditions, the amplified results can be remarkably significant. You don't have to "change the entire river." You only need, when the first ripple appears, to give a push in the right direction.
The following five micro-actions can be implemented by any teacher starting from their very next class. They require no additional resources, no leadership approval, no textbook changes — only the teacher's awareness of the signals they are releasing and selective adjustment.
The above five actions are the "ideal classroom" full version. In a real large-class setting, you don't need to do all five — pick one, do it for a month, and see if the classroom culture changes. If "letting everyone articulate their erroneous reasoning" is unrealistic, at least achieve letting the whole class see: the teacher's reaction to the first wrong answer is curious, not corrective. You don't need to change everything — you only need to change the nature of that first response — that moment is the starting point of the entire class's "variation switch."
Three Subject Scenarios · Pick One to Start
▶ Mathematics: Publicize Errors
When returning a unit test, don't start with the correct answers. Use a projector to display three of the most typical wrong problems (anonymized). Ask the entire class: "What do these three students' approaches have in common? At which step did things go off track?" If the first student who raises their hand says something wrong, don't say "wrong" — say "Interesting — where do you think they got stuck?" This action only takes 5 minutes after returning the test, but the signal it releases is: errors here are not a source of shame; they are a public teaching resource.
▶ Chinese/Literature: Respond to Errors with Curiosity
When teaching reading comprehension, a student misses the main idea of the passage. Don't immediately correct them. First ask: "Can you find a sentence in the text that made you think this?" If the student points one out, the whole class looks at that sentence together — then continue probing: "Then what do the title and the ending of the passage suggest?" The student discovers their own deviation through comparison, rather than being corrected by you. Your role throughout is only to ask questions, not to judge.
▶ English: Ask Why
During a dictation exercise, a student spells "receive" as "recieve." Don't tell them to "copy it ten times." First ask them: "What rule did you use to spell it this way?" The student might say "'i before e' rule..." — then you know, they weren't being careless; they have a gap in their understanding of English spelling rules. Next, have the whole class discuss: "In which cases does the 'i before e' rule not apply?" A single spelling error unearths an entire gap in understanding of a grammar rule.
Summary
Classroom culture is the core of the meso level — it is neither the micro "individual cognition" nor the macro "institutional design." It lies in the middle, the emergent layer between the two.
This chapter's core message can be condensed into three propositions:
Proposition One: Classroom culture is not designed — it emerges. It is not written in the opening-week class meeting script — it "grows" naturally through every teacher-student interaction, via BVSR's selection algorithm.
Proposition Two: The teacher cannot control the direction of emergence, but can design the conditions for emergence. The three selection pressures — time pressure, evaluation pressure, power pressure — are the most direct "condition-designing tools" in the teacher's hands.
Proposition Three: Emergence is sensitive to initial conditions. The first response to the first wrong answer on the first day of school is not a "detail" — it is the initial condition for the entire trajectory of classroom culture. To change the culture, you don't need to change the entire river — you only need to change the direction of the first ripple.
But here is a question: if classroom culture "emerges," what about school culture?
Those three selection pressures in classroom culture — time pressure, evaluation pressure, power pressure — where do they come from? Why are they so similar across so many teachers? What shapes the teacher's "way of responding to the first wrong answer"?
The answers to these questions all connect to a larger ecology — the school-level rules, evaluation systems, and cultural inertia. The "emergence space" of classroom culture is, in fact, bounded by the school as a larger ecosystem. Teachers are not free agents — they teach within a school's institutional environment.
This is the level we will enter in the next chapter: the macro level — School as Ecosystem.
The meso level reveals how classroom culture emerges through every interaction —
But the classroom does not exist in isolation.
It is embedded in a larger system.
That system — the school —
How does it determine where the classroom's "selection pressures" come from?
How do institutional-level rules, culture, and evaluation systems
become a higher-level selective environment?
Next Chapter: Chapter 11 — "The Macro-Level: School as Ecosystem"
Sawyer, R. K. (2004). Creative teaching: Collaborative discussion as disciplined improvisation. Educational Researcher, 33(2), 12–20. (Classroom emergence and improvisational teaching)
Rowe, M. B. (1986). Wait time: Slowing down may be a way of speeding up! Journal of Teacher Education, 37(1), 43–50. (Classic wait time research)
Mazur, E. (1997). Peer Instruction: A User's Manual. Prentice Hall. (Peer instruction methodology)
Kulik, J. A., & Kulik, C.-L. C. (1988). Timing of feedback and verbal learning. Review of Educational Research, 58(1), 79–97. (Feedback timing meta-analysis)
Chi, M. T. H., et al. (1989). Self-explanations: How students study and use examples in learning to solve problems. Cognitive Science, 13(2), 145–182.
In the previous chapter, we established a core proposition: classroom culture is not designed — it "emerges" through every teacher-student interaction. The teacher cannot control the direction of emergence, but can create favorable conditions by adjusting three forces — how time is allocated, how evaluation is phrased, and who gets the chance to speak.
But one question has been hanging in the air: where do those "adjustment tools" in the teacher's hands — how long to wait before calling the next student, how to respond to a wrong answer, whom to give the speaking opportunity to — come from?
Why do so many teachers respond so similarly to the first wrong answer? Why does a teacher who dares to let students make mistakes at School A become cautious and hesitant after transferring to School B? Why, no matter how elegantly institutional documents are written, does the atmosphere of the school remain unmoved?
The answers to these questions do not lie inside the classroom. They point to a larger ecology — the school. The classroom is the frontier of innovation, but the institutional system is the logistical system behind the front line. The school's rules, evaluation systems, cultural inertia, and time structure — these together form an entire invisible selective environment that determines what teachers "can do" in the classroom and "what is worth their while to do."
In the language of CLE: the classroom is the "front line" of innovation; the school is the "rear" that exerts decisive influence. Whatever the rear rewards and punishes, the front line will grow to match that style.
In this chapter, we enter the third analytical level — the macro level. Here, the school is not a machine designed from a blueprint; it is a living ecosystem. Reform cannot simply replace one bolt; it must simultaneously adjust four interlocking components.
Macro-Level: Institution as Selective Environment
When I was a principal, I experienced a moment that troubled me for a long time.
The reform stopped. Not the kind of stop that comes from failure — no opposition, no protest, no one saying "Principal, you're wrong." It just silently stopped, like a stone thrown into a pond: after the initial ripples faded, the water returned to stillness.
What I had been pushing was "rebuilding the relationships between people in education" — breaking the binary opposition between managers and the managed, between transmitters and receivers, so that everyone would become an equal learner. I led by example: I listened seriously to dissenting opinions, crouched down to speak with students, stood in the hallway and chatted with them. The effects were real — middle management said "Principal Wang is different from other principals," teachers said "at Yingjian you can speak honestly," students said "the principal actually listens to us." I almost believed the culture had been fully established.
But that was where it stopped. When middle management dealt with teachers, the old patterns returned — "This is what the principal requires; you have to execute it." When teachers dealt with students, the old patterns returned — "The standard answer is this; don't overthink it." The students also kept to the old patterns — exercising the power of their "positions." They enjoyed the equality others gave them, but when dealing with others, they identified with power as more efficient.
That lesson from failure cut deeper than any success. It made me truly understand one thing: culture is not something a single person can "build."
Schools were not designed by a sage sitting at a desk. The academies of ancient Greece took shape as Plato walked and talked with his students beneath the olive trees. The private schools of China gradually crystallized when a village elder who could read a few characters was invited by neighbors to teach — slowly settling into form over time. The modern subject-divided school emerged from the Industrial Revolution's need for large numbers of literate workers — society forced a new form into being. Every morphological change followed the same cycle: the old ways stopped being enough → try something new → the new ways stabilized → they too stopped being enough. No one sat in God's seat to draw up a "blueprint for schools" that humanity then constructed.
Schools are not designed — they evolve. Institutions, culture, teaching paradigms — every layer of structure in the school ecosystem formed gradually through history via the cycle of "try new ideas → accepted or eliminated by the environment → what survives becomes institutionalized."
The practical significance of this perspective: the object of reform is not "fixing a machine," but "changing the direction in which a forest grows." When a machine breaks, you can replace a part. But after replacing the part, the forest will not immediately reorganize itself to match the new component — it needs time, it needs conditions, it needs change throughout the entire forest.
Institution vs. Culture
The relationship between institution and culture is often misunderstood. Many people believe institution sits on top and culture on the bottom — once institutions are set, culture follows. But the truth is the reverse. Culture is a deeper operating system than institution. Institution is the application layer; culture is the kernel layer.
You can install a Mac-style skin on a Windows computer, but the underlying system is still Windows. Run it long enough, and everything will readapt to the logic of Windows. This explains why institutions can be rewritten beautifully while culture remains unmoved — you are trying to change a kernel-layer variable at the application layer.
I have seen more than one school where the administration established very detailed "encourage innovation" systems — monthly innovative teacher awards, innovation bonuses, innovation sharing sessions. But the core of the school's evaluation system was the end-of-term standardized test rankings. Teacher promotions were tied to rankings, and parent chat groups compared scores every day. How long do you think the "encourage innovation" system lived? After a brief gust of wind, the teachers returned to their familiar operating patterns. Not because they didn't want to innovate — but two opposing forces were pulling at them simultaneously, and the deeper institution — evaluation — won.
The fast variable is always dragged back to the origin by the slow variable. Teaching methods are fast variables. Evaluation systems are slow variables. If you change the classroom but not the assessment, the fast variable will be pulled back by the slow one. This is why reforms keep going in circles. An analogy: you keep a fish tank; you swap in a new fish, but the water temperature, water quality, and filtration system haven't changed — the new fish won't live long. The classroom is that new fish; the evaluation system is the water temperature and quality. Change only the fish without changing the water, and the new fish will sooner or later be bred to death.
Two Clocks
As a principal, I carried two clocks in my mind. One watches the days — lesson progress, project developments, administrative paperwork, daily management. The other watches the seasons — the cognitive transformation of teachers, the cultural evolution of classrooms, the sedimentation of school values. The rhythms of these two clocks are entirely different.
The source of frustration for many principals is using the clock of days to measure seasonal matters. A new teaching method is introduced at the start of the semester; midterms show no effect and anxiety rises; by end-of-term it is abandoned when results still haven't appeared. Then a new method is introduced. Teachers, having just begun to get a feel for the new approach before it could become internalized, are informed that "we're not doing that anymore; we're doing this now."
In schools where direction changes frequently, teachers are forever starting from zero.
From the CLE perspective, the conflict between two timescales boils down to a simple question: you are using the rhythm of planting vegetables to measure how much a tree has grown. A teacher learning a new teaching method (equivalent to having a "new idea") can complete this within weeks, but for a new practice to be accepted, recognized, and become habit at the school level — that requires spanning multiple cycles. When higher-level evaluation applies pressure on a "monthly" timescale, you are essentially using a short-term metric to measure a change that needs a long time to stabilize — it's like dumping the bacteria out of the petri dish every few weeks and replacing them, then wondering why the bacterial colonies never form.
Conditions for Cultural Growth
Culture cannot be commanded, but it can be cultivated. Just as a gardener cannot command a plant to grow but can create conditions suitable for growth. From the CLE perspective, the core mechanism of cultural growth is: create a stable environment where good behaviors have ample opportunity to be repeatedly practiced, seen, and retained.
A set of ordinary systems with clear direction, consistently executed for five years, far surpasses five sets of "perfect" systems changed annually. The core of culture is "making things predictable" — teachers and students need to know "what this school truly values," and that this valuing is sustained and credible. If this year emphasizes scores, next year emphasizes holistic development, and the year after reverts to scores, everyone will enter "waiting for the trend to pass" mode.
Consistency creates safety. Safety is the soil in which culture grows. From the CLE perspective, the essence of consistency is simple: if you want people to follow a certain approach, give them enough time to follow it. If the school's demands point one direction today and another tomorrow, everyone is interrupted just as they begin to adapt and can never form stable habits.
Culture is not the principal's solo work. It is what the entire school community, through daily interactions, slowly precipitates — like a coral reef. If the institution is "formulated" by the principal's office, "communicated" at faculty meetings, and "implemented" by teaching-research groups — it will forever remain an institution and never become culture.
Let teachers participate in the construction process of institutions. Not performative consultation — genuine co-design. When teachers feel that "this rule was arrived at through our collective deliberation," compliance transforms from "obedience" to "identification." From the CLE perspective, the logic is simple: if only the principal is generating ideas for the institution, the pace of progress is capped at one person. If all the school's teachers participate, there are dozens of people generating ideas together — the pace of progress becomes entirely different.
In the process where old culture dissolves and new culture forms, there is a phase of minimum efficiency. Old practices have been questioned, but new practices are not yet proficient. During this time, the school's operations appear "messy." Many principals panic at this stage and hurriedly restore the old system — "stability above all."
But this is precisely what cuts off the possibility of the new culture taking shape. The inefficiency of the transition period is the growing pain of gaining height. You cannot negate the entire process just because results are not immediate.
CLE translates this problem more bluntly: any systemic change must pass through a period where new ideas burst forth in concentration — before the old patterns are eliminated and the new ones are settled, there is inevitably a phase where various new practices are being tried and the system appears "unstable." This is precisely evidence that the system is growing, not that something has gone wrong. Suppressing these new ideas is equivalent to depriving the system of the possibility of generating better practices.
The Interlocking of Four Components
Having understood that schools "grow," the next question becomes unavoidable: why do reforms always go in circles and return to where they started?
Over a hundred years ago, the American Progressive Education movement failed. Not because the ideas were wrong — on the contrary, they worked. Student-centered learning, project-based learning, collaborative inquiry — these are common sense today. But in 1957, a satellite was launched into space, national panic ensued, and everything was overturned. Nearly a hundred years later, generation after generation of principals still repeats the same pattern: begin with burning passion, achieve some effects, and then be consumed by larger forces.
It's not just one principal. For twenty years as a principal myself, every teaching reform attempt, large or small, began with the same passion and ended with the same regret — the classroom did indeed change, but after a while, everything slowly slid back to the starting point.
Why? You only changed the classroom, but everything surrounding the classroom remained untouched.
CLE theory provides a systematic diagnostic framework: the classroom is not an isolated island within the school. It is a "source of innovation" within the school's ecosystem. Any new practice it generates must pass through the layer-by-layer screening of four components — if any one of these components does not recognize the practice, the new approach will be eliminated.
These four components are not four unrelated domains — they are tightly interlocked. Their relationship can be understood as a chain:
Classroom teaching produces new practices (a new teaching method, a new interaction pattern) → The evaluation system determines whether this new practice can be recognized (does the exam reward this kind of ability?) → Teacher development determines whether teachers have the capacity to sustain the new method (is there training? Is there space for trial and error?) → Parent cognition determines whether the school can withstand external pressure (do parents support or oppose?) — and parents, through social pressure, feed back into the evaluation system and classroom teaching, forming a complete loop.
If any one component doesn't move, the innovations of the other three will be dragged back to the starting point by it.
This is not theoretical deduction. This is the real dynamic unfolding in every school that attempts reform. Classroom teaching changes, but the evaluation system still rewards old behaviors — the teacher is pulled by two opposing forces and ultimately chooses to "do a bit of both," then slowly slides back to the old pattern. Teacher development has no supporting measures — the teacher wants to change but doesn't know how and doesn't dare to; new methods are stiffly grafted onto old habits. Parent cognition hasn't caught up — parents see the new methods as "not looking like proper studying"; pressure flows in from outside; the school can't withstand it, and reform stops.
Fullan (2007), in his long-term tracking of educational reform, found that whether reform can deepen and sustain depends on synchronous changes across four dimensions: teacher practice, evaluation systems, school culture, and external policy. Changing any single dimension alone almost never succeeds. The easiest things to change are teaching materials and classroom activities, but what truly determines whether reform can last — teachers' educational beliefs and the school's institutional culture — is far harder to change than most people imagine.
Tyack & Cuban (1995), in their systematic review of over a century of American educational reform history, found that most reforms stop at the level of "tinkering" — changing the classroom without changing the evaluation, changing the method without changing the institution. They discovered: the deepest structures of schools — what time classes are held, how students are grouped, how they are evaluated — are extraordinarily stubborn, because they are not decided by a single principal but are supported by the entire social system.
Once we understand the interlocking relationship of the four components, the path of reform becomes clear. Not grander — but smarter.
Step One: Don't try to overturn the entire system at once. First, run through the process in one class or one teaching-research group. The purpose of small-scale piloting is to accumulate experience in "how to operate in a real environment." One successful case is worth more than a hundred pages of reform plans.
Step Two: Move at least two of the four components in sync. If resources are limited, at minimum ensure classroom teaching and the evaluation system move together. Classroom changes but evaluation doesn't — the teacher knows how to run the new classroom but knows it won't add points at the end of term; motivation drops to zero. Evaluation changes but classroom doesn't — teachers still use old methods but the evaluation requires new indicators; anxiety doubles. Only when both move together can teachers "change with peace of mind."
Step Three: State clearly from the beginning that "this will take three years." Most reform failures happen not because effects didn't appear, but because support was withdrawn before effects had time to appear. Communicate clearly with all stakeholders from the start — while also demonstrating intermediate results so that external parties see reform is underway. Managing expectations is one of the most important tasks of a reformer.
Finland's educational reform provides a more complex reference. Many people imagine it was a linear path from "classroom to institution," but the actual process was multi-dimensional synchronous advancement: the comprehensive school reform eliminated ability tracking, dramatically raised teacher education standards (all teachers must hold a Master's degree), granted high local curricular autonomy, and simultaneously eliminated standardized external examinations. These measures were not completed in a "classroom first, then institution" sequence — institutional reform and classroom practice operated simultaneously at multiple nodes. As Sahlberg pointed out in Finnish Lessons (2011), the uniqueness of Finland's reform lay in the high trust and bidirectional penetration between institutional design and grassroots practice.
It took over twenty years, not two.
Finland's case reveals one core lesson: don't try to change the entire system at once. Start by changing the rules at one point, and let the change transmit outward, link by link, between adjacent components.
The Trap of Progress
The Progressive Education movement (1919–1957) provides an extraordinarily expensive lesson — Chapter 13 will unfold this case study in complete historical detail. Here I extract only the conclusions directly relevant to the CLE framework: Progressive Education proved that classrooms can produce critical thinking and autonomous learning ability (the Eight-Year Study had empirical data), but it failed — because the exam system couldn't see these outputs, parents exerted counter-pressure from outside, and when a social crisis hit, everything reverted to zero. Change the classroom without changing the institution, and new practices will be rejected by the old system.
The lesson of Progressive Education points to a key proposition:
In any educational system, only those things that are "seen" by formal measurement standards (exams, evaluations, promotions) have a chance of being systematically retained and scaled. Things that are not seen — no matter how educationally valuable — cannot survive at the institutional level.
Institution as Selection Pressure
If the Visibility Principle of Selection holds, then an even more fundamental question surfaces: institutions are not "managing" behavior — they are "selecting" behavior.
This shift in perspective is critically important. The implicit assumption of most educational reform is: institutions are neutral management tools — once the institution is set, behavior will move in the expected direction. If behavior does not move correctly, then either the institution is not detailed enough or the execution is not strict enough.
CLE theory proposes a different perspective: institutions are not neutral management tools; they are themselves a set of reward-and-punishment rules. They determine what practices within the school will be "retained" and what will be "eliminated."
Every institutional design is answering an implicit question: "In this school, what is worth doing?"
If the sole criterion for teacher promotion is student end-of-term average scores — this institution effectively releases a clear signal: "Anything you do in the classroom unrelated to raising scores falls under 'not worth it.'" Teachers are not unwilling to innovate. They are making the most rational choice within a rational environment.
If exams test only standard answers — the signal this institution releases is: "Proposing different ideas is 'not worth it.' Writing the standard answer is 'worth it.'" Students are not unwilling to think. They are making the most adaptive response within a system that rewards "standard" rather than "unique."
This is the power of institutions: an institution is like a "reward-and-punishment map" — on this map, certain behaviors are marked "doing this brings benefits," and certain behaviors are marked "doing this brings no benefits" or even "doing this brings trouble." Everyone, in their own position, naturally gravitates toward the path marked "benefits." This is not opportunism — it is the normal response of any sentient system to its environment.
Cuban (1984), in How Teachers Taught, conducted a systematic historical analysis of over a century of American classroom teaching. His finding is thought-provoking: educational theories and policies changed again and again, but the core of the classroom — teacher lectures, students listen, textbook-driven — was astonishingly stable.
CLE theory offers a simple explanation: it's not that teachers are stubborn; it's that reward-and-punishment rules have shaped the same behaviors generation after generation. Every generation of teachers faces the same evaluation system, the same class sizes, the same time structure, the same textbook regime — each, in their own separate classroom, "independently" made very similar choices. This looks like "inertia," but in reality everyone walked the same path under the same rules.
Once we clarify that "institutions are reward-and-punishment rules," institutional design transforms from a question of "how to manage people" into a question of "how to design the rules of the game." The following three principles form the institutional design framework proposed by CLE theory:
Every school has "what it says with its mouth" (mottos, visions, slogans) and "what it does with its hands" (promotion criteria, bonus distribution, career advancement paths). When the two are inconsistent, the school will absolutely operate according to "what the hands do" — not because people lack integrity, but because what truly determines people's actions are tangible benefits, not elegant slogans. The first step of institutional design is not writing better slogans, but seeing clearly: in the current system, which behaviors actually bring people real benefits?
If a school's goal is "cultivating creative thinking," but the sole reward criterion is end-of-term standardized test ranking — then without any analysis whatsoever, one can predict the goal will not be achieved. This is not a methodological problem — it's that rewards naturally steer toward the direction being rewarded. What institutional designers need to do is not add a pile of new titles like "Encourage Innovation," but check: are the existing reward-and-punishment rules consistent with the direction you want to head toward? When they are inconsistent, adding more new institutions is just noise.
This connects back to the "Two Clocks" framework from Section 11.1. Reward-and-punishment rules are not a notice posted once and done — they need continuous operation. The culture a school genuinely wants is something that can only settle through at least 3–5 years of stable, consistent reward-and-punishment rules. A system that changes today and swaps tomorrow is like walking east today and west tomorrow — the people in the system can never form stable habits.
By defining "what behavior brings benefits, what behavior brings costs, what behavior is irrelevant," institutions draw a "reward-and-punishment map" for everyone in the school. People are not "managed" by institutions — rather, on this map, they find the most advantageous path for themselves.
Reform fails to advance not usually because people don't agree with the reform vision — but because the path indicated by the reform vision (the long-term benefits of new practices) still looks insufficiently clear and insufficiently certain on this map, compared to the safety of the old path.
From this perspective, the core task of institutional design is not "telling people what to do" — it is "drawing this map well, so that good practices naturally become the path everyone most wants to take."
The Panorama of the Macro Level
In Section 11.1, we established that schools are not design blueprints — they "grew" generation by generation through the process of "trial and error → see what works → keep it." We distinguished the relationship between institution (the surface operation interface) and culture (the underlying operating system), proposed the framework of Two Clocks, and the three conditions for culture to grow naturally.
In Section 11.2, we translated the dilemma of reform into an interlocking chain — four components (classroom teaching / evaluation system / teacher development / parent cognition) linked together; if any one doesn't move, the other three will be dragged back to the starting point.
In Section 11.3, through the historical case of the Progressive Education movement, we revealed a painful truth: only things that can be "seen" by measurement standards have a chance of truly being retained. This is not about right or wrong — it is an objective law of how systems operate.
In Section 11.4, we redefined institutions as "reward-and-punishment maps" — institutions are not managing people; they are pricing every person's behavior. Changing culture is not about writing better institutional texts — it's about redrawing this map so that good practices naturally become the "most worthwhile" path.
The first three sections analyzed the internal operating mechanisms of the school institution as a selective environment — how the four components lock each other in place. But on a larger timescale, schools face a new kind of external selection pressure: one that does not come from Ministry of Education documents, parent complaints, or international comparisons — it comes from a foundational technology that has changed the mode of information production.
The traditional school rests on three pillars. For the past 150 years, these three pillars have supported the legitimacy of the concept of "school." But AI is dismantling them one by one:
First Pillar: Knowledge Distribution. The oldest function of schools — transmitting information from those who know to those who don't. The teacher stands at the podium, textbooks lie open on desks, the library is a building. This pillar was built on "information asymmetry." Now a student pulls out a phone and asks AI — the density and accuracy of the knowledge obtained may exceed what they absorb in an entire class. Information asymmetry has been demolished by technology. Knowledge distribution is no longer the school's core value.
Second Pillar: Standardized Testing. Schools need a fair, comparable measurement tool — who can advance, who should repeat a year, which teacher is "teaching well." Standardized testing is currently the only large-scale viable solution. But GPT-4 scoring 130 on the gaokao math section exposes a fatal fact: a system that has no idea what "understanding" means can score higher on this measurement tool than most humans can. This shows that the exam is not measuring understanding — it is measuring brute-force coverage ability within a low-dimensional closed space. When the measurement tool is not measuring what you think it's measuring, this pillar begins to wobble.
Third Pillar: The Physical Classroom. We lock a group of people in the same room not because "co-presence" has any educational magic — but because technological conditions made it the only option. For one teacher to teach fifty students simultaneously, they had to be placed in the same physical space. But AI has already proven: personalized learning does not require co-presence. Physical space is no longer a prerequisite for learning.
With each pillar that falls, the school's legitimacy drains away. This is not AI attacking the school — it is AI dismantling the physical foundations of the "old school."
When information no longer needs distributing, exams no longer measure what matters, and classrooms are no longer necessary —
what remains of the school?
In the language of CLE: what remains is precisely the school's irreplaceable nature as a "macro-level selective pressure environment." AI can replace knowledge transmission — but AI cannot replace a classroom culture in which the teacher genuinely looks forward to seeing each student. AI can generate perfect-score answers — but AI cannot replace the argument and epiphany sparked between classmates by a wrong answer. AI can let you learn anywhere — but AI cannot replace the feeling, when you walk into a classroom, that "the people here care about what I care about."
This is not a crisis for schools. On the contrary — the three pillars AI is dismantling were never the school's true value. They were merely the form the school had to grow into during an era of technological constraints. By dismantling these forms, AI forces the school back to its irreplaceable essence: a social and cultural cognitive ecosystem.
CLE Perspective: School is not a venue — it is a selective pressure environment. AI has taken away the "venue" part. It has returned the "selective pressure environment" part to the school. This is the greatest return to essence in the history of education.
But this chapter leaves an equally profound question: Micro, meso, macro — the analytical framework of three levels is now complete. But how are they interwoven in a real educational ecology? How does one student's cognitive shift influence school institutions through classroom atmosphere? How do the school's reward-and-punishment rules, in turn, shape every teacher's attitude toward students?
These three levels are not independent of one another — they are entangled, forming a complete cycle from individual to classroom to school. In the next chapter, we will weave the three levels together and examine how this cycle spins — when it spins well, when it gets stuck, and at which nodes we can give it a push.
The micro level examines how individuals learn —
The meso level examines how classroom culture grows —
The macro level examines how school institutions reward and punish —
But these three levels do not operate in isolation.
They are entangled, forming a
complete cycle from individual to classroom to school.
How does this cycle spin?
Under what conditions does it spin better and better?
Under what conditions does it get stuck?
And at which points can we give it a push?
Next Chapter: Chapter 12 — "The Three-Level Cycle: How the Classroom Influences the School, and How the School Influences the Classroom"
In Chapters 9 through 11, we disassembled the three levels of CLE theory and examined each in turn. The micro level (Chapter 9) revealed how individual learning "varies, selects, and retains" in every cognitive interaction; the meso level (Chapter 10) revealed how classroom culture grows out of teacher-student interaction; the macro level (Chapter 11) revealed how school institutions, as a reward-and-punishment environment, define what is "worth doing" for teachers in the classroom.
But one question has been set aside: how do these three levels interlock?
They are not independent of one another. A new idea sparked by a student in class seeps along teacher-student interaction into the classroom atmosphere; the patterns of that classroom atmosphere travel along teacher evaluation and parent pressure to the institutional design of the entire school; and the school's institutions deliver the final verdict — determining whether that student, in the next class, "still dares" to spark another new idea.
This is a cycle with neither beginning nor end. The three levels are not three segments of pipe joined together — they are three intermeshing gears.
Under what conditions does it trend toward a virtuous cycle that spins smoother and smoother — individuals learning more boldly, classrooms becoming more open, schools growing more vibrant?
And under what conditions does it sink into a vicious cycle that locks ever tighter — individuals growing more conservative, classrooms becoming more lifeless, schools growing more rigid?
More importantly: at which node can we leverage it?
This chapter answers these three questions.
A Systems-Thinking View of CLE's Levels
The previous three chapters respectively established the analytical frameworks for the micro (Chapter 9), meso (Chapter 10), and macro (Chapter 11) levels. I will not repeat the definitions here, only distill the most crucial structural relationships:
Micro Level — Individuals construct their knowledge through "try new ideas → see what works → retain it." Meso Level — Classroom culture grows from micro-level interactions, regulated by how teachers allocate time, how they evaluate, and how they distribute speaking opportunities. Macro Level — School institutions set the reward-and-punishment boundaries, determining what is "worth doing" for teachers in the classroom.
But CLE's true leverage lies not in the layering itself — but in how these three levels intermesh and spin together. They are not an architectural diagram of tiers. They are a mutually nested cycle: each level is both the source of new material for the next level and the environment that determines whether the previous level's new material survives.
In the language of systems thinking (Senge, 1990), this is a system where the components interlock but transmission always involves time lags: bottom-up movement is gradual accumulation — one new classroom idea will not directly change school institutions, but a hundred times, a thousand times, it will; top-down movement is gradual permeation — a new institutional policy will not change classroom culture the next day, but sustained, consistent reward-and-punishment rules will, over a semester or an academic year, slowly reshape the behavior of teachers and students.
In a multi-level learning ecosystem, what each level "can do" is jointly determined by the reward-and-punishment rules of the level above it and the output of new ideas from the level below it. The lower level provides the "raw materials" for the higher level — without abundant micro-level new ideas, both meso-level culture and macro-level institutional innovation lack sources. The higher level sets the "rules" for the lower level — institutions determine what is worth doing and what is not, directly influencing everyone's direction of action.
The direct consequence: any "problem" you observe at one level often has its root cause at another level. A student "dares not make mistakes" — on the surface this looks like an individual psychological issue, but the root cause may lie in the macro-level evaluation system. A school "can't push reform through" — on the surface this looks like a problem with institutional documents, but the root cause may be the lack of supporting teacher development at the meso level.
Changes at the micro level can be very fast — when a student's unconventional idea is affirmed in a single class discussion, their mode of thinking may begin adjusting within minutes. Changes at the meso level require weeks to months — a new teacher-student interaction pattern must be repeated enough times to become part of the "classroom atmosphere." Changes at the macro level require years — the real impact of a new evaluation system takes an entire semester or even an academic year to become visible.
This difference in timescales is a major reason why reforms "take so long to show results." You changed the institution (macro), but classroom behavior (meso) and individual performance (micro) will not follow in sync — they need transmission time.
From Description to Dynamics: Two Cycle Paths
Having understood the nested structure of the three levels, let us examine the first possible pathway: the positive cycle — the kind that spins smoother and smoother.
A positive cycle can begin at any level — but the most natural starting point is the micro level.
Step One: A student raises a "non-standard" answer in class — not the conventional line of reasoning, but an alternative valid way of thinking. At the micro level, this is one person producing a new idea.
Step Two: The teacher does not immediately negate or correct it, but pauses for three seconds (regulation of time pressure), then responds in a neutral, curious tone: "That's an interesting line of thought — how did you arrive at it?" (regulation of evaluation pressure). This response releases a signal: here, different ideas are welcomed.
Step Three: Other students receive this signal. The next time, not just one but two or three students are willing to propose non-standard ideas. When this interaction pattern repeats within the classroom, it begins to upgrade from individual student behavior to class-level culture — this class "is a place where you can express yourself freely, without fear of being wrong."
Step Four: The class's positive performance begins to be noticed by school administrators — students are more engaged, the classroom atmosphere is lively, and even midterm exam scores haven't dropped because of "error tolerance." If the school's institutions are sufficiently flexible, administrators will begin to reflect: Is the current evaluation system too rigid? Should teachers be given more autonomy? Once institutional adjustments are made — for example, shifting from "only looking at exam rankings" to "multi-dimensional evaluation that includes classroom participation and innovative performance" — that is change at the macro level.
Step Five: The new institutions release clearer signals: innovative teaching is rewarded, error-tolerant culture is supported. This signal in turn reinforces teachers' "sense of safety" in the classroom, empowering them to wait longer, respond more inclusively, and distribute speaking opportunities more equitably in more contexts. More students therefore dare to produce new ideas. The cycle accelerates.
In this positive cycle there is one critical turning point: the "leap" from individual behavior to institutional change. Not every vibrant classroom automatically drives institutional change. Too many classrooms are "islands of innovation" — there is one outstanding teacher in the school whose classroom is brilliant, but the school's evaluation system is entirely unmoved. This is because transmission from micro to macro requires an intermediate bridge: either enough classes simultaneously demonstrate similar positive patterns (critical mass), or a particular class's positive performance catches institutional attention (precedent effect), or teachers proactively translate classroom experience into institutional recommendations through formal channels (institutional pathway).
Condition One: Sufficient safety to initiate. The first step of the positive cycle — students daring to voice that "non-standard" answer, teachers daring to wait those extra three seconds — requires a sense of safety. In systems lacking safety, new ideas are suppressed from the start, and the positive cycle cannot even begin.
Condition Two: Institutional pathways to transmit "micro-level success" upward. Positive meso-level culture needs mechanisms to be "seen" — teacher debriefings, teaching case sharing, student feedback, administrators' classroom observations. Without institutional pathways, no amount of positive micro-level behavior will be anything more than "isolated excellent classrooms."
Condition Three: Institutions need to respond fast enough. If an innovative teaching practice has already proven its effectiveness, but the school needs two to three years to adjust its evaluation system, that delay itself will strangle the cycle. Teachers will say: "Forget it — at the end of the year they still just look at rankings." — The cycle breaks.
The positive cycle is the ideal pathway of educational improvement described by CLE theory. But in reality, the other path is far more common.
Two Cycles, One Structure
Now let's look at the second path.
We still begin at the micro level.
A student writes a non-standard but correct problem-solving sequence. They hesitate slightly before raising their hand, because they know this line of reasoning doesn't quite look like the "standard answer." The teacher glances at it and says: "This method won't work on exams — just follow the standard procedure." — No waiting, no exploration, direct negation. The signal the student receives: different ideas are not encouraged. The safe choice is "follow the standard."
When this signal repeats over a day, a week, a semester, the students in this class learn one thing: don't produce "strange" ideas in class. Classroom interaction becomes increasingly predictable — teacher asks, standard-answer-style response, next question. No more "different" ideas emerge. The classroom atmosphere shifts from a place of free expression to "a place where only standard answers count."
This classroom atmosphere is transmitted to the administration. When school administrators walk into the classroom, what they see is quiet, orderly, with students' answers clean and crisp. Administrators are satisfied — "This class has good discipline; their scores are stable." In the absence of multi-dimensional evaluation indicators, this "orderliness" is easily mistaken for "a good classroom." Administrators have no motivation to modify the evaluation system — after all, the current system "works."
Worse: this system will in turn affect teachers' career development. If a teacher tries innovation and tolerates "non-standard" lines of thought, their classroom may appear "messy" — students discussing heatedly, but less organized than standard-response uniformity. In an evaluation environment that only considers discipline and scores, this teacher will suffer in evaluations and promotions. So, they too begin to "return to the standard."
One teacher returns to the standard, two teachers return to the standard... eventually, all teachers move in the same direction. The school becomes an efficient standard-answer production machine. Students' new ideas are systematically suppressed. The micro level no longer produces "different" ideas — because "different" has been eliminated at every preceding checkpoint.
This is what CLE theory calls the negative cycle (or getting stuck): the three levels form a set of mutually reinforcing, ever-narrowing loops. It's not that "there is no reform" — it's that reform is dissolved by the system's own reward-and-punishment logic.
Note the cruelest thing about this negative cycle: it looks "good" locally. Student exam scores are stable, classroom order is impeccable, parents are satisfied, administrators are untroubled. It has no visible "problems" — its problem is precisely that there is "no problem." The system is in a state that appears stable but is in fact highly fragile. In the short term, this state looks highly efficient — but it has no redundancy, no resilience, no capacity whatsoever to withstand changes in the environment.
Tyack & Cuban (1995), in Tinkering Toward Utopia, described the school's "grammar of schooling" — subject division, grade segmentation, class grouping, grading — as having barely changed over the past century despite countless waves of reform washing over it. This is the classic manifestation of systemic lock-in. Not because "no one tried to reform" — but because every reform attempt was pulled back to the starting point by the system-level negative cycle.
Even more worth reflecting on is Fullan's (2007) finding: most educational reforms touch only the surface (curriculum content, textbook editions, class-hour allocations); very few reforms reach the system's deepest layers — teachers' educational beliefs and the evaluation system. Because touching the deep layers triggers the system's "resistance to reform" — not human stubbornness, but the system's own self-stabilizing mechanism.
Under the CLE framework, this "resistance to reform" has a more precise explanation: the system's macro level (school institutions) has already formed a stable set of reward-and-punishment rules that continuously "reward" safe, standard, predictable practices while "punishing" risky, innovative, unpredictable ones. Any attempt to change practices, without simultaneously changing these reward-and-punishment rules, will be pulled back to the original state by the system's own force.
Understanding this allows us to understand why "reformers burn with passion while the front line remains unmoved." Not because the front line "doesn't want to change" — but because the system level has a solid set of reward-and-punishment rules saying "not changing is safer." Unless these rules change, any new practice will be eliminated by the system.
Understanding the System to Find the Leverage Point
At this point, the three-level analysis of CLE theory has revealed a disquieting picture:
The positive cycle requires a series of preconditions (safety, institutional pathways, response speed) to initiate and sustain. Meanwhile, the negative cycle — getting stuck — is almost the default path: it requires no conditions; it only requires the system to "keep operating the way it always has."
This is not a pessimistic conclusion. It is a clear-eyed recognition of systemic inertia. The greatest mistake reformers make is not aiming at the wrong target — it is underestimating the system's own capacity for self-stabilization.
So, where is the most effective point of intervention?
Donella Meadows (1999), in her classic paper "Leverage Points: Places to Intervene in a System," ranked twelve types of intervention points from least to most effective. The most superficial interventions (e.g., "extend exam time by 30 minutes") have almost no effect; the deepest interventions (changing the system's core goal — e.g., "shifting from controlling learning to inspiring learning") are most powerful but also hardest.
Drawing on Meadows' framework, we can classify interventions into three tiers of power. But note: these three tiers do not simply correspond to micro/meso/macro. As we said earlier, the three levels form a mutually nested cycle — change at any level, through the cycle, affects the other levels. So to judge whether an intervention has power, do not look at which level it falls on; look at whether it can produce a chain reaction within the cycle.
This is the most natural intervention — and the most common. "Train teachers in wait time," "encourage students to ask questions," "give children an extra sticker as a reward" — none of these are difficult to do, but their common problem is isolation: when a teacher returns from training to find the school's evaluation system unchanged, their new skills have no room to survive in the original reward-and-punishment environment. "The child learned to ask questions in class, but the exam still tests memorization" — the new practice is "not worth it" within the system, and is quickly eliminated.
This is not to say micro-level interventions have no value — rather, whether a single-point change can take effect depends on whether it can cause adjacent links to change in response. Doing just one is like asking a transplanted sapling to solve drought, pests, and poor soil on its own.
More effective than moving a single point is changing the flow of information within the system — the speed and direction of information transmission.
Concretely, interventions at this level include:
The beauty of this level of intervention: it does not require changing the institution itself — only changing how information flows within the institution. When teachers and administrators can see the effects of their own behavior more quickly and more comprehensively, the system's self-regulating capacity naturally improves.
The most powerful, but also hardest to execute, intervention is changing the system's reward-and-punishment rules — what Meadows called "the rules of the system" and "the structure of the system." In the CLE framework, this corresponds to the three institutional design principles proposed in Chapter 11:
Principle One: See clearly what is truly being rewarded. What is the school's evaluation system actually "saying"? "Innovation brings benefits" or "stability is safer"? If a school says with its mouth "encourage innovation," but promotions depend solely on years of teaching and exam scores — then the real reward signal is the latter. The first thing to do is identify the gap between "what is said to be rewarded" and "what is actually rewarded," and narrow it.
Principle Two: Reward exactly what you want to see. If a school's educational goal is "cultivating critical thinking," but the sole measurement standard is end-of-term standardized test ranking — then without any analysis whatsoever, one can predict the goal will not be achieved. Either adjust the goal, or adjust the evaluation. When neither can be moved, at minimum let teachers and students be aware of this contradiction — awareness itself is an important intervention.
Principle Three: Give new reward-and-punishment rules enough time. Changing reward rules is not a one-time event. A new evaluation system needs at least 2–3 years of continuous operation to truly begin producing effects. Rewarding group discussion today, rewarding individual performance tomorrow — will only leave everyone "not knowing which direction to head," and they will all just stop in place. Institutional instability is itself a negative signal.
For interventions at this level, there is one core strategy in execution: changing the system's core goal is the most powerful intervention. When the school's core goal shifts from "raising exam rankings" to "cultivating students' capacity to construct" — the entire reward-and-punishment map will adjust accordingly. But this is also the hardest, because it involves the reconfiguration of values and educational beliefs.
Drawing on Meadows' (1999) leverage points framework, CLE proposes three criteria for judging whether an intervention in a school system has power:
The optimal strategy is not to fixate on the single most powerful lever — but to act simultaneously at multiple levels: start prying from feasible points, and let the change transmit outward, link by link, between adjacent components.
This leads to CLE's layered advancement strategy for three-tier intervention:
Step One: Pilot at small scale. Within the existing institutional framework, find a "safe space" where teachers and students can experiment with new interaction patterns. Not full-scale reform — just try, in one class, one subject, one semester, waiting a few extra seconds, accepting more non-standard answers, letting students lead more discussions. The purpose at this stage is to produce a "successful" case as evidence, not to change the entire system.
Step Two: Make the success visible. Disseminate and discuss the positive cases accumulated in Step One among the teacher body — through teaching case sharing, peer observations, teaching-research group reflections, allowing more teachers to "see" the feasibility of an alternative teaching model. The purpose at this stage is not to persuade everyone, but to let the system "see" that another possibility exists.
Step Three: Institutionalize the new practice. When enough positive evidence has accumulated — when enough teachers and classes have demonstrated positive behavior — institutional-level adjustment becomes "a natural next step" rather than "a radical gamble." At this point, changing the evaluation system, adjusting the focus of teacher training — is no longer a lonely gamble, but confirmation and consolidation of good changes already happening.
Step One: Pilot — Produce a good case within a safe space; accumulate evidence that "another way can also work."
Step Two: Amplify — Spread the good case through sharing, exchange, etc.; let more people "see" different possibilities.
Step Three: Institutionalize — When enough positive evidence has accumulated, turn the new practice into an institution; transform "one teacher's success" into "the school's standard practice."
Not: change the institution first, then change the classroom — but: change the classroom first, then let the classroom's success drive the institution.
From Analysis to Action
From micro to meso to macro, CLE theory uses a unified framework to explain how learning happens at the individual level, the classroom level, and the school level. But these three levels are not the entirety of education — they are situated within an even deeper background: culture and history.
School institutions did not emerge from thin air. They grew in specific cultural soil, shaped by the social expectations, economic structures, and political environments of particular historical periods. Why does a Chinese school place such weight on standardized testing? Why is "good classroom discipline" considered the primary criterion for a "good classroom"? The deep roots of these institutions lie not inside the school — but in the larger cultural system.
CLE theory provides an analytical framework, but analysis is not the same as action. When we apply this framework to real educational systems, we encounter a thorny question: since we know what should be done, why, after all these years of reform, has the core pattern of the classroom barely changed?
In the next chapter, we confront this question head-on: The Progressive Education movement rose and fell; Finland succeeded while others could not replicate; China's educational reforms come one wave after another — why are good ideas so difficult to implement?
Next Chapter
Chapter 13: Why Reforms Keep Going in Circles
In 1919, Dewey and his colleagues launched an educational revolution in the United States. They believed that learning should be child-centered, experience-based, and inquiry-driven.
Over the next thirty-eight years, the Progressive Education Movement established thousands of experimental schools, produced a large body of research, and generated genuine success stories. In 1942, the Eight-Year Study, which tracked thousands of students over eight years, demonstrated that students who had received a progressive education significantly outperformed their traditionally educated peers in critical thinking, social skills, and college-level academic ability.
Then, on October 4, 1957, the Soviet Union launched a satellite.
The entire movement effectively collapsed within two years. Not because the ideas were wrong. Because the environment gave it no room to survive.
This story is one of the most heartbreaking episodes in the history of educational reform. But it is not an isolated case.
Around the world, every decade or so a new wave of educational reform surges forth: constructivism, project-based learning, the flipped classroom, STEM education, personalized learning... Each wave is backed by theoretical grounding, supported by experimental data, and accompanied by inspiring case studies. Yet most reforms quietly fade within three to five years, the classroom returns to its previous state, and all that remains are a few academic papers and a handful of solitary pioneers still holding on.
It is not that the people are incapable. It is that the system entraps everyone within it.
Let us examine the history of the Progressive Education Movement in its entirety. Not to mourn, but to understand.
From a systemic perspective, the Progressive Education Movement made one fatal mistake: it changed only classroom instruction (a fast variable) without simultaneously transforming the evaluation system (a slow variable).
When the assessment standard of the entire society remained "standard answers," any pedagogical approach advocating "inquiry processes" was vulnerable in the face of a crisis. Because there was simply no way to prove on a standardized test that your students were better at "critical thinking."
The Sputnik Crisis was merely the straw that broke the camel's back. The true failure of the Progressive Education Movement was sealed the moment it refused to confront the evaluation regime head-on.
Fast variables are always dragged back to the origin by slow variables. Classroom instruction is a fast variable. The examination and assessment regime is a slow variable. If you change only the fast variable, the fate of your reform is already written.
In the 1970s, Finland's basic education ranked in the lower-middle tier among OECD nations. By the year 2000, when the first PISA results were released, Finland had become number one in the world.
Everyone wanted to know: what did Finland do?
Many people thought the answer was "fewer exams" or "more play." That answer is incomplete. What Finland did, in the language of CLE, is this: it reformed all four dimensions simultaneously, rather than changing only what happened in the classroom.
| Dimension | What Finland Did | CLE Correspondence |
|---|---|---|
| Teacher Selection | Teachers are selected from the top 10% of university graduates; teacher education is a five-year master's program requiring original educational research. | Macro-level selective pressure: ensuring that those entering the system possess the capacity to "construct cognitive space." |
| Assessment System | Only one national external examination exists across the entire basic education cycle (the matriculation exam at the end of upper secondary); formative assessment dominates, with teacher-led autonomous evaluation. | Reform of the slow variable: assessment standards are aligned with classroom philosophy, allowing "critical thinking" to germinate in the classroom. |
| Teacher Autonomy | The state provides a core curricular framework; specific content is determined by teachers themselves. There are no standardized textbooks. | Teacher as Selective Pressure Designer: the macro-level system grants teachers sufficient cognitive niche space. |
| Parental Perception | The culture deeply trusts teachers; parents do not use test-score rankings to evaluate school quality. | Soil conditions: the external ecology does not exert countervailing selective pressure. |
But here is an honest statement from Finland's own reform leader:
I am asked over and over: can we take the Finnish model with us? My answer is always: you can take all the ideas, but you cannot take the culture. Finland was able to do this not only because of what we did, but because of who we are—a society that deeply trusts teachers and is not afraid of educational diversity.
—Pasi Sahlberg, Finnish Lessons, 2011
This is not to say that the Finnish model cannot be replicated. It is to say: copying a reform blueprint is not enough. You must simultaneously change the cultural soil in which the reform takes root. This is precisely what Chapter 11 described as "the school as an ecosystem": you cannot transplant a single tree. You must bring some of the soil with it.
The Progressive Education Movement and the Finnish case reveal the same pattern: when selective pressure is misaligned across the three levels, reform cannot take hold. But "misalignment" is not precise enough. We need to specify exactly where things go wrong.
The core contradiction can be stated in one sentence: change at the classroom level is a fast variable. Change in the assessment system is a slow variable. When these two point in opposite directions, the classroom can never outrun assessment.
This is not a uniquely Chinese problem, nor is it a problem of any particular era. It is a structural feature of all educational systems. Why? Because in any system—whether a school, a corporation, or a society—change close to the implementation layer is always easier to initiate than change near the institutional layer. You want to change the way you teach (the micro level)? You can start next period. You want to change your school's exam system (the meso level)? You need faculty meetings, pilot programs, convincing parents. You want to change how an entire society defines talent (the macro level)? That is measured in decades.
Micro Level (Classroom): A teacher can change questioning style, wait time, or error-correction strategy within a single class period. The cycle from variation generation to selection outcome: one class period to one week.
Meso Level (School): Reform requires revising evaluation criteria, adjusting teaching-research systems, reallocating resources. From initiation to visible effect: one semester to one year.
Macro Level (Society): The gaokao or high school entrance exam system, society's definition of "good education," parental trust in test scores. From initiation to genuine change: ten years or more.
Here lies the problem: when the micro level has already shifted in one direction while macro-level evaluative standards remain frozen, micro-level variation is ineffective variation in the face of macro selective pressure—it will be eliminated.
This explains the fate of the Progressive Education Movement. Teachers made radically different attempts in their classrooms—rich micro-level variation. But when the Sputnik crisis hit, American society's evaluative standard snapped back decisively to "we need more scientists"—a sudden shift in macro selective pressure. Before classroom-level variation could be retained, the macro-level selective pressure sentenced it to death before it had a chance to prove itself.
Finland's success confirms the same principle from the opposite direction: Finland did not just change classrooms. It simultaneously reformed teacher selection (the macro entry point), the assessment system (the meso anchor), and social trust (the macro soil). The time lag across the three levels was compressed to the greatest extent possible.
Reform keeps spinning in circles not because people do not want to change. It is because the system contains a structural asymmetry: changes closest to the classroom are easiest to start, yet also easiest to be crushed by selective pressure from higher levels.
Based on the mechanistic analysis above, what follows is not a "how-to" operations manual—that is the work of Chapter 15. It is CLE theory's structural diagnosis of how to break the reform cycle.
For reform to stop spinning in circles, one condition must be met: change across all three levels must move in the same direction, and the speed of micro-level change must not exceed the tolerance threshold of the meso and macro levels by too great a margin.
From this, three concrete paths of action can be derived.
Micro-level changes (classroom experiments) must produce evidence visible to the meso level. Not "I feel like students enjoy inquiry more now"—that is not enough. It means collecting three months of data: student questioning frequency rose from 2 times per class to 8 times; student error types shifted from "blank, silent" to "partially correct, willing to try"; unit test averages did not drop, while the top-score band increased by 15%. This kind of data is what enables the meso level (principals, department heads) to willingly adjust evaluative criteria in parallel.
This mechanism corresponds to the micro-to-meso positive cycle discussed in Chapter 12: effective variation at the micro level must be able to "emerge upward" into institutional change; otherwise, it is simply absorbed by the system.
When you are a decision-maker at the meso or macro level, the most effective reform lever is not "require all teachers to change their teaching methods"—this is precisely the approach most likely to provoke backlash. A far better approach is: change the selective pressure that teachers face. For instance, revise the lesson-observation rubric from "good classroom discipline, neat blackboard writing" to "frequency and quality of student questions." Teachers' classroom behavior will automatically follow the observation rubric. You no longer need to train them one by one. This is the "selective pressure design" discussed in Chapter 5—you do not need to control the direction of variation; you only need to control what kind of variation gets selected.
This is the most easily overlooked path. Every reform, in its early implementation phase, goes through a stage where it "appears less effective than the old method." This is the plateau period discussed in Chapter 7, and the "one-semester law" discussed in Chapter 2. During this phase, what reform needs most is not more methodology; it is not to be prematurely eliminated. Specifically: while launching the reform, begin cultivating the "soil"—communicate with parents in advance, experiment on a small scale, give teachers room for trial and error, explicitly tell everyone that "the first three months are a trial-and-error period; we are not in a hurry to see score changes." This is, in effect, creating a protective selective environment: allowing new variation to survive until it is robust enough to withstand mainstream selective pressure.
Path One: let micro-level variation be seen from above.
Path Two: let meso-level institutions release the right selective pressure downward.
Path Three: give the first two paths enough time to complete a full cycle.
When all three paths operate simultaneously, this becomes the three-level cycle described in Chapter 12—micro to meso to macro. The three levels are no longer fighting separate battles; they form one positive cycle after another. Reform no longer spins in circles, because change at each level creates conditions for change at the other levels, instead of canceling each other out.
Chapter 15 will provide more concrete classroom operational tools. But what needs saying here is something more urgent: the root cause of failed reform is not that the methods are insufficiently good. It is that the structure does not support them. Structure can be changed—as long as you move all three levels simultaneously, and give the slow variables enough time.
The greatest enemy of reform is not a conservative system. It is the reformer's own impatience—declaring "this method doesn't work" before the slow variables have had time to change.
We understand why reform is difficult. But before we discuss what constructivism can do, there is something more honest that must be said:
Constructivism has boundaries. It is not a panacea. There are things it cannot do; situations where it is not suitable; criticisms that it needs to take seriously.
→ Chapter 14: What Constructivism Is Not
Some teachers are bound to ask: "What about multiplication tables? How would you constructivists teach those?"
That is a good question. A very good question.
If a theory cannot answer the question "What can I not do?", then it is not an honest theory.
In this chapter, I am going to do something uncommon in educational theory books: systematically and rigorously lay out what the constructivist view of education—specifically, Evolutionary Pedagogy (CLE)—cannot do, where it is not suitable, and how it should not be misapplied.
Negation is not the goal. Honesty is. A hammer cannot turn a screw—this is not a defect of the hammer; it is an honest description of it.
There is a recurring lesson in evolutionary biology: species that become hyper-specialized are often the first to vanish when the environment changes. Trilobites ruled the oceans for three hundred million years—and when the meteor struck, nothing remained. The panda's digestive tract is a perfect bamboo-processing machine, but eating only bamboo means that if a bamboo forest withers on a large scale, the panda has no backup plan. It is not that specialization is wrong—it is that refusing to acknowledge the boundaries of specialization is the problem. A theory is not mature until it has the courage to honestly say, "I am not suited for this." This chapter is CLE's statement of honesty.
This is the most common misunderstanding. And the one most in need of correction.
Every time I say "errors are raw material" or "don't give the answer right away," someone asks: then should our children still memorize poetry? Should they still memorize multiplication tables? Should they still practice writing Chinese characters?
My answer is: yes. Absolutely yes.
Let us return to multiplication tables. Memorizing multiplication tables is about making "7 x 8 = 56" an automatic response—occupying no cognitive resources in the brain. This way, when you are solving a complex application problem, your working memory does not need to allocate itself to basic computation; it can be fully invested in understanding the problem, building a model, and designing a strategy.
This is precisely what Chapter 6 (True Learning vs. Pseudo-Effort) described: automation is not the endpoint of learning. It is the infrastructure for higher-level construction.
Principle: first establish automation through high-frequency repetition, then deepen understanding through constructivist approaches. The two are sequential, not mutually exclusive.
More precisely, CLE's critique is not directed at "repetitive practice itself." It is directed at the practice of "using repetition as a substitute for understanding"—treating memorization as comprehension, treating fluency as depth. That is the problem. One clarification: doing 100 practice problems can also generate genuine construction. A student who, on the 70th problem, suddenly realizes that "completing the square is faster than the quadratic formula under certain conditions"—that is cognitive variation generated during practice. CLE does not oppose practice volume. What it opposes is practice with zero cognitive variation—mechanical repetition without ever pausing to ask, "why am I doing this?"
The same 100 practice problems: one person generates 3 moments of cognitive variation ("Oh, so that's how it works"), another generates 0. The difference is not in quantity—it is in whether, between problems, the brain is comparing, reflecting, and reconnecting. This frequency is what we call cognitive variation density—a new metric that measures the quality of practice, not the quantity.
A child who has not eaten breakfast can complete only limited cognitive variation in the classroom. Working memory has a resource ceiling. When basic safety, belonging, and esteem needs are unmet, the brain's cognitive resources are first consumed by anxiety and hypervigilance, not by the exploration of new knowledge.
CLE does not solve poverty. It does not solve family background. It does not solve the neurodevelopmental damage caused by community violence, household trauma, or malnutrition. When discussing constructivist education, we must remain soberly aware of this reality.
Telling a hungry child to "please actively construct the meaning of knowledge"—that is Marie Antoinette saying "let them eat cake." The prerequisite for constructivism is that the child's basic needs have already been met.
This was touched upon in the previous chapter, but it deserves to be stated clearly on its own.
The establishment of deep understanding requires multiple cycles of variation → selection → retention, and every cycle takes time. The punctuated equilibrium phenomenon discussed in Chapter 7 means that before a genuine breakthrough in understanding arrives, students may experience a pronounced "plateau period"—possibly even accompanied by a temporary decline in scores (because the "surface fluency" accumulated under traditional teaching is receding, while genuine understanding has not yet fully formed).
This is not evidence that constructivism has failed. It is a signal that constructivism is working. But it places enormous strain on teachers, students, and parents alike.
The following situations require caution in applying constructivist methods:
This is another common misunderstanding—and it often comes from constructivism's supporters, not its critics.
There is a "child-centered" narrative that holds: if the teaching method is good enough, and the learning environment ideal enough, learning should be effortless and natural—like a child playing a game.
That is a beautiful image. But it misunderstands the nature of constructivism.
Bjork (1994) introduced an important concept: "desirable difficulties"—those arrangements that make learning appear slower and harder (spaced practice, retrieval practice, interleaved practice) are precisely the ones that produce the most durable memory and the most transferable understanding.
The feeling of ease is often a signal that you think you understand something you do not truly understand. Truly effective constructivist learning frequently makes students feel the right amount of confusion, struggle, and effort—and then, insight. The struggle that precedes that insight is not a problem to be eliminated. It is a process to be protected.
Constructivism is learning with a "cost." Its cost is greater cognitive effort. Its return is deeper understanding, more durable memory, and stronger transfer ability. This trade-off is worthwhile—but the cost must be paid honestly.
Over twenty years of educational practice, I have heard a great deal of criticism directed at constructivism. In this section, I want to respond seriously—without evasion—to several of the most important objections.
This criticism is valid—if you understand constructivism as "pure discovery learning, where the teacher says nothing and students explore entirely on their own." That extreme form is indeed very difficult to implement in a real classroom. The critique by Kirschner et al. (2006) of "minimally guided instruction" has merit.
But CLE is not extreme discovery learning. Its core claim is: with effective scaffolding, allow students to experience a sufficient number of complete cognitive cycles. The teacher does not disappear. The teacher's role shifts from "information transmitter" to "environment designer." This is not more difficult to implement than traditional teaching—it is simply implemented differently.
Constructivism is not idealistic—traditional teaching is the one that is idealistic. It assumes that "explaining clearly means learning has occurred." That assumption cannot stand in the face of cognitive science evidence.
This is a question that cannot be dodged. And it requires a precise answer.
If "exam-oriented education" refers to an examination system that "only looks at scores and only tests memorized repetition," then there is indeed a fundamental contradiction between constructivism and that system—but it is not a contradiction of philosophy. It is a contradiction of time scales.
CLE theory can articulate this contradiction with clarity. Chapter 7 (Punctuated Equilibrium) tells us: deep understanding, during its formation, must pass through a "plateau period." During that time, student performance may not improve—it may even decline—because old "surface fluency" is being deconstructed while new structures of understanding have not yet fully formed.
The problem is: the Gaokao does not wait for your plateau period. The Gaokao only cares about outcomes, not processes.
This does not mean that "constructivism is ineffective in the face of the Gaokao." Quite the opposite: once construction is complete, deep understanding will absolutely outperform rote memorization on higher-order exam items (synthesis problems, application problems, novel problem types). The problem is the word "once"—this time lag is the most genuine contradiction between constructivism and exam-oriented education.
So the answer is not the mild compromise of "constructivism can also raise scores." It is: if you have only three months before a purely memory-based exam, constructivism is not the optimal strategy. If you have a year and a half to two years, constructivism is more efficient than any memory-training approach—because through multiple cycles of variation → selection → retention, it builds transferable structures of understanding.
This is not a problem unique to constructivism. It is a shared feature of all deep learning. When learning a language, you cannot say anything for the first three months. When learning a musical instrument, you cannot play a complete piece for the first six months. You cannot dismiss the method itself because of this "initial silence."
The real question is: is our assessment system willing to pay for this initial silence?
There is a misunderstanding here that needs to be clarified: construction is not a "higher-order ability" that only clever students can perform.
All students are constructing—the question is what they are constructing. A student who listens silently in class is also constructing: constructing the self-concept that "I don't understand math," or the attribution pattern that "I just have a bad memory." These constructions are equally deep and equally durable. They are simply destructive rather than constructive.
CLE's claim is: create conditions that allow all students to generate effective cognitive variation—not "more difficult tasks," but "appropriate challenges within their current zone of proximal development." This applies equally to students with weaker learning abilities. The scaffolding just needs to be designed with greater precision.
This objection is the most honest one. My answer is: you do not need to do all of it. You only need to start doing part of it.
Chapter 15 will provide ten specific, actionable principles that you can start using tomorrow. These principles do not require you to redesign your entire curriculum. They do not require you to overturn your existing teaching methods. They are small, cumulative adjustments—such as extending wait time from 1 second to 3 seconds, or asking students to articulate their thinking process before you correct them.
Constructivism is not an all-or-nothing choice. It is a direction. You go as far as you can, making the best design possible within the scope of your capacity.
Having written this chapter to this point, I want to say something about theories and honesty.
I believe that Evolutionary Pedagogy describes something real. I believe that the variation → selection → retention algorithm is indeed the deep mechanism of human learning. I believe that the role of the teacher as a Selective Pressure Designer comes closer to the essence of education than the role of information transmitter.
But I believe equally firmly: any theory that cannot honestly state its own boundaries is harming practitioners. When a teacher tries constructivist methods in an under-resourced classroom and encounters setbacks—if she has not been told that "these constraints are real"—she may attribute her failure to "not having done it well enough." That is unfair. And it is dishonest.
What CLE can do is this: under appropriate conditions, provide understanding that is deeper, more durable, and more transferable than traditional teaching. That is enough. We do not need it to be a panacea. We only need it to be real.
The honesty of constructivist education does not lie in claiming omnipotence. It lies in acknowledging: education is, first and foremost, something that happens inside the learner's own mind. No external force—teacher, textbook, exam—can substitute for the learner's active construction.
And precisely because of this, it returns the sovereignty of learning to the student. And it returns professional dignity to the teacher—you are a designer, an observer, a conversation partner, knowing when to reach out and when to hold back. This is difficult. Precisely because it is difficult, it is called professionalism.
We now know what constructivism can do, and what it cannot do.
So: what can a teacher standing at the lectern today do differently tomorrow?
Not three years from now. Not after reform arrives. Tomorrow. The first class. The first question. The first moment of waiting.
→ Chapter 15: The Teacher as Selective Pressure Designer
This book has reached its final chapter. We have discussed Darwin's theory of natural selection and Campbell's BVSR algorithm, Piaget's assimilation and accommodation, punctuated equilibrium, cognitive niches, the emergence of classroom culture, and the systemic dilemmas of school reform.
Now it is time to return to the simplest question:
In your very first class tomorrow, what can you do differently?
Not three years from now. Not after school reform arrives. Not after you have read more constructivist literature. Tomorrow. The first class. The first question. The first moment of waiting. The first time you choose not to immediately correct an error.
Trying to activate all ten principles at once is a recipe for chaos. What follows is a classroom-tested minimal startup sequence—add one thing each week, not ten things in one day.
Week One: Only "Wait 3 Seconds." After every question you ask, silently count to three before calling on anyone. Change nothing else. You are simply extending wait time from 0.7 seconds to 3 seconds. After one week, you will notice more hands going up—not because students became smarter, but because you gave their BVSR cycle time to start.
Week Two: Add "Ask Why." A student gets an answer right. Now ask: "Why?" Not because you suspect them—because you are forcing their cognitive system to descend from "got it right" to "why did I get it right?" This step pulls automation back into conscious awareness, transforming retention from surface fluency into transferable understanding structures.
Week Three: Introduce "Error Publicization." Set up a "Today's Best Error" corner on the blackboard. Each class, select the single most instructive error, write down the name and the reasoning. Not to criticize—to let the entire class see that an erroneous line of reasoning is still a line of reasoning, and that exposing it is more useful than hiding it. This step expands the selective pressure from "teacher as sole judge" to "the whole class as collective observers."
After three weeks, your classroom will simultaneously run three CLE mechanisms: waiting creates space for variation, "why" activates deep selection, and error publicization extends the feedback loop from one person to the entire class. The ten principles that follow are refined elaborations of these three things.
These ten principles—every one of them—can be put into practice starting tomorrow. They are not theoretical propositions. They are classroom-tested operational tools. After each one, I note its corresponding mechanism within the CLE framework and the empirical research that supports it.
You do not need to do all ten at once. Choose one or two. Start next Monday. Do them consistently for six full weeks. Observe what changes in your classroom. Accumulate your own data.
Theories and principles truly exist only when they happen in a real classroom. The story that follows comes from a transformation trajectory I have witnessed repeatedly over twenty years as a school principal—it is not a specific individual, but every detail is drawn from genuine classroom observation.
The transformation did not happen overnight. For the first four weeks, every time she resisted the urge to correct an error, her hand gripping the chalk turned white. In the fifth week, a student who had never spoken up in class wrote a completely wrong approach on the board. She was about to speak—when another student raised a hand first. Not to tattle. To propose an improvement.
In that moment, she said, she felt "loss of control" in the classroom for the first time—but not in the fearful sense. It was the realization that she was no longer the only "knower" in the room.
In the eighth week, she began proactively categorizing and displaying student errors. "Look—today our class generated three different pathways. Let's analyze them one by one. Which one comes closest to the answer?" The students did not feel humiliated—because they saw that the teacher's gaze had changed: not a "let me correct you" gaze, but a "let's look at this together" gaze.
At the end of the semester, this class's scores were not significantly higher than those of the parallel classes—but the number of questions they asked was three times that of the parallel classes. One semester later, their higher-order problem-solving accuracy began to pull ahead.
I asked her: what was the hardest thing this semester?
She said: "The hardest thing was not designing questions. It was holding back—waiting 3 extra seconds when every nerve in my body was screaming 'correct him now.'"
Those 3 seconds, she said, were the most important professional decision of her twenty-year teaching career.
Evolutionary Pedagogy (CLE) is a theory under construction. Its core logic already has substantial empirical support, but it still has large gaps that need filling. Here are what I consider the most important directions for future research:
This book began with a single class taught by Harvard physics professor Eric Mazur. In that class, he discovered that after twenty years of teaching, his students' physical intuition had not genuinely changed—they had simply learned to solve problems using formulas, without truly understanding the concept of force.
That discovery changed his entire approach to teaching.
This book ends not with an answer, but with an invitation:
You can also do what Mazur did—not at Harvard, but in your own classroom, using the subject you know best. Conduct a serious diagnosis: do my students truly understand, or have they simply learned to write down the right answer?
This diagnosis may make you uncomfortable. Mazur was very uncomfortable that year. But it was precisely that discomfort that triggered his variation, his change.
You are the reader of this book. You are also a learner. If this book has any meaning—that meaning is something you constructed yourself, not something I placed inside it. Someone else's chewed-up bread has no flavor.
Now, put the book down. Go teach that class.
This is not a rhetorical game of "East meets West." It is not a cultural performance of "dialogue between Eastern and Western thought."
These thinkers, in different languages, different cultures, and different eras, independently touched the same thing. That fact itself is evidence that this "thing" is real.
When an observer deep in the mountains of Asia, a scientist on the Galápagos Islands, and a psychologist on the shores of Lake Geneva—each independently—describe the same structure, that structure is very likely real.
And then AI arrived.
It did not overturn any of the above. It precisely proved all of it. GPT-4 proved that a closed space can be exhausted by computational power. It proved that "memorization" and "understanding" do indeed collapse into the same thing within a low-dimensional space. It told everyone, with a Gaokao math score of 130: those things you spent twelve years training yourself to do—a few weeks of training data can cover them.
So the goal of education in the post-AI era is not to cultivate people who can outrun algorithms within a closed space—it is to cultivate people whom algorithms cannot outrun. These are people who do not turn their brains into potato fields. They can reorient themselves when the scene changes. They can sense the existence of a problem before it has been properly defined. They can ask the next question when AI has already provided every answer.
This glossary collects the core concepts of Evolutionary Pedagogy, arranged in the logical order of their first appearance. Each entry includes the Chinese and English names, a definition, and the chapter index where the term receives its first detailed treatment.
The design principle of this observation form: it is not for evaluating teachers—it is for recording the cognitive processes that actually occur in the classroom. Based on CLE's three core levels (micro / meso / macro) and three core verbs (variation / selection / retention), it provides a structured lens for teaching-research activities, peer observation, and self-reflection.
Observation focus: Are individual students experiencing a complete "variation → selection → retention" cycle?
| # | Observation Indicator | Evidence Collection | Rating (1-5) |
|---|---|---|---|
| 1 | Quantity of Variation: Did students generate enough different lines of reasoning / attempts? | Count the number of distinct solution methods or response types that appeared in class | |
| 2 | Quality of Variation: Did these attempts go beyond repeating the standard answer? | Record whether student responses included "their own words," "analogies," "questioning," "connecting to prior knowledge" | |
| 3 | Wait Time: Did the teacher provide ≥3 seconds of thinking time after posing a question? | Use a stopwatch to record wait time for 3 randomly selected questions; take the median | |
| 4 | Error Exposure: Were students' erroneous lines of reasoning allowed to be fully expressed? | Record the number of erroneous answers heard in full vs. immediately interrupted | |
| 5 | Self-Explanation: Were students asked to explain "what were you thinking?" | Count the frequency with which the teacher asked "why" or "what were you thinking?" | |
| 6 | Retrieval Practice: Was there a design element requiring students to recall before opening the book? | Record whether a closed-book recall segment occurred and its duration |
Observation focus: What kind of classroom culture is emerging from teacher-student interaction patterns?
| # | Observation Indicator | Evidence Collection | Rating (1-5) |
|---|---|---|---|
| 7 | Error Handling: How did the teacher respond to the first incorrect answer? | Record the teacher's exact words and facial expression when the first error appeared | |
| 8 | Error-Tolerant Atmosphere: Did students dare to express views differing from the teacher or textbook? | Count the number and tone of "differing-opinion" contributions in class | |
| 9 | Peer Selective Pressure: Was there peer evaluation, discussion, or debate among students? | Record the proportion of time spent on peer interaction and the quality of interaction (explanatory vs. telling) | |
| 10 | Type of Teacher Feedback: Was the feedback evaluative (right/wrong) or explanatory (why)? | Randomly sample 10 instances of teacher feedback; calculate the ratio of evaluative to explanatory | |
| 11 | Distribution of Speaking Rights: Which students were speaking? Was there a silent majority? | Draw a seating-chart frequency distribution of student contributions | |
| 12 | Improvisational Responsiveness: Did the lesson flow adjust based on students' genuine responses? | Compare the degree of deviation between the lesson plan's preset flow and the actual flow |
Observation focus: Is the selective pressure at the institutional level aligned with classroom goals?
| # | Observation Indicator | Evidence Collection | Rating (1-5) |
|---|---|---|---|
| 13 | Assessment Alignment: Are the goals of this lesson consistent with the school / grade-level assessment standards? | List the lesson's instructional objectives; cross-reference corresponding entries in the end-of-term / standardized exam syllabus | |
| 14 | Time Structure: Did the time allocation for this lesson leave space for student construction? | Record the proportion of time spent on: teacher lecture / independent student inquiry / peer discussion / practice | |
| 15 | Process Visibility: Were students' construction processes recorded and displayed? | Check for process-based records such as learning journals / error notebooks / knowledge graphs | |
| 16 | Parental Cognitive Alignment: Do parents understand the instructional philosophy of this lesson? | Spot-check 2-3 parent interviews or review home-school communication records for articulation of instructional philosophy |
After completing the observation, use the following framework for structured analysis:
| Diagnostic Dimension | Core Question |
|---|---|
| Variation Scarcity (Micro 1-2 low scores) | Does the classroom overemphasize the "standard answer"? Is the openness of questions insufficient? Is wait time too short? |
| Selection Misalignment (Micro 3-6 + Meso 7-10 low scores) | Are the feedback loops effective? Were student errors "seen" and "used"? |
| Retention Deficit (Micro 6 low score) | Is there a design for spaced repetition and retrieval practice? Or only one-shot explanation? |
| Negative Emergence (Meso 7-12 low scores) | Did the response to the first error send a "don't make mistakes" signal? Is there an absence of error-tolerant culture? |
| Institutional Mismatch (Macro 13-16 low scores) | Is the assessment system in conflict with classroom goals? Is the time structure squeezing out space for construction? |
This comprehensive bibliography is organized alphabetically by author surname and collects the major works cited across all fifteen chapters. Each entry is followed by the chapter(s) of first or primary citation.
Why "grinding practice problems" is the optimal strategy within the Gaokao framework—and why this optimal strategy is precisely the worst strategy for university adaptation
Let P be the problem space of the Gaokao mathematics examination. Based on an analysis of problem types across the past decade of national examinations:
Conclusion: dim(P) ≤ 6. Eighty percent of the available points are concentrated within 3 dimensions. This is not speculation—it is a fact that any student who has completed a full year of senior-three training can intuitively perceive.
For a finite discrete space with dim(P) ≤ 6, a strategy exists—"practice every problem type exhaustively"—such that as the number of training samples N → ∞, the probability of correct solution approaches 1. This is not "learning" mathematics. This is achieving exhaustive coverage within a 6-dimensional discrete space.
GPT-4's achievement of a 130+ score on the Gaokao mathematics examination is empirical confirmation of this theorem. A hundred-billion-parameter model has indirectly covered this low-dimensional space within its training data. When a language model can cover it, the significance of "grinding practice problems" as a human learning strategy has been rendered moot.
Within the closed space of dim(P) ≤ 6, the following two strategies are nearly indistinguishable in terms of scores:
Four points. The resolution of the Gaokao cannot distinguish between these two strategies. But Strategy A produces Buchnera—perfectly adapted to an ecological niche that is about to vanish. Strategy B preserves cognitive diversity—and in the infinite-dimensional space of the university, the difference in adaptation speed is night and day.
The more optimized Strategy A becomes within the space of dim(P) ≤ 6, the more fragile it becomes in the real world where dim(P) = ∞. This is not a matter of learning ability—it is the negative correlation between adaptive value and adaptive range. In the language of evolutionary biology: the Gaokao is a bottleneck of strong directional selection—at the very moment you pass through it, it shapes you into the form it requires. But the environment you need to enter next requires precisely the opposite form.
This paradox carries no moral implication. It is not to say that students who grind practice problems are "wrong"—they made a perfectly rational choice within the given rules of the game. The problem lies not with the choosers. It lies with the designers of the selective pressure.
Automation is not irreversible—but it comes at a cost. Myelinated neural pathways do not spontaneously disappear—new BVSR cycles can only operate within the constraints of existing pathways. Muller's Ratchet applies here as well: reversing highly automated cognitive structures is more effortful than constructing them from scratch.
Under what conditions, then, is automation reversible?
The task of education is not to prevent automation—it is to design reversible automation.
▾ Next: Bringing These Principles into Real Classrooms
Appendix D has provided a mathematical proof—demonstrating the structural relationships among the Gaokao system, grinding strategies, and cognitive fragility. Returning to the main text, we will discuss how these patterns can be translated into actionable principles in everyday teaching.
Wang Sai is the founder of Xuzhou Yingjian Education Group and the principal of an international school. He has spent twenty years in school leadership, working deeply at the front lines of basic education and international education.
He is also the founder of TeachClaw, an AI-powered teaching assistant platform, dedicated to bringing artificial intelligence into authentic instructional contexts. Many of the classroom cases in this book are drawn from real teaching practice at his school.
Over twenty years, he has lived through round after round of teaching reforms that "began with passion and ended with disappointment." It was precisely this refusal to accept defeat that drove him to find, in Darwin's theory of evolution, a new answer to "how learning actually happens"—and that is how this book came to be.