Education in the AI Age: a Teacher Rethinks Learning & Purpose, w/ Johan Falk of Graspable AI
Summary
- AI makes education’s labor-market mandate newly contingent: if society moves toward UBI, Johan Falk says supplying employers with competence could become “basically irrelevant.” Schools would then need to emphasize citizenship, cooperation, well-being, arts, and personal growth—yet Falk’s children are seven and 10, with roughly another decade in a system that takes about five years merely to change a curriculum.
- The binding constraint on education technology is institutional speed, not model capability. Falk points to South Korea, Singapore, China, and especially agile, data-ready Estonia as early movers, while warning that research arrives roughly two years late and can rely on tiny, poorly generalizable samples. His policy call is deliberately asymmetric: establish national direction, fund teacher competence, define data boundaries, and act because “non-action is a great risk.”
- AI tutoring could democratize what was once available only to royalty, but motivation remains the scarce input. A conversational expert costing perhaps $20—or even $2—a month may accelerate learning dramatically, yet Khanmigo’s early experience showed self-directed students racing ahead while already-confused students remained stuck. Nathan Labenz’s formulation survives the caveat: “There’s never been a better time to be a motivated learner,” but never an easier time to fool yourself about learning.
- The near-term opportunity is teacher augmentation, while mandatory classroom deployment is premature. AI can turn four hours of research into 20 minutes, produce differentiated exercises, organize notes, and help teachers reclaim time; Falk’s practical target is to “save five minutes every day.” The investable category is workflow leverage with human review—not an AI silver bullet imposed uniformly across classrooms.
- Assessment is both a major automation market and a governance trap. Falk rejects AI detectors outright and opposes unstructured AI grading because teachers may “fall asleep at the wheel,” but he accepts tightly governed second opinions—and, under pressure, a system of explicit rubrics, subgroup bias tests, transparency, and appeals. Nathan adds inference-time redundancy: multiple models or prompts, consensus thresholds, and mandatory human review when grades diverge.
- AI literacy is more urgent than AI-mediated instruction, especially around learning self-deception and synthetic relationships. Falk would prohibit AI companions below age 18, while hedging that “90%, 95%” of interactions may cause no concern; the danger is concentrated in dependency, emotional harm, and products optimized for retention, ads, or commerce. Tutor and companion will blur because relationships drive learning, making incentives and institutional ownership as consequential as model quality.
- Falk thinks “the age of grades is coming to an end,” because continuous AI observation makes a single letter or standardized test informationally crude. Grades also create the incentive to cheat, whereas a student motivated by learning receives mostly upside from AI. The deeper transition is from schools designed for the first industrial revolution toward systems that cultivate agency, authentic expression, relationships, and worthwhile lives—even if AI eventually outperforms most people economically.
Deep dive
1. AI has reopened the question of what education is for
Falk divides education into three purposes: supplying competence to the labor market, forming functional and responsible citizens, and helping people grow through arts, knowledge, and experiences that are “good for you” without being economically necessary. AI destabilizes the first purpose while making the contested values inside the other two impossible to avoid.
Sweden weights those purposes differently from the United States. Falk describes a social-democratic culture with high taxes, broadly available healthcare, and free higher education, producing less educational competition and a stronger conception of the common good; Labenz notes that employability has clearer external validation than either “good citizen” or “good person.”
Falk’s uncertainty is personal: his children are seven and 10 and will spend roughly 10 more years in school, yet he has “no idea” what society will look like when they finish. Reading, writing, self-understanding, and cooperation remain durable; quadratic equations and additional languages might be useful, but he no longer assumes they are necessary.
2. A five-year curriculum cycle cannot track fast-moving AI
Falk identifies agility as the system’s most important missing capability. Changing a curriculum in five years counts as relatively fast, but “five years in the AI world” contains massive change; educational policy can therefore become obsolete before implementation, just as Labenz has watched startup features become outdated during a development cycle of only months.
The institutional drag is both emotional and structural. Organizations cling to sunk costs and familiar paradigms, while genuinely new approaches have not been stress-tested; large public systems add procurement and organizational inertia that can make the mismatch “an order of magnitude harder, maybe two.”
Falk finds no jurisdiction that has solved the problem, but sees useful pieces: Singapore combines top-down adoption with education-data standards; South Korea couples deployment with limits on using learning data for unrelated purposes; Estonia pairs national agility with strong digital infrastructure. China has long developed AI curriculum and offers Squirrel AI, while U.S. deployment includes Khanmigo’s planned expansion toward one million students and teachers, reportedly with Microsoft support.
His national playbook is concise: declare AI strategically important, resource teacher training, state what is permitted, and draw firm data boundaries. Countries hesitate because the right move is unknowable, but Falk’s judgment is that “doing something and improving along the way is better than doing nothing”; otherwise, they risk widening the digital divide.
3. Falk’s incomplete cube separates four different education markets
What initially looked like an undifferentiated mass eventually became “four different sides of a cube” for Falk: using AI for learning, using it to reduce teachers’ non-classroom workload, teaching students about AI, and preparing for system-level changes to schools, curricula, and teachers’ roles.
The missing two faces are deliberate: “we don’t have the full picture yet.” The framework prevents leaders from treating every AI-and-education question as tutoring or cheating while acknowledging that new categories may emerge as products and institutions evolve.
Counterintuitively, Falk calls the two most visible categories—student learning tools and teacher productivity—the least important at the moment. Teaching AI competence is more urgent because children already use these systems independently, while systemic consequences may ultimately reshape the entire purpose and structure of schooling.
4. Thin evidence argues for experimentation, not universal mandates
The research points in opposing directions: some studies suggest students can learn twice as much or halve the required time, while others report learning losses. Falk’s objection is not that the findings are useless, but that one might rest on 25 adults in Nigeria and another on 18 people instructed to use a chatbot for an essay—weak foundations for a Swedish middle-school mandate.
Research also reaches decision-makers roughly two years after the intervention, by which time language-based AI may have changed substantially. Falk distinguishes these fast-moving “thinking machines” from slower, traditional applications of neural networks to educational data; pooling both under “AI” can make ostensibly evidence-based policy misleading.
His recommendation is classroom-level discretion: interested, competent teachers should try AI where it fits their students and goals, but systems should not yet require everyone to use it as a learning tool. Labenz accepts that hedge while warning that “not proven for everyone” can become an excuse for most educators to avoid learning anything at all.
Falk’s answer is to build teacher competence without declaring a silver bullet. Teaching a class is “bizarrely difficult,” involving perhaps 2,000 decisions daily; AI may help one child practice German, another confront mathematics, and another become curious, while the correct intervention for a restless student may simply be running outside for 10 minutes.
5. AI tutoring amplifies agency more reliably than it creates it
Falk sees enormous potential in conversing with an expert on almost anything—an experience previously available mainly to “some kind of royalty” with personal tutors. At perhaps $20 or even $2 per month, such access could matter even more in poorer countries; Labenz notes Khan Academy’s U.S. retail price around $4 and the generosity of ChatGPT’s free tier.
The constraint is whether the learner can articulate a goal and keep moving. Salman Khan reportedly observed some children immediately “run” with Khanmigo while others remained confused and stuck; teachers recognized the same students who, without AI, could not explain what they were doing or what help they needed.
Falk’s former math classes contained advanced students who were understimulated and 16-year-olds still struggling with negative numbers. Adapting work to actual knowledge could restore interest, but the broader teacher role becomes motivation, initiation, and persistence—the human work Alpha School assigns to mentors, guides, and coaches rather than content presenters.
Labenz’s own AI-assisted reading of technical and biology papers supports the booster case, yet Alpha School’s two-hour learning model reportedly uses no chatbots at all, instead combining internally built and licensed applications. The contrast matters: strong results attributed broadly to “AI education” do not establish conversational tutors as the operative mechanism.
6. The best classroom uses are still small, situated experiments
Falk imagines sending a student fascinated by black holes to converse with Gemini, then asking for a Thursday presentation plus three difficult questions classmates would genuinely want answered. The teacher chooses the tool, subject, deadline, and social output; the chatbot expands available depth without replacing educational judgment.
His strongest anecdote began when a visiting child grew bored and reached for a phone. Falk generated an age-adapted interactive story with choices; half an hour later, the child’s mother said she had rarely seen him read so closely or intently. The point is exploratory, not causal proof: experimentation reveals affordances that standardized policy cannot anticipate.
Falk also describes an “AI wizard” acquaintance rapidly building tools for dyslexia, adapting them for non-native Swedish speakers, and possibly creating a tool to interpret shaky handwriting associated with Parkinson’s disease. AI-written code lowered the cost of moving between adjacent needs, illustrating how individual teachers and parents may discover niches before formal procurement processes can name them.
Yet children are already conducting uncontrolled experiments at home. Some use AI to deepen learning; others “fool themselves into believing that they’re learning stuff while they’re actually not,” fall behind, and accumulate academic and social consequences. That is why instruction about using—and deliberately not using—AI cannot wait for consensus about classroom tutoring.
7. Teacher workflow gains are immediate but require tolerance for imperfection
Outside the classroom, Falk expects AI to help teachers convert rough notes into parent communications, digest large information sets, plan lessons, and investigate unfamiliar student needs. Research and brainstorming that once took four hours might take 20 minutes, freeing attention for students rather than merely increasing administrative output.
A teacher could feed in an old mathematics test and request three variants, including one themed around soccer, then edit the results. Falk treats generated material as inspiration rather than automatically classroom-ready content; the more important personalization is usually matching difficulty to knowledge, not wrapping every exercise in basketball references.
His operating threshold is pragmatic: material that is 90% or 95% good may contain exercises that fail, but the net outcome can still improve if the saved time enables conversation and follow-up. Sometimes an experiment will even be net negative while yielding learning—a tolerance schools need if they expect teachers to discover responsible practices.
Learning-style doctrine is the wrong justification for multimodality: Falk says the science is clear that fixed visual or auditory learning styles “don’t exist in that way,” though the subjective feeling does. Audio on a bus, images or text where useful, and movement between media can still help; he cautiously recalls evidence favoring blended modalities, adding, “don’t quote me on that.”
8. Interactivity is a new primitive, not proof of better learning
Falk expects interactivity to support engagement through anticipation and physiological mechanisms involving dopamine, and he recalls language-learning research attributing gains to that feature. He remains appropriately uncertain about a general causal claim, but AI dramatically expands the number of contexts in which students can ask, respond, branch, and revise.
Two voices are often more engaging than a monologue, which helps explain NotebookLM-style conversational audio. Labenz already gives such synthetic discussions a minority—but growing—share of his listening because they can cover a paper or small collection for which no human podcast exists; the relevant comparison is not perfection but reading and falling asleep.
NotebookLM’s narrated presentations suggest another convergence of text, slides, speech, and interaction. Erik says that, in AI-and-education topics he knows deeply, he can produce better content, but has seen teachers and other presenters who do not do a better job; Labenz jokes to wait another six months.
9. AI assessment needs governance before it needs scale
Falk’s verdict on AI detectors is categorical: “They don’t work. Dead end.” On grading, his initial answer is also no, because a teacher who accepts the model’s first assessment may “fall asleep at the wheel,” reproducing hidden disadvantages for non-native speakers or other atypical groups.
A second opinion is more defensible: grade the work first, ask AI independently, and revisit large discrepancies. Falk’s experience overseeing Swedish national mathematics tests complicates human exceptionalism—evaluations of oral-test scorers found reliability comparable to “tossing dice and sometimes worse.”
AI can be more consistent, but one LLM applying one bias to everyone may be worse than a thousand teachers whose biases partly dissolve into noise. Detected model bias can theoretically be corrected more easily than human bias; Falk nevertheless warns that individual teachers should not casually deploy generic chatbots for consequential grades.
Pressed by Labenz’s hypothetical teacher who will automate regardless, Falk proposes decomposed rubrics with numeric subscales, documented aggregation, subgroup comparisons, transparency, and a route for students or parents to appeal. Labenz adds multiple models or prompts: accept strong consensus, but require human review for a 3–2 split or evaluations separated by two points on a seven-point scale.
10. AI literacy now includes relationships, incentives, and power
The first urgent competency is recognizing when AI supports learning and when it substitutes for it. Children have inhabited a world of chatbots for almost three years, Falk says, yet most schools still do little to explain the technology, the learner’s responsibilities, or the distinction between productive assistance and self-deception.
The second is AI companionship. Falk would prohibit companions for anyone under 18; he has heard of cases involving severe harm and warns of risks including suicide as well as emotional and social harms. He simultaneously preserves the denominator, estimating that “90%, 95%” of interactions may create no concern. Education should teach warning signs rather than imply that every attachment is catastrophic.
Labenz’s pushback is structural: relationships are among the strongest contributors to learning, so the best tutor may necessarily resemble a friend. A long-lived relationship may also become an application moat because users do not abandon a friend merely upon meeting someone smarter; educational rapport, emotional dependency, and commercial lock-in can therefore emerge from the same product feature.
Falk’s partial safeguard is incentive design. A nonprofit optimizing education has better chances of avoiding dark patterns than a subscription business maximizing retention, while Labenz worries even more about free, ad-funded companions steering children toward engagement or purchases. “We still lack the map for navigating this terrain,” Falk admits.
11. Outdated mental models cause institutions to underestimate AI
Falk sees many users treating ChatGPT as “Google but from OpenAI,” entering “vegan pancake recipe” and missing its roles as collaborator, critic, assistant, and problem solver. Six- or 12-month-old experience is also dangerous evidence in a field whose capabilities may change before a school completes a policy review.
Labenz argues that the hallucination narrative was formed around weaker systems: GPT-5 or Claude 4 may not be less accurate than Wikipedia for elementary-school material, although neither speaker claims infallibility at the fringes. The corrective has swung too far when occasional error becomes reassurance that educators can ignore the technology.
“It just predicts the next token” is similarly incomplete. Labenz points to language-independent internal representations of shared concepts and to reinforcement learning that rewards correct final outcomes rather than any particular token sequence; research into hidden reasoning traces has even found compressed, nonstandard shorthand unlike ordinary training text.
Falk adds the institutional implication: AI is not traditional software whose decisions can be recovered by reading source code. Researchers can illuminate fragments only with major effort; the systems are “essentially black boxes grown rather than built,” demanding empirical oversight even when outputs look reliable.
12. Continuous evidence makes grades and standardized tests look primitive
If an AI continuously observes what a student attempts, learns, struggles with, and excels at, Falk says reducing that record to A–F “makes no sense.” Standardized tests remain legible to incumbent institutions—one reason Alpha School still uses them—but they are a thin snapshot beside longitudinal, task-level evidence.
Falk’s most radical claim is that “the age of grades is coming to an end.” AI-assisted cheating is chiefly a problem when the grade or exam supplies the motivation; for someone driven by learning itself, AI is overwhelmingly a tool. He wants education to move from extrinsic ranking toward the desire to learn, while admitting, “I don’t know how to make that transition.”
Abolishing grades would hurt and create new allocation problems: universities would need admission tests, lotteries, or another mechanism. Falk does not pretend every consequence improves, only that grade-centered motivation may already do more harm than good and will deteriorate “for every six months that passes.”
The historical framing is intentionally large: today’s educational system was built for the first industrial revolution nearly 200 years ago. AI tutors, persistent assessment, disappearing entry-level work, and changing concepts of knowledge make it doubtful that the same standardized architecture fits the present revolution.
13. Authentic expression may survive after economic superiority does not
Labenz proposes replacing minimum-length essays with stringent caps—even tweet-length assignments asking students to state something they fully endorse and could read aloud in 15 seconds. Falk offers a companion exercise: write about a beloved hockey team or band, where an AI’s first draft will omit details the student considers essential and therefore provoke genuine revision.
Falk resists defining human purpose as beating AI. A person may enjoy an instrument without becoming the best professional; the reward is development, expression, and expanded capability. That may offer a path forward even when machines can perform economically useful tasks better.
Labenz asks whether schools should instead find the top 3% capable of Einstein-level paradigm shifts. Falk sees experimentation value but doubts society should run an entire system around that screening problem when AI may gain the capability five years later; meanwhile, it still must create meaningful lives for the other 97%.
Falk is less worried that people will complain simply because they lack jobs than that a small minority will capture all resources and leave everyone else without UBI. Personally, he could happily play board games, read, and spend time with family; the challenge would be helping people feel they are living well rather than merely being unemployed and useless.
14. Teachers should model learning by making more mistakes
Labenz’s preferred culture is “teach and learn every day”: adults and children are on the same AI timeline, regardless of age, so teachers should experiment, share failures, and model adaptation. Falk sharpens that into “make more mistakes”—after identifying the costly ones, schools should celebrate ordinary failures as evidence that learning is occurring.
Falk’s practical mantra is for every teacher to use AI to “save five minutes every day.” Repeated utility builds intuition about what AI can do, what students need to understand, where it belongs in teaching, and where it fails, all while returning rather than consuming scarce time.
Science fiction can widen institutional imagination because its useful form asks what society becomes when one or two premises change; dismissing scenarios as “just science fiction” ignores how much former fiction is already present. Labenz extends that into a student assignment: imagine a different future and practice the agency required to steer toward it.
Falk’s closing division of responsibility is calming but firm. Teachers should focus on students and learn enough AI to decide what their subject requires; principals and school owners must fund professional development; districts and national leaders must plan across three-, five-, and 10-year scenarios. The individual starting point remains deliberately small: “save five minutes a day using AI.”