There are a lot of "AI" EdTech tools on the market right now. A quick search turns up dozens of platforms promising to revolutionize instruction, personalize learning, and save teachers hours every week. The marketing is compelling. The demos look impressive. And adoption is accelerating at a pace that few predicted even two years ago.
But behind the polished interfaces and the bold claims, most of these tools are doing essentially the same thing: prompting a large language model to generate lesson content, wrapping it in a classroom-friendly format, and calling it innovation.
That raises a question the K-12 community should be asking much more aggressively than it currently is: does this approach actually help students learn more?
The national data suggests the answer is no. And understanding why requires looking past the technology itself and into the learning science (or lack thereof) that underpins it.
The Explosion of AI-Powered Lesson Generators
AI adoption in education has moved from early experimentation to mainstream integration in a remarkably short window. By 2025, roughly 69% of high school teachers were using generative AI in some capacity, with 42% of elementary teachers and 29% of all educators using AI specifically for lesson plan creation. Teachers who use AI tools regularly report saving an average of about six hours per week, time that gets redirected toward grading, parent communication, and administrative tasks that have long consumed instructional planning time.
On the surface, this looks like progress. Educators are overworked. Planning periods are too short. The demands on a single teacher’s time have expanded far beyond what the profession was designed to accommodate. If AI can absorb some of that burden, the logic goes, teachers can focus on what matters most: connecting with students and delivering instruction.
The problem is that the content these tools produce has a fundamental design flaw. Not a technical one. A pedagogical one.
What LLM-Generated Lessons Actually Produce
The vast majority of AI EdTech tools on the market today use large language models (the same foundational technology behind ChatGPT, Claude, Gemini, and other general-purpose AI systems) to generate lesson content inside their software. Some add subject-matter guardrails. Some align output to specific state standards. Some format the result into slides, handouts, or interactive modules. But the core process is the same: a prompt goes in, and a block of standards-adjacent instructional text comes out.
In other words, these tools produce nearly the same thing you would get from asking any LLM directly: "Create a California standards-aligned lesson for 6th grade math about understanding and calculating unit rates."
Like magic, a polished-looking lesson appears. It has a learning objective. It has an introduction. It walks through the concept with examples and maybe a few practice problems. It looks like instruction. It reads like instruction. But it is missing virtually every mechanism that learning science has identified as essential for moving information from a student’s short-term awareness into durable, retrievable understanding.
What’s typically absent from these generated lessons tells the real story:
- No built-in checks for understanding. The lesson presents information but never pauses to verify whether students are actually comprehending it before moving forward. Without real-time formative assessment embedded throughout instruction, misconceptions go undetected and compound as the lesson progresses.
- No structured practice sequences. Effective instruction moves students through a deliberate progression: teacher modeling, guided practice with immediate feedback, and independent practice only after demonstrated readiness. LLM-generated content typically jumps from explanation to practice with no scaffolded transition.
- No corrective feedback loops. When a student answers incorrectly, what happens next matters enormously. Research on corrective feedback consistently shows that how errors are addressed (immediately, specifically, and with re-teaching) determines whether the student learns from the mistake or simply moves past it. Static content cannot do this.
- No scaffolded vocabulary development. Academic vocabulary is one of the strongest predictors of student achievement across all content areas. Effective vocabulary instruction requires explicit teaching of word meaning, multiple exposures in context, and structured opportunities for students to use new terms. An LLM might define a word in passing; it doesn’t systematically build comprehension of it.
- No retrieval practice or spaced repetition. Decades of cognitive science research have established that the act of retrieving information from memory, not simply re-reading it, is what strengthens retention. Spaced repetition, interleaving, and low-stakes retrieval opportunities are among the most robustly supported findings in learning science. They are almost entirely absent from AI-generated lesson content.
- No connection to peer-reviewed instructional strategies. The structure of a lesson matters as much as its content. Research-backed methodologies like Explicit Direct Instruction have been validated across hundreds of studies and nearly 4,000 measured effects, consistently producing statistically significant improvements in student achievement across subjects and grade levels. LLM-generated content follows no such methodology.
The National Data Tells a Stark Story

If AI-generated instructional content were effective at improving learning outcomes, we would expect to see some measurable signal from the scale of adoption that has already occurred. Millions of teachers are now using these tools. Billions of dollars in venture capital have flowed into AI EdTech companies. The market is projected to grow from roughly $7 billion in 2025 to over $112 billion by 2034.
Yet student performance is not improving. It is getting worse.
These trends predated the pandemic but have accelerated since, and no countervailing signal from AI tool adoption has emerged.
As Lesley Muldoon, executive director of the National Assessment Governing Board, put it when the results were released: these students are entering adulthood with fewer skills and less knowledge than their predecessors a decade ago, at a time when society demands more of them, not less.
This is not an argument against technology in education. It is an argument that the kind of technology matters enormously, and that the current dominant approach (using general-purpose language models to generate instructional content) is not addressing the problem it claims to solve.
Why Content Generation Is Not the Same as Instructional Design
The distinction is critical, and it is one that gets lost in most conversations about AI in education.
Content generation answers the question: What information should be presented to students? An LLM is genuinely good at this. It can produce accurate, well-organized explanations of virtually any academic concept at any grade level, aligned to any state standard.
Instructional design answers a fundamentally different question: How should that information be structured, sequenced, practiced, assessed, and reinforced so that students actually retain it and can apply it?
These are not the same question. They are not even close. And the gap between them is where student learning lives or dies.
Consider an analogy outside education. A large language model can generate a comprehensive, medically accurate description of how to perform a particular surgical procedure. That does not make the output a surgical training program. A training program would include simulation sequences, progressive skill-building, assessment checkpoints, error correction protocols, and supervised practice, all grounded in evidence about how motor skills and clinical judgment are actually developed. The information might be the same. The design is entirely different. And the design is what determines whether someone can actually perform the procedure.
What the Research Actually Says About Effective Instruction
The evidence base for effective instructional practice is not speculative. It is one of the most extensively studied areas in education, and the findings are remarkably consistent.
John Hattie’s meta-analysis of over 50,000 studies involving more than 80 million students identified the instructional strategies with the highest measurable impact on student achievement. Among the most powerful: teacher clarity (effect size 0.75), formative assessment, direct instruction, scaffolded practice, and structured feedback. These are not philosophical preferences. They are empirical findings replicated across decades of research, thousands of classrooms, and millions of students.
A separate meta-analysis published in the Review of Educational Research examined 328 studies involving nearly 4,000 measured effects over a half century of Direct Instruction research. The findings showed consistently positive, statistically significant effects across reading, math, language, and spelling, with effect sizes that were educationally meaningful and comparable in magnitude to the performance gaps between advantaged and disadvantaged students.
Barak Rosenshine’s influential work on principles of instruction synthesized findings from cognitive science, classroom research, and the practices of master teachers into a coherent framework: begin with a review of previous learning, present new material in small steps, ask questions and check for understanding, provide models and worked examples, guide student practice, check for student understanding frequently, obtain high success rates, provide scaffolds for difficult tasks, require and monitor independent practice, and engage students in weekly and monthly review.
These principles do not describe what most AI lesson generators produce. They describe something fundamentally more structured, more interactive, and more responsive to how the human brain actually acquires and retains knowledge.
The Problem with the "Easy Button"
There is an understandable appeal to tools that reduce teacher workload quickly and visibly. Teachers are under enormous pressure. Planning time is inadequate. Class sizes are too large. Administrative demands are relentless. When a tool promises to generate a week’s worth of lesson plans in minutes, it addresses a real and urgent pain point.
But the framing matters. If the goal is to reduce the time a teacher spends creating documents, then LLM-powered lesson generators accomplish that goal. If the goal is to improve what students actually learn, then the tool needs to do something fundamentally different, and the current generation of AI EdTech, for the most part, does not.
The risk is that widespread adoption of tools optimized for teacher convenience creates a false sense of instructional improvement. Lessons look more polished. Materials are produced faster. But if the underlying instructional design is no better than what a teacher could produce by prompting ChatGPT directly, the investment is not moving the needle on the metric that actually matters: student achievement.
This is especially consequential for the students who need effective instruction the most. Students who are behind grade level, English learners, students with disabilities, students who have missed significant instructional time due to chronic absenteeism: these are the students for whom instructional quality is not a nice-to-have. It is the determining factor in whether they catch up or fall further behind.
The Chronic Absenteeism Connection
The relationship between instructional quality and student attendance is often discussed as separate problems, but they intersect in ways that matter for this conversation about AI in education.
Post-pandemic chronic absenteeism rates remain elevated across the country, with 25–30% of students in many districts missing 10% or more of the school year. Every absence creates a gap in instructional continuity. When students return, they need instruction that is structured enough to help them re-engage with missed content, not a static text block that assumes continuous attendance.
For districts grappling with this challenge, the quality of the instruction delivered during attendance recovery sessions and independent study programs is particularly important. These are contexts where students are working to recover lost instructional time, often outside the regular school day, sometimes with less direct teacher support than they would receive in a typical classroom. The instructional materials used in those settings need to do more of the heavy lifting, providing the structure, the checks for understanding, the scaffolding, and the corrective feedback that a live teacher would normally provide in real time.
Generic AI-generated content is especially poorly suited for these use cases, precisely because it lacks the interactive instructional architecture that struggling and recovering students need most.
What "AI in Education" Should Actually Mean
None of this is an argument against using artificial intelligence in education. AI is a genuinely powerful technology with enormous potential to improve how instruction is designed, delivered, and personalized. But that potential is only realized when AI is used to implement what learning science has already proven works, not to generate text that looks like instruction but functions like a digital textbook.
The distinction between these two approaches is worth making explicit:
❌ Approach 1: AI as Content Generator
A large language model receives a prompt describing a topic, grade level, and standard. It generates text that explains the concept, provides examples, and includes practice problems. The output is formatted to look like a lesson. The teacher saves time. The student reads content. No structural mechanism ensures learning occurs.
✅ Approach 2: AI as Instructional Design Engine
AI constructs lessons according to a deterministic, research-backed instructional framework. Every lesson follows a defined pedagogical sequence. Checks for understanding are embedded at specific intervals. Practice is scaffolded. Vocabulary is explicitly taught. Question formats mirror standardized assessments. The lesson is an interactive experience engineered to produce measurable learning.
Approach 2 is harder to build. It requires deep expertise in instructional methodology, not just prompt engineering. It requires treating learning science as a set of engineering specifications, not optional enhancements. And it requires accepting that the hard problem in education technology was never "how do we generate content faster?" It was always "how do we design instruction that reliably produces learning?"
What Educators and Administrators Should Be Asking
As AI tools continue to proliferate across K-12, the evaluation criteria that districts, schools, and educators apply to these products will determine whether the technology actually serves students or simply serves procurement cycles.
Here are the questions worth asking of any AI-powered instructional tool:
The Standard Our Students Deserve
The edtech market is growing at 36% annually. AI adoption among educators has crossed the threshold from experimental to ubiquitous. Hundreds of millions of dollars are flowing into tools that promise to transform instruction.
But transformation is not measured by adoption rates, revenue growth, or hours saved on lesson planning. It is measured by whether students learn more, retain more, and can demonstrate more. By that standard (the only one that ultimately matters) the current generation of AI-powered lesson generators has not delivered.
The technology exists to do this differently. AI can be used to implement the full architecture of research-backed instruction: the modeling, the guided practice, the checks for understanding, the scaffolded vocabulary, the corrective feedback, the retrieval practice, the standards-aligned assessment preparation. It can be done at scale, across every grade level and content area, aligned to every state’s standards.
But it requires building something fundamentally more sophisticated than a prompt wrapper around a language model. It requires treating instructional design as an engineering discipline, grounded in decades of peer-reviewed research, and holding the output to the standard that our students deserve.
Our students, especially those who are already behind, who have missed instructional time, who are working to recover from chronic absence, who are completing coursework through independent study, do not need another easy button.
They need instruction that is designed, from the ground up, to produce learning. That is what AI in education should have been from the start.
