This practitioner article presents an evidence-informed developmental architecture for learning with AI. The architecture is proposed as a synthesis and testable research direction, not as a claim to have invented its established component traditions.
A student is struggling to construct an argument. She turns to AI.
The AI could write the argument for her. It could explain how arguments work.
It could suggest evidence. It could model a strong response. It could ask her questions. It could challenge her assumptions. Or it could deliberately wait until she has tried. Every one of these could reasonably be described as using AI for learning. But they do not create the same learning. In one case, AI performs the important thinking. In another, it supports the learner while the learner thinks. In another, it deliberately provokes deeper reasoning. And in another, doing nothing may be the most educationally useful thing the AI can do.
This is why one of the most common questions schools ask about generative AI — How much should students use it? — is increasingly inadequate.
A learner can use AI extensively and think deeply. Another can use it briefly and outsource the most important thinking in the task. Removing AI does not automatically create active learning, just as adding AI does not automatically create dependency.
What thinking does this learner need to do, what role should AI play in that thinking now, and how should that relationship change as the learner develops? For educators, that final question may matter most.
A Better Product Does Not Necessarily Mean a More Capable Learner
Generative AI has made an old educational problem much harder to ignore.
A strong performance does not necessarily mean that the underlying capability is equally strong.
A student might produce an excellent essay because she has developed the ability to construct a sophisticated argument — or because AI constructed much of it. A student might solve a difficult problem because his understanding has developed — or because AI supplied the crucial reasoning. A language learner might produce sophisticated English because her language capability has improved — or because much of the language belongs to the system.
This distinction is increasingly explicit in AI-education research. Rowe’s Gradual Release of AI Technology (GRAIT) model distinguishes AI-supported performance from durable learning and proposes staged AI access linked to demonstrated development (Rowe, 2026). Recent assessment-validity work goes further: when generative AI changes how a performance is produced, it can change what that performance can reasonably be taken to mean (Weidlich, 2026).
The quality of the product cannot, by itself, tell us the quality of the learning. But there is an equally important mistake waiting on the other side. The solution is not simply to remove AI.
Less AI Is Not the Goal
Imagine two students.
The first does not use AI. He waits for the teacher to explain what to do, follows the instructions, completes the activity, and rarely questions his reasoning.
The second uses AI frequently. She develops an initial interpretation herself, asks AI to generate a strong counterargument, evaluates that counterargument, rejects one of AI’s claims because the evidence is weak, revises her position, and then asks AI to attack the revised argument. Which learner is carrying more meaningful cognitive responsibility?
AI usage alone cannot answer that question.
This territory already has substantial theoretical neighbors. Extended Executive Cognition describes strategic allocation of cognitive effort, delegation, and orchestration across human-AI systems (Sidorkin, 2025). Epistemic co-agency emphasizes learning to reason not only with AI, but also through and against it (Samuel, 2026). These perspectives make an important point: sophisticated AI use can itself be a developmental outcome rather than evidence of dependence.
The better question is not simply: How much AI? It is: Where should the thinking happen?
A Developmental View of Learning With AI
The figure below brings these ideas together in a deliberately simplified form. It does not prescribe a fixed sequence of AI use. Instead, it asks a developmental question: given what this learner is developing toward, where should the learning-critical cognition happen now — and what evidence would justify changing that configuration later?
Figure 1. Simplified view of how educational purpose, evidence, cautious interpretation, and reconfiguration can interact as a learner develops. AI is one part of the learning configuration rather than the starting point.
Start With the Learning-Critical Cognition
Suppose I am teaching students to evaluate evidence.
AI could retrieve information, translate difficult vocabulary, explain unfamiliar background knowledge, or organize sources. Those forms of assistance might allow a learner to devote more cognitive effort to evaluating evidence. But if AI decides which evidence is credible, relevant, and sufficient while the learner simply accepts its judgment, AI has performed the very cognition I intended the learner to develop.
Now change the objective. Suppose I want students to evaluate AI-generated judgments. AI making the initial judgment is no longer inappropriate. It is part of the task.
There can therefore be no universal list of cognitive activities that AI should or should not perform. The learning purpose changes the answer. Protect the cognition that constitutes the intended learning.
Strategically support, share, or delegate other cognition when doing so enables deeper engagement with that learning.
This principle is not a new learning theory. It builds on established work in scaffolding, constructive alignment, cognitive support, distributed cognition, self-regulated learning, dynamic assessment, and the assistance dilemma. The AI-era complication is that the scaffold can now perform many of the same cognitive functions the learner is supposed to acquire. AI can explain. Reason. Question. Generate. Evaluate. Critique. Revise. Plan. Decide.
The teacher’s challenge is therefore no longer merely determining how much assistance to provide. It is increasingly determining what kind of cognitive relationship is developmentally appropriate.
Even Good AI Questioning Can Hide a Problem
One attractive response is to turn AI from an answer machine into a questioner.
Instead of saying, Your evidence doesn’t support your conclusion, AI might ask: Which piece of evidence most strongly supports your conclusion?
Instead of providing another interpretation, it might ask: What alternative explanation could also fit the evidence?
Instead of correcting reasoning, it might ask: What evidence would make your conclusion wrong?
This can produce much richer learner thinking. But consider what happens after months of excellent AI questioning. The learner evaluates evidence when AI asks. The learner generates alternatives when AI asks. The learner identifies assumptions when AI asks. Now remove the questions.
Does the learner spontaneously initiate those processes?
Perhaps. Perhaps not.
This is not a newly discovered educational problem. Self-questioning, metacognitive scaffolding, adaptive tutoring, and fading have long addressed movement from supported toward increasingly self-regulated cognition. Koedinger and Aleven (2007) described the assistance dilemma as the challenge of balancing the giving and withholding of assistance in ways that optimize learning.
There is a meaningful developmental difference between performing a cognitive process when prompted and recognizing when that cognitive process needs to be initiated.
So the educational question eventually becomes: Can the learner think when AI asks the right question? And then: Does the learner know which question needs asking?
Development Does Not Necessarily Mean AI Fading Away
A tempting response is to imagine good AI-supported learning as a simple progression: lots of AI -> less AI -> no AI. That is too crude.
Consider a learner developing the ability to evaluate evidence. Early on, AI might model the process. Later, AI might prompt it. Later still, AI deliberately waits. The learner begins initiating evaluation independently. Eventually the learner evaluates unfamiliar evidence without assistance. At that point, AI might return. But it does not return to evaluate the evidence for the learner. Instead it says: Here is evidence that appears to contradict your conclusion. Now AI is creating a harder cognitive problem.
Later still, the learner might decide: I want AI to generate the strongest evidence against my position before I finalize it.
AI involvement has increased again. But dependence has not necessarily increased. The learner has changed, and therefore the educationally appropriate function of AI may also have changed.
The individual ingredients here are not new constructs. GRAIT already proposes staged AI access (Rowe, 2026). Extended Executive Cognition already describes strategic allocation and orchestration (Sidorkin, 2025). Epistemic co-agency already emphasizes critical engagement with AI as an epistemic partner (Samuel, 2026).
The narrower proposition is this: evidence relevant to learner development may justify a qualitative change in the cognitive function AI performs, not merely a change in the amount of AI assistance.
The important transition may therefore be: one AI function -> withdrawal from that function -> learner control -> AI return in a different function that creates a higher developmental demand.
I use developmental reconfiguration descriptively for this process. It is a proposed mechanism and research direction, not a claim that reconfiguration, scaffolding, fading, or cognitive allocation themselves are new.
A Developmental Chain
To make this practical, begin with a familiar instructional logic:
Educational Purpose -> Target Development -> Cognitive Demand -> Learning Configuration -> Learning Experience.
Constructive alignment has long connected intended outcomes, learning activities, and assessment (Biggs, 1996). Formative and dynamic approaches similarly use evidence to inform subsequent support. AI introduces an additional question inside that familiar logic: Who or what is performing the learning-critical cognition through which the intended development is supposed to occur?
The answer may include the learner, peers, the teacher, AI, other tools and resources, or some combination of them. I use learning configuration simply as a practical description of that arrangement.
The important question is whether the configuration is appropriate for the development we intend.
Finding a Broken Chain
Suppose the purpose is to develop critical reasoning. The target capability is evaluating evidence. The required cognition includes judging relevance, credibility, and sufficiency.
The activity looks sophisticated. Students investigate a difficult question using AI. But AI identifies the relevant evidence, assesses source credibility, compares the evidence, and recommends the strongest conclusion. The students then create polished presentations.
Everything looks successful. Except the cognition we intended to develop was largely performed elsewhere.
The problem is not that AI was used. The problem is that Target Development -> Cognitive Demand -> Learning Configuration became misaligned.
Now consider another failure. A student independently evaluates scientific evidence extremely well but struggles to explain the evaluation in English. If the system interprets weak English performance as weak scientific reasoning, the problem lies somewhere else:
Performance -> Evidence -> Interpretation. Or imagine that the system correctly identifies that a learner rarely generates alternative explanations. It responds by having AI automatically generate three alternatives every time. Again something goes wrong: Interpretation -> Reconfiguration.
The diagnosis may be correct. The response removes precisely the cognition that needs developing.
Gap One: What Are We Justified in Believing?
This is why the chain cannot simply be continuous. There are places where we should deliberately stop. The first lies between what we observe and what we believe it means. A student produced an excellent argument. That is an observation.
Therefore the student can independently construct excellent arguments. That is an inference. Those are not the same thing.
Epistemic Gap: What does this evidence actually justify us in believing about the learner?
This gap is firmly grounded in existing assessment-validity reasoning rather than being a new invention. Weidlich (2026) makes the GenAI-specific problem particularly clear: AI can affect attribution of performance, construct representation, extrapolation beyond the assessment event, and whether the evidence is adequate for a proposed use.
AI makes this issue relevant far beyond formal assessment. The same discipline matters whenever classroom evidence is used to update our understanding of a learner.
Gap Two: What Are We Justified in Changing?
There is another gap.
Suppose we have good evidence that a learner can independently evaluate evidence but still requires prompting to generate alternative explanations. It does not automatically follow that we know what intervention should come next.
Pedagogical Gap: What does our current understanding justify us in changing about the learner’s next experience?
These are different questions: What are we justified in believing? and What are we justified in changing?
AI makes the distinction especially important because an adaptive system can compress observe -> infer -> recommend -> personalize -> intervene into an almost invisible sequence. The smoother that sequence becomes, the easier it is to forget that every arrow contains assumptions.
Two Warranted Crossings
Before evidence becomes a claim about the learner, ask: What exactly happened? Under what conditions? What did the learner do? What did AI do? What did peers and the teacher do? What alternative explanation fits? What contradictory evidence exists? What does the learner think happened? What remains uncertain?
Then we can move from Evidence -> warranted provisional interpretation. But before that interpretation changes the learner’s environment, ask again: What development are we trying to produce? Why should this particular change help? What cognition will it give to or take from the learner? What happens if our interpretation is wrong? Is the change reversible? Does the learner have a meaningful voice? Does professional judgment support it? What evidence would tell us whether the change helped?
Then we can move from Understanding -> warranted reconfiguration.
The idea of warrants is inherited from established validity and professional-reasoning traditions. The proposal here is not that evidence needs interpretation or that interventions need justification. The practical point is that, in AI-mediated learning, both crossings need to remain visible.
Evidence should not automatically become a learner claim, and a learner claim should not automatically become an intervention.
Either crossing can legitimately stop. Sometimes the appropriate conclusion is simply: We don’t know yet.
Continue, Scaffold, Climb, Branch — or Investigate
Once a change is justified, several routes are possible. These are navigation choices, not separate learning theories.
Continue
The current developmental pathway remains appropriate. Keep going.
Scaffold
The learner needs temporary support to access worthwhile cognition. A student may have sophisticated scientific reasoning but insufficient academic English to express it. Scaffold the language. Do not unnecessarily lower the science.
Climb
Evidence shows that the current challenge is no longer sufficient. Do not simply give the learner more of the same work. Increase the quality of the cognitive demand. A learner who already evaluates evidence independently might instead construct competing explanations and determine what evidence would discriminate between them.
Branch
A different route may better reveal or develop the intended learning. Written explanation may not be the best route. Perhaps oral defense, investigation, modeling, collaboration, or another representation is more appropriate.
Investigate
Sometimes the evidence does not justify a consequential reconfiguration. Do not personalize merely because the system can. Investigate.
These routes can also combine. A learner might Climb conceptually while receiving a Scaffold for academic language. Development is multidimensional; learners should not be reduced to a single global level.
Sometimes the Next Learning Experience Should Test Our Interpretation
Imagine a learner consistently generates alternative explanations after AI asks: What else could explain this?
There are at least two plausible interpretations. Hypothesis A: The learner cannot yet generate alternatives without support.
Hypothesis B: The learner can generate alternatives but has not been given an opportunity to initiate the process independently.
Five more AI-prompted activities may produce five more apparently successful performances without distinguishing A from B. So change the condition. Do not prompt.
Observe.
The next learning experience now serves two purposes: it provides another opportunity for development, and it produces evidence capable of challenging our current interpretation of the learner. When uncertainty matters, do not always collect more of the same evidence. Design a condition capable of distinguishing between competing interpretations.
This matters especially in AI-supported learning because an intervention can otherwise help manufacture evidence that confirms the assumption that produced the intervention.
The Learner Is More Than the Model
This becomes critical if schools eventually use AI-supported systems that connect learning evidence across lessons, subjects, teachers, and time.
Such a system could be enormously useful. It might notice: In English, evidence evaluation is independently initiated. In history, source reliability is independently questioned, but competing interpretations remain prompted. In science, evidence interpretation appears dependent on language support. That could be much more educationally useful than three isolated grades. But the system still does not know the learner. It maintains a bounded representation based on evidence.
Open and negotiated learner-model research has long shown the value of allowing learners to inspect, discuss, and correct representations of their learning (Bull, 2016). In an AI-rich system, that principle becomes even more important.
A responsible learner representation should preserve where evidence came from, when it occurred, under what conditions, what support was present, what contradicts it, how uncertain the interpretation remains, and whether an older interpretation has been superseded. The learner should be able to say: That isn’t true anymore. The teacher should be able to say: The system has misunderstood this evidence. And the system must be able to say: There is not enough evidence to know. The representation is not the learner.
The Full Architecture
The complete architecture makes the distinction visible. The learner and our representation of the learner are not the same thing. Learning experiences can change the learner; performance provides an evidence window into that development. Evidence must then cross two different inferential gaps before it should influence future learning conditions: first, what are we justified in believing about the learner? Second, what are we justified in changing?
The architecture therefore connects learner development with a second process of interpretation and reconfiguration without collapsing the two.

Figure 2. Full evidence-informed architecture for learning with AI. The learner develops through experience; performance provides an evidence window. Evidence can inform warranted interpretation and, where justified, reconfiguration of learning-critical cognition across the learner, teacher, peers, AI, and other resources.
Experience generates evidence. Evidence informs warranted understanding. Warranted understanding can justify reconfiguration. Reconfiguration changes experience. The learner — not the model — develops through those experiences.
Two Loops Are Running at the Same Time
The architecture can be understood as two coupled processes.
The Learner Development Loop asks: What is the learner becoming?
Purpose leads to a developmental target. The target implies cognitive demands. Those demands inform a learning configuration. The configuration produces experience.
Experience can change the learner and also generate performance and evidence. Over time we examine development, transfer, persistence, and new possibilities.
The Learning Support and Reconfiguration Loop asks: What are we learning about how to support this learner’s development?
Evidence is interpreted. Interpretations are challenged. A bounded learner representation is updated. Possible changes are considered. A new configuration is selected. The change is implemented. Its consequences become new evidence.
Human-AI shared-regulation research already conceptualizes humans and AI as interacting regulatory systems (Järvelä, Nguyen, & Hadwin, 2023). The distinction here is more specific: improvement in the support system is not itself evidence that the learner has developed.
A learner is learning. The educational system is learning about how to support that learning. Those are related processes, not identical ones.
Across Subjects, Preserve the Context
Now imagine the architecture operating beyond one lesson.
An English teacher observes that a learner independently evaluates textual evidence. A science teacher finds that the same learner requires substantial prompting to evaluate experimental evidence. A history teacher finds strong source criticism but little spontaneous generation of competing interpretations. A simplistic system might calculate: Critical Thinking: 78%. That would throw away much of what matters.
A better representation might say: Across several contexts, the learner repeatedly demonstrates independent evaluation of evidence. Generation of competing explanations remains more dependent on prompting, and transfer into unfamiliar scientific contexts is not yet sufficiently evidenced.
Now the next learning opportunity can be more intelligent. Perhaps another teacher deliberately creates a novel task requiring competing explanations before either teacher or AI prompts them.
Cross-subject evidence then becomes more than data aggregation. It becomes an opportunity to investigate transfer.
Development Is Not Always Upward
Another danger is imagining development as a staircase. It isn’t.
A learner may independently demonstrate a capability and later struggle with it. A capability may be newly acquired, fragile, support-dependent, consolidating, stable, transferring, unused, weakened, or recovered.
AI makes this especially important. A learner might become exceptionally skilled at working with AI while gradually exercising a previously independent capability less often.
Performance could remain excellent. AI orchestration could improve. Independent capability could weaken. That is why some capabilities should occasionally be reconfirmed under informative conditions.
Not because students should constantly be tested without AI. But because if we make claims about independent capability, we occasionally need evidence relevant to independent capability.
The More Advanced Learner May Sometimes Use More AI
Imagine two learners. Learner A needs AI to identify weaknesses in an argument.
Learner B identifies those weaknesses independently and then uses AI to construct the strongest possible adversarial case against the argument. Learner B uses more AI. But AI is performing a different function. That learner may be less cognitively dependent while more technologically augmented.
The amount of AI use is a poor proxy for educational quality. The developmental question is: What is AI doing, what is the learner doing, and why?
Eventually, the Learner Should Help Navigate
Initially, the teacher may make most of these decisions. Continue. Scaffold. Climb. Branch. Wait. Use AI. Do not use AI yet. Ask AI for a challenge.
Try independently first. But if the learner always depends on the teacher or AI to configure learning, something remains unfinished. Over time, the learner should increasingly participate.
A learner might eventually say:
I want to try this independently first because I need to know whether I can still do it.
I understand the science, but the English is stopping me from explaining it. I need language support, not an easier problem.
I’ve already done several examples like this. Give me something where the evidence conflicts. Don’t tell me what’s wrong with my argument. Give me the strongest evidence against it.
Or even:
I’m not sure whether I actually understand this or whether AI has been carrying too much of the reasoning. I need to test myself.
This overlaps substantially with self-regulated learning, learner agency, epistemic co-agency, and Extended Executive Cognition. The claim is therefore not that learners managing their relationship with AI is a new educational outcome.
The more specific research question is: Can participation in evidence-informed reconfiguration help learners become better at navigating future learning configurations for themselves? That is an empirical question. It should be tested rather than assumed.
What This Architecture Is — and Is Not
This is not a claim to have invented scaffolding, formative assessment, adaptive learning, learner modeling, metacognition, distributed cognition, socially shared regulation, assessment validity, student agency, strategic cognitive offloading, human-AI orchestration, or staged AI use. All of those traditions contribute important pieces.
The proposal is narrower: When an intelligent system can perform many of the cognitive functions education is trying to develop, evidence relevant to learner development should inform not only how much support is provided, but potentially which learning-critical functions are performed by the learner, teacher, peers, AI, and other resources. Those changes should then be tested through subsequent evidence rather than assumed to be beneficial.
The approach is therefore not a fixed sequence of AI use. It is an evidence-informed developmental architecture.
Evidence does not automatically determine what happens next. It informs professional and learner judgment. Claims about learners remain bounded and provisional. Changes to learning conditions require justification. AI may withdraw from one cognitive function and later return in another. Different capabilities can require different configurations simultaneously. A learner can need scaffolding in one dimension while being ready to climb in another. And sometimes the appropriate response is: We do not know yet. Let’s create a better opportunity to find out.
The Teacher’s Role Becomes More Consequential
None of this removes the teacher. It makes professional judgment more consequential. The teacher increasingly becomes a designer of developmental conditions.
The teacher asks: What development matters here? What thinking would produce it? Who is currently doing that thinking? What does the evidence actually show? What might we be wrong about?
What should change? What should remain difficult? When should AI help? When should it wait? When should it challenge? When should it return in a different role? And eventually, how can the learner increasingly make these judgments too?
AI can help us reason about some of these questions. It should not quietly take ownership of them.
Beyond Personalization
Much of the conversation around educational AI focuses on personalization.
Give the learner the right content. At the right difficulty. At the right time. With the right feedback. Those things matter. But generative AI creates another possibility.
What if adaptation concerns not only what the learner receives, but also where learning-critical cognition occurs?
Not merely: What should AI give this learner next? But: What should this learner increasingly become capable of doing, and what configuration of learner, teacher, peers, resources, and AI is most likely to support that development now?
Then, as the learner changes, change the configuration. And test again. That is a more demanding vision of AI-enhanced learning. But it is also a more human one. Because the ultimate objective is not to build an AI system that claims to know exactly what every learner needs. It is to create educational conditions in which learners continue to develop — and increasingly become capable of making good judgments about learning themselves.
The Question Behind the Question
So when AI does the thinking, what should the learner do?
There is no single answer. Sometimes the learner should observe. Sometimes attempt. Sometimes struggle.
Sometimes collaborate. Sometimes question. Sometimes evaluate what AI produced.
Sometimes reject it. Sometimes use AI to go beyond what unaided cognition could reasonably achieve. Sometimes work without it. Sometimes ask it to remain silent. And sometimes deliberately invite it back. The important thing is that these decisions should not be accidental. They should be connected to what the learner is becoming.
Perhaps that is the deeper educational challenge of generative AI. Not preserving an artificial world in which humans perform every cognitive task unaided. And not constructing a frictionless world in which intelligent systems perform every difficult cognitive task for us. But helping learners develop the knowledge, capability, judgment, agency, and self-understanding to navigate intelligently between the two. The future learner may not be defined by independence from AI.
The more demanding goal may be a learner who increasingly understands what they need to learn, recognizes what thinking they need to do, knows when support is useful, knows when challenge is needed, recognizes when they are ready to climb or take another route, and can decide when AI should assist, question, challenge, extend — or get out of the way. That is not less thinking in the age of AI.
It requires us to become much more deliberate about where thinking happens, why it happens there, and how that relationship should change as a learner develops.
A Research Proposition, Not a Finished Theory
This architecture is a developing research proposition. Its distinctive claims require empirical testing.
One useful comparison would place four conditions side by side: static AI support, simple fade-only support, performance-adaptive AI, and developmental cognitive reconfiguration.
Outcomes could include immediate performance, later independent capability, spontaneous initiation of target cognitive processes, transfer, persistence, strategic AI delegation, and learner judgment about when to recruit or resist AI.
The architecture would be weakened if qualitative reconfiguration adds complexity without developmental benefit, if simple fading performs just as well, or if learners cannot meaningfully distinguish appropriate configurations. That is a feature, not a defect. A useful educational proposal should be capable of being wrong.
References
Biggs, J. (1996). Enhancing teaching through constructive alignment. Higher Education, 32, 347-364. DOI: 10.1007/BF00138871.
Bull, S. (2016). Negotiated learner modelling to maintain today’s learner models. Research and Practice in Technology Enhanced Learning, 11, DOI: 10.1186/s41039-016-0035-3.
Järvelä, S., Nguyen, A., & Hadwin, A. F. (2023). Human and artificial intelligence collaboration for socially shared regulation in learning. British Journal of Educational Technology, 54(5), 1057-1076. DOI: 10.1111/bjet.13325.
Koedinger, K. R., & Aleven, V. (2007). Exploring the assistance dilemma in experiments with Cognitive Tutors. Educational Psychology Review, 19(3), 239-264. DOI: 10.1007/s10648-007-9049-0.
Rowe, L. (2026). From gradual release of responsibility to gradual release of technology: A case for the staged use of AI in formal education. Education Sciences, 16(8), 1291. DOI: 10.3390/educsci16081291.
Samuel, A. (2026). Learning with machines: Toward a theory of epistemic co-agency. Computers and Education: Artificial Intelligence, 10, Article 100573. DOI: 10.1016/j.caeai.2026.100573.
Sidorkin, A. M. (2025). Extended executive cognition, a learning outcome for the AI age. Computers and Education Open, 9, Article 100294. DOI: 10.1016/j.caeo.2025.100294.
Weidlich, J. (2026). Which inference is at risk? Assessment validity reasoning and generative AI. Assessment & Evaluation in Higher Education. Advance online publication, September 20, 2026. DOI: 10.1080/02602938.2026.2734795.
Positioning note. This article proposes an evidence-informed developmental architecture and a testable reconfiguration mechanism. It synthesizes established traditions rather than claiming invention of scaffolding, formative assessment, assessment validity, learner modeling, self-regulated learning, distributed cognition, staged AI access, or human-AI cognitive orchestration.
