By
Assistant Professor at aivancity
Director of the MSc in Generative and Agent-Based Artificial Intelligence
Zainab Assaghir
University Professor
Faculty of Sciences, Lebanese University
Libraries of educational prompts are proliferating, but they describe what we ask generative AI to produce, rarely the educational function of that request. Our study proposes a taxonomy of five pedagogical functions of prompts and shows that it can be reliably applied by two evaluators. It also shows that it is difficult to deduce, based solely on the text of a prompt, the level of cognitive complexity it will require.
Generative artificial intelligence is becoming increasingly integrated into educational practices, to the point that UNESCO and the OECD have now developed recommendation frameworks dedicated to it (Miao and Holmes, 2023; OECD, 2026). Teachers use it to prepare lessons, design activities, or create teaching materials; students use it to get explanations, practice, or receive feedback on their work. In all these uses, interaction occurs via a prompt—that is, the written instruction given to a generative AI model.
Behind this diversity of uses lies a more fundamental question: What role does a prompt actually play in education?
Does it simply ask the AI to generate an answer, does it gradually guide the learner, does it require the learner to provide reasoning, or does it encourage the learner to evaluate their own reasoning?
Our study , “From Engineering to Pedagogy: A Functional Taxonomy for Teacher and Student-Facing Educational Prompts,” published in the proceedings of the EC-TEL 2026 conference (Yaacoub and Assaghir, 2027), addresses this question. In it, we propose classifying educational prompts not solely based on the task they require students to perform, but rather based on the pedagogical function they seek to fulfill.
After calibration, two raters who applied this classification to 50 previously unseen prompts achieved a high level of agreement (88%, κ = 0.80). However, the text of a prompt, taken in isolation, does not allow for a reliable determination of the level of cognitive complexity it will actually require. The corpora, coding grid, and annotation workbook are available as open access on the OSF platform.
Why Look Beyond "Prompt Engineering"?
Libraries of educational prompts are growing rapidly. Recent literature reviews highlight the rise of large language models in education, the diversity of prompt-writing practices, and the lack of a shared analytical framework (Chen et al., 2024; Shi et al., 2026). These libraries are most often organized around the task at hand: preparing a lesson, generating a quiz, creating an assessment rubric, or explaining a concept.
This organization is useful in practice, but it provides little insight into the intended pedagogical mechanism and makes it difficult to compare prompts across different contexts.
Two prompts that appear to ask for the same task can, in fact, organize the cognitive work very differently.
One provides an answer directly; another guides the user step by step; a third asks the user to explain their reasoning, consider a counterargument, or evaluate their own work.
The question, therefore, is no longer just what we ask AI to produce, but how the prompt structures the cognitive activity of the teacher or learner.
This distinction is important in education, where the quality of an interaction with an AI depends not only on the quality of the response it generates, but also on what the user is prompted to do on their own.
Five Pedagogical Functions for Analyzing Prompts
Our taxonomy is based on the theory of scaffolding (Wood, Bruner, and Ross, 1976) and on research on self-regulated learning (Zimmerman, 2002). It distinguishes five functions, and each prompt is assigned a single code: that of its dominant function. The English terms listed in parentheses are those used in the published taxonomy.
1. Scaffolding: guiding without thinking for the learner
The " Support " category includes prompts that offer steps, templates, cues, or forms of gradual support.
The idea is to provide enough structure to help the user move forward, while still leaving the user responsible for the core reasoning.
For example, one prompt in our corpus asks the AI to create an educational simulation led by a “game master” who structures the flow of the activity.
Distinguishing this category from others requires precision. A series of questions is not automatically Socratic: if the questions are used primarily to gather the information needed to complete a task, they fall under the category of scaffolding. This is one of the decision rules established during calibration.
2. Socratic questioning: asking questions rather than giving the answer
In this category, the AI is explicitly instructed to ask questions that bring out, clarify, or challenge the user’s reasoning, and to record the direct response.
The goal is not to gradually gather information, but to use questioning to stimulate critical thinking.
This definition avoids labeling any dialogue in which the AI asks several questions in succession as “Socratic.”
3. Cognitive Enforcement: imposing tasks that require deeper thinking
This category refers to prompts that introduce constraints requiring the user to engage in certain cognitive processes: critical thinking, synthesis, counterargumentation, analysis of trade-offs, or evaluation using a rubric.
The challenge is to go beyond merely carrying out a task in a superficial way.
Simply asking the AI to generate a text in several steps is not enough; asking it to examine a counterargument, compare options based on explicit criteria, or critique a piece of work using a rubric adds an additional constraint. One prompt in our corpus, for example, asks the AI to critique a lesson plan using a rubric based on Universal Design for Learning (UDL).
4. Metacognition: observing one's own reasoning
This feature applies to prompts that explicitly guide users through planning, monitoring, self-assessment, error analysis, or review based on defined criteria.
The presence of reflexive vocabulary is not enough: learners must be asked to take concrete steps to check their own work—for example, by using a checklist to review their work during a tutoring session.
This category refers to the self-regulatory mechanisms of learning: planning one's activities, monitoring one's progress, and evaluating the results achieved (Zimmerman, 2002).
5. Other/Unclear: producing content without an identifiable instructional protocol
This last category mainly includes prompts that request the generation of an artifact—such as an assessment rubric or course materials—without a specific teaching protocol.
However, the creation of a teaching resource is not automatically classified in this category: if the prompt includes a teaching sequence, an investigative process, or another identifiable function, that function takes precedence.

How was this taxonomy tested?
The study was conducted in two phases.
An initial development corpus of 50 prompts in English was compiled from four public libraries consulted on February 9, 2026: More Useful Things (11 prompts), Wharton Interactive (6), Microsoft prompts-for-edu (6), and AI for Education (27). The prompts were selected to cover the full range of categories offered by each repository.
This corpus reflected a characteristic of today's public libraries: it was overwhelmingly intended for teachers (46 out of 50 prompts).
The two authors of the study, who are researchers in educational technology, independently coded each prompt along two dimensions.
The first was its primary educational function.
The second was based on the SOLO (Structure of the Observed Learning Outcome) taxonomy, which classifies the complexity of a response into successive levels: unstructured, multistructured, relational, and extended abstract (Biggs and Collis, 1982; Yaacoub et al., 2025). It was used here on an exploratory basis to determine whether the targeted level of cognitive complexity could be inferred from the text of a prompt alone.
To measure inter-rater reliability, the study uses the percentage of agreement and Cohen’s kappa (κ). The latter estimates the agreement between two raters by accounting for the portion of agreement that could occur by chance. According to the most common interpretation scale, a κ between 0.41 and 0.60 corresponds to moderate agreement, between 0.61 and 0.80 to substantial agreement, and above that to near-perfect agreement (Landis and Koch, 1977). These thresholds are arbitrary and vary depending on the author (McHugh, 2012).
An initial, imperfect classification, followed by a marked improvement after calibration
During the initial coding of the development corpus, classification based on pedagogical function achieved 58% agreement, with a κ of 0.41 (95% confidence interval: [0.24; 0.57]), indicating moderate agreement.
The disagreements center on three areas: Scaffolding/Other (4 cases), Scaffolding/Socratic Questioning (3), and Scaffolding/Metacognition (3).
This finding is instructive in and of itself: categories do not become reliable simply because they have been defined; their boundaries must be precise enough for multiple evaluators to apply them consistently.
A calibration session was therefore conducted on ten cases involving significant disagreement. It resulted in seven decision rules, focusing in particular on the distinction between Socratic questioning and justification, and on how to handle prompts related to the generation of artifacts.
The coding scheme was then finalized. A new coding of the same corpus yielded 88% agreement and a κ of 0.82 ([0.67; 0.94]).
However, this improvement—achieved using prompts that had already been discussed—was not enough to demonstrate that the rules worked on new prompts.
Validation using 50 previously unseen prompts
The taxonomy was therefore evaluated on a second corpus of 50 prompts, which was compiled after the coding scheme had been established and excluded all prompts from the development corpus.
This corpus was intentionally balanced to include 25 prompts for teachers and 25 prompts for students, drawn from Microsoft prompts-for-edu (6 and 8) and AI for Education (19 and 17).
Only the pedagogical function was coded here, in two independent runs conducted blind with a fixed grid, with no discussion between runs prior to calculating the agreement.
The main result is an 88% agreement, with κ = 0.80 and a 95% confidence interval of [0.64, 0.93].
The results are similar for both groups: 92% agreement and κ = 0.81 for the prompts intended for teachers, and 84% agreement and κ = 0.78 for those intended for students.
However, these two subsamples each contain only 25 prompts: these figures are for illustrative purposes only, and the difference between them has not been statistically tested.
Under the conditions of this study, these results indicate that the taxonomy retains substantial reliability for prompts not used during its development, whether they are intended for teachers or students. However, not all discrepancies have been resolved. The six remaining discrepancies concern Scaffolding/Cognitive Constraint (2), Scaffolding/Other (2), Scaffolding/Socratic Questioning (1), and Scaffolding/Metacognition (1). All involve scaffolding, making it, in our corpus, the category whose boundaries with other functions remain the most difficult to define.

Not all functions are represented in the same way
In the validation corpus, scaffolding is the dominant function, with 28 out of 50 prompts: 18 intended for teachers and 10 for students.
The cognitive task consists of 12 prompts, divided equally between the two groups.
The six metacognitive prompts, on the other hand, are all intended for students, as are the two Socratic prompts. The “Other/Undefined” category includes two prompts.
This distribution should be interpreted with caution. The corpus was deliberately constructed to be equally divided between the two audiences, using only two libraries: its distribution describes this sample, not the frequency of these functions across all educational prompts. It does, however, suggest a hypothesis to be tested on larger corpora: certain functions, such as metacognition and Socratic questioning, may be more prevalent in prompts directly intended for learners.

An insightful negative result: the limitations of SOLO when using only the prompt
One of the study's key findings does not concern what worked, but rather what did not work consistently.
During the initial coding, the use of SOLO to infer cognitive complexity from the text of the prompts yielded only 36% inter-rater agreement, with κ = 0.05—a level close to chance.
An ordinal weighting, which takes into account the distance between levels, did not improve the results. Given this limitation, the SOLO coding method was not applied to the validation corpus.
The discrepancy with the pedagogical function can be explained by the nature of what is encoded. The functions in our taxonomy often correspond to mechanisms explicitly stated in the text: “proceed step by step,” “use a checklist,” “offer a counterargument,” “check one’s work.”
Cognitive complexity, on the other hand, requires inferring the actual level of understanding required—information that the prompt, taken on its own, rarely provides.
The same text can lead to different interactions depending on the context, the user, the model used, and the successive responses generated during the exchange.
This finding calls for methodological caution: when analyzing cognitive complexity using a framework such as SOLO, the relevant unit of analysis might be the prompt-response pair—or even the entire interaction—rather than the prompt alone.
Designing an educational prompt also involves deciding where the cognitive effort lies
This taxonomy can serve as a framework for interpretation, but also as a design tool. The following recommendations do not stem from the study’s findings, which do not identify any learning outcomes; rather, they stem from its theoretical framework and align with existing recommendations on the responsible use of generative AI in education (Miao and Holmes, 2023; OECD, 2026).
First, they urge us to avoid a systematic “response-first” approach, in which AI immediately provides the expected result.
On the contrary,scaffolding and Socratic questioning can help preserve some productive effort by requiring intermediate steps or justifications.
Cognitive constraints make it possible to explicitly incorporate critical analysis, synthesis, counterarguments, and evaluation.
Finally, metacognitive prompts can help make the stages of planning, monitoring, and self-assessment more visible.
In modules and activities that require writing code, I often observe the same pattern: students paste their code and error message, then ask the AI to correct it. The prompt receives a response, but the student misses the diagnostic insight. So I’ve reworded the instructions I give them: the assistant should not provide corrected code, but should ask one question at a time to guide the student toward formulating a hypothesis about the source of the error, and then verifying it. The task remains the same. What changes is who is doing the reasoning.
However, no category is inherently superior to the others. The appropriate approach depends on the educational objective, the context, the learner, and how the interaction with AI is actually implemented.
These design choices are not merely technical. Deciding how much of the reasoning process to entrust to AI also means deciding how much autonomy to grant the learner: a prompt that systematically provides the answer can relieve the learner of that responsibility, whereas a prompt that withholds it leaves that responsibility to the learner. This decision calls for transparency, because a student who does not understand why the assistant refuses to give them the solution may perceive this withholding as a limitation of the tool rather than as a pedagogical choice. It also raises a question of inclusion: the prompts studied come from English-language libraries designed for specific educational contexts, and their adaptation for French-speaking or multilingual learners is not straightforward. Finally, these prompts run on models whose behavior evolves with each new version; a pedagogical function carefully embedded in a prompt therefore also depends on technical choices beyond the teacher’s control.
What this study allows us to affirm, and what it does not yet allow us to conclude
The study shows that, after calibration, the five proposed functions can be distinguished with a high degree of reliability on the validation dataset examined.
However, it does not demonstrate that one type of prompt leads to better learning outcomes than another.
The analysis focuses on the text of the prompts: it does not examine the interactions that follow, the models’ responses, the work produced by the learners, or their learning outcomes.
However, the educational function embedded in a prompt can change when it is applied in a real-world situation.
A next step in the research, therefore, is to link the categories of prompts to the interactions and outcomes observed: for example, to determine whether scaffolding prompts or Socratic questioning prompts elicit more “productive struggle”—the effort through which the learner must search, reason, and justify—than prompts aimed at the direct generation of a result.
There are also several limitations that must be taken into account.
Both corpora are small in size and come from English-language public libraries; therefore, the results cannot be generalized without caution to multilingual environments, less structured libraries, or real-world tutoring situations.
Furthermore, the development and validation corpora come from the same broad categories of sources. Finally, the validation was conducted in two independent rounds using the same fixed coding scheme, rather than by a new team of annotators. It should therefore be interpreted as evidence of the coding scheme’s out-of-sample stability, rather than as a comprehensive demonstration of its generalizability to other teams.
From the prompt as an instruction to the prompt as a teaching tool
The development of generative AI in education raises a question that goes beyond simply mastering prompt engineering—that is,the art of formulating effective instructions for a model.
An educational prompt is not merely an instruction designed to produce a better output: it also determines who is reasoning, when, under what constraints, and with what degree of autonomy.
The proposed taxonomy provides a vocabulary for analyzing these differences.
This is only the first step. It remains to be seen how these functions play out in real-world interactions and what effects they have on learning processes and outcomes.
The focus then shifts: it is no longer just a matter of writing prompts that enable AI to provide better responses, but of designing interactions in which AI supports learning without taking over the cognitive work that gives it its value.
References
[1] Biggs, J. B., and Collis, K. F. (1982). Evaluating the Quality of Learning: The SOLO Taxonomy (Structure of the Observed Learning Outcome). New York: Academic Press.
[2] Chen, E., et al. (2024). A Systematic Review on Prompt Engineering in LLMs for K-12 STEM Education. arXiv:2410.11123. View the preprint
[3] Landis, J. R., and Koch, G. G. (1977). The Measurement of Observer Agreement for Categorical Data. Biometrics, 33(1), 159–174.
[4] McHugh, M. L. (2012). Interrater Reliability: The Kappa Statistic. Biochemia Medica, 22(3), 276–282. View the scientific article
[5] Miao, F., and Holmes, W. (2023). Guidance for Generative AI in Education and Research. Paris: UNESCO. View the guide
[6] OECD (2026). OECD Digital Education Outlook 2026. Paris: OECD Publishing. View the report
[7] Shi, Y., Yu, K., Dong, Y., and Chen, F. (2026). Large Language Models in Education: A Systematic Review. Computers and Education: Artificial Intelligence, 10, 100529. View the study
[8] Wood, D., Bruner, J. S., and Ross, G. (1976). The Role of Tutoring in Problem Solving. Journal of Child Psychology and Psychiatry, 17(2), 89–100. View the scientific article
[9] Yaacoub, A., and Assaghir, Z. (2027). From Engineering to Pedagogy: A Functional Taxonomy for Teacher and Student-Facing Educational Prompting. In J. Weidlich et al. (eds.), Mindful TEL: Learning Technologies Shaped with Intention. EC-TEL 2026. Lecture Notes in Computer Science, vol. 16906, pp. 255–260. Cham: Springer. View publication
[10] Yaacoub, A., Assaghir, Z., and Da-Rugna, J. (2025). Cognitive Depth Enhancement in AI-Driven Educational Tools via SOLO Taxonomy. In ACR’25, LNNS 1346, pp. 14–25. Springer. View publication
[11] Open Science Framework (OSF). Open materials: corpus, coding grid, and annotation workbook. View the research materials
