Blog

Why ChatGPT is the wrong AI for education


Most universities are still having the wrong debate about AI. The question in faculty meetings and policy committees tends to be binary: should we allow students to use ChatGPT, or should we ban it? Neither answer is particularly useful. Banning it is unenforceable, since students are already using it and detection tools are unreliable. Permitting it without guidance leaves students using a tool that's poorly suited to education in ways that may actively harm their learning.

The better question, the one that most policy conversations haven't reached yet, is which kind of AI is appropriate for education and which kind is counterproductive.


Three problems with general-purpose chatbots in education

ChatGPT and similar general-purpose AI tools have three fundamental limitations when used for studying.

They hallucinate. General-purpose language models generate plausible-sounding text regardless of whether it's accurate. In academic contexts, this means students can receive confident, detailed explanations that are partially or entirely wrong. A student asking ChatGPT about a concept from their political science module might get an answer that sounds authoritative but misattributes a theory, cites a paper that doesn't exist, or presents a contested claim as settled consensus. The student, who turned to the AI because they didn't understand the material, has no way to evaluate whether the answer is correct, which means they may learn and reproduce errors.

They don't know the curriculum. ChatGPT answers from its training data, which is the internet at large. It doesn't know what's on the student's syllabus, which readings were assigned, what the lecturer emphasised, or how the course frames a particular debate. When a student asks about the concept of sovereignty in international relations, ChatGPT gives a generic answer drawn from its training corpus. But the student's course might frame sovereignty through a specific theoretical lens, assign specific readings that argue a particular position, and expect the student to engage with that framing in their essay. The generic answer is at best unhelpful and at worst misleading, because it doesn't match what the student is actually being asked to learn.

They don't cite sources the student can verify. When ChatGPT provides a reference, it's often fabricated or approximate. The student can't click through to the original passage in their assigned reading, check the claim against their lecture slides, or verify the citation in their course materials. This is a serious problem for academic work, where the ability to trace an argument back to its source is fundamental to the entire enterprise.


What a course-grounded AI tutor does differently

An AI tutor grounded in the student's own course materials works on a fundamentally different model. The student uploads their syllabus, their lecture recordings (which are transcribed and searchable), their assigned readings as PDFs or ebooks, and their lecture slides. The AI indexes this specific content and answers questions from it.

When the student asks about sovereignty in international relations, the answer comes from their assigned readings and their lecturer's explanation. The response cites the specific page in the specific reading, or the specific timestamp in the specific lecture, where the concept was discussed. The student can click the citation and see the original text or hear the original explanation.

If the student's materials don't contain the answer, the AI says so rather than generating one from its general training data. This is a constraint by design: the tutor is scoped to the student's course materials, which means it can't produce the kind of plausible hallucination that makes general-purpose chatbots unreliable in academic settings.

The distinction matters for three reasons.

Accuracy. Every answer is grounded in materials that the institution itself has assigned. The AI is retrieving and synthesising from known-good sources rather than generating from a vast, unverifiable training corpus.

Relevance. The answers match the course's framing, terminology, and theoretical approach, because they're drawn from the course's actual materials. A student preparing for an essay gets help that's directly applicable rather than generically adjacent.

Verifiability. Every claim includes a citation that points back to the original source. The student can check the AI's answer against the reading or the lecture. Professors can trust the tool because every answer is traceable.


The academic integrity question

Faculty concern about AI in education is largely, and understandably, about academic integrity. The worry is that students will use AI to generate work they submit as their own, bypassing the learning that the assessment was designed to produce.

This concern is valid when applied to general-purpose chatbots that generate essays, produce answers to exam questions, and create original text that the student passes off as their own thinking. These uses actively undermine learning, because the student is outsourcing the cognitive work that the assessment exists to develop.

A course-grounded AI tutor operates differently. It helps students navigate and understand the materials they've already been assigned. It explains concepts using the lecturer's framing. It finds the relevant passage for an essay they're writing themselves. It generates practice questions so the student can test their understanding before an exam. It connects ideas across readings and lectures to deepen comprehension.

This is closer to what a really good search engine over the student's personal library does than to what ChatGPT does. The tutor helps the student understand their sources better. It doesn't replace those sources with generated text. The student still has to do the thinking, the analysis, and the writing, but they do it with better comprehension of the material they're working with.

A useful framework for thinking about this: would a professor be comfortable if a student used a human tutor to do exactly the same things? If the answer is yes, the AI version is equally appropriate. Human tutors explain concepts, help students find relevant sources, and quiz them on material. That's what a course-grounded AI tutor does, at a fraction of the cost and available at any hour.


The positive incentive loop

There's a secondary effect worth noting for faculty members who curate their own course materials. Because a course-grounded AI tutor is only as good as the materials it's trained on, the quality of the tutoring scales with the quality of the curriculum.

A course with carefully selected readings, well-structured lectures, and clear supporting materials produces an AI tutor that gives excellent, well-grounded answers. A course with vague readings and disorganised slides produces a tutor that reflects that disorganisation. This creates a positive incentive loop: professors who invest in their materials see direct evidence of that investment in the quality of the AI support their students receive.

For departments thinking about AI adoption, this reframes the conversation from "will AI undermine our teaching?" to "does AI give us a new reason to invest in the quality of our teaching materials?"


The policy implication

Universities that treat all AI tools as equivalent in their policies are missing the most important distinction in the landscape: the difference between a general-purpose chatbot that generates from its training data and a course-grounded tutor that retrieves from assigned materials.

A policy that says "students may not use AI" bans both, losing the benefits of the second because of the risks of the first. A policy that says "students may use AI" permits both, failing to distinguish between appropriate and inappropriate uses.

The more useful policy framework distinguishes between AI that generates work the student submits as their own (clearly inappropriate), AI that helps students understand their assigned materials (clearly appropriate), and a grey area that requires explicit guidance. Course-grounded AI tutoring sits firmly in the second category, and making that distinction explicit protects both learning outcomes and academic integrity.


Frequently asked questions

Can students just copy the AI tutor's answers into their essays? The tutor's answers are explanations and references grounded in course materials, not polished essay prose. Copying an explanation of a concept is no different from copying a textbook passage, and institutions already have plagiarism detection for that. The tutor is designed to help students understand, not to produce submittable text.

What if a student uses both ChatGPT and a course-grounded tutor? Many students will, and a clear AI policy should address this directly. The key distinction: using the course-grounded tutor to understand assigned readings is studying. Using ChatGPT to generate an essay is not. The tools are different, the uses are different, and the policy should reflect that.

How do we know the AI tutor's answers are accurate? Because every answer is retrieved from materials the institution itself assigned, and every claim includes a citation pointing to the specific source document and passage. The accuracy is bounded by the quality of the course materials, which the institution controls.

Does this work for all subjects? It works for any subject where the learning involves engaging with texts, lectures, and source materials, which covers the vast majority of higher education. Subjects with heavy practical components (lab work, clinical placements, studio practice) benefit from the AI tutor for the theoretical content alongside the practical work.


Related reading: Bloom's two sigma problem, Why students drop out, An AI that only answers from your course materials, A practical AI policy for universities. Related pages: AI tutor, For students, Fabric vs ChatGPT, Fabric vs Khanmigo.

The workspace that thinks with you.

Ready when you are.

The workspace that thinks with you.

Ready when you are.

The workspace that thinks with you.

Ready when you are.