Blog

How to run an AI tutor pilot at your university


If you're reading this, you probably already believe that AI tutoring grounded in course materials could benefit your students. The problem you're solving now is a different one: how to propose, structure, and run a pilot that produces evidence convincing enough for the people who control the budget.

This post is designed to be forwarded to a committee. It covers the scope, the metrics, the timeline, the setup, and the analysis, in enough detail that a proposal can be drafted from it directly.


Step 1: Choose the right cohort

A pilot needs to be small enough to manage but large enough to produce meaningful data.

Size: 100-200 students in the treatment group, with a matched comparison group of similar size. This is large enough for statistical significance on key metrics while being small enough to set up and monitor within a single department.

Selection: Choose a course that has historically had challenges with student comprehension, engagement, or retention. First-year courses are particularly suitable because the compounding gap that drives dropout typically begins in the first year, and first-year cohorts are large enough to provide a comparison group.

Comparison group: Either a parallel section of the same course taught by the same lecturer (if available), or the previous year's cohort taking the same course with the same materials. The comparison should control for as many variables as possible: same lecturer, same materials, same assessment structure.

Subjects: The AI tutor works for any subject where students engage with texts, lectures, and source materials, which covers the majority of higher education. Subjects with particularly strong reading lists (humanities, social sciences, law) tend to produce the most immediately visible results, because the tutor's ability to help students navigate and understand readings is most obviously valuable where the reading load is highest.


Step 2: Define what you're measuring

Decide on your metrics before the semester starts, not after. The committee will want to know what you expected to find, not just what you did find.

Primary metrics:

Grade distribution. Compare the distribution of final grades between the treatment and comparison groups. A shift in the grade distribution (fewer D/F grades, more B/C grades, or a higher mean) is the most direct evidence of academic impact. Even a small shift in the fail rate is significant, because each student who passes rather than failing is a student who doesn't need to retake the module.

Retention rate. Compare the percentage of students who complete the course in the treatment versus comparison groups. If the course has historically had a withdrawal rate above 5%, this metric can be informative. For lower-withdrawal courses, the sample may not be large enough for statistical significance on this metric alone.

Student satisfaction. A brief survey at the end of the semester asking about the AI tutor's usefulness, frequency of use, and perceived impact on their understanding. Include questions about confidence, feelings of academic support, and willingness to recommend the tool to other students.

Secondary metrics:

Engagement frequency. How often students used the AI tutor, measured by session count, queries per week, and the time distribution of usage (do students use it most before assessments? Late at night? During reading weeks?). These patterns are useful for understanding how the tool fits into student behaviour.

Material coverage. Which course materials students uploaded, how completely they used the tutor across different topics, and whether usage correlated with the topics they found most difficult.

Qualitative feedback. Open-ended responses from students about what they found most and least useful, and structured interviews with a sample of heavy users and non-users to understand adoption and non-adoption patterns.


Step 3: Set up the pilot

The setup should require minimal effort from faculty and students. If the pilot creates significant administrative overhead, it undermines the case for scale.

Student onboarding. Students in the treatment group are given access to Fabric and instructed to upload their course materials: syllabus, lecture slides, assigned readings, and textbook chapters. If the course includes recorded lectures, students can record and transcribe these directly. The onboarding session should take 15-30 minutes and can be incorporated into the first tutorial or seminar.

Faculty involvement. The faculty member teaching the course doesn't need to change their teaching, materials, or assessment. The AI tutor works from whatever materials the students upload. Faculty should be briefed on what the tool does (and what it doesn't do), and should be invited to test it themselves with their own course materials so they can see what students experience.

Data handling. Student usage data should be anonymised and aggregated for the pilot analysis. Individual student usage should not be visible to faculty or linked to grades. Student data is encrypted and not used to train AI models. Clarify this in the student consent process to address privacy concerns.

Ethical approval. If you intend to publish the results or use the data for research purposes, obtain ethical approval before the semester begins. If the pilot is purely for institutional evaluation, your institution's standard evaluation procedures may be sufficient, but check with your ethics committee.


Step 4: Run the semester

Once the pilot is live, the operational demands are minimal.

Mid-semester check-in. Halfway through the semester, review the engagement data to ensure the tool is being used and to identify any adoption issues. If usage is lower than expected, a brief reminder or a guided session showing students specific use cases (asking for explanations, generating practice questions, finding connections between readings) can boost engagement.

No intervention on the comparison group. The comparison group should receive exactly the same teaching and support they would have received without the pilot. The point is to measure the incremental impact of the AI tutor, not the impact of extra attention.

Document everything. Keep a brief log of any unusual events during the semester (a module being restructured, a lecturer being replaced, an assessment being changed) that might affect the comparison. These will be important when interpreting the results.


Step 5: Analyse and present the results

At the end of the semester, compile the data and prepare a presentation for the committee that controls the budget.

Lead with the retention and grade data. These are the metrics that matter most to institutional decision-makers. Present the grade distributions side by side. Highlight any differences in fail rates, withdrawal rates, or the proportion of students achieving upper grades.

Include the financial context. Connect any improvements to the financial case for retention. If the pilot retained even a handful of students who would otherwise have withdrawn, calculate the revenue implications. At around $10 per student per month, the tool pays for itself if it retains a very small number of students.

Show the engagement patterns. The committee will want to know whether students actually used the tool. Show session counts, usage patterns over the semester, and the correlation between usage and grades if one exists.

Include student voices. Quantitative data makes the case. Qualitative data makes it compelling. Selected quotes from student surveys about how the tool helped them understand difficult material or catch up after falling behind are powerful supporting evidence.

Present the scale-up proposal. If the results are positive, include a clear proposal for expanding the pilot: which departments, how many students, what the cost per student would be at scale, and what the expected return on retention would be based on the pilot data.


The proposal template

For the person who needs to send a brief proposal to the committee before the pilot is approved, here's the core structure:

The problem. Student dropout costs the institution £X / $X per student per year. Current retention interventions are reactive and human-dependent. The compounding gap that drives dropout begins with small, early academic difficulties that current support can't catch at scale.

The intervention. An AI tutor, grounded in the course's own materials, available to students at any hour, that explains concepts, answers questions from the assigned readings, and generates practice questions. Every answer cites the source, distinguishing it from general-purpose chatbots.

The evidence. Bloom's research on personal tutoring. The 2025 Harvard RCT showing AI tutoring produced twice the learning rate of active-learning classrooms. The measured effects of adaptive learning systems across multiple studies.

The pilot. One course, 100-200 students, one semester. Measuring grades, retention, satisfaction, and engagement. Cost: around $10 per student per month for the duration of the pilot.

The ask. Approval to run the pilot and a commitment to evaluate the results fairly. If the data shows improvement, a follow-up discussion about broader deployment.


Frequently asked questions

How much does a pilot cost? At around $10 per student per month for 200 students over a nine-month academic year, the total pilot cost is around $18,000. For most institutions, this falls within departmental discretionary spending and doesn't require a major budget approval.

What if the pilot shows no effect? That's a valid and useful outcome. It might mean the tool wasn't used enough (adoption issue), the course wasn't one where the tool adds value (scope issue), or the comparison wasn't well-matched (methodology issue). Understanding why a pilot didn't show results is nearly as valuable as understanding why it did, and the cost of finding out is minimal.

Can we run pilots in multiple departments simultaneously? Yes, and this is often more informative than a single-department pilot, because it reveals which kinds of courses and which student populations benefit most. Coordinate the metrics and timeline across departments so the results are comparable.

Who should own the pilot internally? Ideally, a faculty member who teaches the course and a member of the student success or academic support team. The faculty member provides academic credibility and can speak to the tool's appropriateness. The student success team can connect the pilot to institutional retention data and the broader support infrastructure.


Related reading: Why students drop out, The cost of student dropout, AI tutoring vs ChatGPT, Bloom's two sigma problem. Related pages: AI tutor, For students, AI tutoring use case.

The workspace that thinks with you.

Ready when you are.

The workspace that thinks with you.

Ready when you are.

The workspace that thinks with you.

Ready when you are.