Confusion-Calibrated
Tutoring
Tutors optimise in-session performance, not durable learning
LLM tutors resolve student confusion immediately — they explain and reveal answers, which maximises how well the session feels and how quickly the current problem is solved.
Cognitive science shows durable learning requires a controlled degree of struggle — desirable difficulty and productive failure. By removing all friction, current tutors produce strong in-session performance but weak retention and transfer — the learning–performance paradox.
No deployed tutor measures a student's confusion or decides how much struggle to preserve. Confusion is unobserved, and teaching strategy is therefore heuristic.
Objectives
- 1Measure per-turn confusion as a calibrated, uncertainty-aware signal grounded in cognitive science.
- 2Ask questions with an adaptive question system — difficulty and the next graph-linked question chosen from the learner's confusion and mastery.
- 3Hold confusion inside a productive-struggle band so the student truly learns the concept, not just the answer.
- 4Train the tutor policy on delayed transfer — retention measured after a time gap, not in-session comfort.
A control layer around the language model
A multi-head DistilBERT classifier over the student turn: one shared encoder with two heads producing a confusion score and an uncertainty score. Scores are calibrated so confusion is a usable probability.
Deterministic control against a per-skill band. An uncertainty gate suppresses low-confidence signals, and hysteresis prevents band-status jitter across adjacent turns; the band adapts slowly per student.
The PPO tutor policy is pre-trained on the simulated dataset to a usable, sub-optimal start, then keeps learning online from each user's conversations — adapting continuously to become more personalized.
Beyond the initial dataset, the system adapts to each learner from real input — the band and question difficulty shift with their confusion and recall as the policy keeps updating on live conversations.
A novel, niche solution — trained, not prompted
A gap no one else fills
The niche: tutors that deliberately preserve confusion to build retention, instead of removing it for a smoother session.
The novelty: CCT is the first to compose all four — a trained tutor policy, a calibrated per-turn confusion signal, an explicit productive-struggle band, and a delayed-transfer objective — the paradox PEARL and TASA name but never train on.
Scored from the graph: when a related question is later answered correctly, the concept counts as retained
Both too-easy and too-confusing turns cost reward, holding the learner inside the productive-struggle band.
Revealing the answer before the student has struggled erases the very learning we want to keep.
Implementation and tooling
Backend & Serving
- Python3.11
- FastAPIREST API
- UvicornASGI server
- python-multipartuploads
Machine Learning & RL
- PyTorchmodels + PPO
- TransformersDistilBERT
- scikit-learn / joblibcalibration
- PPO + GAEin-house
Data, Config & Frontend
- PyYAMLconfiguration
- LangChain splitters, pypdfingestion
- Next.js, React, TypeScriptproduct UI
- NumPyRL policy
Working end-to-end pipeline
- 4Components integrated — estimator, controller, policy, and adaptive question generation run end-to-end via the session runner.
- <1 msController latency — deterministic band decision per turn, no GPU required.
- TrainedConfusion estimator — DistilBERT artifact versioned and served offline.
- PPOPolicy training — raises delayed transfer over the untrained policy on the simulated dataset (0.573 → 0.745).
- AppWeb application — a Next.js frontend with a Python API backend serving live tutoring sessions.
Where CCT differs
We optimise for capability, not task completion
AI made intelligence abundant — but if an engineer still asks the same questions six months later, organisations aren't building capability, they're renting it. CCT optimises for capability growth, not task completion.
Faster onboarding and more independent task completion — employees who grow monthly, not more dependent on AI.
Shift from assignment completion to measurable understanding and long-term retention.
Success measured by interview performance and job readiness, not course completion.
Most AI companies collect conversation data. We collect learning trajectories — time-to-help, hints that unlock understanding, recurring misconceptions, transfer, retention.
Independence Score — how well a learner solves problems without help over time. As it rises, AI dependency falls.
Confusion estimate, band decision, and strategy run offline with zero LLM tokens — at enterprise scale, millions saved vs prompt-engineered tutors.
A tutor that learns when to help
Summary
- CCT manages confusion within an adaptive productive-struggle band instead of removing all difficulty.
- Source documents become a graph-based conceptual map of connected questions with chunk-level provenance.
- The real-time platform integrates confusion estimation, adaptive control, PPO strategy selection, and tutor response in one learning loop.
- The policy starts from synthetic training and is wired to continue adapting from completed learner conversations.
Next Steps
- Research publication. Submit the end-to-end system and human-centered evaluation to ACM CHI 2027.
- Native application. Extend the web platform into a dedicated mobile learning experience.