CCT Confusion-Calibrated Tutoring
NIT Hackathon · 2027 H26EDU05
Team H26EDU05 · PSG College of Technology

Confusion-Calibrated
Tutoring

Team ID H26EDU05
Project Title Confusion-Calibrated Tutoring
College PSG College of Technology
P
Paramaguru HTeam Lead
M
Mohamed AklamaashTeam Member
Measure Decide Train CCT
Problem Statement

Tutors optimise in-session performance, not durable learning

LLM tutors resolve student confusion immediately — they explain and reveal answers, which maximises how well the session feels and how quickly the current problem is solved.

Cognitive science shows durable learning requires a controlled degree of struggle — desirable difficulty and productive failure. By removing all friction, current tutors produce strong in-session performance but weak retention and transfer — the learning–performance paradox.

No deployed tutor measures a student's confusion or decides how much struggle to preserve. Confusion is unobserved, and teaching strategy is therefore heuristic.

Objectives

  • 1Measure per-turn confusion as a calibrated, uncertainty-aware signal grounded in cognitive science.
  • 2Ask questions with an adaptive question system — difficulty and the next graph-linked question chosen from the learner's confusion and mastery.
  • 3Hold confusion inside a productive-struggle band so the student truly learns the concept, not just the answer.
  • 4Train the tutor policy on delayed transfer — retention measured after a time gap, not in-session comfort.
System Design & Algorithms

A control layer around the language model

Per-turn loop
Question Served from the adaptive map
→
1 Confusion Estimator Confusion (0–1) and uncertainty per turn
→
2 Budget Controller Band status: below / inside / above
→
3 Tutor Policy Strategy: probe / verify / scaffold
→
Response Language model renders the reply
Student replies — the turn repeats until the question is solved
→ Question solved — the next question is served from the adaptive map, and the loop begins again.
A1Confusion Estimation

A multi-head DistilBERT classifier over the student turn: one shared encoder with two heads producing a confusion score and an uncertainty score. Scores are calibrated so confusion is a usable probability.

DistilBERT2 headscalibration
A2Budget Control

Deterministic control against a per-skill band. An uncertainty gate suppresses low-confidence signals, and hysteresis prevents band-status jitter across adjacent turns; the band adapts slowly per student.

band [b, b̄]hysteresisuncertainty gate
A3Policy Optimisation

The PPO tutor policy is pre-trained on the simulated dataset to a usable, sub-optimal start, then keeps learning online from each user's conversations — adapting continuously to become more personalized.

PPOsimulated → personalizedonline update
A4Adaptive Learning

Beyond the initial dataset, the system adapts to each learner from real input — the band and question difficulty shift with their confusion and recall as the policy keeps updating on live conversations.

per-learner banddifficulty adaptsreal data
Research Impact

A novel, niche solution — trained, not prompted

MetricUntrainedTrainedDirection
Delayed Transfer Accuracy 0.573 0.745 ↑ good
Turns-in-band fraction 0.574 0.727 ↑ good
Premature-answer rate (sampled) 0.100 0.052 ↓ good
Niche & novelty

A gap no one else fills

The niche: tutors that deliberately preserve confusion to build retention, instead of removing it for a smoother session.

The novelty: CCT is the first to compose all four — a trained tutor policy, a calibrated per-turn confusion signal, an explicit productive-struggle band, and a delayed-transfer objective — the paradox PEARL and TASA name but never train on.

The reward is our novelty — the RL is grounded only on these three terms
R= α·DTA − β·hinge − γ·scaffold
+ α · DTADelayed transfer

Scored from the graph: when a related question is later answered correctly, the concept counts as retained

− β · hingeBand violation

Both too-easy and too-confusing turns cost reward, holding the learner inside the productive-struggle band.

− γ · scaffoldPremature help

Revealing the answer before the student has struggled erases the very learning we want to keep.

Technology Stack

Implementation and tooling

Backend & Serving

  • Python3.11
  • FastAPIREST API
  • UvicornASGI server
  • python-multipartuploads

Machine Learning & RL

  • PyTorchmodels + PPO
  • TransformersDistilBERT
  • scikit-learn / joblibcalibration
  • PPO + GAEin-house

Data, Config & Frontend

  • PyYAMLconfiguration
  • LangChain splitters, pypdfingestion
  • Next.js, React, TypeScriptproduct UI
  • NumPyRL policy
Results & Demonstration

Working end-to-end pipeline

demo · session walkthrough
Demo recording
  • 4Components integrated — estimator, controller, policy, and adaptive question generation run end-to-end via the session runner.
  • <1 msController latency — deterministic band decision per turn, no GPU required.
  • TrainedConfusion estimator — DistilBERT artifact versioned and served offline.
  • PPOPolicy training — raises delayed transfer over the untrained policy on the simulated dataset (0.573 → 0.745).
  • AppWeb application — a Next.js frontend with a Python API backend serving live tutoring sessions.
Comparison with Existing Solutions

Where CCT differs

Approach Trained tutor policy Measured confusion Productive-struggle band Delayed-transfer objective
Prompt-only tutors (Khanmigo, ChatGPT) ✗✗✗✗
Socratic-form RL (PEARL, 2026) ✓✗✗✗
Single-conversation RL (Nam, 2025) ✓✗✗✗
Forgetting-conditioned (TASA, 2025) ✗✗✗✓
CCT (this work) ✓✓✓✓
Novelty. CCT is the first system to combine a trained tutor policy, a calibrated per-turn confusion signal, an explicit productive-struggle band, and a delayed-transfer training objective. It operationalises the learning–performance paradox that prior work names but does not address.
Business Impact

We optimise for capability, not task completion

AI made intelligence abundant — but if an engineer still asks the same questions six months later, organisations aren't building capability, they're renting it. CCT optimises for capability growth, not task completion.

Enterprises

Faster onboarding and more independent task completion — employees who grow monthly, not more dependent on AI.

Universities

Shift from assignment completion to measurable understanding and long-term retention.

Bootcamps & training

Success measured by interview performance and job readiness, not course completion.

Our moat

Most AI companies collect conversation data. We collect learning trajectories — time-to-help, hints that unlock understanding, recurring misconceptions, transfer, retention.

North Star

Independence Score — how well a learner solves problems without help over time. As it rises, AI dependency falls.

Zero-token control

Confusion estimate, band decision, and strategy run offline with zero LLM tokens — at enterprise scale, millions saved vs prompt-engineered tutors.

Conclusion & Next Steps

A tutor that learns when to help

Summary

  • CCT manages confusion within an adaptive productive-struggle band instead of removing all difficulty.
  • Source documents become a graph-based conceptual map of connected questions with chunk-level provenance.
  • The real-time platform integrates confusion estimation, adaptive control, PPO strategy selection, and tutor response in one learning loop.
  • The policy starts from synthetic training and is wired to continue adapting from completed learner conversations.

Next Steps

  • Research publication. Submit the end-to-end system and human-centered evaluation to ACM CHI 2027.
  • Native application. Extend the web platform into a dedicated mobile learning experience.