
.jpg)
As part of the committee organising the Festival of Learning 2026 in Seoul, we came away feeling both energised and humbled. The conference spanned seven days and brought together around 1,200 people from the AIED, EDM, and Learning at Scale communities under one roof, all focused on a shared question: how do we improve education through technology?
AI in education has become a very popular but also a sometimes polarising topic in mainstream and social media. This event was a refreshing break from the constant AI newscycle and recognition that there are no silver bullets or easy answers. The Festival of Learning is a coming together of careful thinkers trying to measure what actually works. Every paper represented months, or years, of serious work, bravely presented to peers who are also experts in the field of education technology to provide praise and (hopefully constructive) criticism.
And indeed, the reactions were often humbling. Even predicting what a student will do next is far from solved: one benchmark had the strongest foundation models (like GPT-OSS-120B, Qwen3-Next-80B-Thinking) scoring only around 56% at forecasting student responses. It was a reminder of how early the field still is, and how much progress will depend on many disciplines moving together.
Across the week a handful of threads kept popping up which are worth naming, because they are the questions the whole field is concentrating on.
When 1,200 people are all trying to solve the same problem, it forces you to ask where your own contribution sits.
We have attended the Festival of Learning and its constituent AIED, EDM and L@S conferences for several years. But this was the first year we arrived not only as participants, but as contributors with our own research to present.
Our role is not simply to go deep on one narrow research problem. Our edge is being the end-to-end connector, working across the whole chain rather than owning any single link in it.
A lot of academic work goes deep on one vertical. That depth is valuable. But schools and students do not experience isolated research questions. They experience complete systems. At Eedi, we care about how data analysis, model optimisation, deployment, benchmarking, evaluation and classroom use all fit together. That systems view feels increasingly important.
This was also the strongest Eedi presence we have had at the conference so far. We contributed our own papers across topics including knowledge tracing (where we showed that smaller, specialised models beat general-purpose LLMs at predicting student responses, and do so faster, cheaper and more accurately) and misconception diagnosis from tutoring dialogues. We also contributed through direct collaborations, and it was exciting to see wider use of Eedi datasets in other research.
We also co-organised two workshops: the third Impact and Responsibility in AI Supported Education (RAISE) workshop and the inaugural Pedagogical Evaluation of Automated Feedback (PEAF) workshop.
That breadth matters. We have grown from being known mainly for sharing datasets and collaborating with others, to also conducting publication-worthy research of our own and taking that work toward rigorous real-world evaluation.
Some of the papers that stayed with us most were the ones that felt immediately useful, so we have each picked out a couple.
Simon’s picks.
One was the keystroke dynamics paper To What Extent Do Keystroke Dynamics Differentiate Correct from Incorrect Problem Solving?, by Oscar Deho, who completed it during an internship with Eedi. Oscar ran a counterfactual analysis suggesting that around 30% of the time a student would have answered correctly if they had taken a little longer. Whilst Oscar was hesitant about reading too much into the results, it does at least suggest that it would be worth providing a little behavioural nudge when we detect a student is rushing more than usual.

The other was Capture-Calibrate-Coach: A Graph-Based Framework for Knowledge Monitoring Estimation and Adaptive Feedback presented by Gen Li, which offered compelling evidence for a deceptively simple idea: helping students judge how well they really understand something, then coaching them based not only on what they know but their perception of what they know. The study was in higher education, but it raises exciting questions about how well the approach might transfer into K–12 settings.
Prarthana’s picks.
Prarthana was drawn to the work on measurement and modality. Modality Matters (Carnegie Learning) looked at how text, audio and video shape engagement and performance with AI tutors.
She was also struck by FoundationalASSIST, the knowledge tracing benchmark behind the 56% figure above, where models predicted correct answers far better than incorrect ones. In other words, they were weakest at spotting the very students who are struggling.
A theoretical paper which uses circuit complexity to quantify how hard knowledge tracing is as a task was striking in the insights it surfaced for practitioners. The authors suggest knowledge-tracing systems should supervise intermediate concepts and add depth through iterative, recurrent, or hybrid structured models.
The investigation from Khan Academy on improving AI tutor quality showed the value of hill climbing: they conducted 50 live experiments over eight months, and improved next-item correctness by 15% and cognitive engagement by 14%.
Across these papers a common thread emerged: progress depends not just on smarter models, but on better ways of measuring learning, understanding behaviour and grounding interventions in evidence.
A few directions feel especially important for us coming out of the conference.
First, knowledge tracing. There is a sense that research here has stagnated, and some pessimism about how much further it can go. We are more optimistic, we believe that with the right data and inductive biases there is real opportunity.
Second, turning research into live studies. The hard part is doing something with the research, collaborating with learning designers and product developers to run real classroom studies that demonstrate genuine learning gains.
Third, multimodality. We expect audio and other modalities to become much more important, especially as AI tutors need to work for younger learners and students with lower literacy.
And fourth, model sovereignty and control. Questions of trust, privacy, bias, performance drift and dependence on external models are becoming central. If we care about durable educational systems, we also need to care about how much control we really have over the models inside them.
The big takeaway for us is less about any single area of research.
We want to connect research to practice: through dataset sharing, collaboration, product development, trials and evidence. We want to help ensure that educators and learners are exposed not just to impressive demos, but to solutions that are rigorously built, grounded in research, carefully evaluated and genuinely useful in classrooms.
This is exactly what our diagnostic engine does: it bundles everything we've learnt from our research and data into a simple service other edtech providers can use to keep their own solutions cutting-edge.
That’s Eedi’s role: weaving together research, product and the classroom into a single loop. Learning, experimenting and evidencing. Looking forward to sharing how we are getting on at the Festival of Learning 2027! 🦘