Education’s dark matter
There is no evidence to inform most of what we do
For the reasonable price of just $5 AUD per month or $50 per year, you can be a full subscriber to Filling the Pail with complete access to my extensive archive on a range of topics, as well as access to the unabridged version of my weekly Curios posts.
I have often wondered whether the outlook of Richard Dawkins would be different if he had studied physics rather than biology. In biology, there is a unifying theory and only small gaps in our understanding into which Dawkins has famously argued that the need for a God retreats ever further. In contrast, physics is a looking-glass world with huge chasms rather than gaps. We only know what about 5% of the universe is made from, for instance, with the rest being mysterious forms of ‘dark’ energy and matter.
In this sense, education is more like physics than biology.
I often receive emails from teachers who are being subjected to some new and crazy scheme, asking if I know of any evidence supporting or refuting it. Most of the time, there is no direct evidence. It won’t have been researched in any meaningful way. The best we can do is look for something a little similar, opening up scope for different interpretations and arguments.
Most of what we do every day in schools is not ‘evidence-based’ and probably never can be. Instead, it is based on craft knowledge. This is why the subtitle of my first book used ‘evidence-informed’ instead and why this is a key theme of my upcoming book, How to Improve a School.
The difference between my position and that of education academics who take a relativist stance is that where evidence does exist, I believe we have a responsibility to be informed by it. This is not ‘positivism’, a boo-word misused by some to dismiss all evidence. Positivism insists that only empirical evidence matters. It is impossible to run a school as a genuine positivist. The leaders would be valueless and indecisive.
Still, education has a history of grand empirical claims. Advocates for Cognitive Acceleration and for Philosophy for Children, for example, have claimed ‘far transfer’ effects. This means they have claimed that training in one subject area can improve outcomes in a different one. Early research suggested Cognitive Acceleration improved performance in a range of GCSE subjects and Philosophy for Children has supposedly shown an impact on reading and mathematics. In both cases, the research basis for these claims has been criticised—including by me here and here—and when the UK’s Education Endowment Foundation subsequently ran more rigorous trials, the claims collapsed (see here and here).
That is why randomised controlled trials are important. No, they certainly cannot tell us everything, but they do tell us something.
What other evidence may teachers draw upon? The field is limited. Direct tests of educational programs are expensive and hard to do. We can draw further inferences from indirect evidence, but this opens up scope for further debate.
I am taken with the process-product research tradition that peaked in the 1960s. This is also sometimes known as ‘teacher effectiveness’ research and involved observing teachers and looking for correlations between teacher behaviours and learning gains. It is, by nature, correlational, but it is nonetheless suggestive and some attempts to train teachers in the behaviours that were associated with larger gains seem to have been successful. Process-product research is what Brophy and Good summarise in an excellent 1980s monograph and forms the spine of Rosenshine’s Principles.
The other place I would look is to lab-based studies of the kind I completed for my PhD. The term ‘lab-based’ is ambiguous. Sometimes these studies are indeed conducted in education faculties in actual laboratory settings, but more often, like my research, they take place in schools and universities and they mostly target educationally relevant objectives. The ‘lab’ nature comes from the creation of artificial conditions. These are not typical learning programs and the advantage of this artificiality is that it lets researchers properly control all the variables. The disadvantage is that people can then claim any findings may not translate to real classrooms. I am sceptical of this claim because I don’t see why subjects’ minds would change that much between environments.
Lab-based studies have included research into spaced practice, retrieval practice and other ‘desirable difficulties’. They have formed much of the evidence base for cognitive load theory.
I believe this evidence to be robust, but it does not tell me how to intervene pastorally with a student. We can draw tentative inferences from, say, the effectiveness of cognitive behavioural therapy and so align a school’s approach around Stoic principles, but we can never really know whether that is valid. Ultimately, it is more about our values. No study will ever definitely tell us whether school uniform is somehow effective—at what? It may be able to tell us one small thing about presenting information to students, but that will inform one decision out of the many a teacher makes when planning a lesson.
To me, this is more an engineering question than a scientific one. At Clarendon, we iterate learning materials, use engineering-type processes to try to identify what works best and then repeat the cycle.
Does it work? How do we know if a school is effective?
In Australia, we can look at a school’s NAPLAN results. That allows us to compare students’ performance in one school with ‘similar’ student elsewhere, but it is hardly definitive. We can also look at VCE results—the examinations taken by students at 18.
The problem with VCE results is that the figure published in the papers is flawed. It reports the median study score. But when used for university entry, study scores are further adjusted to reflect the cohort that takes the subject. To be brief, a school could game this system by entering more students for subjects with study scores that are adjusted downwards. We don’t do that at Clarendon. We have large and expanding numbers of students doing challenging subjects such as mathematical methods, chemistry, physics and so on, but you would have to take my word for it. There is no independent way of checking and so the VCE ranking that is published every December does not tell us much and is more an index of socioeconomic status than anything else.
In England, there is a superior evaluation system based on the value a school adds to student progress. Students are tested before entering secondary school and then again at 16. This is why I am fairly certain schools such as Michaela Community School in London are doing something right. But when I visited, I observed the interaction of many different choices and practices. It would be impossible to disentangle whether some, many or all of these are needed to have the same effect. So, I still don’t really know.
In one important sense, therefore, education is like physics. It has chasms rather than gaps. Unlike physics, people point to these chasms in education to argue away the strong evidence we do have. That is silly and just one of the many reasons why education research is held in such low esteem.




I agree with everything in this article except perhaps the choice of analogy. One way physics is very disanalogous to education academia is the rigor and overwhelming agreement about the fundamental principles.
If you try to be a serious physicist with your own idiosyncratic, vibes-based definition of relativity or quantum mechanics that departs even slightly from the very solid and rigorous existing understanding of these concepts you are dismissed as a crank. Not so in education. You can get very far by insisting that whole language teaches phonics correctly, or that inquiry learning involves explicit teaching, or that desirable difficulty means not telling novices the answers to questions, or [choose your favourite example here]
Scepticism to new popular ideas would do a lot to reduce harm.
A proposed change to approach should have evidence to support it.
Much as Russel’s tea pot (proposed to be in orbit without evidence) can be safely ignored, a proposed change should need evidence to back its effectiveness before being imposed on teachers or students.
In medicine, prescribing a drug that lacks clinical trial evidence and causes patient illness is iatrogenesis. What’s the educational equivalent ?