‘Let’s Talk Method’ interview with Xavier D’Haultfoeuille, lecturer at CREST and ENSAE
The difference-of-differences method is a modern estimation method that is currently gaining popularity. Do you recall, however, that it is not entirely new and that it was first used as far back as the 1840s – that is, before the invention of the turbine, the rubber tyre, anaesthesia or oil drilling?
Exactly! It is surprising to see such early use of this method, even though it remained very much a niche practice. The first applications of this method, by Ignaz Semmelweis and John Snow, were in the medical field, at a time when the scientific method had not yet been fully adopted in medicine: Claude Bernard’s seminal essay on the experimental method was not published until 1865. Ignaz Semmelweis’s conclusions were, in fact, ridiculed by his peers.
However, this early adoption can be understood in light of the highly intuitive nature of this method – arguably more so than the instrumental variables method, for example – and the fact that, in its simplest form at least, it requires neither computer technology nor complex statistical tools.
How can we explain its rise in the fields of economics and public policy evaluation? Is this, for example, at the expense of other methods that have fallen somewhat out of favour?
It seems to me that its rise is due to two factors. On the one hand, the institutional context is often conducive to this: public policies are very often targeted, which means that one can establish ‘treatment’ groups, which benefit from the policy from a certain date onwards, and ‘control’ groups, which do not. This framework therefore makes it possible, in principle at least, to apply the method. On the other hand, it is possible, to a certain extent, to test its underlying assumptions.
I shall now address your second question. The two main competing methods – namely the use of instrumental variables and discontinuity regression – have indeed suffered as a result of this recent surge. The main criticism levelled at instrumental variables is that it is difficult to find a credible instrument. And even when one is found, their use often leads to very imprecise estimates. The discontinuity regression method, on the other hand, holds up better against the recent wave of differences-in-differences analyses. It is often highly credible but suffers from a significant drawback compared with differences-in-differences: effects are estimated only for a very small sub-population, which is not always the population of interest.
What is commonly referred to as the ‘credibility revolution’ is now accompanied by a very high level of technical sophistication in the conduct of evaluations. Consequently, if the work is to be carried out properly, must evaluations not necessarily be conducted by researchers with in-depth expertise in both statistics and economics?
In economics, as in other scientific fields, research has become specialised. Evaluation is carried out by experts in their respective fields (labour, education, taxation, etc.), who very rarely develop the statistical tools they use. Instead, it is econometricians who develop and provide these tools. It is, however, important that evaluators have a good understanding of the tools they use, so that they can select the most appropriate ones and be aware of their limitations.
Could this complexity have a negative effect by making it more difficult for funders, policy-makers, journalists and members of the public to understand the methods … and therefore the results?
In fact, it seems to me that differences of differences, just like regression across a discontinuity, are very intuitive. The main results can be visualised directly. From this perspective, this method strikes me as more transparent and accessible than other traditional methods such as linear regression. Of course, interpreting the corresponding graphs can be tricky, but on the whole these methods seem to me to be a step towards making empirical research more accessible to a wider audience.
Do you understand the political criticism regarding the lack of robustness in certain studies or the difficulty in interpreting their results, even though these studies are based on the most advanced methods from an academic perspective?
I’m not sure I understand this criticism. Is it the idea that the effects of a policy can be multidimensional, and that it is difficult to aggregate these different dimensions to draw a single conclusion? If so, I would say that this argument reinforces, rather than undermines, the need for evaluation. For before deciding how these different dimensions should be weighted – which is partly the role of the politician – it is necessary to understand the effects themselves, in each of these dimensions.