Skip to product information
1 of 1

Causal Inference with Differences-in-Differences

Regular price $49.95
Sale price $49.95 Regular price $49.95
Sale Sold out
A comprehensive, rigorous introduction to modern differences-in-differences (DiD) estimators, covering both standard practices and alternativesDifferences-in-differences (DiD) is one of the most wi...
Read More
  • Format:
  • Publication Date: 08 December 2026
  • ISBN: 9780691264189
  • Pages: 368
  • Imprint: Princeton University Press

View Product Details

A comprehensive, rigorous introduction to modern differences-in-differences (DiD) estimators, covering both standard practices and alternatives

Differences-in-differences (DiD) is one of the most widely used methods for impact-evaluation in economics and the social sciences. The key idea behind DiD is to compare outcomes trends for treated and control groups, allowing researchers to estimate the effects of policies or interventions when randomized experiments are not feasible. This book provides a clear and rigorous guide to modern DiD methods, covering both classical approaches and newer estimators developed for complex real-world settings. Designed for advanced undergraduate students, graduate students, and applied researchers, it explains when standard methods are reliable, when they can mislead, and how alternative approaches can provide more credible results. Throughout, theoretical discussion is paired with empirical applications, exercises using real datasets, and practical recommendations for implementation.

The book offers:
• Discussion of all designs in which DiD apply: classical designs, staggered adoption designs, designs with variation in treatment dose, and staggered first switch designs
• Study of standard estimators, estimators without parallel trends (e.g., synthetic controls), and heterogeneity-robust estimators
• 5 lists of dos and don’ts for practitioners
• 150 exercises to understand the theory, and 50 practical exercises to apply it to real empirical examples in Stata and R

files/i.png Icon
Price: $49.95
Pages: 368
Publisher: Princeton University Press
Imprint: Princeton University Press
Publication Date: 08 December 2026
ISBN: 9780691264189
Format: Paperback
Clement de Chaisemartin is professor of economics at Sciences Po, Paris. Xavier D’Haultfoeuille is Professor of Economics at CREST-ENSAE.
  • Preface
  • I Introduction and Setup
    • 1 Introduction
      • 1.1 The classical DID design: Chapters 3 and 4
        • 1.1.1 Definition of a classical DID design
        • 1.1.2 Potential outcomes and target parameter
        • 1.1.3 Three possible estimators: treated versus control, before–after, and DID
        • 1.1.4 The parallel-trends assumption
        • 1.1.5 Parallel-trends is weaker than the assumptions underlying treated-versus-control and before–after comparisons
        • 1.1.6 The parallel-trends assumption remains a strong assumption, whose plausibility should be assessed
        • 1.1.7 Relaxations of the parallel-trends assumption
        • 1.1.8 Application to the effect of the 2009 VAT reduction in French restaurants on profits, prices, and wages
      • 1.2 Beyond the classical DID design: Chapters 5 to 8
        • 1.2.1 Two-way fixed effects regressions
        • 1.2.2 What does a TWFER estimate outside of the classical DID design?
        • 1.2.3 Heterogeneity-robust DID estimators
          • 1.2.3.1 Staggered adoption designs
          • 1.2.3.2 Heterogeneous adoption designs
          • 1.2.3.3 Staggered first switch designs
    • 2 Data, Notation, and Assumptions
      • 2.1 Data requirements: group-level panel data
      • 2.2 Treatment and potential outcomes
      • 2.3 Identifying assumptions
        • 2.3.1 Exclusion restrictions
        • 2.3.2 Parallel trends
      • 2.4 The book’s perspective on statistical inference
        • 2.4.1 A “model-based” perspective on statistical inference
        • 2.4.2 Assuming independent outcomes between groups does not rule out common shocks*
        • 2.4.3 At what level should we cluster standard errors?*
  • II The Classical Design
    • 3 The Classical DID Design
      • 3.1 Target parameters
      • 3.2 Two-way fixed effects regressions
        • 3.2.1 Static two-way fixed effects regressions
          • 3.2.1.1 Assuming randomized treatment rather than parallel-trends?
        • 3.2.2 Event-study two-way fixed effects regressions
          • 3.2.2.1 Estimating event-study effects
          • 3.2.2.2 Pretrend tests
          • 3.2.2.3 Application to the effect of compulsory licensing on innovation
      • 3.3 Inference
        • 3.3.1 Asymptotically valid CIs with many treated and control groups
        • 3.3.2 When are G0 and G1 large enough to rely on CIs with HC2 standard errors and Bell and McCaffrey’s critical values? Some simulations
        • 3.3.3 What can researchers do when HC2-BM CIs seem unreliable?*
          • 3.3.3.1 Few treated groups but many control groups
          • 3.3.3.2 Few treated and control groups
          • 3.3.3.3 Recommendations for practitioners
      • 3.4 Limitations of pretrend tests
        • 3.4.1 We can test for parallel-trends before but not after treatment
          • 3.4.1.1 Concomitant shocks
          • 3.4.1.2 Other policies
        • 3.4.2 Pretrend tests often lack power
        • 3.4.3 Do pretrend tests lead to a pretesting problem?*
      • 3.5 An alternative, imputation estimator
        • 3.5.1 Imputation estimator
        • 3.5.2 Three numerical equivalences
        • 3.5.3 Comparing the properties of β̂fe and β̂imp
        • 3.5.4 Application to the effect of compulsory licensing on innovation
      • 3.6 Estimating heterogeneous effects
        • 3.6.1 Estimating the correlation between treatment effects and some covariates
        • 3.6.2 Estimating the variance of group-specific effects*
        • 3.6.3 Estimating the distribution of group-specific effects*
      • 3.7 Nonlinear DID
        • 3.7.1 Limited dependent variables
        • 3.7.2 Sensitivity to functional form*
          • 3.7.2.1 The parallel-trends assumption may not be invariant to functional form
          • 3.7.2.2 A necessary and sufficient condition to have parallel-trends for any functional form
          • 3.7.2.3 When should researchers worry about sensitivity to functional form?
        • 3.7.3 Estimating quantile treatment effects*
      • 3.8 Instrumental-variable DID estimators
      • 3.9 Further topics
        • 3.9.1 Imbalanced panels
        • 3.9.2 Weighting
        • 3.9.3 Accounting for and estimating spillover effects
      • 3.10 Practitioners’ do’s and don’ts
      • 3.11 Appendix*
        • 3.11.1 The Frisch–Waugh–Lovell theorem
        • 3.11.2 Proof of (3.14)
    • 4 Alternatives to parallel-trends
      • 4.1 TWFE and DID estimators with control variables
        • 4.1.1 Conditional parallel-trends
        • 4.1.2 TWFERs with control variables
        • 4.1.3 DID estimators with control variables
          • 4.1.3.1 Parametric estimators
          • 4.1.3.2 Nonparametric estimators
          • 4.1.3.3 Matched DIDs
          • 4.1.3.4 Controlling for group-specific linear trends
          • 4.1.3.5 Triple-difference estimators*
          • 4.1.3.6 When to include control variables in the estimation?
        • 4.1.4 Controlling for the lagged outcome?
          • 4.1.4.1 AR(1) model for the outcome without treatment*
          • 4.1.4.2 Self-selection, and parallel-trends conditional on the baseline outcome
        • 4.1.5 Computing DID estimators with controls in Stata and R
        • 4.1.6 Application to the effect of compulsory licensing on innovation
      • 4.2 Interactive fixed effects, synthetic controls, and synthetic DID
        • 4.2.1 Interactive fixed effects
        • 4.2.2 Synthetic control and synthetic DID
          • 4.2.2.1 Synthetic control
          • 4.2.2.2 Synthetic DID
        • 4.2.3 Application to the effect of compulsory licensing on innovation
      • 4.3 Bounded differential trends
      • 4.4 Practitioners’ do’s and don’ts
      • 4.5 Appendix*
        • 4.5.1 Proof of Theorem 4.1
        • 4.5.2 Comparison of variances of DID estimators with and without covariates
        • 4.5.3 Details on the computation of the TWFE-IFE estimator
  • III Beyond the Classical Design
    • 5 The pifalls of TWFE estimators
      • 5.1 A decomposition of β̂fe
      • 5.2 β̂fe may be biased for ATT
      • 5.3 β̂fe may not estimate a convex combination of effects
      • 5.4 Decompositions of related estimators
      • 5.5 Stata and R commands to compute the implicit weights of TWFERs and first-difference regressions
      • 5.6 Application to the effect of newspapers on turnout in US elections
        • 5.6.1 Basic TWFER
        • 5.6.2 TWFER with state–year FEs
        • 5.6.3 First-difference regression with state–year FEs
      • 5.7 Next steps
    • 6 Staggered Adoption Designs
      • 6.1 Target parameters
      • 6.2 Two-way fixed effects estimators
        • 6.2.1 Static two-way fixed effects estimator
          • 6.2.1.1 Decomposition of β̂fe
          • 6.2.1.2 Application to the effect of unilateral divorce laws on divorces
          • 6.2.1.3 The origin of the negative weights
          • 6.2.1.4 Assuming randomized treatment timing instead of parallel trends?*
        • 6.2.2 Event-study TWFE regressions
          • 6.2.2.1 Decomposition of ES-TWFERs
          • 6.2.2.2 Application to the effect of unilateral divorce laws on divorces
        • 6.2.3 Local projection regressions*
      • 6.3 Heterogeneity-robust estimators
        • 6.3.1 Target parameters
        • 6.3.2 DID estimators
          • 6.3.2.1 Estimators
          • 6.3.2.2 Extensions
          • 6.3.2.3 Inference
          • 6.3.2.4 Numerical equivalences with regression coefficients
          • 6.3.2.5 Computation in Stata, R, and Python
        • 6.3.3 Imputation estimators
          • 6.3.3.1 Some numerical equivalences
          • 6.3.3.2 Extensions
          • 6.3.3.3 Inference
          • 6.3.3.4 Computation in Stata and R
        • 6.3.4 A comparison of heterogeneity-robust estimators
          • 6.3.4.1 Variance
          • 6.3.4.2 Confidence intervals coverage
          • 6.3.4.3 Bias
          • 6.3.4.4 Implementation in Stata, R, and Python
        • 6.3.5 Application to the effect of unilateral divorce laws on divorces
      • 6.4 TWFE and HR estimators in 13 SADs in political science
      • 6.5 Estimating heterogeneous treatment effects
      • 6.6 Nonlinear DID
        • 6.6.1 Limited dependent variables
          • 6.6.1.1 Application to the effect of free-trade agreements on trade
        • 6.6.2 Estimating quantile treatment effects*
      • 6.7 Further topics*
        • 6.7.1 Using always-treated groups as controls?
        • 6.7.2 Imbalanced panels
        • 6.7.3 Weighting
        • 6.7.4 Accounting for and estimating spillover effects
      • 6.8 Practitioners’ do’s and don’ts
      • 6.9 Appendix*
    • 7 Heterogeneous Adoption Designs
      • 7.1 Target parameters
        • 7.1.1 The conditional-average-slope function
        • 7.1.2 Two unconditional averages of slopes
      • 7.2 TWFERs in heterogeneous-adoption designs
        • 7.2.1 Parallel-trends assumption
        • 7.2.2 β̂fe may not identify a convex combination of slopes
        • 7.2.3 The origin of the negative weights
        • 7.2.4 Assuming randomized treatment dose rather than parallel trends*
      • 7.3 Heterogeneity-robust estimators
        • 7.3.1 Parallel-trends assumptions and a fundamental decomposition
        • 7.3.2 Designs with stayers or quasi-stayers
          • 7.3.2.1 Identification
          • 7.3.2.2 Estimation with stayers
          • 7.3.2.3 Estimation without stayers but with quasi-stayers
        • 7.3.3 Designs without stayers or quasi-stayers
        • 7.3.4 Testing the null that there are quasi-stayers
        • 7.3.5 Application to the effect, on US employment, of eliminating a potential tariffs’ spike on Chinese imports
      • 7.4 Practitioners’ do’s and don’ts
      • 7.5 Appendix*
        • 7.5.1 Proof of Theorem 7.1
        • 7.5.2 Proof of Theorem 7.2
        • 7.5.3 Proof of Theorem 7.3
    • 8 General Designs
      • 8.1 Static TWFER
        • 8.1.1 Decomposition of TWFERs under the usual parallel-trends assumption
        • 8.1.2 Decomposition of TWFERs under a parallel-trends assumption in a counterfactual where groups’ treatment does not change
        • 8.1.3 Decomposition of TWFERs with randomly assigned treatments*
        • 8.1.4 Extensions
          • 8.1.4.1 TWFERs with several treatments
          • 8.1.4.2 Two-stage-least-squares TWFERs, and Bartik regressions*
      • 8.2 Distributed-lag TWFER
      • 8.3 HR estimators for staggered first switch designs
        • 8.3.1 Staggered first switch designs
        • 8.3.2 Parallel-trends assumption
        • 8.3.3 Target parameters and estimators
          • 8.3.3.1 Building blocks: the group-specific actual-versus-status-quo effects
          • 8.3.3.2 A generalization of the ATT event-study effects to staggered first switch designs
          • 8.3.3.3 Path-specific event-study effects
          • 8.3.3.4 Normalized event-study effects
          • 8.3.3.5 Distributed-lag regressions allowing for heterogeneous effects across groups
          • 8.3.3.6 A generalization of ATT to staggered first switch designs
          • 8.3.3.7 Pretrend estimators
          • 8.3.3.8 Testing if lagged treatments affect the outcome
        • 8.3.4 Inference
        • 8.3.5 Extensions
          • 8.3.5.1 Continuous treatment
          • 8.3.5.2 Estimators with control variables
          • 8.3.5.3 Estimating heterogeneous effects
          • 8.3.5.4 Estimators with several treatments*
          • 8.3.5.5 Instrumental-variable DID estimators*
          • 8.3.5.6 The initial-conditions problem in designs where the treatment varies at period one*
        • 8.3.6 Application to the effect of newspapers on turnout in US elections
      • 8.4 Heterogeneity-robust estimators, without dynamic effects
        • 8.4.1 Switchers and stayers designs
        • 8.4.2 Parallel-trends assumption
        • 8.4.3 Target parameters and estimators
          • 8.4.3.1 Target parameters
          • 8.4.3.2 Estimators
          • 8.4.3.3 Pretrend estimators
          • 8.4.3.4 Estimators robust to dynamic effects up to a prespecified number of lags
        • 8.4.4 Imputation estimators
        • 8.4.5 Extensions
          • 8.4.5.1 Continuous treatment
          • 8.4.5.2 Estimators with several treatments*
          • 8.4.5.3 Instrumental-variable DID estimators*
        • 8.4.6 Computation in Stata and R
        • 8.4.7 Application to the effect of newspapers on turnout in US elections
      • 8.5 Conclusion
        • 8.5.1 TWFE and HR estimators in general designs
        • 8.5.2 HR estimators in designs without stayers?
          • 8.5.2.1 Using quasi-stayers
          • 8.5.2.2 Changing the treatment’s definition
          • 8.5.2.3 Functional-form assumptions
      • 8.6 Practitioners’ do’s and don’ts
      • 8.7 Appendix*
        • 8.7.1 Proof of Theorem 8.1
        • 8.7.2 Proof of Theorem 8.2
        • 8.7.3 Proof of Theorem 8.3
        • 8.7.4 Proof of Theorem 8.4
        • 8.7.5 Proof of Theorem 8.5
        • 8.7.6 Proof of Theorem 8.6
  • Bibliography
  • Index