인과추론 섹션의 참고 편입니다. 앞선 다섯 편 뒤에 있는 책·논문·소프트웨어·강의를 모았습니다.
This is the reference post for the Causal Inference section — the books, papers, software and courses behind the five posts that precede it.
읽기 목록은 순서가 붙어야 쓸모가 있습니다. 목록 자체보다 먼저, 어디서 출발하느냐에 따른 경로를 제안합니다.
A reading list is only useful with an order attached, so before the list itself, here is a suggested route depending on where you are starting from.
어디서 시작할 것인가Where to start
- 책 한 권만 읽는다면. Scott Cunningham의 Causal Inference: The Mixtape. 매칭·IV·RD·DiD·합성통제까지 도구 전체를 R과 Stata 코드와 함께 다루면서도 진지한 책 중 가장 접근하기 좋습니다. 온라인 무료입니다.
- 대학원 표준 참고서를 원한다면. Mostly Harmless Econometrics(Angrist·Pischke)가 이 분야의 공통 어휘로 남아 있습니다. 밀도가 높고 가끔 가볍지만 그만한 값어치가 있습니다. 같은 저자들의 Mastering Metrics가 더 부드러운 판본이고 수학이 낯설면 첫 책으로 더 낫습니다.
- 식이 아니라 그래프로 생각한다면. 직관은 Pearl의 The Book of Why, 엄밀함은 Hernán·Robins의 Causal Inference: What If(온라인 무료). DAG 전통과 잠재적 결과 전통은 다른 경로로 같은 결론에 도달합니다. 둘 다 아는 것이 실제로 쓸모 있습니다. 각각이 다른 것을 자명하게 만들기 때문입니다. 충돌부는 그래프에서 훨씬 명확하고 LATE는 잠재적 결과에서 훨씬 명확합니다.
- 금요일까지 돌아가야 한다면. 책은 건너뛰십시오.
DoWhy문서를 읽으십시오. 분석을 모형화·식별·추정·반박 구조로 짜 줍니다. 두 편 앞의 여덟 단계 틀과 깔끔하게 맞아떨어집니다. 그다음 추정량은EconML입니다.
- If you want one book. Causal Inference: The Mixtape by Scott Cunningham. It covers the whole toolkit — matching, IV, RD, DiD, synthetic control — with code in R and Stata, and it is the most approachable serious treatment available. Free online.
- If you want the standard graduate reference. Mostly Harmless Econometrics (Angrist and Pischke) remains the field shared vocabulary. Dense, occasionally glib, and worth the effort. Their Mastering Metrics is the gentler version and a better first book if the mathematics is unfamiliar.
- If you think in graphs rather than equations. Pearl The Book of Why for the intuition, then Hernán and Robins Causal Inference: What If (free online) for the rigour. The DAG tradition and the potential-outcomes tradition reach the same conclusions by different routes, and knowing both is genuinely useful because each makes different things obvious. Colliders are much clearer in graphs; LATE is much clearer in potential outcomes.
- If you need it working by Friday. Skip the books. Read the
DoWhydocumentation, which structures an analysis as model / identify / estimate / refute and maps cleanly onto the framework two posts back. ThenEconMLfor the estimators.
원문으로 읽을 값어치가 있는 논문The papers worth reading in the original
이 분야 대부분은 교과서로 흡수할 수 있지만 몇 편은 직접 읽을 값어치가 있습니다. 다른 모두가 인용하는 게 그 논문들의 틀 자체이기 때문입니다.
Most of the field can be absorbed from textbooks, but a handful of papers repay direct reading because their framing is what everyone else is citing.
| 논문 | 왜 읽는가 |
|---|---|
| Rosenbaum & Rubin (1983) | 성향점수. 고차원 매칭을 다룰 수 있게 만든 결과입니다. |
| LaLonde (1986), Dehejia & Wahba (1999) | 도전과 응답. 이 섹션의 한 편 전체가 이 주제입니다. |
| Imbens & Angrist (1994) | LATE. 효과가 이질적일 때 IV가 실제로 추정하는 것. |
| Bertrand, Duflo & Mullainathan (2004) | DiD 표준오차. 짧고 파괴적이며 여전히 일상적으로 무시됩니다. |
| Goodman-Bacon (2021) | 엇갈린 DiD. 상당량의 양방향 고정효과 추정치를 무효화한 논문입니다. |
| Chernozhukov et al. (2018) | 이중 머신러닝. 추론을 깨뜨리지 않고 유연한 ML을 쓰는 법. |
| Paper | Why |
|---|---|
| Rosenbaum and Rubin (1983) | The propensity score — the result that makes high-dimensional matching tractable. |
| LaLonde (1986), Dehejia and Wahba (1999) | The challenge and the response — the subject of an earlier post in this section. |
| Imbens and Angrist (1994) | LATE — what IV actually estimates under heterogeneous effects. |
| Bertrand, Duflo and Mullainathan (2004) | DiD standard errors. Short, devastating, and still routinely ignored. |
| Goodman-Bacon (2021) | Staggered DiD — the paper that invalidated a large body of two-way fixed effects estimates. |
| Chernozhukov et al. (2018) | Double machine learning — how to use flexible ML for nuisance estimation without breaking inference. |
마지막 둘은 충분히 최근이라 상당량의 출판된 연구가 그보다 앞섭니다. 오래된 응용 논문을 읽을 때 기억해 둘 값어치가 있습니다.
The last two are recent enough that a great deal of published work predates them, which is worth remembering when reading older applied papers.
소프트웨어와 각각의 용도Software, and what each is for
이 도구들은 서로 대체 가능하지 않습니다. 적합성이 아니라 익숙함으로 고르는 것이 흔한 낭비입니다.
The tools are not interchangeable, and choosing by familiarity rather than by fit is a common waste of effort.
DoWhy(Microsoft) — 틀 계층입니다. 가정을 명시적으로 부호화하고 반박 검정을 제공합니다. 분석을 구조화하는 출발점으로 가장 좋습니다.EconML(Microsoft) — 이질적 처치효과, 이중 ML, 인과 숲. 평균이 아니라 부분집단별 효과가 필요할 때 갈 곳입니다.linearmodels— 올바른 표준오차가 붙은 IV·패널·GMM 추정량. IV 편의 작업용 도구입니다.statsmodels— 나머지 전부, 그리고 이 시리즈의 회귀표 출처입니다.- R의
did,fixest,Synth— 현대적 엇갈린 DiD 추정량은 R에 먼저 도착했고 일부는 아직 성숙한 파이썬 대응물이 없습니다. 건너갈 값어치가 있습니다.
DoWhy(Microsoft) — the framework layer. Encodes assumptions explicitly and provides refutation tests. Best starting point for structuring an analysis.EconML(Microsoft) — heterogeneous treatment effects, double ML, causal forests. Where to go when you need effects by subgroup rather than an average.linearmodels— IV, panel and GMM estimators with correct standard errors. The workhorse for the IV post.statsmodels— everything else, and the source of the regression tables in this series.- R
did,fixest,Synth— the modern staggered DiD estimators landed in R first and some still have no mature Python equivalent. Worth crossing over for.
목록을 쓰는 법How to use a list like this
두 가지를 제안합니다. 첫째, 넓이보다 깊이를 택하십시오. 코드를 열어 놓고 책 한 권을 끝까지 하는 편이 다섯 권을 훑는 것보다 낫습니다. 이 분야의 방법들은 서술하기는 어렵지 않고 제대로 적용하기가 어렵습니다. 그 간격은 직접 해 봐야만 좁혀집니다.
Two suggestions. First, prefer depth to breadth. Working through one book with the code open beats skimming five. The methods in this field are not hard to describe and are hard to apply correctly, and that gap only closes by doing.
둘째, 논문만큼 연구자에게 주의를 기울이십시오. 인과추론은 실무 표준이 몇 년마다 옮겨 가는 빠른 분야입니다. 엇갈린 DiD 문헌이 가장 최근의 명확한 예입니다. 몇몇 사람을 팔로우하는 것이 어떤 정적인 읽기 목록보다 — 이것을 포함해서 — 최신 상태를 유지하는 더 믿을 만한 방법입니다.
Second, pay attention to the researchers as much as the papers. Causal inference is a fast-moving field where the practical standard shifts every few years — the staggered DiD literature is the clearest recent example. Following a handful of people is a more reliable way to stay current than any static reading list, including this one.
이 섹션을 닫으며Closing the section
그래서 표준오차는 실제 불확실성을 과소진술합니다. 때로는 심하게 그렇습니다. 설계가 주어졌을 때의 잡음을 재는 것이지 설계가 유효했는지는 아무 말도 하지 않습니다. 민감도 분석이 선택적 강건성 부록이 아니라 결과의 핵심 부분인 이유입니다. 글쓰기가 추정만큼 중요한 이유이기도 합니다.
The standard error is therefore an understatement of your actual uncertainty, sometimes a severe one. It measures noise in the estimate given the design; it says nothing about whether the design was valid. This is why sensitivity analysis is not an optional robustness appendix but a core part of the result, and why the write-up matters as much as the estimation.
작업 방식에 주는 실무적 함의는 이렇습니다. 추정량이 아니라 설계와 가정에 시간을 쓰십시오. 깨끗한 설계에 단순한 추정량을 얹는 편이 훼손된 설계에 정교한 추정량을 얹는 것보다 매번 낫습니다. LaLonde 편이 그 실증입니다. 어떤 방법론적 정교함도 비교 가능하지 않은 비교군을 구해 내지 못했습니다.
The practical consequence: spend your time on the design and the assumption, not on the estimator. A clean design with a simple estimator beats a sophisticated estimator applied to a compromised design, every time. The LaLonde post is the empirical demonstration — no amount of methodological sophistication rescued a comparison group that was not comparable.
뒤따르는 머신러닝 섹션은 질문을 바꿉니다. 여기까지는 식별 — 계수 하나를 제대로 맞히는 것 — 이었습니다. ML 편들은 예측과 차원을 다룹니다. 감당할 수 있는 것보다 변수가 많을 때 무엇을 할 것인가, 그리고 통제변수 수가 늘어나는 동안 인과 질문을 시야에 붙들어 두는 방법. 시리즈의 마지막 편인 Double LASSO에서 두 갈래가 만납니다.
The Machine Learning section that follows changes the question. Everything up to here has been about identification — getting one coefficient right. The ML posts are about prediction and dimensionality: what to do when you have more variables than you can handle, and how regularisation lets you keep the causal question in view as the number of controls grows. Double LASSO, the final post, is where the two threads meet.