Tech PostsTech Posts
인과 추론Causal Inference · 14 / 24
14 인과 추론Causal Inference

정답을 미리 아는 데이터셋The Dataset Where We Know the Answer in Advance

무작위 대조군을 일부러 버리고 관측 데이터로 대체한 뒤, 그것으로 실험 결과를 되찾을 수 있는지 물었습니다.Throw the randomised control group away, substitute an observational one, and ask whether the experimental answer can still be recovered.

LaLonde 데이터셋은 관측 인과추론에서 가장 통제된 시험에 가깝습니다. 이유가 특이합니다. 정답을 미리 알기 때문입니다.

The LaLonde dataset is the closest thing observational causal inference has to a controlled test, and the reason is unusual: we know the right answer in advance.

Robert LaLonde는 1986년 논문에서 National Supported Work 실증사업을 가져다가 무작위 대조군을 버렸습니다. 실제로 무작위 배정된 직업훈련 프로그램이고 연간 소득 약 1,800달러 증가라는 실험 추정치가 있었습니다. 버린 자리에는 전국 서베이 데이터에서 뽑은 비실험 비교군을 넣었습니다. 그러고는 당시의 관측 기법들을 돌려 실험 숫자를 되찾을 수 있는지 물었습니다.

Robert LaLonde 1986 paper took the National Supported Work demonstration — a genuinely randomised job training programme with an experimental estimate of roughly $1,800 in additional annual earnings — and threw the randomised control group away. In its place he substituted a non-experimental comparison group drawn from national survey data, ran the observational methods of the day, and asked whether they could recover the experimental number.

되찾지 못했습니다. 설정에 따라 강한 음수에서 믿기 어려운 양수까지 나왔습니다. 관측 프로그램 평가의 신뢰성에 던지는 심각한 도전이었습니다. 응용 경제학에 신빙성 혁명이 일어난 큰 이유이기도 합니다.

They could not. Estimates ranged from strongly negative to implausibly positive depending on the specification. It was a serious challenge to the credibility of observational programme evaluation, and a large part of why the credibility revolution in applied economics happened.

Dehejia와 Wahba가 1999년에 성향점수 방법으로 다시 다뤘습니다. 중첩과 균형에 주의를 기울이면 실험 기준값을 되찾을 수 있다는 주장이었습니다. 그 주장 자체도 논쟁 대상입니다. 이 데이터셋은 그 이후로 이 분야의 시험장이 되었습니다.

Dehejia and Wahba revisited it in 1999 using propensity score methods and argued that with careful attention to overlap and balance the experimental benchmark could be recovered. That claim has itself been contested. The dataset has been the field proving ground ever since.

좋은 연습인 이유Why this makes a good exercise

거의 모든 응용 인과분석에는 같은 문제가 있습니다. 숫자를 하나 내놓는데 그게 맞는지 아무도 말해 줄 수 없습니다. 여기서는 확인할 수 있습니다. 실험 추정치가 목표값입니다. 그 값을 빗나갔다면 보통 연구에서라면 눈에 띄지도 않았을 무언가를 잘못한 것입니다.

Almost every applied causal analysis has the same problem: you produce a number and nobody can tell you whether it is right. Here you can check. The experimental estimate is the target, and any method that misses it is doing something wrong that would be invisible in an ordinary study.

특히 두 가지가 눈에 들어옵니다.

Two things in particular become visible.

중첩 실패는 기술적 세부사항이 아니라 사건의 본체입니다. 서베이에서 뽑은 비교군은 처치군과 엄청나게 다릅니다. 연령 분포도, 소득 이력도, 고용 패턴도 다릅니다. 공변량 공간의 넓은 영역에서는 처치 단위에 짝지을 비교 가능한 대조군이 아예 없습니다. 원자료의 순진한 평균 차이는 형편없이 틀립니다. 추정량의 어떤 미묘함과는 상관없습니다. 순전히 이것 때문입니다.

Overlap failures are the main event, not a technicality. The comparison group drawn from survey data differs enormously from the treated group — different age profiles, earnings histories, employment patterns. Large parts of the covariate space contain treated units with no comparable controls at all. The naive difference in means is wildly wrong, and it is wrong because of this, not because of any subtlety in the estimator.

설정 선택이 답을 크게 움직입니다. 성향 모형에 어떤 공변량이 들어가는지, 소득을 수준으로 넣을지 로그로 넣을지, 캘리퍼 폭, 복원추출 여부 — 각각이 추정치를 옮깁니다. 기준값이 없는 연구에서는 그중 하나를 보고하고 영영 모를 것입니다.

Specification choices move the answer a lot. Which covariates enter the propensity model, whether earnings enter in levels or logs, the caliper width, whether you match with or without replacement — each shifts the estimate. In a study with no benchmark you would report one of them and never know.

무엇을 볼 것인가What to watch for

작업은 표준 순서를 따릅니다. 읽을 때는 점추정치보다 진단이 중요합니다.

The work follows the standard sequence. As you read it, the diagnostics matter more than the point estimates.

  1. 먼저 원시 평균 차이. 형편없이 틀려 보여야 합니다. 이 연습 전체가 고치려는 기준선입니다.
  2. 성향점수 추정, 그다음 집단별 분포. 중첩 확인입니다. 점수가 아주 낮은 비교 단위가 크게 뭉쳐 있을 것으로 예상하십시오. 애초에 프로그램에 들어올 일이 사실상 없었던 사람들입니다. 효과를 판단할 정보가 여기엔 없습니다.
  3. 공통 지지 영역으로 자르기. 중첩 밖의 단위를 버려야 남은 비교가 의미를 갖습니다. 동시에 추정 대상이 바뀝니다. 이제 비교가 가능한 부분모집단에 대한 효과입니다.
  4. 균형 평가. 전후 표준화 차이. 조정이 통했는지를 가리는 기준은 성향 모형 자신의 적합도가 아니라 이것입니다.
  5. 층화·매칭 추정치. 약 1,800달러라는 실험 기준값과 비교하십시오.
  1. The raw difference in means first. It should look badly wrong. That is the baseline the whole exercise is trying to fix.
  2. Propensity score estimation, then its distribution by group. This is the overlap check. Expect a large mass of comparison units with very low scores — people who would essentially never have entered the programme and carry no information about its effect.
  3. Trimming to common support. Discarding units outside the overlap region is what makes the remaining comparison meaningful. Note that it also changes the estimand: an effect for the subpopulation where comparison is possible.
  4. Balance assessment. Standardised differences before and after. This, not the propensity model own goodness of fit, is the criterion for whether the adjustment worked.
  5. Stratification and matching estimates. Compare against the experimental benchmark of roughly $1,800.
python
import numpy as np
import pandas as pd
from sklearn.linear_model import LogisticRegression

# NSW treated units vs a survey comparison group (CPS/PSID style)
covariates = ["age", "education", "black", "hispanic", "married",
              "nodegree", "re74", "re75"]

X = df[covariates].to_numpy()
D = df["treat"].to_numpy()
Y = df["re78"].to_numpy()

naive_gap = Y[D == 1].mean() - Y[D == 0].mean()
print(f"naive difference in means: {naive_gap:>9,.0f} USD")
print(f"experimental benchmark   : ~1,800 USD")

# propensity score, then the overlap check that decides everything
ps = LogisticRegression(max_iter=2000).fit(X, D).predict_proba(X)[:, 1]

lo, hi = ps[D == 1].min(), ps[D == 1].max()
support = (ps >= lo) & (ps <= hi)
print(f"\ncontrols on common support: {(support & (D==0)).sum():,}"
      f" of {(D==0).sum():,}")
설정 탐색은 이 노트북의 결함이 아닙니다. 이 방법이 실제로 요구하는 바를 정직하게 묘사한 것입니다.The specification search is not a defect of the exercise — it is the honest depiction of what this method requires.

이 연습이 가르치는 것What the exercise teaches

자르면 질문이 바뀝니다. 중첩되지 않는 단위를 버린 뒤에 추정하는 것은 정책이 아니라 데이터가 정의한 부분모집단의 효과입니다. 여전히 쓸모 있을 수 있지만 이 추정치가 누구에게 해당하는지는 글에서 말해야 합니다. 잘라 낸 표본의 ATT는 프로그램이 실제로 서비스할 모집단의 ATE가 아닙니다.

Trimming changes the question. After discarding non-overlapping units you are estimating an effect for a subpopulation defined by the data rather than by policy. That may still be useful, but the write-up must say who the estimate applies to. An ATT on a trimmed sample is not the ATE for the population the programme would actually serve.

관측 가능한 것에 의한 선택은 어떤 진단도 확인할 수 없는 가정 위에 섭니다. 균형표가 확인해 주는 것은 관측된 공변량이 균형 잡혔다는 사실뿐입니다. 의욕, 건강, 가족 상황, 그 밖에 등록과 소득을 함께 움직이면서 측정되지 않은 무엇에 대해서는 아무 말도 하지 않습니다. LaLonde의 원래 지적이 유효합니다. 비교군이 무작위 배정된 게 아니라 구성된 것이라면 논증의 부담은 분석가에게 있습니다.

Selection on observables rests on an assumption no diagnostic can check. Balance tables demonstrate that observed covariates are balanced. They say nothing about motivation, health, family circumstances or any other unmeasured driver of both enrolment and earnings. LaLonde original point stands: when the comparison group is constructed rather than randomised, the burden of argument is on the analyst.

실무에서In practice

  • 순진한 추정치를 조정된 추정치와 나란히 보고하십시오. 둘 사이의 간격이 식별 전략이 얼마나 많은 일을 하는지 독자에게 알려 줍니다. 간격이 크면 그 자체가 분석의 본체입니다. 숨길 문제가 아닙니다.
  • 균형은 t검정보다 표준화 차이로 보십시오. 매칭이 유효 표본 크기를 바꾸므로 p값은 균형과 검정력을 뒤섞습니다. 0.1 아래가 통상적 문턱입니다. 평균뿐 아니라 분산비도 확인하십시오.
  • 성향점수 자체 대신 그 로짓으로 매칭하십시오. 로짓이 선형에 가깝고 꼬리에서 더 잘 행동합니다. 매칭 문제의 대부분이 꼬리에 삽니다.
  • 캘리퍼를 쓰십시오. 없으면 좋은 짝이 없는 처치 단위가 아무리 멀어도 가장 가까운 것과 짝지어집니다. 로짓 점수 표준편차의 0.2가 흔한 기본값입니다. 짝을 못 찾아 버려진 단위 수를 보고하십시오. 그 수 자체가 결과입니다.
  • 부트스트랩은 조심해서 쓰십시오. 매칭 절차의 표준오차는 성향점수가 추정되었다는 사실을 반영하지 않습니다. Abadie와 Imbens는 특히 최근접이웃 매칭에서 통상적 부트스트랩이 유효하지 않음을 보였습니다. 그들의 해석적 분산추정량이나 방법에 맞는 설계를 쓰십시오.
  • 가능하면 이중 강건 추정량을 쓰십시오. AIPW와 TMLE는 성향 모형과 결과 모형 중 어느 한쪽만 올바르게 설정되어 있어도 일치성을 유지합니다. 둘 다 요구하지 않습니다. 매칭 단독보다 엄격히 안전하고 현대의 기본값입니다.
  • 더 나은 설계가 있는지 생각해 보십시오. LaLonde의 교훈은 결국 추정이 아니라 설계입니다. 패널 데이터가 있으면 이중차분이 각 단위를 자기 자신의 대조군으로 써서 횡단면 비교 가능성 문제를 아예 피합니다. 임의의 문턱이나 그럴듯한 도구변수가 있으면 그 설계들이 더 약한 가정을 씁니다. 다음 두 편이 그것입니다.
  • Report the naive estimate alongside the adjusted one. The gap between them tells the reader how much work the identification strategy is doing. A large gap is not a problem to hide — it is the substance of the analysis.
  • Prefer standardised differences over t-tests for balance. Matching changes the effective sample size, so p-values conflate balance with power. Under 0.1 is the convention, and check variance ratios too, not just means.
  • Match on the logit of the propensity score, not the score itself. The logit is closer to linear and behaves better in the tails, where most matching problems live.
  • Use a caliper. Without one, a treated unit with no good match gets matched to whatever is nearest, however far away. A caliper of 0.2 standard deviations of the logit score is the common default. Report how many units were dropped for lack of a match — that count is itself a result.
  • Bootstrap with care. Standard errors from a matching procedure do not account for the fact that the propensity score was estimated. Abadie and Imbens showed the ordinary bootstrap is invalid for nearest-neighbour matching specifically; use their analytic variance estimator or a design appropriate to the method.
  • Prefer doubly robust estimators where you can. Augmented IPW and TMLE remain consistent if either the propensity model or the outcome model is correctly specified, rather than requiring both. Strictly safer than matching alone, and the modern default.
  • Consider whether a better design is available. LaLonde lesson is ultimately about design, not estimation. With panel data, differences-in-differences uses each unit as its own control and avoids the cross-sectional comparability problem entirely. If there is an arbitrary threshold or a plausible instrument, those designs make weaker assumptions. Those are the next two posts.

노동경제학·보건경제학·개발경제학의 프로그램 평가가 정확히 이 문제에 지배됩니다. 산업에서도 다른 이름으로 끊임없이 나타납니다. 가입한 고객에 대한 로열티 프로그램 효과, 방문을 받은 계정에 대한 영업 효과, 채택한 사용자에 대한 기능 효과. 모든 경우에 처치군이 처치받기를 선택했습니다. LaLonde의 경고가 그대로 적용됩니다.

Programme evaluation in labour, health and development economics is dominated by exactly this problem. It also appears constantly in industry under different names: the effect of a loyalty programme on customers who opted in, of a sales call on accounts that received one, of a feature on users who adopted it. In every case the treated group chose to be treated, and the LaLonde warning applies directly.