앞의 두 편은 모든 교란변수를 측정했다는 데 기댔습니다. 이 편은 그것을 포기합니다. 도구변수와 회귀 불연속은 교란변수를 관측하지 않고도 작동합니다. 처치의 변동 중 미관측 요인과 무관한 원천을 찾아 그것만 씁니다.
The previous two posts relied on having measured every confounder. This one gives that up. Instrumental variables and regression discontinuity both work without observing the confounders — they find a source of variation in treatment that is unrelated to the unobservables, and use only that.
대가는 더 이상 모두에 대한 효과를 추정하지 못한다는 점입니다. 그 특정한 변동이 움직인 사람들에 대한 효과를 추정합니다. 그게 정확히 누구인지 알아내는 일이 어려운 부분입니다.
The price is that you no longer estimate the effect for everyone. You estimate it for whoever that particular source of variation happens to move. Understanding exactly who that is turns out to be the hard part.
도구변수Instrumental variables
도구변수 Z는 처치를 움직이면서 결과로 가는 다른 경로가 없는 변수입니다. 세 조건이 정의합니다.
An instrument Z shifts treatment but has no other route to the outcome. Three conditions define it.
- 관련성. Z가 실제로 D에 영향을 줍니다. 이건 검정 가능합니다. 1단계 회귀입니다.
- 배제. Z는 D를 통해서만 Y에 영향을 줍니다. 검정 불가능하고, 도구변수의 생사가 여기서 갈립니다.
- 독립성. Z가 잠재적 결과를 놓고 보면 무작위 배정이나 다름없습니다.
- Relevance. Z actually affects D. This is testable — it is the first-stage regression.
- Exclusion. Z affects Y only through D. Not testable, and this is where instruments live or die.
- Independence. Z is as good as randomly assigned with respect to the potential outcomes.
이 조건들이 성립하면 2단계 최소제곱이 D의 변동 중 Z에서 온 부분을 분리해 그것만으로 Y에 대한 효과를 추정합니다. 교란변수는 여전히 관측되지 않습니다. 가정상 Z와 무관하므로 상관없습니다.
Given these, two-stage least squares isolates the part of the variation in D that comes from Z and uses only that. The confounders remain unobserved and it does not matter, because they are unrelated to Z by assumption.
고전적 예들이 감을 줍니다. 교육 연수에 대한 출생 분기(의무교육법이 태어난 시기에 따라 다르게 구속함), 대학 진학에 대한 최근접 대학까지의 거리, 군 복무에 대한 징병 추첨 번호. 각 경우에 도구변수가 결과와 무관해 보이는 이유로 처치를 움직입니다.
Classic examples give the flavour: quarter of birth as an instrument for years of schooling (compulsory schooling laws bind differently depending on when you were born), distance to the nearest college for university attendance, draft lottery number for military service. In each case the instrument moves treatment for reasons plausibly unrelated to the outcome.
약한 도구변수가 실무적 실패 양상입니다. 1단계가 약하면 2SLS가 OLS 쪽으로 편향되고 표준오차가 심하게 과소평가됩니다. 자신 있게 틀린 답을 얻습니다. 잘못 설정된 모형 편과 같은 방식입니다. 통상적 스크린은 1단계 F > 10인데, 최근 연구(Lee, McCrary, Moreira, Porter)는 그 문턱이 너무 관대하고 유효한 추론에는 F > 100이 필요한 경우가 많다고 봅니다. 1단계를 보고하십시오. 항상.
Weak instruments are the practical failure mode. If the first stage is weak, 2SLS is biased toward OLS and its standard errors are badly understated — a confidently wrong answer, in the same way as the misspecification post. The conventional screen is a first-stage F above 10, though recent work (Lee, McCrary, Moreira and Porter) argues that is far too lenient and valid inference often requires F above 100. Report the first stage. Always.
import numpy as np
from linearmodels.iv import IV2SLS
import statsmodels.api as sm
rng = np.random.default_rng(4)
n = 5_000
TRUE = 1.0
confounder = rng.normal(size=n) # unobserved, drives D and Y
z = rng.normal(size=n) # the instrument
d = 0.8 * z + 1.5 * confounder + rng.normal(size=n)
y = TRUE * d + 2.0 * confounder + rng.normal(size=n)
ols = sm.OLS(y, sm.add_constant(d)).fit()
iv = IV2SLS(y, np.ones((n, 1)), d, z).fit()
first = sm.OLS(d, sm.add_constant(z)).fit()
print(f"true effect : {TRUE:.3f}")
print(f"OLS (confounded) : {ols.params[1]:.3f}")
print(f"IV / 2SLS : {iv.params[0]:.3f}")
print(f"first-stage F : {first.fvalue:,.0f}")IV2SLS를 쓰십시오.Running the two stages by hand gives the right point estimate and the wrong standard errors. Use IV2SLS.true effect : 1.000
OLS (confounded) : 1.788
IV / 2SLS : 0.994
first-stage F : 1,199
OLS는 교란 때문에 79% 위로 편향되어 있습니다. IV는 교란변수를 한 번도 보지 않고 참값을 되찾습니다.
OLS is biased 79% upward by the confounder. IV recovers the truth without ever observing it.
IV가 실제로 추정하는 것. 효과가 이질적이면 2SLS는 국소평균처치효과를 되찾습니다. 순응자, 그러니까 도구변수가 실제로 처치 상태를 바꾼 단위들에 대한 효과입니다. 항상 받는 사람도, 절대 안 받는 사람도 아닙니다. 대학까지의 거리가 도구변수라면 "가까우면 진학하고 아니면 안 하는 사람"의 교육 수익률을 배웁니다. 특수하고 대표성이 없을 수도 있는 집단입니다. LATE는 진짜 추정 대상이지만 정책이 묻는 대상인 경우는 드뭅니다. 그 간격은 명시적으로 말해야 합니다.
What IV actually estimates. Under heterogeneous effects, 2SLS recovers the local average treatment effect — the effect for compliers, the units whose treatment status the instrument actually changed. Not always-takers, not never-takers. If distance to college is your instrument, you learn the return to schooling for people who would attend if nearby and not otherwise: a specific and possibly unrepresentative group. LATE is a real estimand, but it is rarely the one policy asks about, and the gap deserves to be stated explicitly.
회귀 불연속Regression discontinuity
RDD는 다른 우연을 이용합니다. 연속적인 배정변수에 있는 임의의 문턱입니다. 시험 점수 상위에 주는 장학금, 소득선 아래 가구에 주는 급여, 등록 인원에서 발동되는 학급 규모 규칙 같은 것들입니다.
RDD exploits a different accident: an arbitrary threshold in a continuous running variable. Scholarships above a test score cutoff, benefits below an income line, class-size rules triggered at an enrolment count.
논리는 이렇습니다. 문턱 바로 위와 바로 아래의 단위는 사실상 동일합니다. 69점과 70점의 차이는 잡음입니다. 그런데 한쪽은 처치받고 한쪽은 안 받습니다. 그러니 문턱에서 나타나는 결과 도약이 곧 처치효과입니다. 교란변수는 문턱을 매끄럽게 통과하는 반면 처치만 도약하므로 처리됩니다.
The logic: units just above and just below the cutoff are essentially identical — the difference between scoring 69 and 70 is noise — but one group gets treated and the other does not. So the jump in outcomes at the threshold is the treatment effect, and confounders are handled because they vary smoothly through the cutoff while treatment jumps.
RDD는 관측 설계 중 가장 신뢰할 만한 것으로 평가받습니다. 핵심 가정이 검정 가능에 가깝기 때문입니다. 두 종류가 있습니다. 샤프 RD는 문턱을 넘는 것이 처치를 정확히 결정하고 퍼지 RD는 처치 확률을 바꿉니다. 후자는 문턱 지시변수를 도구변수로 쓰는 IV이고 문턱에서 성립하는 LATE를 추정합니다.
RDD is generally regarded as the most credible observational design, because its key assumption is close to testable. Two flavours: sharp RD, where crossing the threshold determines treatment exactly, and fuzzy RD, where crossing changes the probability of treatment. The latter is IV with the threshold indicator as instrument, estimating a LATE at the cutoff.
주된 튜닝 선택은 대역폭입니다. 문턱에서 얼마나 멀리까지 포함할 것인가. 좁으면 신빙성은 높고 관측치가 적습니다. 넓으면 정밀하지만 더 이상 비교 가능하지 않은 단위를 끌어들입니다. 범위에 걸친 민감도를 보고하십시오. 임의 선택보다 Calonico–Cattaneo–Titiunik 최적 대역폭과 편향보정 추론을 쓰십시오.
The main tuning choice is bandwidth: how far from the cutoff to include. Narrow gives credibility and few observations; wide gives precision and pulls in units that are no longer comparable. Report sensitivity across a range, and prefer the Calonico-Cattaneo-Titiunik optimal bandwidth with bias-corrected inference over an arbitrary choice.
실무에서In practice
- 1단계를 부록 표가 아니라 주요 결과로 보고하십시오. 계수, F통계량, 부호. 이것 없이는 독자가 IV 추정치를 평가할 수 없습니다. 약한 1단계는 하류 전체를 무효화합니다.
- OLS와 IV를 명시적으로 비교하십시오. 비슷하면 내생성이 크지 않고 IV는 주로 정밀도를 깎은 것입니다. IV가 훨씬 크면 왜인지 물으십시오. 약한 도구변수는 OLS 쪽으로 편향시키므로 약한 1단계와 함께 나타난 큰 간격은 발견이 아니라 경고입니다.
- RD는 그리십시오. 그래프가 곧 분석입니다. 배정변수를 구간화해 구간별 평균 결과를 찍고 양쪽 적합선을 보이십시오. 그 그림에서 보이지 않는데 회귀에서만 나타나는 불연속은 거의 항상 함수 형태의 인공물입니다.
- 없어야 할 곳의 도약을 검정하십시오. 처치 전에 측정된 공변량은 문턱에서 연속이어야 합니다. 기저 특성이 도약하면 설계가 깨진 것입니다. 진짜 문턱에서 떨어진 가짜 문턱에서 돌리는 위약 검정이 또 다른 표준 확인입니다.
- RD에서 고차 다항식을 피하십시오. Gelman과 Imbens는 3차 이상의 전역 다항식이 문턱에서 먼 점들에 좌우되는 불규칙한 추정치를 낳는다는 것을 보였습니다. 합리적인 대역폭 안의 국소 선형 회귀가 현재 표준입니다.
- 도구변수가 여럿이면 무엇을 평균 내는지 주의하십시오. 여러 도구변수를 쓴 2SLS는 도구변수별 LATE들의 가중평균을 내는데, 그 가중치는 당신이 고르지 않았고 마음에 안 들 수도 있습니다. 각 도구변수의 정확식별 추정치를 따로 보고하는 편이 대개 더 유익합니다.
- Report the first stage as a headline result, not an appendix table. The coefficient, the F-statistic, the sign. A reader cannot evaluate an IV estimate without it, and a weak first stage invalidates everything downstream.
- Compare OLS and IV explicitly. If they are similar, endogeneity may be modest and IV mostly cost you precision. If IV is much larger, ask why — weak instruments bias toward OLS, so a large gap alongside a weak first stage is a warning rather than a discovery.
- Plot the RD. The graph is the analysis. Bin the running variable, plot mean outcomes per bin, show the fitted lines on each side. A discontinuity invisible in that plot but present in the regression is almost always a functional form artefact.
- Test for jumps where there should be none. Covariates measured before treatment should be continuous at the cutoff. If baseline characteristics jump, the design is broken. Placebo tests at fake cutoffs away from the real one are the other standard check.
- Avoid high-order polynomials in RD. Gelman and Imbens showed global polynomials of order three and above produce erratic estimates driven by points far from the cutoff. Local linear regression within a sensible bandwidth is the current standard.
- With multiple instruments, be careful what you are averaging. 2SLS with several instruments produces a weighted average of instrument-specific LATEs, with weights you did not choose and may not like. Reporting just-identified estimates from each instrument separately is often more informative.
가격은 수요에 반응해 정해지므로 수량을 가격에 OLS 회귀하는 건 가망이 없습니다. 투입가·환율·농산물 날씨 같은 비용 이동 요인이 표준 도구변수입니다. 플랫폼 경제학에서는 무작위 홀드아웃을 실제 노출에 대한 도구변수로 쓰는데, 이게 정확히 실험 편의 불응 사례입니다. 배정이 실제 처치의 도구변수가 됩니다. 급여·보조금·신용점수·세율 구간의 자격 문턱은 행정 데이터 어디에나 불연속을 만듭니다. 특히 신용점수 문턱은 소비자 금융 연구의 큰 축을 떠받쳐 왔습니다.
Price is set in response to demand, so OLS of quantity on price is hopeless — cost shifters (input prices, exchange rates, weather for agricultural inputs) are the standard instruments. In platform economics, randomised holdouts instrument for actual exposure, which is exactly the non-compliance case from the experiments post: assignment instruments for treatment received. Eligibility cutoffs for benefits, subsidies, credit scores and tax brackets create discontinuities throughout administrative data, and credit scoring cutoffs in particular have supported a large body of consumer finance research.
2021년 노벨상이 Angrist, Imbens, Card에게 간 것은 대체로 이 작업 때문입니다. 구체적으로는 효과가 이질적일 때 이 설계들이 무엇을 식별하는지를 명확히 한 공로입니다. 그 명확화, 즉 LATE가 응용 논문에서 가장 자주 생략되는 부분이자 가장 제대로 다룰 값어치가 있는 부분입니다.
The 2021 Nobel Prize went to Angrist, Imbens and Card largely for this body of work — specifically for clarifying what these designs identify when effects are heterogeneous. That clarification, LATE, is the part most often skipped in applied write-ups and the part most worth getting right.