Tech PostsTech Posts
기초와 계량경제학Basics & Econometrics · 08 / 24
08 기초와 계량경제학Basics & Econometrics

분산을 성가신 것이 아니라 예측 대상으로 보기Treating Variance as the Thing to Forecast, Not a Nuisance

금융 수익률의 분산은 움직입니다. 그것도 예측 가능하게 움직입니다. 그러면 보정할 게 아니라 모형화할 대상입니다.The variance of financial returns moves, and it moves predictably. That makes it something to model, not something to correct for.

금융 수익률에는 성질이 하나 있습니다. 보통의 회귀는 그것을 다루도록 만들어지지 않았습니다. 분산이 움직입니다. 조용한 주가 함께 뭉치고 격렬한 주도 함께 뭉칩니다. 만델브로가 1963년에 짚었습니다. "큰 변화는 부호와 무관하게 큰 변화 뒤에 오고, 작은 변화는 작은 변화 뒤에 온다."

Financial returns have a property that ordinary regression is not built for: their variance moves. Calm weeks cluster together, and so do violent ones. Mandelbrot noticed it in 1963: "large changes tend to be followed by large changes, of either sign, and small changes by small changes."

이것이 변동성 군집이고 OLS의 핵심 가정을 깹니다. 등분산성은 오차 분산이 일정하다고 말합니다. 여기서는 분명히 아닙니다. 게다가 일정하지 않기만 한 게 아니라 예측 가능합니다. 그러면 보정할 대상이 아니라 모형화할 대상입니다.

That is volatility clustering, and it breaks a core OLS assumption. Homoskedasticity says the error variance is constant. Here it plainly is not — and the variance is not just non-constant but predictable, which makes it something you can model rather than merely correct for.

로버스트 표준오차로는 안 되는 이유Why robust standard errors are not enough

이분산을 다루는 통상적 대응은 White나 HAC 표준오차입니다. 변하는 분산이 성가신 방해물일 때는 그게 맞는 수입니다. 추론이 고쳐지고 넘어가면 됩니다.

The usual response to heteroskedasticity is White or HAC standard errors. That is the right move when the changing variance is a nuisance: it fixes your inference and you move on.

ARCHARCH

Engle의 1982년 통찰은 오늘의 분산이 어제의 충격에 의존하도록 두는 것이었습니다.

Engle 1982 insight was to let today variance depend on yesterday surprises.

σ²_t = ω + α₁ ε²_(t−1) + ... + α_q ε²_(t−q)

지난 기의 큰 충격이 — 제곱으로 들어가므로 부호와 무관하게 — 이번 기의 분산을 올립니다. 군집이 곧바로 재현됩니다.

A large shock last period raises the variance this period — of either sign, because it enters squared. That reproduces clustering directly.

약점은 금융 변동성의 지속성이 길다는 데 있습니다. 시차 제곱오차만으로 그것을 담으려면 시차가 많이 필요하고 시차마다 추정할 모수가 붙습니다.

The weakness is that persistence in financial volatility is long, and capturing it with lagged squared errors alone needs many lags, each with a parameter to estimate.

GARCHGARCH

Bollerslev의 1986년 일반화는 우변에 시차 분산을 더합니다.

Bollerslev 1986 generalisation adds lagged variance to the right-hand side.

σ²_t = ω + α₁ ε²_(t−1) + β₁ σ²_(t−1)

이 항 하나가 모형을 실용적으로 만들었습니다. σ²_(t−1)이 다시 σ²_(t−2)에 의존하므로 이 재귀는 과거 충격의 무한한 가중 이력을 품습니다. 가중치는 기하급수적으로 감소합니다. 지수평활과 같은 수법입니다. 모수 세 개짜리 GARCH(1,1)이 시차 열 개짜리 ARCH를 대개 이깁니다.

That single extra term is what made the model practical. Because σ²_(t−1) itself depends on σ²_(t−2), the recursion embeds an infinite weighted history of past shocks with geometrically declining weights — the same trick exponential smoothing uses. GARCH(1,1) with three parameters typically outperforms an ARCH model with ten.

적합된 GARCH에서 반드시 읽어야 할 두 값이 있습니다.

Two quantities are worth reading off any fitted GARCH.

  • α + β는 지속성입니다. 충격이 사라지는 데 걸리는 시간이고 일별 주식 수익률에서는 경험적으로 0.95~0.99에 자리 잡습니다. 변동성 충격이 며칠이 아니라 몇 주에 걸쳐 사그라진다는 뜻입니다. α + β ≥ 1이면 비정상이고 장기분산이 정의되지 않습니다. 대개는 진짜로 폭발적인 변동성이라기보다 표본 안에 구조 변화가 있다는 신호입니다.
  • ω / (1 − α − β)는 장기분산입니다. 예측은 이 수준으로 평균회귀합니다. 지속성이 0.98이면 그 회귀에 오랜 시간이 걸립니다. 위기 이후 몇 주 동안 GARCH 예측이 높게 유지되는 이유입니다.
  • α + β is persistence. How long shocks take to decay; empirically this lands around 0.95-0.99 for daily equity returns, meaning volatility shocks die out over weeks rather than days. If α + β >= 1 the model is non-stationary and its long-run variance is undefined — usually a sign of a structural break in the sample rather than genuinely explosive volatility.
  • ω / (1 - α - β) is the long-run variance. Forecasts mean-revert toward this level. With persistence at 0.98 that reversion takes a long time, which is why GARCH forecasts stay elevated for weeks after a crisis.

기억할 진단 하나The one diagnostic to remember

제곱 관측치의 자기상관함수를 그리는 것이 ARCH 효과의 표준 검정입니다. 논리를 말해 둘 값어치가 있습니다. 수익률 자체는 예측 불가에 가까우니 그 ACF는 백색잡음처럼 보입니다. 그런데 제곱은 부호를 없애고 크기만 남깁니다. 크기는 강하게 자기상관되어 있습니다.

Plotting the autocorrelation function of the squared observations is the standard test for ARCH effects, and the logic is worth stating. The returns themselves are close to unpredictable, so their ACF looks like white noise. But squaring removes the sign and leaves the magnitude, and magnitudes are strongly autocorrelated.

python
import numpy as np
from arch import arch_model
from statsmodels.graphics.tsaplots import plot_acf
import matplotlib.pyplot as plt

rng = np.random.default_rng(11)

# a series whose variance grows over time - the effect we want to detect
T = 2_000
scale = np.linspace(0.5, 3.0, T)
returns = rng.normal(scale=scale) * 100

fig, axes = plt.subplots(1, 2, figsize=(12, 3.5))
plot_acf(returns, lags=25, ax=axes[0], title="ACF of returns")
plot_acf(returns**2, lags=25, ax=axes[1], title="ACF of SQUARED returns")
왼쪽은 백색잡음처럼 보이고 오른쪽이 조건부 이분산의 서명입니다.The left panel looks like white noise. The right one is the signature of conditional heteroskedasticity.

제곱 잔차의 ACF에 유의한 막대가 서 있으면 조건부 이분산입니다. 그 그림이 평평하면 GARCH가 필요 없습니다.

Significant spikes in the ACF of squared residuals are the signature. If that plot is flat, you do not need GARCH.

python
split = int(T * 0.8)
train, test = returns[:split], returns[split:]

model = arch_model(train, mean="Constant", vol="GARCH", p=1, q=1, dist="t")
res = model.fit(disp="off")

alpha = res.params["alpha[1]"]
beta  = res.params["beta[1]"]
omega = res.params["omega"]

print(f"persistence  alpha+beta = {alpha + beta:.4f}")
print(f"long-run variance       = {omega / (1 - alpha - beta):.4f}")
print(res.forecast(horizon=5).variance.iloc[-1].round(3).to_string())
분포를 t로 두는 것은 기본값이 아니라 의도입니다 — 아래 참조.Choosing the t distribution is deliberate rather than default. See below.

실무에서In practice

  • 분산보다 평균을 먼저 모형화하십시오. GARCH는 어떤 평균식의 잔차 분산을 기술합니다. 평균이 잘못 설정되어 있으면 — 빠진 추세, 모형화하지 않은 시차 — 그 구조가 잔차로 새어 들어가 ARCH 효과를 부풀립니다. 평균을 먼저 맞추고 잔차에 계열상관이 없음을 확인한 다음 그 제곱을 보십시오.
  • 가격이 아니라 수익률을 쓰고 100을 곱하십시오. 가격은 비정상입니다. 로그 차분을 모형화하고 100배로 스케일하십시오. arch 패키지의 최적화기는 데이터가 0.001 규모일 때 힘들어합니다. 흔한 "수렴 실패"의 정체가 스케일 문제인 경우가 많습니다.
  • 정규 혁신은 꼬리를 과소평가합니다. 가우시안 오차를 쓴 GARCH는 군집은 잡지만 극단적 움직임은 여전히 덜 예측합니다. 변동성을 조건부로 준 뒤에도 수익률이 첨도가 높기 때문입니다. Student-t로 적합하는 편이 거의 항상 더 나은 기술이고 99% 수준의 VaR 숫자를 실질적으로 바꿉니다.
  • 비대칭을 고려하십시오. 기본 GARCH는 충격이 제곱으로 들어가므로 5% 하락과 5% 상승을 똑같이 다룹니다. 주식 변동성은 하락에 훨씬 크게 반응합니다. 레버리지 효과입니다. GJR-GARCH와 EGARCH가 이 항을 더합니다. 주식 데이터에서 비대칭 항은 거의 항상 유의합니다.
  • 무작위 훈련/테스트 분할을 절대 쓰지 마십시오. 변동성 모형은 순차적이고 분할은 시간순이어야 합니다. 교차검증 편의 TimeSeriesSplit 이야기이고 여기서 그 어느 곳보다 중요합니다.
  • 제곱 수익률이 아니라 실현변동성으로 평가하십시오. 하루치 제곱 수익률은 그날 분산의 대리치로는 극도로 잡음이 커서 좋은 모형이 형편없어 보일 수 있습니다. 일중 데이터가 있으면 5분 수익률로 만든 실현변동성이 훨씬 나은 기준입니다. 없으면 RMSE 대신 표본 외 로그우도로 모형을 비교하십시오.
  • Model the mean before the variance. GARCH describes the variance of residuals from some mean equation. A misspecified mean — an omitted trend, an unmodelled lag — leaks into the residuals and inflates the apparent ARCH effect. Fit the mean first, confirm the residuals are serially uncorrelated, then look at their squares.
  • Use returns, not prices, and scale by 100. Prices are non-stationary. Model log differences and multiply by 100: the arch package optimiser struggles when the data are on the order of 0.001, and a common "convergence failure" is nothing more than a scaling problem.
  • Normal innovations understate the tails. GARCH with Gaussian errors captures clustering but still under-predicts extreme moves, because returns are leptokurtic even after conditioning on volatility. Fitting with a Student-t is nearly always a better description and materially changes value-at-risk at the 99% level.
  • Consider asymmetry. Plain GARCH treats a 5% fall and a 5% rise identically, since shocks enter squared. Equity volatility responds far more to falls — the leverage effect. GJR-GARCH and EGARCH add a term for this, and for equity data it is almost always significant.
  • Never use a random train/test split. Volatility models are sequential and the split must be chronological. This is the TimeSeriesSplit point from the cross-validation post, and it matters more here than almost anywhere.
  • Evaluate against realised volatility, not squared returns. A single squared return is an extremely noisy proxy for that day variance, so a good model can look terrible against it. With intraday data, realised volatility from five-minute returns is a far less noisy benchmark. Otherwise compare by out-of-sample log-likelihood rather than RMSE.

어디서 마주치게 되는가Where this shows up

  • 리스크 관리. VaR과 기대손실은 조건부 분산 예측입니다. 군집을 무시한 모형은 위험이 가장 높은 시기에 정확히 위험을 과소평가합니다.
  • 옵션 가격. Black–Scholes는 변동성이 일정하다고 가정합니다. 변동성 스마일은 시장이 그에 동의하지 않는다는 표시입니다. GARCH 옵션 가격 모형은 분산이 진화하도록 둡니다.
  • 포지션 사이징. 변동성 타기팅은 예측 변동성에 반비례해 노출을 조정합니다. 가장 널리 쓰이는 응용 중 하나이고 과거 평균이 아니라 예측이 필요합니다.
  • 거시경제학. 인플레이션과 성장률의 불확실성도 같은 방식으로 모형화합니다. Engle의 1982년 원논문은 자산 수익률이 아니라 영국 인플레이션을 다뤘습니다.
  • Risk management. VaR and expected shortfall are conditional-variance forecasts. A model that ignores clustering understates risk precisely during the periods when risk is highest.
  • Option pricing. Black-Scholes assumes constant volatility, and the volatility smile is the market disagreeing. GARCH option pricing models let the variance evolve.
  • Position sizing. Volatility targeting — scaling exposure inversely to forecast volatility — is one of the most widely used applications, and it needs a forecast rather than a historical average.
  • Macroeconomics. Inflation and output growth uncertainty are modelled the same way. Engle original 1982 paper was about UK inflation, not asset returns.

Engle은 2003년에 이 작업으로 노벨상을 받았습니다. 이유는 틀 자체에 있습니다. 그는 분산을 없애 버려야 할 통계적 방해물이 아니라 그 자체로 예측할 값어치가 있는 경제적 대상으로 다뤘습니다.

Engle received the Nobel Prize in 2003 for this work. The reason is in the framing: he treated variance not as a statistical nuisance to be corrected away, but as an economic object worth forecasting in its own right.