RFS-使用机器学习和大数据预测期权价格-英-55页_2mb
报告摘要
ABSTRACT / INTRODUCTION
This paper examines the predictability of individual U.S. equity option returns (delta-hedged) using machine learning and big data techniques. Drawing from over 12 million observations between 1996 and 2020, the study finds that accounting for nonlinearities significantly boosts out-of-sample predictive performance. Nonlinear machine learning models generate statistically and economically significant profits after accounting for transaction costs, even if the data set is large. While option-based characteristics are the most important standalone predictors, stock-based measures add substantial incremental predictive power.
MACHINE LEARNING MODELS AND PREDICTION
Key findings regarding ML prediction:
- Nonlinear models outperform linear models: Techniques including gradient-boosted trees (GBR, Dart), random forests, and neural networks significantly improve predictive power compared with penalized regression models (Lasso, Ridge) and dimensionality reduction methods (PCR, PLS).
- Best performing models: Dart gradient boosting achieved the highest out-of-sample R² of 2.26%, followed by random forests at 1.96%. Ensembles (L-En for linear, N-En for nonlinear) combine predictions to improve accuracy, with N-En generating an R² of 1.92%.
- Cross-sectional predictability: Prediction quality varies by option bucket (calls, puts, short/long-term, moneyness), but nonlinear models consistently improve predictability. During the COVID-19 period, nonlinear models performed particularly better.
FINDINGS: TRANSACTION COSTS & ECONOMIC SIGNIFICANCE
Investigation on transaction costs showed that even high costs don't invalidate profitability from nonlinear predictions. After accounting for effective spreads (25%-100%) and margin requirements, the nonlinear ensemble consistently yielded significant outperformance relative to linear counterparts. Key trading outcomes include:
- High-minus-low portfolio returns: N-En generated $2.04% monthly with Sharpe ratio 1.28.
- Transaction costs impact: Nonlinear models perform better across cost scenarios, while linear models only yield significant returns at low costs.
FEATURE IMPORTANCE AND PREDICTABILITY MECHANISMS
Important characteristics identified through SHAP analysis:
- Top characteristics: Implied volatility (iv), underlying stock's bid-ask spread (baspread), industry momentum, maturity-specific at-the-money implied volatility, variance risk premium (ivrv).
- Information frictions: Higher informational frictions and mispricing in options and underlying stocks correlate with higher predictability. Delta-hedged options on stocks with high information frictions (quintile 5) show the highest predictability at R²=5.32%.
INFORMATION AND MISPRICING EFFECTS
- Predictability increases with informational frictions, suggesting limits to arbitrage in delta-hedged options.
- Mispricing proxies (composite score) show predictability is higher for mispriced options; high-mispricing options yielded 4.09% monthly R².
REFERENCES
Sourced from the original article's bibliography.
CONCLUSION
The study successfully demonstrates that machine learning techniques capture predictable patterns in equity option returns, especially through nonlinear interactions and effects from multi-level characteristics. These findings have substantial practical relevance for investors heavily using options for hedging or speculation, and provide insights into markets characterized by information frictions and option mispricing.
试读结束,高清完整版pdf/doc/ppt,请点下载