亚开行-在卫星影像上应用人工智能编制贫困统计数据(英文)-2020.12-28页_4mb
报告摘要
Summary of "Applying Artificial Intelligence on Satellite Imagery to Compile Granular Poverty Statistics"
Core Content
This paper explores the use of artificial intelligence (AI), specifically convolutional neural networks (CNNs) and ridge regression, to improve the spatial granularity of poverty statistics using satellite imagery. The study focuses on the Philippines and Thailand, where small area poverty estimates are available, and aims to provide more detailed poverty maps than those traditionally produced by government surveys.
The primary motivation is to address the challenges of data granularity and timeliness in poverty statistics, as conventional household surveys often lack the necessary sample size to produce reliable estimates at the subnational level, and the release of updated data can be delayed due to the high cost and time required for data collection.
Main Views and Key Information
1. Spatial Granularity and SDGs
- The Sustainable Development Goals (SDGs) require poverty data to be disaggregated by geographic, ethnic, gender, and income-related dimensions.
- Traditional methods using household surveys are not sufficient for such granularity, especially in developing countries.
- AI-based methods offer a cost-effective alternative for generating more detailed poverty estimates.
2. Use of Satellite Imagery
- Daytime satellite images (e.g., Landsat and Sentinel) and nighttime light intensity data (e.g., from VIIRS) are used as inputs.
- Nighttime light intensity is used as a proxy for economic activity and wealth.
- Daytime images are used to train the model to extract features that can be used for poverty prediction.
3. Methodology Overview
- The method involves:
- Training a CNN to predict nighttime light intensity using daytime satellite images.
- Extracting features from the last layer of the CNN.
- Averaging features across SAE-level areas (municipalities, cities, tambons).
- Using ridge regression to find the relationship between these features and ground truth poverty data.
- Predicting poverty at grid-level using only daytime images.
4. Data and Tools
- Satellite data used:
- Landsat 8 (15-meter resolution)
- Sentinel 2 (10-meter resolution)
- Nightlight data:
- DMSP-OLS and SNPP-VIIRS
- Ground truth data:
- Philippines: SAE-level poverty estimates from the Philippine Statistics Authority (PSA) for 2012, 2015, and 2018.
- Thailand: SAE-level poverty estimates for 2013, 2015, and 2017.
- Shapefiles are used to define the geographical boundaries of the areas being studied.
5. Key Findings
- CNN performance:
- The model was trained using ResNet34 and weighted cross entropy loss to handle class imbalance.
- Data augmentation was used to improve model generalization and reduce overfitting.
- Validation accuracy for CNN predictions:
- Thailand: 85.79% (2013), 85.22% (2015), 86.43% (2017)
- Philippines: 94.15% (2012), 93.50% (2015), 92.91% (2018)
- Full dataset accuracy:
- Thailand: 86.52%, 87.28%, 86.98%
- Philippines: 94.66%, 93.86%, 90.56%
- Ridge regression:
- Used to relate image features to ground truth poverty data.
- The model was validated against published poverty rates and showed alignment after calibration.
- Robustness of the model was tested using different data splitting strategies and machine learning algorithms.
6. Robustness Assessment
- The study tested the robustness of the model to:
- Algorithmic parameters (e.g., learning rate, number of epochs)
- Model specifications (e.g., using ridge regression versus random forest)
- Data splitting strategies (e.g., 90% training, 10% validation)
- It was found that changes in parameters and strategies did not significantly affect the model's performance.
- Harmonizing AI-based predictions with published poverty rates showed that the model could produce reliable estimates when properly calibrated.
Conclusion
The study provides a computational framework for generating granular poverty maps using AI and satellite imagery, which can be used to support development planning and target resources more efficiently. The results suggest that publicly available satellite data, even with medium resolution, can be used to produce accurate poverty estimates when combined with household survey data and machine learning techniques.
The methodology can be applied in other countries and regions with similar data availability, and it offers a potential solution to the challenges of data granularity and timeliness in poverty statistics.
References
- Chen, J., & Nordhaus, W. (2011)
- Henderson, J., Storeygard, A., & Weil, D. (2012)
- Jean, S., et al. (2016)
- Xie, Y., et al. (2015)
- Elvidge, C. D., et al. (2013)
- Deng, J., et al. (2009)
- Perez, A., & Wang, J. (2017)
- Tingzon, M., et al. (2019)
- Yeh, S., et al. (2020)
- ADB (2020)
Tables and Figures
Tables
- Table 1: Prediction accuracy of CNNs for different country-year combinations.
- Table 2: Root mean square error for poverty rate estimates.
- Table 3: Prediction accuracy of CNNs using alternative data splitting strategies.
- Table 4: Comparison of predictive performance of ridge regression and random forest.
- Table 5: Root mean square error for poverty rate (validation) and survey level.
Figures
- Figure 1: Methodology for predicting poverty using satellite imagery.
Authors and Affiliations
- Martin Hofer (Vienna University of Economics and Business)
- Tomas Sako (Freelance Data Scientist)
- Arturo Martinez Jr. (Asian Development Bank)
- Joseph Bulan (Asian Development Bank)
- Mildred Addawe, Ron Lester Durante, Marymell Martillan (Consultants at ADB)
Publication Information
- ADB Economics Working Paper Series No. 629
- Date: December 2020
- License: Creative Commons Attribution 3.0 IGO
- ISSN: 2313-5867 (print), 2313-5875 (electronic)
- Stock No.: WPS200432-2
- DOI: http://dx.doi.org/10.22617/WPS200432-2
Keywords
- Big data
- Computer vision
- Data for development
- Machine learning algorithm
- Official statistics
- Poverty
- SDG
JEL Codes
- C19
- D31
- I32
- O15
试读结束,高清完整版pdf/doc/ppt,请点下载