2024-06-10-世界银行-全球劳动力数据库用户手册_全球劳动力数据库的理解_使用和交互指南(英)_143页_2mb
报告摘要
Global Labor Database (GLD) User Manual Summary
Core Content of the GLD
The Global Labor Database (GLD) is a World Bank initiative aimed at harmonizing labor force surveys (LFS) and household surveys with labor modules. It provides a platform for researchers and analysts to access, use, and expand harmonized labor data for cross-country and cross-time comparisons.
Main Objectives
- Create a harmonized database of labor surveys with reliable and comprehensive data for analytical work.
- Allow users to customize and expand the harmonization by providing full access to the codes and documentation, enabling deeper analysis.
Key Features
- Open and transparent approach, allowing users to trace and deviate from the standard harmonization.
- Complements I2D2 and GMD by focusing more on labor market details and including industry and occupation classifications using ISIC and ISCO codes.
- Sustainability through collaboration and community engagement on GitHub.
Main Audience
The intended users of GLD include:
- Researchers and data analysts
- Practitioners in international development
- Statistical offices
- Ministries of labor, economy, and planning
- Government agencies analyzing labor data
Users can leverage GLD in two ways:
- As-is harmonization: Using the harmonized data files directly.
- Amended or hacked harmonization: Editing the harmonization code or adding new variables to suit specific research needs.
Principles of GLD
1. Coverage and Expansion
- As of April 2024, GLD includes 345 surveys from 24 countries (1 high-income, 9 upper middle-income, 11 lower middle-income, and 9 low-income).
- The countries covered are: ARM, BGD, BRA, CHL, COL, EGY, ETH, GEO, IDN, IND, LKA, MEX, MNG, NPL, PAK, PHL, RWA, SLE, THA, TUR, TUN, TZA, ZAF, ZMB, ZWE.
- Surveys are updated to include the latest data (from the previous four years) to ensure timeliness.
- Coverage is balanced across regions and income groups, though limitations may exist due to data availability and NSO restrictions.
2. Transparency and Data Access
- All harmonization steps, including documentation, code, and survey details, are transparent and traceable.
- Outputs such as harmonization code and survey documentation are freely shared on GitHub.
- Raw and harmonized microdata are access-restricted based on data privacy rules and NSO policies.
- Access to raw data is limited to World Bank colleagues or via NSO portals in some cases (e.g., Egypt).
3. Data Quality and Validation
- GLD ensures high-quality data to support reliable cross-country comparisons.
- Three main validation tools are used:
- Manual validation with NSO and country office colleagues.
- Automated checks for data integrity and coherence with external sources (e.g., ILO, WDI).
- Automated checks for consistency across survey series over time.
- Users can report issues through GitHub or direct communication with the GLD team, which will then update the harmonization as needed.
Data Types and Storage
1. Raw Microdata
- Individual-level data as received from NSOs, aggregators, or colleagues.
- May be in formats like CSV, TXT, SPSS, etc.
- Converted to Stata .dta format for consistency.
- Access rights depend on the data sharer and are categorized into:
- Public domain: Free for all users.
- World Bank official use: Accessible within the World Bank.
- Limited release: Access requires permission from the data sharer.
2. Harmonized Microdata
- Output of the harmonization process.
- Contains only variables that have been successfully coded.
- Stored in Stata .dta files.
- Access rights are inherited from raw data.
3. Harmonization Code
- Instructions to convert raw data to harmonized data.
- Stored as Stata .do files.
- Shared under the MIT License.
- Includes comments explaining the rationale behind coding decisions.
4. Other Code
- Includes:
- Raw data conversion code
- Quality check code
- Code templates
- Ecosystem tools code (e.g., for ISIC/ISCO code conversion)
- All code is shared on GitHub and under the MIT License.
5. Survey and Documentation
- Includes:
- Questionnaires
- Enumerator manuals
- Reports
- Stored in PDF or spreadsheet format.
- Country Survey Details document changes in survey definitions (e.g., employment classification) and are accessible to all.
Data Storage Structure
- All data is stored on the GLD Server, organized by country and survey vintage.
- The GLD WB Staff Server is a subset:
- Only includes surveys with freely shareable raw data.
- Only contains the latest harmonized version.
- Folder structure example for Armenia:
ARM_2014_LFS(country and year)ARM_2014_LFS_V01_M(master data version)ARM_2014_LFS_V01_M_V01_A_GLD.dta(harmonized output)ARM_2014_LFS_V01_M/Doc(contains questionnaires, reports, and technical documents)ARM_2014_LFS_V01_M/Programs(contains harmonization code)
Access and Collaboration
- Access to the GLD Server is restricted to the GLD team and World Bank staff under specific conditions.
- The GLD WB Staff Server is accessible to all World Bank staff and is used for easier access to the latest versions.
- GitHub is the primary platform for:
- Sharing code and documentation
- Collaborating on harmonization and data expansion
- Receiving user feedback and suggestions for improvements
- Users can contribute to GLD by:
- Sharing new raw data
- Providing harmonization code for new surveys
- Collaborating on new or expanded harmonizations
- Correcting or expanding existing harmonizations and tools
展开完整摘要
试读结束,高清完整版pdf/doc/ppt,请点下载