APL-冠状病毒检测的操作分析(英)-2021-56页_9mb
报告摘要
Summary of Operational Analysis for Coronavirus Testing
Core Content
This document provides an operational analysis of coronavirus testing, emphasizing the importance of understanding the relationship between surface positivity (observed positive test results) and the true incidence rate of the disease, while accounting for test errors (false negatives and false positives). It builds on a companion report that outlines methods for translating surface positivity into incidence rate estimates, then into risk assessments for groups of varying sizes.
The analysis is grounded in statistical modeling and simulation, using mathematical tools such as the binomial distribution and Gaussian approximation to derive accurate and precise estimates of the incidence rate and associated risks.
Main Purpose
The main purpose of this report is to support the recommendations from the companion report by:
- Demonstrating how to compute the incidence rate from surface positivity and test error rates.
- Showing how to derive a range of possible incidence rates (the test range) based on the number of tests.
- Analyzing the accuracy and precision of incidence rate estimates under different test error assumptions.
- Introducing a certainty equivalent (CE) approximation to handle uncertainty in test error rates.
- Addressing the case where surface positivity is zero, which is common as the pandemic wanes.
Key Concepts and Equations
1. Surface Positivity and Incidence Rate
- Surface Positivity: Observed fraction of positive tests, $ p = \frac{P}{T} $, is not equal to the true incidence rate $ f_t $ due to test errors.
- Expected Positivity Rate:
$$
p_+(f_t) = f_t(1 - p_{FN}) + (1 - f_t)p_{FP}
$$
This equation shows how the expected positivity rate is influenced by both true positives and false positives.
2. Incidence Rate Estimate
- The estimate of the incidence rate $ \hat{f} $, assuming test error rates $ p_{FN} $ and $ p_{FP} $, is given by:
$$
\hat{f} = \frac{P/T - p_{FP}}{1 - p_{FN} - p_{FP}}
$$
This is the maximum likelihood estimate (MLE) of the incidence rate.
3. Test Range
- The test range, which defines the range of incidence rates compatible with the data and model, is:
$$
\text{Range}(\hat{f}) = 3.92 \sqrt{\frac{p_+(\hat{f})(1 - p_+(\hat{f}))}{T(1 - p_{FN} - p_{FP})^2}}
$$
This provides a symmetric range around the estimated incidence rate $ \hat{f} $, with lower and upper bounds defined as:
$$
\hat{f}{lower} = \hat{f} - 0.5 \cdot \text{Range}(\hat{f}), \quad \hat{f}{upper} = \hat{f} + 0.5 \cdot \text{Range}(\hat{f})
$$
4. Risk Assessment
-
The risk $ \mathcal{R}(g, \hat{f}) $ that a group of size $ g $ contains at least one infected individual is:
$$
\mathcal{R}(g, \hat{f}) = 1 - (1 - \hat{f})^g
$$
This equation is used to calculate the risk for different group sizes based on the estimated incidence rate. -
The maximum group size $ g(\hat{f}, \mathcal{R}{acc}) $ that is consistent with a specified level of acceptable risk $ \mathcal{R}{acc} $ is:
$$
g(\hat{f}, \mathcal{R}{acc}) = \frac{\ln(1 - \mathcal{R}{acc})}{\ln(1 - \hat{f})}
$$
5. Certainty Equivalent (CE) Approximation
- The CE approximation replaces the stochastic surface positivity with its mean, allowing for more accurate and precise estimates of the incidence rate.
- This approach is used to account for uncertainty in test error rates and to derive the CE estimate for the incidence rate:
$$
\hat{f}{CE} = \frac{p{CE} - p_{FP}}{1 - p_{FN} - p_{FP}}
$$
where $ p_{CE} $ is the mean of the surface positivity.
Main Findings
- Surface positivity is not an accurate indicator of the true incidence rate due to test errors.
- False negatives lead to overestimation of the incidence rate, while false positives lead to underestimation.
- The risk of a group is not an either-or situation but depends on the incidence rate, group size, and acceptable risk level.
- The MLE for the incidence rate is accurate for even a modest number of tests, but its precision is low due to high variance.
- The standard deviation of the MLE and test range declines as $ 1 / \sqrt{T} $, so with sufficient tests, both accuracy and precision can be achieved.
- Oversampling can occur when test ranges are too narrow, leading to inefficient allocation of testing resources.
- When no positive tests are observed, the maximum incidence rate consistent with this result can be calculated, and this is used in risk assessments.
- The CE approximation confirms simulation results and allows for more robust risk calculations when test errors are uncertain.
Additional Recommendation
- Avoid relying solely on surface positivity to assess the true incidence rate. Always consider test error rates and use the MLE or CE approximation for more accurate estimates.
Supporting Analysis
- The report uses simulation and analytical methods to validate the MLE and test range formulas.
- 27 additional sets of parameters are used to test the robustness of the findings, confirming that the qualitative patterns remain consistent across different scenarios.
- The binomial distribution and its likelihood are used to model test outcomes, with the Gaussian approximation providing a close match for large numbers of tests.
Conclusion
This report provides a comprehensive framework for interpreting coronavirus test data, accounting for test errors, and deriving accurate and precise estimates of the incidence rate and associated risks. It supports the recommendations in the companion report and offers a practical approach for policy makers and health care managers to use when making decisions based on test data.
试读结束,高清完整版pdf/doc/ppt,请点下载