Engineering Papers⌕ Search

DOE OSTI · 1982504

Private Tabular Survey Data Products through Synthetic Microdata Generation

Abstract

We propose two synthetic microdata approaches to generate private tabular survey data products for public release. We adapt a pseudo posterior mechanism that downweights by-record likelihood contributions with weights ∈[0,1] based on their identification disclosure risks to producing tabular products for survey data. Our method applied to an observed survey database achieves an asymptotic global probabilistic differential privacy guarantee. Our two approaches synthesize the observed sample distribution of the outcome and survey weights, jointly, such that both quantities together possess a privacy guarantee. The privacy-protected outcome and survey weights are used to construct tabular cell estimates (where the cell inclusion indicators are treated as known and public) and associated standard errors to correct for survey sampling bias. Through a real data application to the Survey of Doctorate Recipients public use file and simulation studies motivated by the application, we demonstrate that our two microdata synthesis approaches to construct tabular products provide superior utility preservation as compared to the additive noise approach of the Laplace Mechanism. Moreover, our approaches allow the release of microdata to the public, enabling additional analyses at no extra privacy cost.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Hu, Jingchen, Savitsky, Terrance D., Williams, Matthew R. (ORCID:0000000188941240). 2022-03-03. Private Tabular Survey Data Products through Synthetic Microdata Generation. https://doi.org/10.1093/jssam%2Fsmac001

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related reports

Direction of impact for explainable risk assessment modeling

Abstract Several graphical indicators have been recently introduced to help analysts visualize the marginal effects of inputs in complex models. The insights derived from such tools may help decision‐makers and risk analysts in designing interventions. However, we know little about the adequacy and consistency of different indicators. This work investigates popular marginal effect indicators to understand whether they yield indications consistent with the properties of the quantitative model under inspection. Specifically, we examine the notions of monotonicity, Lipschitz, and concavity consistency. Surprisingly, only PD functions satisfy all these notions of consistency. However, when selecting the indicators, in addition to consistency, analysts need to consider the risk of model extrapolation. For situations where such risk is under control, we utilize individual conditional expectations together with PD plots. Two applications, on a NASA space risk assessment model and a susceptible exposed infected recovered (SEIR) model for the COVID‐19 pandemic illustrate the insights obtained from these indicators.

Mathematical Methods In Social Sciences↗

Enhancing risk and crisis communication with computational methods: A systematic literature review

Abstract Recent developments in risk and crisis communication (RCC) research combine social science theory and data science tools to construct effective risk messages efficiently. However, current systematic literature reviews (SLRs) on RCC primarily focus on computationally assessing message efficacy as opposed to message efficiency. We conduct an SLR to highlight any current computational methods that improve message construction efficacy and efficiency. We found that most RCC research focuses on using theoretical frameworks and computational methods to analyze or classify message elements that improve efficacy. For improving message efficiency, computational and manual methods are only used in message classification. Specifying the computational methods used in message construction is sparse. We recommend that future RCC research apply computational methods toward improving efficacy and efficiency in message construction. By improving message construction efficacy and efficiency, RCC messaging would quickly warn and better inform affected communities impacted by current hazards. Such messaging has the potential to save as many lives as possible.

Mathematical Methods In Social Sciences↗

On the compatibility of established methods with emerging artificial intelligence and machine learning methods for disaster risk analysis

Abstract There is growing interest in leveraging advanced analytics, including artificial intelligence (AI) and machine learning (ML), for disaster risk analysis (RA) applications. These emerging methods offer unprecedented abilities to assess risk in settings where threats can emerge and transform quickly by relying on “learning” through datasets. There is a need to understand these emerging methods in comparison to the more established set of risk assessment methods commonly used in practice. These existing methods are generally accepted by the risk community and are grounded in use across various risk application areas. The next frontier in RA with emerging methods is to develop insights for evaluating the compatibility of those risk methods with more recent advancements in AI/ML, particularly with consideration of usefulness, trust, explainability, and other factors. This article leverages inputs from RA and AI experts to investigate the compatibility of various risk assessment methods, including both established methods and an example of a commonly used AI‐based method for disaster RA applications. This article utilizes empirical evidence from expert perspectives to support key insights on those methods and the compatibility of those methods. This article will be of interest to researchers and practitioners in risk‐analytics disciplines who leverage AI/ML methods.

Mathematical Methods In Social Sciences↗