Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “multivariate data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Development of a prediction model for radiotherapy response among patients with head and neck squamous cell carcinoma based on the tumor immune microenvironment and hypoxia signature

Abstract Introduction The immune system and hypoxia are major factors influencing radiosensitivity in patients with different cancer types. This study aimed at developing a model to predict radiotherapy response in patients with head and neck squamous cell carcinoma (HNSCC) based on the tumor immune microenvironment and hypoxia signature. Materials and Methods We first evaluated the hypoxia status and tumor immune microenvironment in the Cancer Genome Atlas (TCGA) cohort by using transcriptomic data. Differentially expressed genes (DEGs) were identified between the “high immunity and low hypoxia” and “low immunity and high hypoxia” groups and those DEGs significantly associated with disease‐specific survival in the univariate Cox regression analysis were selected as the prognostic DEGs. We selected the immune hypoxia–related genes (IHRGs) by intersecting prognostic DEGs with immune and hypoxia gene sets. We used the IHRGs to train a multivariate Cox regression model in the TCGA cohort, based on which we calculated the IHRG prognostic index (IHRGPI) for each patient and validated its efficacy in predicting radiotherapy response in the Gene Expression Omnibus cohorts. Furthermore, we explored potential mechanisms and effective combinational treatment strategies for different IHRGPI groups. Results Five IHRGs were used to construct the IHRGPI, which was used to dichotomize the cohorts. The patients with lower IHRGPI showed a better radiotherapy response across different cohorts and endpoints, including overall survival, progression‐free survival, and recurrence‐free survival ( p < 0.05). Patients with higher IHRGPI showed greater hypoxia and lesser immune cell infiltration. A lower IHRGPI indicated a better immunotherapy response, while a higher IHRGPI indicated a better chemotherapy response. Conclusions IHRGPI is promising for predicting radiotherapy response and guiding combinational treatment strategies in patients with HNSCC.

Zhu, Guang‐Li↗

Investigating Kinetic Mechanisms of Soot Formation in Plasma Pyrolysis of Methane via Active Learning (Final Technical Report)

Plasma pyrolysis of methane is an effective route for zero-carbon hydrogen production. Yet, soot generated from pyrolysis of hydrocarbons is detrimental to the climate and human health. There is ample experimental and theoretical evidence that suggests polycyclic aromatic hydrocarbons (PAHs) are the molecular precursors to soot particles. The reaction pathways of PAH formation are intricately dependent on a multitude of process parameters, whose kinetic mechanisms are not well-understood in plasma pyrolysis. This project aims to leverage advances in the kinetic modeling of soot formation in combustion, as well as in surrogate modeling and active learning, to systematically investigate the effects of process parameter on the kinetics of PAH formation in plasma pyrolysis of methane. To this end, we propose to use the PAH formation kinetics model developed by the PPPL/PU group based on the well-established ABF and HACA mechanisms, coupled with low-temperature plasma models. We will develop an active learning (AL) framework based on Bayesian optimization to systematically and data-efficiently explore the complex and multivariable parameter space of plasma pyrolysis in order to quantify the effects of plasma and feed parameters on the ABF and HACA kinetic pathways. AL is the branch of machine learning concerned with systematically querying samples from a system (experimental or computational) to train a data-driven model that maps design parameters to a performance criterion. We will use the data generated via AL to perform global sensitivity analysis, combined with uncertainty quantification, to elucidate the impact of different reaction pathways on minimizing formation of soot precursors. This study will result in an improved understanding of kinetics of PAH formation in plasma pyrolysis and can pave the way for more advanced mechanistic studies (e.g., soot nucleation mechanisms). Additionally, the findings will be useful for establishing practical strategies for increasing the pyrolysis efficiency and producing high-grade carbon for synthesis of nanomaterials.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Search for non-resonant Higgs boson pair production in the $2b+2\ell +{E}_{\textrm{T}}^{\textrm{miss}}$ final state in pp collisions at $\sqrt{s}$ = 13 TeV with the ATLAS detector

A search for non-resonant Higgs boson pair (HH) production is presented, in which one of the Higgs bosons decays to a b-quark pair ($b\bar{b}$) and the other decays to WW * , ZZ * , or τ + τ – , with in each case a final state with ℓ + ℓ – + neutrinos (ℓ = e, μ). The analysis targets separately the gluon-gluon fusion and vector boson fusion production modes. Data recorded by the ATLAS detector in proton-proton collisions at a centre-of-mass energy of 13 TeV at the Large Hadron Collider, corresponding to an integrated luminosity of 140 fb –1 , are used in this analysis. Events are selected to have exactly two b-tagged jets and two leptons with opposite electric charge and missing transverse momentum in the final state. These events are classified using multivariate analysis algorithms to separate the HH events from other Standard Model processes. No evidence of the signal is found. The observed (expected) upper limit on the cross-section for non-resonant Higgs boson pair production is determined to be 9.7 (16.2) times the Standard Model prediction at 95% confidence level. The Higgs boson self-interaction coupling parameter κ λ and the quadrilinear coupling parameter κ 2V are each separately constrained by this analysis to be within the ranges [–6.2, 13.3] and [–0.17, 2.4], respectively, at 95% confidence level, when all other parameters are fixed.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Causal interaction in high frequency turbulence at the biosphere–atmosphere interface: Structure–function coupling

At the biosphere–atmosphere interface, nonlinear interdependencies among components of an ecohydrological complex system can be inferred using multivariate high frequency time series observations. Information flow among these interacting variables allows us to represent the causal dependencies in the form of a directed acyclic graph (DAG). Here, we use high frequency multivariate data at 10 Hz from an eddy covariance instrument located at 25 m above agricultural land in the Midwestern US to quantify the evolutionary dynamics of this complex system using a sequence of DAGs by examining the structural dependency of information flow and the associated functional response. We investigate whether functional differences correspond to structural differences or if there are no functional variations despite the structural differences. We base our analysis on the hypothesis that causal dependencies are instigated through information flow, and the resulting interactions sustain the dynamics and its functionality. To test our hypothesis, we build upon causal structure analysis in the companion paper to characterize the information flow in similarly clustered DAGs from 3-min non-overlapping contiguous windows in the observational data. We characterize functionality as the nature of interactions as discerned through redundant, unique, and synergistic components of information flow. Through this analysis, we find that in turbulence at the biosphere–atmosphere interface, the variables that control the dynamic character of the atmosphere as well as the thermodynamics are driven by non-local conditions, while the scalar transport associated with CO and H 2 O is mainly driven by short-term local conditions.

58 GEOSCIENCES↗

A machine learning approach to galaxy properties: joint redshift–stellar mass probability distributions with Random Forest

We demonstrate that highly accurate joint redshift–stellar mass probability distribution functions (PDFs) can be obtained using the Random Forest (RF) machine learning (ML) algorithm, even with few photometric bands available. As an example, we use the Dark Energy Survey (DES), combined with the COSMOS2015 catalogue for redshifts and stellar masses. We build two ML models: one containing deep photometry in the griz bands, and the second reflecting the photometric scatter present in the main DES survey, with carefully constructed representative training data in each case. We validate our joint PDFs for 10 699 test galaxies by utilizing the copula probability integral transform and the Kendall distribution function, and their univariate counterparts to validate the marginals. Benchmarked against a basic set-up of the template-fitting code bagpipes, our ML-based method outperforms template fitting on all of our predefined performance metrics. In addition to accuracy, the RF is extremely fast, able to compute joint PDFs for a million galaxies in just under 6 min with consumer computer hardware. Such speed enables PDFs to be derived in real time within analysis codes, solving potential storage issues. As part of this work we have developed galpro 1, a highly intuitive and efficient python package to rapidly generate multivariate PDFs on-the-fly. galpro is documented and available for researchers to use in their cosmology and galaxy evolution studies.

79 ASTRONOMY AND ASTROPHYSICS↗

Quantifying Drivers of Methane Hydrobiogeochemistry in a Tidal River Floodplain System

The influence of coastal ecosystems on global greenhouse gas (GHG) budgets and their response to increasing inundation and salinization remains poorly constrained. In this study, we have integrated an uncertainty quantification (UQ) and ensemble machine learning (ML) framework to identify and rank the most influential processes, properties, and conditions controlling methane behavior in a freshwater floodplain responding to recently restored seawater inundation. Our unique multivariate, multiyear, and multi-site dataset comprises tidal creek and floodplain porewater observations encompassing water level, salinity, pH, temperature, dissolved oxygen (DO), dissolved organic carbon (DOC), total dissolved nitrogen (TDN), partial pressure of carbon dioxide (pCO 2 ), nitrous oxide (pN 2 O), methane (pCH 4 ), and the stable isotopic composition of methane (δ 13 CH 4 ). Additionally, we incorporated topographical data, soil porosity, hydraulic conductivity, and water retention parameters for UQ analysis using a previously developed 3D variably saturated flow and transport floodplain model for a physical mechanistic understanding of factors influencing groundwater levels and salinity and, therefore, CH 4 . Principal component analysis revealed that groundwater level and salinity are the most significant predictors of overall biogeochemical variability. The ensemble ML models and UQ analyses identified DO, water level, salinity, and temperature as the most influential factors for porewater methane levels and indicated that approximately 80% of the total variability in hourly water levels and around 60% of the total variability in hourly salinity can be explained by permeability, creek water level, and two van Genuchten water retention function parameters: the air-entry suction parameter α and the pore size distribution parameter m. These findings provide insights on the physicochemical factors in methane behavior in coastal ecosystems and their representation in local- to global-scale Earth system models.

54 ENVIRONMENTAL SCIENCES↗

Contraceptive Sabotage and Contraceptive Use at the Time of Pregnancy: An Analysis of People with a Recent Live Birth in the United States

Contraceptive sabotage and other forms of intimate partner violence (IPV) can interfere with contraceptive use. We used 2012 to 2015 Pregnancy Risk Assessment Monitoring System data from 8,981 people residing in five states who reported that when they became pregnant, they were not trying to get pregnant. We assessed the relationships between ever experiencing contraceptive sabotage and physical IPV 12 months before pregnancy (both by the current partner) and contraceptive use at the time of pregnancy using multivariable logistic regression. We also assessed the joint associations between physical IPV 12 months before pregnancy and ever experienced contraceptive sabotage with contraceptive use at the time of pregnancy. Few people ever experienced contraceptive sabotage (1.8%; 95% confidence interval [CI]: 1.4, 2.3) or physical IPV 12 months before pregnancy (2.8%; 95% CI: 2.3, 3.3). In models adjusted for age, race/ethnicity, marital status, education, and state of residence, ever experiencing contraceptive sabotage was associated with contraceptive use at the time of pregnancy (adjusted odds ratio [aOR]: 1.73; 95% CI: 1.06, 2.82), but not with physical IPV 12 months before pregnancy (aOR: 0.69; 95% CI: 0.46, 1.02). When examining the joint association, compared to not ever experiencing contraceptive sabotage or physical IPV 12 months before pregnancy, ever experiencing contraceptive sabotage was significantly related to contraceptive use at the time of pregnancy (aOR: 1.72; 95% CI: 1.00, 2.95). However, it was not associated with experiencing physical IPV 12 months before pregnancy (aOR: 0.68; 95% CI: 0.45, 1.04) or with experiencing both contraceptive sabotage and physical IPV 12 months before pregnancy (aOR: 1.21; 95% CI: 0.42, 3.50), compared to not ever experiencing contraceptive sabotage or physical IPV 12 months before pregnancy. Our study highlights that current partner contraceptive sabotage may motivate those not trying to get pregnant to use contraception; however, all people in our sample still experienced a pregnancy.

Huber-Krum, Sarah↗

Kernel-based global sensitivity analysis obtained from a single data set

Results from global sensitivity analysis (GSA) often guide the understanding of complicated input–output systems. Kernel-based GSA methods have recently been proposed for their capability of treating a broad scope of complex systems. In this paper, we develop a new set of kernel GSA tools when only a single set of input–output data is available. Three key advances are made: (1) A new numerical estimator is proposed that demonstrates an empirical improvement over previous procedures. (2) A computational method for generating inner statistical functions from a single data set is presented. (3) A theoretical extension is made to define conditional sensitivity indices, which reveal the degree that the inputs carry shared information about the output when inherent input–input correlations are present. Utilizing these conditional sensitivity indices, a decomposition is derived for the output uncertainty based on what is called the optimal learning sequence of the input variables, which remains consistent when correlations exist between the input variables. Further, while these advances cover a range of GSA subjects, a common single data set numerical solution is provided by a technique known as the conditional mean embedding of distributions. The new methodology is implemented on benchmark systems to demonstrate the provided insights.

42 ENGINEERING↗

Node Distortion as a Tunable Mechanism for Negative Thermal Expansion in Metal–Organic Frameworks

Chemically functionalized series of metal–organic frameworks (MOFs), with subtle differences in local structure but divergent properties, provide a valuable opportunity to explore how local chemistry can be coupled to long-range structure and functionality. Using in situ synchrotron X-ray total scattering, with powder diffraction and pair distribution function (PDF) analysis, we investigate the temperature dependence of the local- and long-range structure of MOFs based on NU-1000, in which Zr 6 O 8 nodes are coordinated by different capping ligands (H 2 O/OH, Cl – ions, formate, acetylacetonate, and hexafluoroacetylacetonate). We show that the local distortion of the Zr 6 nodes depends on the lability of the ligand and contributes to a negative thermal expansion (NTE) of the extended framework. Using multivariate data analyses, involving non-negative matrix factorization (NMF), we demonstrate a new mechanism for NTE: progressive increase in the population of a smaller, distorted node state with increasing temperature leads to global contraction of the framework. The transformation between discrete node states is noncooperative and not ordered within the lattice, i.e., a solid solution of regular and distorted nodes. Density functional theory calculations show that removal of ligands from the node can lead to distortions consistent with the Zr···Zr distances observed in the experiment PDF data. Control of the node distortion imparted by the nonlinker ligand in turn controls the NTE behavior. Furthermore, these results reveal a mechanism to control the dynamic structure of MOFs based on local chemistry.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine Learning–Augmented Laser-Induced Breakdown Spectroscopy for Spectral Discrimination of Iron Oxalates

Enhanced characterization and phase identification of post-PUREX Pu Oxalates (PuOXA) are pivotal for nonproliferation and pre-detonation nuclear forensics. Despite significant advances in the characterization of PuO 2 samples, little is known about the impact of both the chemical structure and oxidation states of PuOXA (i.e., Pu(III) and Pu(IV)) have on optical emission signatures. Here, we demonstrate the analytical capabilities of laser-induced breakdown spectroscopy (LIBS) applied to Fe(II) and Fe(III) oxalate samples as surrogates for PuOXA, highlighting the discriminating features in the LIBS emission spectra arising from differences in the oxidation states within mixed FeOXA samples. We report the enhancement of spectral feature selection using Principal Component Analysis (PCA), which enables the analytical superiority of machine learning algorithms such as Linear Discriminant Analysis (LDA), Quadratic Discriminant Analysis (QDA), Partial Least Squares Regression (PLSR), Support Vector Regression (SVR), and Random Forest Regression (RFR) over conventional univariate techniques for phase discrimination and chemometric analysis. Cluster analysis revealed how both matrix effects and laser ablation influence cluster separability by introducing spectral artifacts that misdirect the maximization of variance. PCA-selected emission lines were used in the regression models, demonstrating that both univariate and multivariate linear regression models (i.e., PLSR and SVR) can achieve acceptable performance, with machine learning models outperforming conventional calibration regressions. Furthermore, the application of non-linearly activated PCA-selected emission lines illustrates how simplifying the data while retaining captured variance enables the use of less complex and more computationally efficient models. Furthermore, this is particularly evident in the underperformance of RFR, which suffers from increased computational costs and overfitting owing to its high complexity.

Oxalates↗

An Exploratory Approach Using Regression and Machine Learning in the Analysis of Mass Absorption Cross Section of Black Carbon Aerosols: Model Development and Evaluation

Mass absorption cross-section of black carbon (MAC BC ) describes the absorptive cross-section per unit mass of black carbon, and is, thus, an essential parameter to estimate the radiative forcing of black carbon. Many studies have sought to estimate MAC BC from a theoretical perspective, but these studies require the knowledge of a set of aerosol properties, which are difficult and/or labor-intensive to measure. We therefore investigate the ability of seven data analytical approaches (including different multivariate regressions, support vector machine, and neural networks) in predicting MAC BC for both ambient and biomass burning measurements. Our model utilizes multi-wavelength light absorption and scattering as well as the aerosol size distributions as input variables to predict MAC BC across different wavelengths. We assessed the applicability of the proposed approaches in estimating MAC BC using different statistical metrics (such as coefficient of determination (R 2 ), mean square error (MSE), fractional error, and fractional bias). Overall, the approaches used in this study can estimate MAC BC appropriately, but the prediction performance varies across approaches and atmospheric environments. Based on an uncertainty evaluation of our models and the empirical and theoretical approaches to predict MAC BC , we preliminarily put forth support vector machine (SVM) as a recommended data analytical technique for use. We provide an operational tool built with the approaches presented in this paper to facilitate this procedure for future users.

54 ENVIRONMENTAL SCIENCES↗

Upscaling Soil Organic Carbon Measurements at the Continental Scale Using Multivariate Clustering Analysis and Machine Learning

Abstract Estimates of soil organic carbon (SOC) stocks are essential for many environmental applications. However, significant inconsistencies exist in SOC stock estimates for the U.S. across current SOC maps. We propose a framework that combines unsupervised multivariate geographic clustering (MGC) and supervised Random Forests regression, improving SOC maps by capturing heterogeneous relationships with SOC drivers. We first used MGC to divide the U.S. into 20 SOC regions based on the similarity of covariates (soil biogeochemical, bioclimatic, biological, and physiographic variables). Subsequently, separate Random Forests models were trained for each SOC region, utilizing environmental covariates and SOC observations. Our estimated SOC stocks for the U.S. (52.6 ± 3.2 Pg for 0–30 cm and 108.3 ± 8.2 Pg for 0–100 cm depth) were within the range estimated by existing products like Harmonized World Soil Database, HWSD (46.7 Pg for 0–30 cm and 90.7 Pg for 0–100 cm depth) and SoilGrids 2.0 (45.7 Pg for 0–30 cm and 133.0 Pg for 0–100 cm depth). However, independent validation with soil profile data from the National Ecological Observatory Network showed that our approach ( R 2 = 0.51) outperformed the estimates obtained from Harmonized World Soil Database ( R 2 = 0.23) and SoilGrids 2.0 ( R 2 = 0.39) for the topsoil (0–30 cm). Uncertainty analysis (e.g., low representativeness and high coefficients of variation) identified regions requiring more measurements, such as Alaska and the deserts of the U.S. Southwest. Our approach effectively captures the heterogeneous relationships between widely available predictors and the current SOC baseline across regions, offering reliable SOC estimates at 1 km resolution for benchmarking Earth system models.

58 GEOSCIENCES↗

Defining Golden Batches in Biomanufacturing Processes From Internal Metabolic Activity to Detect Process Changes That May Affect Product Quality

ABSTRACT Cellular metabolism plays a role in the observed variability of a drug substance's Critical Quality Attributes (CQAs) made by biomanufacturing processes. Therefore, here we describe a new approach for monitoring biomanufacturing processes that measures a set of metabolic reaction rates (named Critical Metabolic Parameters (CMP) in addition to the macroscopic process conditions currently being used as Critical Process Parameters (CPP) for biomanufacturing. Constraint‐based systems biology models like Flux Balance Analysis (FBA) are used to estimate metabolic reaction rates, and metabolic rates are used as inputs for multivariate Batch Evolution Models (BEM). Metabolic activity was reproducible among batches and could be monitored to detect a deliberately induced macroscopic process shift (i.e., temperature change). The CMP approach has the potential to enable “golden batches” in biomanufacturing processes to be defined from the internal metabolic activity and to aid in detecting process changes that may impact the quality of the product. Overall, the data suggested that monitoring of metabolic activity has promise for biomanufacturing process control.

Biotechnology & Applied Microbiology↗

Observation of four top quark production in proton-proton collisions at s = 13 TeV

The observation of the production of four top quarks in proton-proton collisions is reported, based on a data sample collected by the CMS experiment at a center-of-mass energy of 13 TeV in 2016–2018 at the CERN LHC and corresponding to an integrated luminosity of 138 fb − 1 . Events with two same-sign, three, or four charged leptons (electrons and muons) and additional jets are analyzed. Compared to previous results in these channels, updated identification techniques for charged leptons and jets originating from the hadronization of b quarks, as well as a revised multivariate analysis strategy to distinguish the signal process from the main backgrounds, lead to an improved expected signal significance of 4.9 standard deviations above the background-only hypothesis. Four top quark production is observed with a significance of 5.6 standard deviations, and its cross section is measured to be 17.7 − 3.5 + 3.7 (stat) − 1.9 + 2.3 (syst) fb , in agreement with the available standard model predictions.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Representations of Precipitation Diurnal Cycle in the Amazon as Simulated by Observationally Constrained Cloud‐System Resolving and Global Climate Models

Abstract The ability of an observationally‐constrained cloud‐system resolving model (Weather Research and Forecasting; WRF, 4‐km grid spacing) and a global climate model (Energy Exascale Earth System Model; E3SM, 1‐degree grid spacing) to represent the precipitation diurnal cycle over the Amazon basin during the 2014 wet season is assessed. The WRF model coupled with a 3‐D variational data assimilation scheme reproduces the spatial variability of the precipitation diurnal cycle over the basin and the lifecycle of westward propagating MCSs initiated by the coastal sea‐breeze front. In contrast, a single morning peak in rainfall is produced by E3SM for simulations despite the nudging of large‐scale winds toward global reanalysis, indicating precipitation in E3SM is largely controlled by local convection associated with diurnal heating. The role of propagating MCS on the environment are discussed by using a multivariate perturbation analysis. We also find that the advection of moisture perturbations from ocean to inland regions have a higher correlation with the occurrence of MCSs in the Amazon than the intensity of colder air intrusion associated with sea breezes along the coast. Moreover, the presence of large cold pools over the central Amazon basin are responsible for the maintenance of propagating deep convection.

54 ENVIRONMENTAL SCIENCES↗

Latent-Space Dynamics for Prediction and Fault Detection in Geothermal Power Plant Operations

This paper presents a latent-space dynamic neural network (LSDNN) model for the multi-step-ahead prediction and fault detection of a geothermal power plant’s operation. The model was trained to learn the dynamics of the power generation process from multivariate time-series data and the effects of exogenous variables, such as control adjustment and ambient temperature. In the LSDNN model, an encoder–decoder architecture was designed to capture cross-correlation among different measured variables. In addition, a latent space dynamic structure was proposed to propagate the dynamics in the latent space to enable prediction. The prediction power of the LSDNN was utilized for monitoring a geothermal power plant and detecting abnormal events. The model was integrated with principal component analysis (PCA)-based process monitoring techniques to develop a fault-detection procedure. The performance of the proposed LSDNN model and fault detection approach was demonstrated using field data collected from a geothermal power plant.

15 GEOTHERMAL ENERGY↗

Search for charged Higgs bosons decaying into a top quark and a bottom quark at $\sqrt{s}$ = 13 TeV with the ATLAS detector

A search for charged Higgs bosons decaying into a top quark and a bottom quark is presented. The data analysed correspond to 139 fb -1 of proton-proton collisions at $\sqrt{s}$ = 13 TeV, recorded with the ATLAS detector at the LHC. The production of a heavy charged Higgs boson in association with a top quark and a bottom quark, pp → tbH + → tbtb, is explored in the H + mass range from 200 to 2000 GeV using final states with jets and one electron or muon. Events are categorised according to the multiplicity of jets and b-tagged jets, and multivariate analysis techniques are used to discriminate between signal and background events. No significant excess above the background-only hypothesis is observed and exclusion limits are derived for the production cross-section times branching ratio of a charged Higgs boson as a function of its mass; they range from 3.6 pb at 200 GeV to 0.036 pb at 2000 GeV at 95% confidence level. The results are interpreted in the hMSSM and M$_{h}^{125}$ scenarios.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Toward Quantity-of-Interest Preserving Lossy Compression for Scientific Data

Today's scientific simulations and instruments are producing a large amount of data, leading to difficulties in storing, transmitting, and analyzing these data. While error-controlled lossy compressors are effective in significantly reducing data volumes and efficiently developing databases for multiple scientific applications, they mainly support error controls on raw data, which leaves a significant gap between the data and user's downstream analysis. This may cause unqualified uncertainties in the outcomes of the analysis, a.k.a quantities of interest (QoIs), which are the major concerns of users in adopting lossy compression in practice. In this paper, we propose rigorous mathematical theories to preserve four families of QoIs that are widely used in scientific analysis during lossy compression along with practical implementations. Specifically, we first develop the error control theory for univariate QoIs which are essential for computing physical properties such as kinetic energy, followed by multivariate QoIs that are more commonly used in real-world applications. The proposed method is integrated into a state-of-the-art compression framework in a modular fashion, which could easily adapt to new QoIs and new compression algorithms. Experiments on real-world datasets demonstrate that the proposed method provides faithful error control on important QoIs including kinetic energy, regional average, and isosurface without trials and errors, while offering compression ratios that are up to 4x of the compression ratios provided by state-of-the-art compressors.

Jiao, Pu↗