Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Predictive Data analytics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Resource selection functions based on hierarchical generalized additive models provide new insights into individual animal variation and species distributions

Habitat selection studies are designed to generate predictions of species distributions or inference regarding general habitat associations and individual variation in habitat use. Such studies frequently involve either individually indexed locations gathered across limited spatial extents and analyzed using resource selection functions (RSFs) or spatially extensive locational data without individual resolution typically analyzed using species distribution models. Both analytical methodologies have certain desirable features, but analyses that combine individual- and population-level inference with flexible non-linear functions may provide improved predictions while accounting for individual variation. Here, we describe how RSFs can be fit using hierarchical generalized additive models (HGAMs) using widely available software, providing a means to explore individual variation in habitat associations and to generate species distribution maps. We used GPS tracking data from golden eagles Aquila chrysaetos from across eastern North America with four environmental predictors to generate monthly distribution models. We considered three model structures that assumed different amounts of individual variation in the functional relationship between predictors and habitat use and used k-fold cross-validation to compare model performance. Models accounting for individual variability in shape and smoothness of functional responses performed best. Eagles exhibited the least amount of individual variation in response to land cover variables during winter months, with most individuals more closely adhering to the population-level trend. During the summer months, eagles exhibited more substantial individual variation in shape and smoothness of the functional relationships, suggesting some need to account for individual variation in eagle habitat use for both inferential and predictive purposes, during this time of year. Because they allow users to blend flexible functions with random effects structures and are well-supported by a variety of software platforms, we believe that HGAMs provide a useful addition to the suite of analyses used for modeling habitat associations or predicting species distributions.

54 ENVIRONMENTAL SCIENCES↗

Integrated End-to-end Performance Prediction and Diagnosis for Extreme Scientific Workflows

This report details recent progress for the ASCR funded project “Integrated End-to-end Performance Prediction and Diagnosis for Extreme Scientific Workflows”. We refer to the project as IPPD/2, reflecting the 2017 renewal under expanded scope and partners In IPPD/2, we increased our research scope to include data motion. We are focusing on three major aspects: a) observe how data is generated, distributed, and used; b) analyze how data is (repeatedly) consumed with a focus both on repeated patterns and anomalies; and c) explore how to optimize data motion. This new work on data motion will augment and complement IPPD/2’s research that focused on the computational aspects of tasks. We leverage and extend our existing tools and demonstrate our work on the Belle II workflow suite as well as on workflows from NSLS-II. The highlights of our work are as follows: Provenance for Workflows: Provenance is used to provide information enabling quality control, re-run computational workflows, and reproduce results. IPPD/2 has been building a scalable provenance management system that enables the capture of provenance from the high-level workflow through all relevant system levels in one integrated environment. Leveraging this work, our recent efforts have included using provenance as an enabling technique. Workload characterization: Leveraging provenance and analysis, we characterize data movement within network, storage, and memory over a variety of workloads. This characterization enables an understanding by performance analysts and application developers of the range of behaviors that could be expected. Performance Prediction for Workflows: The goal of modeling distributed workflows is to understand performance bottlenecks and enable more intelligent task scheduling to optimize selected metrics of interest (e.g., task throughput or output data rate). IPPD/2 has utilized both analytical and AI/ML modeling methodologies for performance modeling. Advanced Scheduling and Fault Modeling for Workflows: Scheduling of large-scale scientific workflows on geographically distributed resources is a challenging problem. To improve workflow throughput, we combined novel scheduling algorithms with task predictions from performance modeling and fault modeling. Dynamically Alleviating Bottlenecks in Workflows: Exploiting our provenance, analysis, and modeling efforts, we have explored and developed several techniques for dynamically detecting and alleviating bottlenecks in data movement. In particular, we have spent considerable effort demonstrating our techniques on production-like workflow configurations.

97 MATHEMATICS AND COMPUTING↗

Measurements of the groomed and ungroomed jet angularities in pp collisions at $ \sqrt{s}$ = 5.02 TeV

The jet angularities are a class of jet substructure observables which characterize the angular and momentum distribution of particles within jets. These observables are sensitive to momentum scales ranging from perturbative hard scatterings to nonperturbative fragmentation into final-state hadrons. We report measurements of several groomed and ungroomed jet angularities in pp collisions at $\sqrt{s}$ = 5.02 TeV with the ALICE detector. Jets are reconstructed using charged particle tracks at midrapidity (|η| < 0.9). The anti-$k_T$ algorithm is used with jet resolution parameters R = 0.2 and R = 0.4 for several transverse momentum $p^{ch jet}_{T}$ intervals in the 20–100 GeV/c range. Using the jet grooming algorithm Soft Drop, the sensitivity to softer, wide-angle processes, as well as the underlying event, can be reduced in a way which is well-controlled in theoretical calculations. We report the ungroomed jet angularities, λ α , and groomed jet angularities, λ α,g , to investigate the interplay between perturbative and nonperturbative effects at low jet momenta. Various angular exponent parameters α = 1, 1.5, 2, and 3 are used to systematically vary the sensitivity of the observable to collinear and soft radiation. Results are compared to analytical predictions at next-to-leading-logarithmic accuracy, which provide a generally good description of the data in the perturbative regime but exhibit discrepancies in the nonperturbative regime. Moreover, these measurements serve as a baseline for future ones in heavy-ion collisions by providing new insight into the interplay between perturbative and nonperturbative effects in the angular and momentum substructure of jets. They supply crucial guidance on the selection of jet resolution parameter, jet transverse momentum, and angular scaling variable for jet quenching studies.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Observing the onset of pressure-driven K-shell delocalization

The gravitational pressure in many astrophysical objects exceeds one gigabar (one billion atmospheres), creating extreme conditions where the distance between nuclei approaches the size of the K shell. This close proximity modifies these tightly bound states and, above a certain pressure, drives them into a delocalized state. Both processes substantially affect the equation of state and radiation transport and, therefore, the structure and evolution of these objects. Still, our understanding of this transition is far from satisfactory and experimental data are sparse. Here, in this work, we report on experiments that create and diagnose matter at pressures exceeding three gigabars at the National Ignition Facility where 184 laser beams imploded a beryllium shell. Bright X-ray flashes enable precision radiography and X-ray Thomson scattering that reveal both the macroscopic conditions and the microscopic states. The data show clear signs of quantum-degenerate electrons in states reaching 30 times compression, and a temperature of around two million kelvins. At the most extreme conditions, we observe strongly reduced elastic scattering, which mainly originates from K-shell electrons. We attribute this reduction to the onset of delocalization of the remaining K-shell electron. With this interpretation, the ion charge inferred from the scattering data agrees well with ab initio simulations, but it is significantly higher than widely used analytical models predict.

79 ASTRONOMY AND ASTROPHYSICS↗

A database and meta-analysis on the performance of exploding pusher implosions conducted at OMEGA

A database of 222 exploding pusher implosions conducted at the OMEGA Laser Facility is presented. The dataset consists of glass-shell capsules filled with varying pressures of D 2 , T 2 , and 3 He, which were imploded using square laser pulses with intensities ranging from 1 to 1 × 10 15 W/cm 2 . The database includes measurements of bang times, ion temperatures, and yields from the DD, D 3 He, and DT fusion reactions. A semi-analytic exploding pusher model is introduced, which effectively captures the observed trends in the data. This model predicts that the measurements scale according to a power-law relation based on the initial capsule and laser conditions. A generalized power-law scaling relation is directly fit to each dataset, providing a useful interpolation of the entire database. Overall, the database provides a valuable resource to estimating bang times, temperatures, and yields for the design of future experiments. Additionally, it provides a diverse set of data for validating more advanced implosion physics models.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Local Weather Station Design and Development for Cost-Effective Environmental Monitoring and Real-Time Data Sharing

Current weather monitoring systems often remain out of reach for small-scale users and local communities due to their high costs and complexity. This paper addresses this significant issue by introducing a cost-effective, easy-to-use local weather station. Utilizing low-cost sensors, this weather station is a pivotal tool in making environmental monitoring more accessible and user-friendly, particularly for those with limited resources. It offers efficient in-site measurements of various environmental parameters, such as temperature, relative humidity, atmospheric pressure, carbon dioxide concentration, and particulate matter, including PM 1, PM 2.5, and PM 10. The findings demonstrate the station’s capability to monitor these variables remotely and provide forecasts with a high degree of accuracy, displaying an error margin of just 0.67%. Furthermore, the station’s use of the Autoregressive Integrated Moving Average (ARIMA) model enables short-term, reliable forecasts crucial for applications in agriculture, transportation, and air quality monitoring. Furthermore, the weather station’s open-source nature significantly enhances environmental monitoring accessibility for smaller users and encourages broader public data sharing. With this approach, crucial in addressing climate change challenges, the station empowers communities to make informed decisions based on real-time data. In designing and developing this low-cost, efficient monitoring system, this work provides a valuable blueprint for future advancements in environmental technologies, emphasizing sustainability. The proposed automatic weather station not only offers an economical solution for environmental monitoring but also features a user-friendly interface for seamless data communication between the sensor platform and end users. This system ensures the transmission of data through various web-based platforms, catering to users with diverse technical backgrounds. Furthermore, by leveraging historical data through the ARIMA model, the station enhances its utility in providing short-term forecasts and supporting critical decision-making processes across different sectors.

54 ENVIRONMENTAL SCIENCES↗

Development of Explainable, Knowledge-Guided AI Models to Enhance the E3SM Land Model Development and Uncertainty Quantification

Focal Area(s): (2)Predictive modeling using AI techniques and AI-derived model components; use of AI and other tools to design a prediction system comprising of a hierarchy of models. (3) Insight gleaned from complex data (both observed and simulated) using AI, big data analytics, and other advanced methods, including explainable AI and physics- or knowledge- guided AI. Science Challenge: The Energy Exascale Earth System Model (E3SM) is a fully coupled, state-of-the-science Earth system model that uses code optimized for DOE's advanced computers to address the most critical scientific questions facing our nation and society (Golaz et al., 2019). The E3SM Land model (ELM) is designed to understand how the changes in terrestrial land surfaces will interact with other Earth system components and has been used to understand hydrologic cycles, biogeophysics, and ecosystem dynamics. In spite of great successes, the ELM has several known issues that restrain rapid improvements. For example, the ELM uses equilibrium models to simulate dynamic land-climate interactions and it requires long model spin-up time to identify suitable initial conditions for transient simulations. The ELM lacks built-in uncertainty mechanisms that can improve the robustness of model predictions. The ELM is a holistic, deterministic model system with a rigid design, and in many situations, it is hard to modify the ELM system to incorporate new theory/hypothesis and new data across scales to address emerging science problems (such as predicting the impacts of water cycle extremes). In addition, The ELM is technically optimized for traditional CPU-centric computers and it cannot fully utilize the current and incoming leadership computers for model simulations and uncertainty quantification (UQ). The success of artificial intelligence (AI) has inspired scientists to use AI models to discover intrinsic features from simulation data (Chattopadhyay et al., 2020) and observational data (Reichstein et al., 2019) to gain further process understanding of Earth science problems. However, autonomous AI model training through deep learning usually requires a huge amount of annotated data. To overcome the limitations from the data and computing resources, knowledge-guided AI models are necessary where human-knowledge is ingested in model construction (Banino et al., 2018) and training process (Silver et al., 2016) for efficient learning. Herein, we present a new way that leverages the process understanding from the ELM to guide AI model development for the ELM enhancement and UQ. We hope this study can inspire further Earth and environmental system model developments and transformations.

54 ENVIRONMENTAL SCIENCES↗

Operational Analytics Studies for ATLAS Distributed Computing: Data Popularity Forecast and Utilization of the WLCG Centers

Operational analytics is the direction of research related to the analysis of the current state of computing processes and the prediction of future states in order to anticipate imbalances and take timely measures to stabilize a complex system. There are two relevant areas in ATLAS Distributed Computing that are currently the focus of studies: user physics analysis including the forecast of popularity of data samples among users, and evaluating WLCG centers for their readiness to process user analysis payloads. Studying these areas is challenging due to the complexity involved, as it requires a comprehensive understanding of numerous boundary conditions typically found in large-scale distributed computing infrastructures. Forecasts of data popularity are problematic without the categorization of user tasks by their types (data transformation or physics analysis), which do not always appear on the surface but may induce noise, which introduces significant distortions for predictive analysis. Evaluating the WLCG resources by their analysis workloads is also a challenging task as it is necessary to find a balance between the workload of the resource, its performance, the waiting time for jobs on it, as well as the volume of jobs that it processes. This is especially difficult in a heterogeneous computing environment, where legacy resources are used along with modern high-performance machines. We will look at these areas of research in detail and discuss what tools and methods are used in our work, demonstrating results already obtained.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Hybridizing Machine Learning and Physically-based Earth System Models to Improve Prediction of Multivariate Extreme Events (AI Exploration of Wildland Fire Prediction)

Focal Areas: This project responds to two focal areas identified in the DOE Call for AI4ESP White Papers: 1) Predictive modeling through the use of artificial intelligence (AI) techniques, and 2) insights gleaned from complex data using explainable AI and big data analytics. Science Challenge: Large wildland fires (hereafter wildfires) appearing as high-impact compound climate extreme events are closely related to hydroclimate and water cycle extremes that modulate surface fuel supply and combustibility. These compound events have multivariate climatic features (e.g., temperature, precipitation, relative humidity, wind, lightning) and societal drivers (e.g., forest management, land use change, human caused ignitions). Meanwhile, they induce strong feedbacks to the coupled atmosphere, biosphere, and hydrosphere by perturbing regional and global radiation budget as well as ecological, biogeochemical, and water cycles across multiple spatiotemporal scales. The nonlinear interactions between these natural and anthropogenic components of the Earth system are too complex to be completely and adequately represented in today’s Earth system models (ESMs). The inherent stochastic nature of fire activity at all scales further increases the difficulty of its prediction using ESMs that are usually developed from deterministic equations and parameterizations. Besides, concurrence of long-term (decadal to interdecadal) global climate change and fire regime shifts overlapping with short-term (intraseasonal to interannual) variations of regional fire weather and burning activity confound predictability of these compound extreme events. We propose to address the above scientific challenges by using machine learning (ML)-based data-driven modeling techniques to integrate observations and physically-based ESMs’ simulations in a computationally efficient hybrid prediction system. This prediction system is supposed to characterize the wildfire’s sensitivity to climate and exogenous drivers at high resolution (~ 0.25°) on subseasonal to seasonal (S2S) timescales providing improved predictability and explainability. We will use the system to help identify: (1) What are the computational elements of a hybrid system needed to predict compound climate extreme events such as global wildfires? (2) What are the key drivers (either natural or anthropogenic) that modulate short-term variations of multivariate fire weather and burning activity over different regions? How can one take advantage of those driver-response relationships to improve the predictability of large wildfires on S2S time scales? (3) What are the underlying physical mechanisms and sources of improved predictability? Which ML techniques are optimal in revealing and adapting these mechanisms?

54 ENVIRONMENTAL SCIENCES↗

Continental Scale Hydrostratigraphy: Comparing Geologically Informed Data Products to Analytical Solutions

Abstract This study synthesizes two different methods for estimating hydraulic conductivity (K) at large scales. We derive analytical approaches that estimate K and apply them to the contiguous United States. We then compare these analytical approaches to three‐dimensional, national gridded K data products and three transmissivity (T) data products developed from publicly available sources. We evaluate these data products using multiple approaches: comparing their statistics qualitatively and quantitatively and with hydrologic model simulations. Some of these datasets were used as inputs for an integrated hydrologic model of the Upper Colorado River Basin and the comparison of the results with observations was used to further evaluate the K data products. Simulated average daily streamflow was compared to daily flow data from 10 USGS stream gages in the domain, and annually averaged simulated groundwater depths are compared to observations from nearly 2000 monitoring wells. We find streamflow predictions from analytically informed simulations to be similar in relative bias and Spearman's rho to the geologically informed simulations. R ‐squared values for groundwater depth predictions are close between the best performing analytically and geologically informed simulations at 0.68 and 0.70 respectively, with RMSE values under 10 m. We also show that the analytical approach derived by this study produces estimates of K that are similar in spatial distribution, standard deviation, mean value, and modeling performance to geologically‐informed estimates. The results of this work are used to inform a follow‐on study that tests additional data‐driven approaches in multiple basins within the contiguous United States.

54 ENVIRONMENTAL SCIENCES↗

Data for: A hybrid biophysical-machine learning framework for diurnal surface energy flux estimation using proximal sensing

Thermal-based remote sensing of surface energy fluxes has traditionally relied on high spatial resolution satellite data with revisit frequencies on the order of weeks. In this study, we evaluate a biophysics-based analytical surface energy balance model for predicting latent energy (LE) and sensible heat (H) fluxes using proximal sensing observations. The Surface Temperature Initiated Closure (STIC1.2) model has been extensively validated across a wide range of spatial and temporal scales using various satellite-derived thermal datasets. Here we extend this validation by applying STIC at sub-hourly temporal resolution over multiple growing seasons for four distinct agricultural systems. We further develop and evaluate novel STIC variants that incorporate machine learning (ML) techniques to eliminate the need for specific surface energy balance observations, specifically net radiation and soil heat flux, thereby enhancing model applicability in data-sparse settings. The integration of an ML component to estimate surface available energy is shown to have strong predictive performance for both LE (R2 = 0.81-0.94) and H (R2 = 0.46-0.72) across all agricultural systems examined here, demonstrating the potential of hybrid biophysical – machine learning approaches for surface energy balance modeling with minimal data requirements. This study concludes with a novel application of explainable machine learning (exML) to diagnose sources of model error. This exML framework attributes residual prediction errors to both model input variables and environmental drivers not explicitly included in the simulation experiments. This approach provides a new pathway for improving model design and integrating previously overlooked yet influential variables into future model iterations.

Agricultural Sciences↗

A Privacy First Path Analysis using Clickstream Data

In the modern digital economy, data-driven decision making is crucial for effectively meeting the ever-evolving demands of consumer engagement and satisfaction. Clickstream data has become invaluable for understanding customer behavior, yet concerns over privacy and security persist, especially with some internet service providers profiting from its sale. This article introduces an innovative methodology that blends experiential learning with advanced cryptographic techniques, including differential privacy and graph analytics. The core objective of this methodology is to estimate Customer Lifetime Value (CLV) by analyzing clickstream data, achieving an average prediction accuracy of 92.4% in user engagement levels while ensuring user anonymity through Recency, Frequency, and Monetary (RFM) analysis. Our study introduces the concept of a “data depositor” and a privacy manager, employing the composition theorem to merge non-adaptive queries effectively. Privacy budgets (? = 1.0, d = 10-5), sensitivity-specific techniques, and data partitioning were applied. Randomization and noise addition protect data integrity, with special handling for categorical values. This approach, differing from prior studies, offers a 12.6% improvement in privacy-preserving targeting accuracy while maintaining strict confidentiality, presenting a novel path forward in data-driven decision-making.

Frequency and Monetary (RFM) analysis↗

The role of pre-existing heterogeneities in materials under shock and spall

There has been a challenge for many decades to understand how heterogeneities influence the behavior of materials under shock loading, eventually leading to spall formation and failure. Experimental, analytical, and computational techniques have matured to the point where systematic studies of materials with complex microstructures under shock loading and the associated failure mechanisms are feasible. This is enabled by more accurate diagnostics as well as characterization methods. As interest in complex materials grows, understanding and predicting the role of heterogeneities in determining the dynamic behavior becomes crucial. Early computational studies, hydrocodes, in particular, historically preclude any irregularities in the form of defects and impurities in the material microstructure for the sake of simplification and to retain the hydrodynamic conservation equations. Contemporary computational methods, notably molecular dynamics simulations, can overcome this limitation by incorporating inhomogeneities albeit at a much lower length and time scale. This review discusses literature that has focused on investigating the role of various imperfections in the shock and spall behavior, emphasizing mainly heterogeneities such as second-phase particles, inclusions, and voids under both shock compression and release. Pre-existing defects are found in most engineering materials, ranging from thermodynamically necessary vacancies, to interstitial and dislocation, to microstructural features such as inclusions, second phase particles, voids, grain boundaries, and triple junctions. This literature review explores the interaction of these heterogeneities under shock loading during compression and release. Systematic characterization of material heterogeneities before and after shock loading, along with direct measurements of Hugoniot elastic limit and spall strength, allows for more generalized theories to be formulated. Further, continuous improvement toward time-resolved, in situ experimental data strengthens the ability to elucidate upon results gathered from simulations and analytical models, thus improving the overall ability to understand and predict how materials behave under dynamic loading.

36 MATERIALS SCIENCE↗

Bayesian learning with Gaussian processes for low-dimensional representations of time-dependent nonlinear systems

This work presents a data-driven method for learning low-dimensional time-dependent physics-based surrogate models whose predictions are endowed with uncertainty estimates. We use the operator inference approach to model reduction that poses the problem of learning low-dimensional model terms as a regression of state space data and corresponding time derivatives by minimizing the residual of reduced system equations. Standard operator inference models perform well with accurate training data that are dense in time, but producing stable and accurate models when the state data are noisy and/or sparse in time remains a challenge. Another challenge is the lack of uncertainty estimation for the predictions from the operator inference models. Our approach addresses these challenges by incorporating Gaussian process surrogates into the operator inference framework to (1) probabilistically describe uncertainties in the state predictions and (2) procure analytical time derivative estimates with quantified uncertainties. The formulation leads to a generalized least-squares regression and, ultimately, reduced-order models that are described probabilistically with a closed-form expression for the posterior distribution of the operators. The resulting probabilistic surrogate model propagates uncertainties from the observed state data to reduced-order predictions. Furthermore, we demonstrate the method is effective for constructing low-dimensional models of two nonlinear partial differential equations representing a compressible flow and a nonlinear diffusion–reaction process, as well as for estimating the parameters of a low-dimensional system of nonlinear ordinary differential equations representing compartmental models in epidemiology.

Data-driven model reduction↗

AI-Based Integrated Modeling and Observational Framework for Improving Seasonal to Decadal Prediction of Terrestrial Ecohydrological Extremes

Focal Areas: (1) Insight gleaned from complex data (both observed and simulated) using artificial intelligence(AI), big data analytics, and other advanced methods, including explainable AI and physics- or knowledge-guided AI (2) Data acquisition and assimilation enabled by machine learning, AI, and advanced methods including experimental/network design/optimization, unsupervised learning (including deep learning), and hardware-related efforts involving AI (e.g., edge computing).

54 ENVIRONMENTAL SCIENCES↗

Multifidelity Neural Network Formulations for Prediction of Reactive Molecular Potential Energy Surfaces

Here, this paper focuses on the development of multifidelity modeling approaches using neural network surrogates, where training data arising from multiple model forms and resolutions are integrated to predict high-fidelity response quantities of interest at lower cost. We focus on the context of quantum chemistry and the integration of information from multiple levels of theory. Important foundations include the use of symmetry function-based atomic energy vector constructions as feature vectors for representing structures across families of molecules and single-fidelity neural network training capabilities that learn the relationships needed to map feature vectors to potential energy predictions. These foundations are embedded within several multifidelity topologies that decompose the high-fidelity mapping into model-based components, including sequential formulations that admit a general nonlinear mapping across fidelities and discrepancy-based formulations that presume an additive decomposition. Methodologies are first explored and demonstrated on a pair of simple analytical test problems and then deployed for potential energy prediction for C 5 H 5 using B2PLYP-D3/6-311++G(d,p) for high-fidelity simulation data and Hartree–Fock 6-31G for low-fidelity data. For the common case of limited access to high-fidelity data, our computational results demonstrate that multifidelity neural network potential energy surface constructions achieve roughly an order of magnitude improvement, either in terms of test error reduction for equivalent total simulation cost or reduction in total cost for equivalent error.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Establishing performance metrics for quantitative non-targeted analysis: a demonstration using per- and polyfluoroalkyl substances

Abstract Non-targeted analysis (NTA) is an increasingly popular technique for characterizing undefined chemical analytes. Generating quantitative NTA (qNTA) concentration estimates requires the use of training data from calibration “surrogates,” which can yield diminished predictive performance relative to targeted analysis. To evaluate performance differences between targeted and qNTA approaches, we defined new metrics that convey predictive accuracy, uncertainty (using 95% inverse confidence intervals), and reliability (the extent to which confidence intervals contain true values). We calculated and examined these newly defined metrics across five quantitative approaches applied to a mixture of 29 per- and polyfluoroalkyl substances (PFAS). The quantitative approaches spanned a traditional targeted design using chemical-specific calibration curves to a generalizable qNTA design using bootstrap-sampled calibration values from “global” chemical surrogates. As expected, the targeted approaches performed best, with major benefits realized from matched calibration curves and internal standard correction. In comparison to the benchmark targeted approach, the most generalizable qNTA approach (using “global” surrogates) showed a decrease in accuracy by a factor of ~4, an increase in uncertainty by a factor of ~1000, and a decrease in reliability by ~5%, on average. Using “expert-selected” surrogates ( n = 3) instead of “global” surrogates ( n = 25) for qNTA yielded improvements in predictive accuracy (by ~1.5×) and uncertainty (by ~70×) but at the cost of further-reduced reliability (by ~5%). Overall, our results illustrate the utility of qNTA approaches for a subclass of emerging contaminants and present a framework on which to develop new approaches for more complex use cases. Graphical Abstract

Pu, Shirley (ORCID:0000000201223797)↗