Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Error Metrics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Enabling Robust Exoplanet Atmospheric Retrievals with Gaussian Processes

Atmospheric retrievals are essential tools for interpreting exoplanet transmission and eclipse spectra, enabling quantitative constraints on the chemical composition, aerosol properties, and thermal structure of planetary atmospheres. The James Webb Space Telescope (JWST) offers unprecedented spectral precision, resolution, and wavelength coverage, unlocking transformative insights into the formation, evolution, climate, and potential habitability of planetary systems. However, this opportunity is accompanied by challenges: modeling assumptions and unaccounted-for noise or signal sources can bias retrieval outcomes and their interpretation. To address these limitations, we introduce a Gaussian process (GP)-aided atmospheric retrieval framework that flexibly accounts for unmodeled features and correlated noise in exoplanet spectra. We validate this method on synthetic JWST observations, and show that GP-aided retrievals reduce bias in inferred abundances and better capture model–data mismatches than traditional approaches. We also introduce the concept of mean squared error to quantify the trade-off between bias and variance, arguing that this metric more accurately reflects retrieval performance than bias alone. We then reanalyze the NIRISS/SOSS JWST transmission spectrum of WASP-96 b, finding that GP-aided retrievals yield broader constraints on CO 2 and H 2 O, possibly alleviating tension between previous retrieval results and equilibrium predictions. Our GP framework provides precise and accurate constraints while highlighting regions where models fail to explain the data. As JWST matures and future facilities come online, a deeper understanding of the limitations of both data and models will be essential, and GP-enabled retrievals like the one presented here offer a principled path forward.

Rotman, Yoav [Arizona State Univ., Tempe, AZ (Unit↗

Characterizing the Reproducibility of Noisy Quantum Circuits

The ability of a quantum computer to reproduce or replicate the results of a quantum circuit is a key concern for verifying and validating applications of quantum computing. Statistical variations in circuit outcomes that arise from ill-characterized fluctuations in device noise may lead to computational errors and irreproducible results. While device characterization offers a direct assessment of noise, an outstanding concern is how such metrics bound the reproducibility of a given quantum circuit. Here, we first directly assess the reproducibility of a noisy quantum circuit, in terms of the Hellinger distance between the computational results, and then we show that device characterization offers an analytic bound on the observed variability. We validate the method using an ensemble of single qubit test circuits, executed on a superconducting transmon processor with well-characterized readout and gate error rates. The resulting description for circuit reproducibility, in terms of a composite device parameter, is confirmed to define an upper bound on the observed Hellinger distance, across the variable test circuits. This predictive correlation between circuit outcomes and device characterization offers an efficient method for assessing the reproducibility of noisy quantum circuits.

97 MATHEMATICS AND COMPUTING↗

Aboveground woody biomass estimation of young bioenergy plantations of Populus and its hybrids using mobile (backpack) LiDAR remote sensing

Woody aboveground biomass (AGB) including short-rotation Populus is used as a feedstock for renewable and carbon-neutral bioenergy. While woody AGB can be estimated with allometric equations requiring labor-intensive field data, remote sensing technologies like mobile terrestrial light detection and ranging (LiDAR) can estimate woody AGB quickly and accurately. Therefore, the goals of this study were to develop a model to predict woody AGB of 2-year-old Populus spp. from three taxa (P. deltoides, P. deltoides × P. maximowiczii and P. deltoides × P. trichocarpa) using allometric (height and diameter at breast height (DBH)) or LiDAR-derived metrics from a mobile terrestrial (backpack) system. Likewise, we sought to compare LiDAR-estimated tree height and DBH with field-measured values. We found that a taxa-specific model containing LiDAR-measured tree height, crown volume, and taxa interactions with the height of the 10 th percentile, and the density of the lowest interval (density metric 0) explained 84 % of the variation in woody AGB with a root mean square error (RMSE) of 28.7 % and performed slightly better than the allometric model. The best model excluding taxa had a slightly higher RMSE but lower bias than the allometric model. LiDAR-derived tree heights were highly correlated with field-measured heights, but DBH could not be estimated accurately. Therefore, terrestrial mobile LiDAR systems can accurately estimate woody AGB and tree height of Populus in short rotation systems to aid in the fast and efficient quantification of woody bioenergy production and renewable energy resources.

AGB↗

Tracing U.S. fuel life-cycle greenhouse gas emissions in a multi-sector dynamics model using LC-GCAM

Model-based analysis of fuel pathways is essential for informing energy and environmental policy. Two major model types are typically used: multi-sector dynamics models, which capture the broader energy-economy, such as GCAM (Global Change Analysis Model), and life cycle assessment models, such as GREET (Greenhouse Gases, Regulated Emissions, and Energy Use in Transportation). Each has distinct strengths and limitations, and recent studies increasingly adopt hybrid approaches to harness the advantages of both. However, such integration is often time-consuming and complicated by inconsistencies in system boundaries and technology definitions. We present LC-GCAM, a new tool that enables estimation of life-cycle greenhouse gas emissions and primary energy use for any fuel pathway represented in GCAM. We apply LC-GCAM to 300 scenarios designed to explore key uncertainties affecting the life-cycle performance of future fuel options in the U.S. freight sector. To evaluate LC-GCAM, we compare its results with those from GREET for nine fuel types in a 2030 reference scenario. When input assumptions are modestly aligned, LC-GCAM and GREET estimates typically agree within 10% (absolute sum-based mean absolute percentage error). LC-GCAM offers a flexible and efficient approach to generating life-cycle metrics within an integrated modeling framework, supporting robust policy analysis across a wide range of interacting energy system uncertainties.

Wolfram, Paul↗

Probing Accuracy-Speedup Tradeoff in Machine Learning Surrogates for Molecular Dynamics Simulations

The performance promise of machine learning surrogates of molecular dynamics simulations of soft materials is significant but generally comes at the cost of acquiring large training datasets to learn the complex relationships between input soft material attributes and output properties. Under the constraint of limited high-performance computing resources, optimizing the size of the training datasets becomes paramount. Using an artificial neural network based surrogate for molecular dynamics simulations of confined electrolytes, we explore the tradeoff between surrogate accuracy and computational gains. Accuracy is assessed by computing the root-mean-square errors between the surrogate predictions and the ground truth results obtained via molecular dynamics simulations. The computational performance is judged by evaluating the speedup which incorporates the training dataset creation time. Improvement in accuracy occurs with a loss of speedup, which scales as the inverse of the training dataset size. Furthermore, the link between surrogate generalizability and the accuracy-speedup tradeoff is assessed by examining the errors incurred in surrogate predictions on unseen, interpolated input variables and developing a net speedup metric to capture the associated gains.

Anions↗

CONSTAX2: improved taxonomic classification of environmental DNA markers

Abstract Summary CONSTAX—the CONSensus TAXonomy classifier—was developed for accurate and reproducible taxonomic annotation of fungal rDNA amplicon sequences and is based upon a consensus approach of RDP, SINTAX and UTAX algorithms. CONSTAX2 extends these features to classify prokaryotes as well as eukaryotes and incorporates BLAST-based classifiers to reduce classification errors. Additionally, CONSTAX2 implements a conda-installable command-line tool with improved classification metrics, faster training, multithreading support, capacity to incorporate external taxonomic databases and new isolate matching and high-level taxonomy tools, replete with documentation and example tutorials. Availability and implementation CONSTAX2 is available at https://github.com/liberjul/CONSTAXv2, and is packaged for Linux and MacOS from Bioconda with use under the MIT License. A tutorial and documentation are available at https://constax.readthedocs.io/en/latest/. Data and scripts associated with the manuscript are available at https://github.com/liberjul/CONSTAXv2_ms_code. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

The Effects of Gamma Ray Integrated Dose on a Commercial 65-nm SRAM Device

Here this work shows that the static random access memory (SRAM) error rate for a commercial 65-nm device in a dose rate environment can be highly dependent upon the integrated dose (dose rate × pulse duration). While the typical metric for such testing is dose rate upset (DRU) level in rad(Si)/s, a series of dose rate experiments at Little Mountain Test Facility (LMTF) shows dependence on the integrated dose. The error rate is also found to be dependent on the core voltage, and the preradiation value of the bits. We believe that these effects are explained by a well charge depletion caused by gamma ray photocurrent.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Numerical Modeling of the Effects of Coating Plates on Terminal Ballistic Performance

This report deals with the development and evaluation of a numerical model to examine applied coating to a metal substrate subjected to a ballistic impact. The numerical model will be used to examine the benefit of the coating in resisting penetration due to the impact. For a detailed examination the Retch-Ipson curve is used as a metric. The numerical data is plotted and then fit to the Retch-Ipson curve and error calculations are used to compare the difference between the numerical output and the experimental data. This initial study is an examination of a few shortcomings of the standard material models used, and demonstrate the future work that is needed to understand the ballistic behavior of materials.

36 MATERIALS SCIENCE↗

Performance Evaluation of Intelligent Solar Control Software Through Hardware-in-the-Loop (CRADA Final Report)

Recent research has highlighted the potential for solar to act as a zero-marginal-cost and zero-emission flexibility resource on the bulk power system when operated with advanced control systems. To increase the performance of these systems, leading technologies, including machine learning (ML) and hierarchical inverter set point allocation, have been developed by Latimer Controls, Inc. to estimate the headroom of large PV plants for grid operation and control; however, these technologies lack comprehensive validation under real-world application scenarios. Latimer Controls, Inc. received two voucher awards for research at a national laboratory from the Department of Energy American Made Solar Prize Round 6. The National Renewable Energy Laboratory (NREL) was selected to collaborate with Latimer staff to conduct a performance evaluation of Latimer PV control software. The NREL team will develop a hardware-in-the-loop (HIL) testbed to perform testing and validation of the Latimer PV control technology in a de-risked yet realistic testbed environment. Latimer and NREL worked together to analyze the test data, draw conclusions from the results, and disseminate the resulting scientific findings. In this CRADA work, we propose to test and validate the real-world application of the Latimer Control solution in an HIL environment. We evaluate the performance of different flexible solar technologies in responding to automatic generation control signals in a closed-loop fashion. In particular, a data-driven potential high limit (PHL) estimation is developed for large solar plants to accurately estimate their headroom so that they have fast and short-time regulation and control capability to participate in grid services and respond to grid signals in real time (e.g., AGC). This PHL estimation algorithm is embedded in a hardware power plant controller (PPC) and tested with an IEEE-39 bus system model developed in RTDS. To account for the varying cloud conditions and diverse inverter dispatches, we developed a 135-MW PV plant with detailed modeling of 27 individual PV modules and inverters using RTDS. The real-world communications used in such big plants, such as ModBus TCP/IP for inverter level and DNP3 for plant level, were developed to emulate the real-world applications in big PV plants. The ML-based PHL estimation method is tested under nine separate weather scenarios against the ‘reference-control’ solution, hereafter referred to as the baseline solution. The baseline method reserves a subset of inverters (reference group) to operate at their PHL at all times and dispatches only the remaining inverters (control group) at curtailed levels to fulfill the flexibility need. Despite being successfully piloted by NREL in California in 2017 and Chile in 2020, there exist two gaps in the state of the art to fully unlock the flexibility of PV plants: a. There is a trade-off between the PHL estimation accuracy and the flexibility range. b. There lacks granularity in the PHL estimation to capture the variation across inverters. The Latimer solution seeks to address these gaps by applying machine learning methods to improve PHL estimation accuracy while accounting for variability at every inverter. Performance metrics were taken from the 2023 Georgia Power CARES utility-scale RFP. The results demonstrate that the ML-based approach outperforms the traditional baseline method in PHL estimation accuracy for 7 of 9 scenarios. The average PHL error across the nine scenarios was 7.40% for the ML-based method, 2.06% less than the 9.46% PHL error average across scenarios that was exhibited by the baseline method. Additionally, the PHL error was below 5% for at least 95% of the testing interval for 3 of 9 tested intervals with the ML approach, whereas it did not achieve this metric for any of the baseline tests. Overall, simulation results indicate the superior performance of an ML-based approach compared to the conventional baseline reference-control approach, showcasing its potential to support grid stability and operational efficiency. This laboratory HIL testing using real PPC, representative power system simulation models in real-time with detailed PV plant and inverter models, and real-world communication protocols gives us confidence that this machine learning based PHL estimation algorithm works well in the hardware PPC and therefore de-risks future field commissioning. The end goal of this project is to advance grid technology to address the grid operation challenges brought by solar plant’s variability and uncertainties in power generation.

14 SOLAR ENERGY↗

EPIsembleVis: A geo-visual analysis and comparison of the prediction ensembles of multiple COVID-19 models

In this work, we present EPIsembleVis, a web-based comparative visual analysis tool for evaluating the consistency of multiple COVID-19 prediction models. Our approach analyzes a collection of COVID-19 predictions from different epidemiological models as an ensemble and utilizes two metrics to quantify model performance. These metrics include (a) prediction uncertainty (represented as the dispersion of predictions in each ensemble) and (b) prediction error (calculated by comparing individual model predictions with the recorded data). Through an interactive visual interface, our approach provides a data-driven workflow for (a) selecting and constructing the COVID-19 model prediction ensemble based on the spatiotemporal overlap of available predictions of multiple epidemiological models, (b) quantifying the model performance using both the uncertainty of each model prediction ensemble, and the error of each ensemble member that represents individual model predictions, and (c) visualizing the spatiotemporal variability in the projection performance of individual models using a suite of novel ensemble visualization techniques, such as the data availability map, a spatiotemporal textured-tile calendar, multivariate rose chart, and time-series leaflet glyph. We demonstrate the capability of our ensemble visual interface through a case study that investigates the performance of weekly COVID-19 predictions, which are provided through the COVID-19 Forecast Hub UMass-Amherst Influenza Forecasting Center of Excellence [47] for the United States and United States Territories. The EPIsembleVis tool is implemented using open-source web technologies and adaptive system design, rendering it interoperable with Elasticsearch and Kibana for automatically ingesting COVID-19 predictions from online repositories, and it is generalizable for analyzing worldwide projections from more epidemiological models.

60 APPLIED LIFE SCIENCES↗

Safeguards Modeling for Advanced Nuclear Facility Design.

Future nuclear fuel cycle facilities will see a significant benefit from considering materials accountancy requirements early in the design process. The Material Protection, Accounting, and Control Technologies (MPACT) working group is demonstrating Safeguards and Security by Design (SSBD) for a notional electrochemical reprocessing facility as part of a 2020 Milestone. The idea behind SSBD is to consider regulatory requirements early in the design process to provide more optimized systems and avoid costly retrofits later in the design process. Safeguards modeling, using single analyst tools, allows the designer to efficiently consider materials accountancy approaches that meet regulatory requirements. However, safeguards modeling also allows the facility designer to go beyond current regulations and work toward accountancy designs with rapid response and lower thresholds for detection of anomalies. This type of modeling enables new safeguards approaches and may inform future regulatory changes. The Separation and Safeguards Performance Model (SSPM) has been used for materials accountancy system design and analysis. This paper steps through the process of designing a Material Control and Accountancy (MC&A) system, presents the baseline system design for an electrochemical reprocessing facility, and provides performance metrics from the modeling analysis. The most critical measurements in the electrochemical facility are the spent fuel input, electrorefiner salt, and U/TRU product output measurements. Finally, material loss scenario analysis found that measurement uncertainties (relative standard deviations) for Pu would need to be at 1% (random and systematic error components) or better in order to meet domestic detection goals or as high as 3% in order to meet international detection goals, based on a 100 metric ton per year plant size.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Evaluation of pressure reconstruction techniques for Model Order Reduction in incompressible convective heat transfer

This paper compares pressure reconstruction strategies in Model Order Reduction for incompressible flows with convective heat transfer. The Navier-Stokes equation are reduced along with the passive scalar transport equation for the temperature using the POD-Galerkin technique. Six different pressure reconstruction methods are evaluated, two of which are novel to the best of the authors’ knowledge. Accurate pressure reconstruction is key to avoid error buildup in when solving for the conservation of linear momentum at the reduced level. The six approaches are compared using Direct Numerical Simulations of convective heat exchange processes in a 3D Backward Facing Step with a heated cylinder. Additionally, when comparing time-averaged metrics, we observe that the reconstruction methods that approximate the reduced pressure field using techniques borrowed from full order models (mechanical analogy, pressure Poisson, and velocity supremizers) yield higher errors than the methods that seek to stabilize the reduced systems (reduced residual stabilization, artificial divergence, and Uzawa operator).

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

AAPM Truth‐based CT (TrueCT) reconstruction grand challenge

Background: This Special Report summarizes the 2022, AAPM grand challenge on Truth-based CT image reconstruction. Purpose: To provide an objective framework for evaluating CT reconstruction methods using virtual imaging resources consisting of a library of simulated CT projection images of a population of human models with various diseases. Methods: Two hundred unique anthropomorphic, computational models were created with varied diseases consisting of 67 emphysema, 67 lung lesions, and 66 liver lesions. The organs were modeled based on clinical CT images of real patients. The emphysematous regions were modeled using segmentations from patient CT cases in the COPDGene Phase I dataset. For the lung and liver lesion cases, 1–6 malignant lesions were created and inserted into the human models, with lesion diameters ranging from 5.6 to 21.9 mm for lung lesions and 3.9 to 14.9 mm for liver lesions. The contrast defined between the liver lesions and liver parenchyma was 82 ± 12 HU, ranging from 50 to 110 HU. Similarly, the contrast between the lung lesions and the lung parenchyma was defined as 781 ± 11 HU, ranging from 725 to 805 HU. For the emphysematous regions, the defined HU values were −950 ± 17 HU ranging from −918 to −979 HU. The developed human models were imaged with a validated CT simulator. The resulting CT sinograms were shared with the participants. The participants reconstructed CT images from the sinograms and sent back their reconstructed images. Further, the reconstructed images were then scored by comparing the results against the corresponding ground truth values. The scores included both task-generic (root mean square error [RMSE] and structural similarity matrix [SSIM]), and task-specific (detectability index [d’] and lesion volume accuracy) metrics. For the cases with multiple lesions, the measured metric was averaged across all the lesions. To combine the metrics with each other, each metric was normalized to a range of 0 to 1 per disease type, with “0” and “1” being the worst and best measured values across all cases of the disease type for all received reconstructions. Results: The True-CT challenge attracted 52 participants, out of which 5 successfully completed the challenge and submitted the requested 200 reconstructions. Across all participants and disease types, SSIM absolute values ranged from 0.22 to 0.90, RMSE from 77.6 to 490.5 HU, d’ from 0.1 to 64.6, and volume accuracy ranged from 1.2 to 753.1 mm3. The overall scores demonstrated that participant “A” had the best performance in all categories, except for the metrics of d’ for lung lesions and RMSE for liver lesions. Participant “A” had an average normalized score of 0.41 ± 0.22, 0.48 ± 0.32, and 0.42 ± 0.33 for the emphysema, lung lesion, and liver lesion cases, respectively. Conclusions: The True-CT challenge successfully enabled objective assessment of CT reconstructions with the unique advantage of access to a diverse population of diseased human models with known ground truth. This study highlights the significant potential of virtual imaging trials in objective assessment of medical imaging technologies.

60 APPLIED LIFE SCIENCES↗

A control-oriented combustion model framework for compression ignition engines operating on low-reactivity fuel

This work focuses on zero-dimensional modeling of the heat release rate in a compression ignition engine operating on gasoline-like fuels. Due to the properties of gasoline, such as high volatility and longer ignition delay than diesel, the injection strategies can vary significantly from the operation with conventional diesel fuel. Different injection strategies are commonly used to achieve varying degrees of in-cylinder stratification in order to shape the combustion event and maximize efficiency. The proposed zero-dimensional combustion model was developed to account for the different stages in combustion caused by the fuel stratification. As the ignition delay model is an integral part of the entire combustion process and significantly affects the prediction accuracy, special attention has been paid to local phenomena influencing ignition delay. A one-dimensional spray model by Musculus and Kattke was employed in conjunction with a Lagrangian tracking approach in order to estimate the local air–fuel ratio within the spray tip, as a proxy for reactivity. The local air–fuel ratio, in-cylinder temperature and pressure were used in an integral fashion to estimate the ignition delay. Heat release rates were modeled using first-order non-linear differential equations. The proposed combustion model was validated against experimental data of a heavy-duty compression ignition engine with up to three injection events at mostly 1038 r/min and 14 bar brake mean effective pressure. Further validation of the model was carried out at other engine loads and speeds. Model prediction errors in CA50 of less than 1 °CA across all conditions were found. Modeling results of other combustion metrics such as combustion duration and indicated mean effective pressure are also highly satisfactory. In addition, the model has been shown to be capable of estimating the ringing intensity for most conditions.

Pamminger, Michael↗

Solution Irregularity Remediation for Spatial Discretization Error Estimation for S N Transport Solutions

The discrete ordinates linear Boltzmann transport equation is typically solved in its spatially discretized form, incurring spatial discretization error. Quantification of this error for purposes such as adaptive mesh refinement or error analysis requires an a posteriori estimator, which utilizes the numerical solution to the spatially discretized equation to compute an estimate. Because the quality of the numerical solution informs the error estimate, irregularities, present in the true solution for any realistic problem configuration, tend to cause the largest deviation in the error estimate vis-a-vis the true error. In this paper, an analytical partial singular characteristic tracking (pSCT) procedure for reducing the estimator’s error is implemented within our novel residual source estimator for a zeroth-order discontinuous Galerkin scheme, at the additional cost of a single inner iteration. Here, a metric-based evaluation of the pSCT scheme versus the standard residual source estimator is performed over the parameter range of a Method of Manufactured Solutions test suite. The pSCT scheme generates near-ideal accuracy in the estimate in problems where the dominant source of the estimator’s error is the solution irregularity, namely, problems where the true solution is discontinuous and problems where the true solution’s first derivative is discontinuous and the scattering ratio is low. In problems where the scattering ratio is high and the true solution is discontinuous in the first derivative, the error in the scattering source, which is not converged by the pSCT scheme, is greater than the error incurred due to the irregularity. Ultimately, a pSCT scheme is judged to be useful for error estimation in problems where the computational cost of the scheme is justified. In the presence of many irregularities, such a scheme may be intractable for general use, but in benchmarks, as an analytical tool, or in problems that have nondissipative discontinuities, the scheme may prove invaluable.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Extreme metrics from large ensembles: investigating the effects of ensemble size on their estimates

Abstract. We consider the problem of estimating the ensemble sizes required to characterize the forced component and the internal variability of a number of extreme metrics. While we exploit existing large ensembles, our perspective is that of a modeling center wanting to estimate a priori such sizes on the basis of an existing small ensemble (we assume the availability of only five members here). We therefore ask if such a small-size ensemble is sufficient to estimate accurately the population variance (i.e., the ensemble internal variability) and then apply a well-established formula that quantifies the expected error in the estimation of the population mean (i.e., the forced component) as a function of the sample size n, here taken to mean the ensemble size. We find that indeed we can anticipate errors in the estimation of the forced component for temperature and precipitation extremes as a function of n by plugging into the formula an estimate of the population variance derived on the basis of five members. For a range of spatial and temporal scales, forcing levels (we use simulations under Representative Concentration Pathway 8.5) and two models considered here as our proof of concept, it appears that an ensemble size of 20 or 25 members can provide estimates of the forced component for the extreme metrics considered that remain within small absolute and percentage errors. Additional members beyond 20 or 25 add only marginal precision to the estimate, and this remains true when statistical inference through extreme value analysis is used. We then ask about the ensemble size required to estimate the ensemble variance (a measure of internal variability) along the length of the simulation and – importantly – about the ensemble size required to detect significant changes in such variance along the simulation with increased external forcings. Using the F test, we find that estimates on the basis of only 5 or 10 ensemble members accurately represent the full ensemble variance even when the analysis is conducted at the grid-point scale. The detection of changes in the variance when comparing different times along the simulation, especially for the precipitation-based metrics, requires larger sizes but not larger than 15 or 20 members. While we recognize that there will always exist applications and metric definitions requiring larger statistical power and therefore ensemble sizes, our results suggest that for a wide range of analysis targets and scales an effective estimate of both forced component and internal variability can be achieved with sizes below 30 members. This invites consideration of the possibility of exploring additional sources of uncertainty, such as physics parameter settings, when designing ensemble simulations.

54 ENVIRONMENTAL SCIENCES↗

Downscaling atmospheric chemistry simulations with physically consistent deep learning

Abstract. Recent advances in deep convolutional neural network (CNN)-based super resolution can be used to downscale atmospheric chemistry simulations with substantially higher accuracy than conventional downscaling methods. This work both demonstrates the downscaling capabilities of modern CNN-based single image super resolution and video super-resolution schemes and develops modifications to these schemes to ensure they are appropriate for use with physical science data. The CNN-based video super-resolution schemes in particular incur only 39 % to 54 % of the grid-cell-level error of interpolation schemes and generate outputs with extremely realistic small-scale variability based on multiple perceptual quality metrics while performing a large (8×10) increase in resolution in the spatial dimensions. Methods are introduced to strictly enforce physical conservation laws within CNNs, perform large and asymmetric resolution changes between common model grid resolutions, account for non-uniform grid-cell areas, super-resolve lognormally distributed datasets, and leverage additional inputs such as high-resolution climatologies and model state variables. High-resolution chemistry simulations are critical for modeling regional air quality and for understanding future climate, and CNN-based downscaling has the potential to generate these high-resolution simulations and ensembles at a fraction of the computational cost.

54 ENVIRONMENTAL SCIENCES↗

Performance and power modeling and prediction using MuMMI and 10 machine learning methods

Energy-efficient scientific applications require insight into how high performance computing system features impact the applications' power and performance. This insight can result from the development of performance and power models. Here, in this article, we use the modeling and prediction tool MuMMI (Multiple Metrics Modeling Infrastructure) and 10 machine learning methods to model and predict performance and power consumption and compare their prediction error rates. We use an algorithm-based fault-tolerant linear algebra code and a multilevel checkpointing fault-tolerant heat distribution code to conduct our modeling and prediction study on the Cray XC40 Theta and IBM BG/Q Mira at Argonne National Laboratory and the Intel Haswell cluster Shepard at Sandia National Laboratories. Our experimental results show that the prediction error rates in performance and power using MuMMI are less than 10% for most cases. By utilizing the models for runtime, node power, CPU power, and memory power, we identify the most significant performance counters for potential application optimizations, and we predict theoretical outcomes of the optimizations. Based on two collected datasets, we analyze and compare the prediction accuracy in performance and power consumption using MuMMI and 10 machine learning methods.

97 MATHEMATICS AND COMPUTING↗