Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Benchmarking Software”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Evaluation of Ocean Biogeochemistry and Carbon Cycling in CMIP Earth System Models With the International Ocean Model Benchmarking (IOMB) Software System

Abstract The International Ocean Model Benchmarking (IOMB) software package is a new community resource that we use here to evaluate surface and upper ocean biogeochemical variables and integrated anthropogenic carbon uptake from earth system models (ESMs) contributing to the 5th and 6th phases of the Coupled Model Intercomparison Project (CMIP5 and CMIP6). IOMB generates graphics and tables for systematically comparing model predictions against multiple datasets. Our analysis reveals some improvement in the multi‐model mean from CMIP5 to CMIP6 for most of the variables we examined. Compared to data‐constrained estimates of ocean anthropogenic carbon uptake for the 1994–2007 period, negative biases exist for many models between 30 and 50°S. Global model estimates of anthropogenic carbon uptake for the same period do not change significantly from CMIP5 to CMIP6, with the combined ensemble mean estimate of 27.8 ± 0.5 Pg C lower than a data‐constrained estimate of 33.0 ± 4.0 Pg C. At the same time, the change in the natural carbon inventory from CMIP is estimated to be a source of 0.7 ± 0.3 Pg C, which is considerably smaller in magnitude than a data‐constrained estimate of 5.0 ± 3.0 Pg C. With chlorofluorocarbon (CFC) predictions available for several models, we demonstrate that negative anthropogenic dissolved inorganic carbon biases coincide with negative biases in CFC concentration, highlighting the importance of weak exchange between the surface and interior ocean in regulating rates of anthropogenic carbon uptake. To examine the robustness of this attribution across the CMIP models, we calculate the global vertical temperature gradient between 200 and 1,000 m as a metric for global stratification and exchange between the surface and deeper waters. We find a linear relationship between the bias of the vertical temperature gradients and the bias in global anthropogenic carbon uptake, consistent with the hypothesis that model biases in anthropogenic carbon uptake are related to biases in surface‐to‐interior exchange by physical processes.

58 GEOSCIENCES↗

DeepSurveySim: Simulation Software and Benchmark Challenges for Astronomical Observation Scheduling

Modern astronomical surveys have multiple competing scientific goals. Optimizing the observation schedule for these goals presents significant computational and theoretical challenges, and state-of-the-art methods rely on expensive human inspection of simulated telescope schedules. Automated methods, such as reinforcement learning, have recently been explored to accelerate scheduling. However, there do not yet exist benchmark data sets or user-friendly software frameworks for testing and comparing these methods. We present DeepSurveySim -- a high-fidelity and flexible simulation tool for use in telescope scheduling. DeepSurveySim provides methods for tracking and approximating sky conditions for a set of observations from a user-supplied telescope configuration. We envision this tool being used to produce benchmark data sets and for evaluating the efficacy of ground-based telescope scheduling algorithms, particularly for machine learning algorithms that would suffer in efficacy if limited to real data for training.We introduce three example survey configurations and related code implementations as benchmark problems that can be simulated with DeepSurveySim.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

DeepSurveySim: Simulation Software and Benchmark Challenges for Astronomical Observation Scheduling

Modern astronomical surveys have multiple competing scientific goals. Optimizing the observation schedule for these goals presents significant computational and theoretical challenges, and state-of-the-art methods rely on expensive human inspection of simulated telescope schedules. Automated methods, such as Reinforcement Learning (RL), have recently been explored to accelerate scheduling. DeepSurveySim provides methods for tracking and approximating sky conditions for a set of observations from a user-supplied telescope configuration.

79 ASTRONOMY AND ASTROPHYSICS↗

Computational Fluid Dynamic Modeling of Dry Cask Simulator with Crosswind

The purpose of this study is to create a STAR-CCM+ model of a Belowground Vertical Dry Cask Simulator (BVDCS) at Sandia National Laboratories (SNL) and validate the model with SNL’s experimental results. The BVDCS consists of a single boiling water reactor assembly fitted with electric heaters encompassed by a containment vessel and shell to represent a belowground spent nuclear fuel (SNF) dry storage system. Blowers are located near the inlet and outlet of the BVDCS to simulate crosswind conditions. In addition to the experimental results, the STAR-CCM+ model developed for this study is compared with a previous computational fluid dynamics (CFD) model in a different software program, which is used as a software-to-software benchmark. The experimental results provide a dataset to compare the STAR-CCM+ model results for a variety of different conditions. The main objective is to validate and improve STAR-CCM+ CFD models for spent nuclear fuel storage systems with explicitly modeled external environments and “wind driven” crossflows. These CFD models aide in the study of external particle deposition in spent nuclear fuel storage systems, which is important to predicting the significance of chloride induced stress corrosion cracking (CISCC). In addition to experimental comparison, a sensitivity analysis study is performed using the STAR-CCM+ model. The sensitivity analysis provides a quantitative assessment of the sensitivity of various parameters. This helps provide information on various parameters that are of particular importance to constructing a model representative of real life systems. The STAR-CCM+ model compared well to the experimental results showing similar responses to changes in cross wind flow, and a number of parameters are identified for model improvement.

Jensen, Ben J.↗

Quantum Application Specifications and Benchmarks

This software describes computational tasks for quantum computers that are derived from LANL basic science research applications such as the modeling of materials, chemicals and compounds at atomic scales. The computations focus on quantum simulation tasks, such as quantum dynamics, thermal state preparation and ground state estimation. The primary function of the software is to develop estimates of the requirements for large-scale fault-tolerant quantum computers to solve these scientific computations. The secondary focus of the software are codes for assessing the limitations of conducting quantum computations on classical computers.

Coffrin, Carleton↗

DOE FAIR Surrogate Benchmarks Supporting AI and Simulation Research (SBI Surrogate Benchmark Initiative) (Final Report)

Computational Science is being revolutionized by integrating AI and simulation and, in particular, by deep learning surrogate models that can replace all or part of traditional large‐scale HPC computations. Such surrogates can achieve remarkable performance improvements, as much as several orders of magnitude, and save both compute time and energy. The Surrogate Benchmark Initiative (SBI) project creates a community repository and FAIR (Findable, Accessible, Interoperable, and Reusable) data ecosystem for HPC application surrogate benchmarks. The SBI team comes from Argonne National Laboratory (ANL), Indiana University (IU), Rutgers University, the University of Tennessee, Knoxville (UTK), and the University of Virginia(UVA). SBI repositories include data, code, and all relevant collateral artifacts, that the science and engineering community needs to use and reuse these data sets and surrogates. SBI repositories generate active research from both participants in SBI and the broader AI and domain science communities. This project develops surrogates that use several different neural nets to learn and quickly infer the results of simulations and data systems and capture them as surrogate benchmarks with a rich set of metadata, covering. Data; Model; Metrics specification; Machine specification; Science, Speed, Power Results, We research FAIR metadata for these benchmarks. We develop application surrogate examples as benchmarks across many fields (ANL, UTK, IU, UVA). We also study non Surrogate benchmarks that have many common features and similar issues regarding FAIRness. We work with MLCommons (UVA, UTK), which is a major machine learning benchmarking activity where we get metadata ontologies, software, and benchmarks, benchmarks have datasets, models, and metadata, and they need a technical framework developed by UTK and Rutgers and deployed by UVA. We study features of Surrogates, including performance, training set size, and uncertainty quantification (Rutgers, UVA and IU).

97 MATHEMATICS AND COMPUTING↗

FAIR Surrogate Benchmarks Supporting AI and Simulation Research (Final Report)

Computational Science is being revolutionized by integrating AI and simulation and, in particular, by deep learning surrogate models that can replace all or part of traditional large‐scale HPC computations. Such surrogates can achieve remarkable performance improvements, as much as several orders of magnitude, and save both compute time and energy. The Surrogate Benchmark Initiative (SBI) project creates a community repository and FAIR (Findable, Accessible, Interoperable, and Reusable) data ecosystem for HPC application surrogate benchmarks. The SBI team comes from Argonne National Laboratory (ANL), Indiana University (IU), Rutgers University, the University of Tennessee, Knoxville (UTK), and the University of Virginia (UVA). SBI repositories include data, code, and all relevant collateral artifacts that the science and engineering community need to use and reuse these data sets and surrogates. SBI repositories generate active research from both the participants in SBI and the broad community of AI and domain scientists. This project develops surrogates that use several different neural nets to learn and quickly infer the results of simulations and data systems and captures them as surrogate benchmarks with a rich set of metadata covering: Data; Model; Metrics specification; Machine specification; and Science, Speed, and Power Results. We research FAIR metadata for these benchmarks. We develop application surrogate examples as benchmarks across many fields (ANL, UTK, IU, UVA). We also study non-Surrogate benchmarks that have many common features and similar issues as regards FAIRness. We work with MLCommons (UVA, UTK), which is a major machine learning benchmarking activity where we get metadata ontologies, software, and benchmarks, Benchmarks have datasets, models, and metadata and they need a technical framework developed by UTK and Rutgers and deployed by UVA. We study features of Surrogates including performance, training set size, and uncertainty quantification (Rutgers, UVA and IU).

97 MATHEMATICS AND COMPUTING↗

Towards Digital and Performance-Based Supervisory HVAC Control Delivery

Upgrading supervisory HVAC control in commercial buildings is one of the most attractive decarbonization tools at our disposal. Modern controls are software programs and can in theory be deployed at scale and with a low up-front carbon "pulse". In practice, however, control delivery is a disjointed and inefficient process, dominated by manual handoffs of imprecise English language documents. A particularly high barrier exists between control implementation and building energy modeling (BEM) which results in control sequences typically not being tested for correctness or performance before implementation. Together with industry partners, DOE and the national labs are developing an ecosystem of tools and standards that can support fully digital performance-based control delivery workflows. This paper describes this ecosystem, which consists of three mutually supportive efforts. Semantic models of buildings and their systems enable automatic configuration and installation of control software. Platform-neutral control descriptions separate control algorithms from control platforms and enable the creation of libraries of reference control implementations. Dynamic whole-building energy-control simulation that can execute physically realistic control sequences makes it possible to test and evaluate the performance of control sequences and then directly compile them for installation and execution in control systems. In addition to digitizing and streamlining project-level control delivery, these standards and related software support benchmarking of control algorithms, both rule-based and optimization-based, and help to both advance the state of the art and to implement ratings and programs that encourage the adoption of high-performance control.

building controls↗

Update to the Microcontroller Benchmark for Radiation Testing

LANL developed a benchmark of software code for radiation testing of microprocessors several years ago, and it was published under an open-source license on GitHub. Publishing the software is necessary for other researchers to adopt and implement this benchmark for radiation testing of other microprocessors to standardize test practices so that test data can be compared across different microprocessors. The original codes have been used several times by other organizations to test a wide range of microcontrollers and microprocessors. After several years of research, LANL is ready to update the benchmark. Changes include: 1. Addition of new codes that allow common software codes to be tested, 2. Addition of new codes that instrument more microprocessor circuitry, 3. Addition of input patterns that allow for a more compressive understanding of how the memory layout affects the sensitivity to radiation-induced faults and better use of automated test pattern generation standards, and 4. Modification of current codes for faster and more resilient detection, reporting and correction of radiation-induced faults. These codes have been tested by LANL researchers over the last few years, which has been published in the open literature. As the code base for the new benchmarks are stable, it is time to release the update to the GitHub repository, where the original codes were released.

Quinn, Heather↗

Uniformly Ordered Binary Decision Algorithm for Benchmark Experiment Correlations in Whisper Validation

When performing a validation exercise for determining the upper subcritical limit of a nuclear criticality safety application, an analyst should select and perform a statistical analysis on a population of benchmark experiments that are neutronically similar to the application. The size of this population should be sufficiently large such that the statistical analysis has a high degree of confidence that the bias plus bias uncertainty (calculational margin) has been accurately quantified. A complication arises because many benchmark experiments share common components, leading to correlations in their measured effective multiplication factors. Correlations between benchmark experiments within the population reduces its predictive power. This motivates the need for methods that consider benchmark experiment correlations and ensure adequate statistical significance of results. The Whisper code is a statistical analysis pack- age that incorporates nuclear data sensitivity coefficients from MCNP to assess benchmark experiment similarity and then performs an extreme-value analysis to estimate the bias plus bias uncertainty. The original methodology in Whisper does not consider the effect of benchmark experiment correlations when making this estimation, and this summary proposes the uniformly ordered binary decision algorithm to address this shortcoming. The original methodology in Whisper computes similarity coefficients ck for an application compared to all benchmark experiments in its library and develops weighting factors for a selected population proportional to the ck values. The methodology can be interpreted as statistically emulating a validation exercise for a particular application where the weighting factors may be viewed as the likelihood that an analyst would include a particular benchmark experiment within the population. The effective sample size of the population is the expected or mean number of benchmark experiments in the population. The uniformly ordered binary decision algorithm identifies clusters of correlated benchmark experiments within the population and then computes adjusted weighting factors based on the magnitude of the correlation coefficients within the cluster to compute a reduced effective sample size accounting for the lower information content because of correlations. Benchmark experiments within the cluster are ordered randomly with equal probability and probabilistic decisions are made as to whether a benchmark. experiment within the cluster should treated as redundant with a previous one; if two redundant benchmark experiments are included, then the conservative worst case bias plus bias uncertainty is used and the pair is counted as a single benchmark experiment in the population. Results are provided for HEU solutions in a research version of the Whisper software using benchmark experiment correlations provided by DICE, the Database for the International Criticality Safety Benchmark Evaluation Project (ICSBEP). These show that there can be a significant increase in the bias plus bias uncertainty because the effective sample size is reduced, and therefore the algorithm, needing to meet sample size requirements, expands the benchmark experiment population by accepting less similar benchmark experiments that would have otherwise not been included.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Tools for unbinned unfolding

Machine learning has enabled differential cross section measurements that are not discretized. Going beyond the traditional histogram-based paradigm, these unbinned unfolding methods are rapidly being integrated into experimental workflows. Here, in order to enable widespread adaptation and standardization, we develop methods, benchmarks, and software for unbinned unfolding. For methodology, we demonstrate the utility of boosted decision trees for unfolding with a relatively small number of high-level features. This complements state-of-the-art deep learning models capable of unfolding the full phase space. To benchmark unbinned unfolding methods, we develop an extension of existing dataset to include acceptance effects, a necessary challenge for real measurements. Additionally, we directly compare binned and unbinned methods using discretized inputs for the latter in order to control for the binning itself. Lastly, we have assembled two software packages for the OmniFold unbinned unfolding method that should serve as the starting point for any future analyses using this technique. One package is based on the widely-used RooUnfold framework and the other is a standalone package available through the Python Package Index (PyPI).

47 OTHER INSTRUMENTATION↗

Bison Verification and Validation Activities for TRISO

Numerical modeling and simulation (M&S) tools play a key role in the research, development, and overall safety assessments of next-generation nuclear energy systems. One such tool, Bison, is a nuclear fuel performance code that is applicable to many fuel forms (e.g., light-water reactor fuel, oxide and metallic fuel for fast reactors, tri-structural isotropic (TRISO) fuel, and plate fuel), and it uses the finite element method to model the thermo- mechanical response of nuclear fuels. One fuel form widely utilized in Generation-IV high-temperature gas-cooled and fluoride- salt-cooled nuclear reactor concepts is TRISO fuel. Recently, Bison’s capabilities were significantly expanded to enable it to model the performance of TRISO particles and compacts. It is important that Bison’s computational results be reliable and predictive, since this code is used to inform high-consequence decisions. The various processes developed to address this issue generally entail two fundamental steps: verification and validation (V&V). Verification ensures that the code functions correctly and is reliable. Code/solution verification, code benchmark, and software quality assurance exercises are examples of verification activities. On the other hand, validation is the process of assessing a code’s capability to accurately model physical problems. Comparisons between code results and experiments quantify the validation level. Application of V&V procedures is crucial to the development of computational tools that are free of coding mistakes and can accurately represent reality. The current study presents an overview of Bison V&V activities relevant to the TRISO fuel concept, which include code/solution verification exercises, CRP-6 Benchmark—a Coordinated Research Program through the International Atomic Energy Agency (IAEA)—exercises, and validation exercises with the Advanced Gas Reactor (AGR)- 1/2/3/4 experiment series.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Noisy intermediate-scale quantum algorithms

A universal fault-tolerant quantum computer that can efficiently solve problems such as integer factorization and unstructured database search requires millions of qubits with low error rates and long coherence times. While the experimental advancement toward realizing such devices will potentially take decades of research, noisy intermediate-scale quantum (NISQ) computers already exist. These computers are composed of hundreds of noisy qubits, i.e., qubits that are not error corrected, and therefore perform imperfect operations within a limited coherence time. In the search for achieving quantum advantage with these devices, algorithms have been proposed for applications in various disciplines spanning physics, machine learning, quantum chemistry, and combinatorial optimization. The overarching goal of such algorithms is to leverage the limited available resources to perform classically challenging tasks. In this review, a thorough summary of NISQ computational paradigms and algorithms is provided. The key structure of these algorithms and their limitations and advantages are discussed. Finally, a comprehensive overview of various benchmarking and software tools useful for programming and testing NISQ devices is additionally provided.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Direct Estimation of Parameters in ODE Models Using WENDy: Weak-Form Estimation of Nonlinear Dynamics

Abstract We introduce the Weak-form Estimation of Nonlinear Dynamics (WENDy) method for estimating model parameters for non-linear systems of ODEs. Without relying on any numerical differential equation solvers, WENDy computes accurate estimates and is robust to large (biologically relevant) levels of measurement noise. For low dimensional systems with modest amounts of data, WENDy is competitive with conventional forward solver-based nonlinear least squares methods in terms of speed and accuracy. For both higher dimensional systems and stiff systems, WENDy is typically both faster (often by orders of magnitude) and more accurate than forward solver-based approaches. The core mathematical idea involves an efficient conversion of the strong form representation of a model to its weak form, and then solving a regression problem to perform parameter inference. The core statistical idea rests on the Errors-In-Variables framework, which necessitates the use of the iteratively reweighted least squares algorithm. Further improvements are obtained by using orthonormal test functions, created from a set of $$C^{\infty }$$ C ∞ bump functions of varying support sizes.We demonstrate the high robustness and computational efficiency by applying WENDy to estimate parameters in some common models from population biology, neuroscience, and biochemistry, including logistic growth, Lotka-Volterra, FitzHugh-Nagumo, Hindmarsh-Rose, and a Protein Transduction Benchmark model. Software and code for reproducing the examples is available at https://github.com/MathBioCU/WENDy .

97 MATHEMATICS AND COMPUTING↗

Benchmarking second and third-generation sequencing platforms for microbial metagenomics

Shotgun metagenomic sequencing is a common approach for studying the taxonomic diversity and metabolic potential of complex microbial communities. Current methods primarily use second generation short read sequencing, yet advances in third generation long read technologies provide opportunities to overcome some of the limitations of short read sequencing. Here, we compared seven platforms, encompassing second generation sequencers (Illumina HiSeq 300, MGI DNBSEQ-G400 and DNBSEQ-T7, ThermoFisher Ion GeneStudio S5 and Ion Proton P1) and third generation sequencers (Oxford Nanopore Technologies MinION R9 and Pacific Biosciences Sequel II). We constructed three uneven synthetic microbial communities composed of up to 87 genomic microbial strains DNAs per mock, spanning 29 bacterial and archaeal phyla, and representing the most complex and diverse synthetic communities used for sequencing technology comparisons. Our results demonstrate that third generation sequencing have advantages over second generation platforms in analyzing complex microbial communities, but require careful sequencing library preparation for optimal quantitative metagenomic analysis. Our sequencing data also provides a valuable resource for testing and benchmarking bioinformatics software for metagenomics.

59 BASIC BIOLOGICAL SCIENCES↗

NMSBA: Continuous Application Benchmarking & Analysis – CABA

The CABA project is a test of a new tool, Survey,developed by Trenza to be used not only for benchmarking or profiling programs but also to allow incorporation of the information provided by Survey to be utilized in a CI, continuous integration,tool such as GitLab CI.Survey is foremost a means of assessing code performance in terms of time and operations which for computer programmers is known as benchmarking.

97 MATHEMATICS AND COMPUTING↗

Mesh independency in topology optimization

The topology optimization community has regularly employed optimization algorithms from the operations research community. However, these algorithms are implemented in the sequence space l2 instead of the proper function space where the design variable resides. In this software, we benchmark the mesh independent implementation of pyMMAopt with three common problems in topology optimization. We show how the volume fraction variable discretization on non-uniform meshes affects the convergence of l2 based optimization algorithms and how pyMMAopt successfully achieves mesh independent convergence.

Salazar De Troya, Miguel↗

CI/CD Efforts for Validation, Verification and Benchmarking OpenMP Implementations

Software developers must adapt to keep up with the changing capabilities of platforms so that they can utilize the power of High-Performance Computers (HPC), including exascale systems. OpenMP, a directive-based parallel programming model, allows developers to include directives to existing C, C++, or Fortran code to allow node level parallelism without compromising performance. This paper describes our CI/CD efforts to provide easy evaluation of the support of OpenMP across different compilers using existing testsuites and benchmark suites on HPC platforms. Our main contributions include (1) the set of a Continuous Integration (CI) and Continuous Development (CD) workflow that captures bugs and provides faster feedback to compiler developers, (2) an evaluation of OpenMP (offloading) implementations supported by AMD, HPE, GNU, LLVM, and Intel, and (3) evaluation of the quality of compilers across different heterogeneous HPC platforms. With the comprehensive testing through the CI/CD workflow, we aim to provide a comprehensive understanding of the current state of OpenMP (offloading) support in different compilers and heterogeneous platforms consisting of CPUs and GPUs from NVIDIA, AMD, and Intel.

Jarmusch, Aaron↗