Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Benchmarking Software”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Assessing and benchmarking the fidelity of posterior inference methods for astrophysics data analysis

In this era of large and complex astronomical survey data, interpreting, validating, and comparing inference techniques becomes increasingly difficult. This is particularly critical for emerging inference methods like Simulation-Based Inference (SBI), which offer significant speedup potential and posterior modeling flexibility, especially when deep learning is incorporated. We present a study to assess and compare the performance and uncertainty prediction capability of Bayesian inference algorithms – from traditional MCMC sampling of analytic functions to deep learning-enabled SBI. We focus on testing the capacity of hierarchical inference modeling in those scenarios. Before we extend this study to cosmology, we first use astrophysical simulation data to ensure interpretability. We demonstrate a probabilistic programming implementation of hierarchical and non-hierarchical Bayesian inference using simulations derived from the DeepBench software library, a benchmarking tool developed by our group that generates simple and controllable astrophysical objects from first principles. This study will enable astronomers and physicists to harness the inference potential of these methods with confidence.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Ranking and Classifying AI Benchmarks

We created a set of standards to efficiently evaluate AI benchmarks through objective means. Although prevalent, especially in recent times, AI benchmarks have no single way to measure their effectiveness. The MLCommons team provided a set of criteria for evaluating benchmarks, although the criteria lacks a clearly defined set of evaluation rules. We created a rubric with preset factors to efficiently and objectively evaluate a benchmark s quality. We created a software framework for processing lists of benchmarks for visualization. The framework and rating system allows researchers to quickly check if their benchmarks are effective.

Shiraishi, Reece C. [Cornell U.]↗

Integrating chromatin conformation information in a self-supervised learning model improves metagenome binning

Metagenome binning is a key step, downstream of metagenome assembly, to group scaffolds by their genome of origin. Although accurate binning has been achieved on datasets containing multiple samples from the same community, the completeness of binning is often low in datasets with a small number of samples due to a lack of robust species co-abundance information. In this study, we exploited the chromatin conformation information obtained from Hi-C sequencing and developed a new reference-independent algorithm, Metagenome Binning with Abundance and Tetra-nucleotide frequencies—Long Range (metaBAT-LR), to improve the binning completeness of these datasets. This self-supervised algorithm builds a model from a set of high-quality genome bins to predict scaffold pairs that are likely to be derived from the same genome. Then, it applies these predictions to merge incomplete genome bins, as well as recruit unbinned scaffolds. We validated metaBAT-LR’s ability to bin-merge and recruit scaffolds on both synthetic and real-world metagenome datasets of varying complexity. Benchmarking against similar software tools suggests that metaBAT-LR uncovers unique bins that were missed by all other methods.

59 BASIC BIOLOGICAL SCIENCES↗

Classifying and rating AI benchmarks

We created a set of standards to efficiently evaluate AI benchmarks through objective means. Although prevalent, especially in recent times, AI benchmarks have no single way to measure their effectiveness. The MLCommons team provided a set of criteria for evaluating benchmarks, although the criteria lacks a clearly defined set of evaluation rules. We created a rubric with preset factors to efficiently and objectively evaluate a benchmark’s quality. We created a software framework for processing lists of benchmarks for visualization. The framework and rating system allows researchers to quickly check if their benchmarks are effective.

Shiraishi, Reece [Fermilab]↗

How simple is software defect detection?

This study benchmarks several FSS techniques and reports several studies where a large set metrics were reduced to a handful with little loss of detection accuracy. This result raises the possibility that software defect detection may be much simpler than previously believed.

feature subset selection software fault model soft↗

Gaussian Process Optimization of Sensitivity-Based Similarity Metrics between New Nuclear Applications and New/Existing Benchmarks

Confidently designing safe, new nuclear criticality experiments requires expert judgement, which could take years of experience. Sensitivity/uncertainty (S/U) analysis can be utilized by less experienced individuals to conservatively estimate uncertainties in important parameters, such as k eff , in newly proposed nuclear applications. This type of analysis relies on matching new nuclear applications with existing benchmark experiments. The Whisper-1.1 software package included in MCNP6.2 ®1 contains more than 1,100 International Criticality Safety Benchmark Evaluation Project (ICSBEP) benchmarks. These benchmarks however rarely match new nuclear applications. The number of benchmarks available to match a given set of materials or geometric configurations varies significantly. Furthermore, recently performed benchmark experiments may not have had enough time to be properly documented and published. Benchmarks are vital for determining the accuracy of nuclear data and can assist nuclear physics and evaluators in improving nuclear data libraries. Exhaustively exploring the parameter space using simulations with software such as MCNP is too computationally expensive. In this work, Gaussian process optimization was implemented to reduce the number of simulations needed for optimization over multiple parameters. This optimization scheme was designed to selectively generate new benchmarks with high sensitivity-based similarity metrics to user-defined nuclear applications. Two benchmark models of spherically nested shells containing plutonium, uranium, tantalum, and water were used in the optimization to match an application containing plutonium plates, stacked in a tantalum reflector, surrounded by water.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Feature review of photovoltaic modeling software utilizing blind performance assessment

While confidence in photovoltaic (PV) modeling software has always been essential, the rapid pace of new PV plant developments makes accuracy and credibility more critical than ever. Independent assessments, particularly through blind modeling comparisons, are therefore necessary to ensure unbiased benchmarking across PV modeling software. Previous studies have been limited by a narrow range of models compared, anonymized results, or system size. This study presents results from the first-ever onymous blind modeling comparison, evaluated using both lab- and utility-scale fixed-tilt, monofacial, south-facing systems at sub-hourly time intervals. Seven commercially used PV software tools were compared: 3E SynaptiQ, PlantPredict, PVsyst, RatedPower, SAM, SolarFarmer, and Solargis Evaluate. Predictions were submitted directly by software representatives, providing unique insights into each software’s implementation and resulting prediction behavior. Notable features, including plane-of-array (POA) transposition model, module temperature model, shading model, and performance model were analyzed and compared. Four summary tables compile these features of the software, serving as a resource to help users understand the methodological differences and select the most suitable software for their applications. The software tools show deviations from mean error in annual yield up to 2.5 % in the lab-scale system, increasing to 6.0 % for the utility-scale system. These differences arise from a combination of user decisions and the inherent behavior of the software, indicating the need for continuous and rigorous validation of modeling methods using these software tools against complex, real-world systems.

14 SOLAR ENERGY↗

DFT-based QM/MM with Particle-Mesh Ewald for Direct, Long-Range Electrostatic Embedding

In this work, we present a DFT-based, QM/MM implementation with long-range electrostatic embedding achieved by direct real-space integration of the particle mesh Ewald (PME) computed electrostatic potential. The key transformation is the interpolation of the electrostatic potential from the PME grid to the DFT quadrature grid, from which integrals are easily evaluated utilizing standard DFT machinery. We provide benchmarks of the numerical accuracy with choice of grid size and real-space corrections, and demonstrate that good convergence is achieved while introducing nominal computational overhead. Furthermore, the approach requires only small modification to existing software packages, as is demonstrated with our implementation in the OpenMM and Psi 4 software. After presenting convergence benchmarks, we evaluate the importance of long-range electrostatic embedding in three solute/solvent systems modeled with QM/MM. Water and BMIM/BF 4 ionic liquid were considered as "simple" and "complex" solvents respectively, with water and p-phenylenediamine (PPD) solute molecules treated at QM level of theory. While electrostatic embedding with standard real-space truncation may introduce negligible error for simple systems such as water solute in water solvent, errors become more significant when QM/MM is applied to complex solvents such as ionic liquids. An extreme example is the electrostatic embedding energy for oxidized PPD in BMIM/BF 4 for which real-space truncation produces severe error even at 2-3 nm cutoff distances. This latter example illustrates that utilization of QM/MM to compute redox potentials within concentrated electrolytes/ionic media requires carefully chosen long-range electrostatic embedding algorithms, with our presented algorithm providing a general and robust approach.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

EQ_phase_detection

The EQ_phase_detection software is designed to scan continuous daily waveforms to detect earthquake phase arrivals from local to regional (150 km) events. The detections are made with a deep learning encoder-decoder model. When the model detects an earthquake in the waveforms, a second model is implemented to classify the first arriving motions. Both deep learning models are trained with the Tensorflow package using publicly available benchmark data sets. The software input is a path to a directory that contains waveforms in mseed format and the associated response files in xml format. The output is a data table of time stamped detections, signal amplitude, signal-to-noise ratio, and softmax probability of the detection in a generic format applicable to post-processing association algorithms for event locations. Additionally, the p-wave and s-wave waveforms are saved in a data table for rapid access when producing improved locations using correlation-based techniques. The software is designed for multiprocessing with multiple GPU’s for rapid processing of large data sets. The configuration file provides flexibility in the trained models implemented and allows access to multiple models trained for different sampling rates or input dimensions. This is particularly useful for regions with multiple networks that do not have the same data parameters.

Johnson, Christopher↗

Preliminary Computational Analysis of the (HIRENASD) Configuration in Preparation for the Aeroelastic Prediction Workshop

This paper presents preliminary computational aeroelastic analysis results generated in preparation for the first Aeroelastic Prediction Workshop (AePW). These results were produced using FUN3D software developed at NASA Langley and are compared against the experimental data generated during the HIgh REynolds Number Aero- Structural Dynamics (HIRENASD) Project. The HIRENASD wind-tunnel model was tested in the European Transonic Windtunnel in 2006 by Aachen University0s Department of Mechanics with funding from the German Research Foundation. The computational effort discussed here was performed (1) to obtain a preliminary assessment of the ability of the FUN3D code to accurately compute physical quantities experimentally measured on the HIRENASD model and (2) to translate the lessons learned from the FUN3D analysis of HIRENASD into a set of initial guidelines for the first AePW, which includes test cases for the HIRENASD model and its experimental data set. This paper compares the computational and experimental results obtained at Mach 0.8 for a Reynolds number of 7 million based on chord, corresponding to the HIRENASD test conditions No. 132 and No. 159. Aerodynamic loads and static aeroelastic displacements are compared at two levels of the grid resolution. Harmonic perturbation numerical results are compared with the experimental data using the magnitude and phase relationship between pressure coefficients and displacement. A dynamic aeroelastic numerical calculation is presented at one wind-tunnel condition in the form of the time history of the generalized displacements. Additional FUN3D validation results are also presented for the AGARD 445.6 wing data set. This wing was tested in the Transonic Dynamics Tunnel and is commonly used in the preliminary benchmarking of computational aeroelastic software.

Chwalowski, Pawel↗

Resonant ultrasound spectroscopy probe for in-situ neutron scattering measurements

Resonant ultrasound spectroscopy (RUS) is an efficient, nondestructive technique to study the elastic properties of solids. A low-temperature (2-300 K) probe has been assembled and tested at the NOMAD beamline at the Spallation Neutron Source at Oak Ridge National Laboratory to assess the probe's neutronic properties, data acquisition system, and compatibility with existing sample environment. A case study on a bulk metallic glass, La 65 Cu 20 Al 10 Co 5, served to benchmark both hardware and software developments. The elastic constants of the metallic glass were determined as a function of temperature between 4 and 300 K and were used to guide neutron diffraction measurements at NOMAD. Tracking of a specific RUS peak center frequency and width enabled live monitoring of the sample temperature evolution that lagged thermometry by upwards of 40 K near room temperature. Assembly of a high-temperature (300-875 K) probe is underway and both probes are scheduled to be available to users by late-2021. Our aim is to provide users with live monitoring of an intrinsic variable at the neutron scattering beamlines, in addition to existing controls, to monitor the state of their samples and its elastic moduli and make informed decisions in real time.

Torres, James↗

VERA BWR progression problems

During the first phase of the Consortium for Advanced Simulation of Light Water Reactors (CASL) program, the Virtual Environment for Reactor Applications (VERA) was developed with a focus on capabilities for high-fidelity, multiphysics simulation of pressurized water reactors (PWRs). During this development effort, a set of progression problems was created ranging from smaller pin cell calculations to larger 3D full-core calculations. These progression problems helped to guide the development of the software and served as benchmarks against which to test VERA. Since 2019, efforts have been made to extend the capabilities of VERA to model boiling water reactors (BWRs). Because BWR simulations come with many unique challenges, a set of BWR progression problems was developed to aid in this new effort. The BWR progression problems range from 2D lattice calculations to 3D mini-core problems, and reference neutronic solutions were computed using continuous-energy Monte Carlo codes. MPACT, one of the neutronics code in VERA, was benchmarked using the BWR progression problems. The code is capable of computing solutions to all problems. The eigenvalues computed by MPACT agree well with the Monte Carlo reference solutions. (authors)

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

CAE for thermal management of aerospace electronic boards using the BETAsoft program

Aerospace electronic boards require special attention to thermal management due to constraints such as their need to be light, small, and maintain high power densities. Also, cooling is mainly through conductive and radiative modes with minor or negligible convective cooling. Due to these particular requirements, thermal design has become an integrated part of the electronic design process in order to avoid expensive repeat prototyping and to ensure high reliability. To achieve high speed simulations, the BETAsoft code uses semi-empirical formulations and an advanced finite difference scheme that incorporates local adaptive grids. Detailed conduction, convection and radiation heat transfer is considered. Various benchmark verifications of the software simulation compared to infrared images typically prove to be within 10% of each other. The thermal analysis of a sample avionic card in a natural convection environment is shown. Then, the individual effects of attaching metal screws to the casing, increasing radiative emissivities of the casing, increasing the conductance of the wedge lock, adding an aluminum core to the board, adding metal strips in board layers, inserting conduction pads under components, and adding heat sinks to components are demonstrated.

Bobish, Kimberly↗

Creating Benchmark Data for Artificial Intelligence and Machine Learning Space Biology Research

To identify an appropriate AI/ML approach for a specific problem, the best practice is to measure algorithm performance through the benchmarking process. A scientific benchmark consists of an AI-ready dataset and a reference implementation on a specific scientific question. The NASA Science Mission Directorate (SMD) has started the “Benchmark Initiative for AI/ML to create scientific benchmark datasets in three applications: 1) scientific benchmarking, which finds the best algorithm for a specific problem; 2) application benchmarking, which measures algorithm performance against a set of parameters; and 3) system benchmarking, which evaluates performance of hardware and software architecture. Currently, there are no standardized datasets available to benchmark AI/ML algorithms in the domain of space biology. In this work, we constructed two AI/ML-ready biological datasets from experiments in space-flown mice: cellular imaging and RNA-seq. First, radiation-exposed immune cells harbor DNA damage foci that can be fluorescently marked to visualize the amount of damage following exposure to ionizing radiation. However, such large datasets are difficult to analyze visually, due to imaging inconsistencies and human bias, and classical image processing approaches can fail on imaging artifacts. AI/ML are therefore exciting alternative, providing the speed of machines and the accuracy of humans. We have made this dataset available at https://registry.opendata.aws/bps_microscopy/. Second, high-throughput nucleic acid sequencing (DNA-seq, RNA-seq) has become widespread in biomedical research due to the growing availability and affordability of these assays. However, most sequencing datasets suffer from high dimensionality and low sample count. In this work, we used a generative adversarial network to synthesize a standardized, AI-ready, publicly available benchmark dataset for space biology RNA-seq data with sufficient space-flown and ground control mouse liver samples from NASA GeneLab. This dataset is available at https://registry.opendata.aws/bps_rnaseq/. These datasets are now fully open the Space Biology community to test their favorite AI/ML approaches.

James Casaletto↗

The TSO Logic and G2 Software Product

This internship assignment for spring 2014 was at John F. Kennedy Space Center (KSC), in NASAs Engineering and Technology (NE) group in support of the Control and Data Systems Division (NE-C) within the Systems Hardware Engineering Branch. (NEC-4) The primary focus was in system integration and benchmarking utilizing two separate computer software products. The first half of this 2014 internship is spent in assisting NE-C4s Electronics and Embedded Systems Engineer, Kelvin Ruiz and fellow intern Scott Ditto with the evaluation of a newly piece of software, called G2. Its developed by the Gensym Corporation and introduced to the group as a tool used in monitoring launch environments. All fellow interns and employees of the G2 group have been working together in order to better understand the significance of the G2 application and how KSC can benefit from its capabilities. The second stage of this Spring project is to assist with an ongoing integration of a benchmarking tool, developed by a group of engineers from a Canadian based organization known as TSO Logic. Guided by NE-C4s Computer Engineer, Allen Villorin, NASA 2014 interns put forth great effort in helping to integrate TSOs software into the Spaceport Processing Systems Development Laboratory (SPSDL) for further testing and evaluating. The TSO Logic group claims that their software is designed for, monitoring and reducing energy consumption at in-house server farms and large data centers, allows data centers to control the power state of servers, without impacting availability or performance and without changes to infrastructure and the focus of the assignment is to test this theory. TSOs Aaron Rallo Founder and CEO, and Chris Tivel CTO, both came to KSC to assist with the installation of their software in the SPSDL laboratory. TSOs software is installed onto 24 individual workstations running three different operating systems. The workstations were divided into three groups of 8 with each group having its own operating system. The first group is comprised of Ubuntus Debian -based Linux the second group is windows 7 Professional and the third group ran Red Hat Linux. The highlight of this portion of the assignment is to compose documentation expressing the overall impression of the software and its capabilities.

TSO Logic↗

Final Report of the NASA Office of Safety and Mission Assurance Agile Benchmarking Team

To ensure that the NASA Safety and Mission Assurance (SMA) community remains in a position to perform reliable Software Assurance (SA) on NASAs critical software (SW) systems with the software industry rapidly transitioning from waterfall to Agile processes, Terry Wilcutt, Chief, Safety and Mission Assurance, Office of Safety and Mission Assurance (OSMA) established the Agile Benchmarking Team (ABT). The Team's tasks were: 1. Research background literature on current Agile processes, 2. Perform benchmark activities with other organizations that are involved in software Agile processes to determine best practices, 3. Collect information on Agile-developed systems to enable improvements to the current NASA standards and processes to enhance their ability to perform reliable software assurance on NASA Agile-developed systems, 4. Suggest additional guidance and recommendations for updates to those standards and processes, as needed. The ABT's findings and recommendations for software management, engineering and software assurance are addressed herein.

Software Assurance↗

A Review of Candidates for a Validation Data Set for High-Assay Low-Enrichment Uranium Fuels

Many advanced reactor concept designs rely on high-assay low-enriched uranium (HALEU) fuel, enriched up to approximately 19.75% 235 U by weight. Efforts are underway by the US government to increase HALEU production in the United States to meet anticipated needs. However, very few data exist for validation of computational models that include HALEU, beyond a few fresh fuel benchmark specifications in the International Reactor Physics Experiment Evaluation Project. Nevertheless, there are other data with potential value available for developing into quality benchmarks for use in data- and software-validation efforts. This paper reviews the available evaluated HALEU fuel benchmarks and some of the potentially relevant benchmarks for fresh highly enriched uranium. It then introduces experimental data for HALEU fuel irradiated at Idaho National Laboratory, from relatively recent irradiation programs at the Advanced Test Reactor. Such data should be evaluated and, if valuable, collected into detailed benchmark specifications to meet the needs of HALEU-based reactor designers.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Benchmark study of a new simplified DFN model for shearing of intersecting fractures and faults

It is challenging to quantitatively predict shearing of intersecting fractures/faults because of dynamic frictional contacts accompanied by possible nonlinear rock deformation. To address such challenges, a new conceptual model—the simplified DFN model—was proposed and validated by Hu et al. 46 to use major paths (MPs) to represent complicated DFNs for calculation of shearing. In this work, we conducted a benchmark study for three examples that involve different levels of complexity of intersecting fractures, and correspondingly different numbers of MPs. The codes and software that were used in the benchmark cover a range of continuum, discontinuum and hybrid numerical methods: NMM (LBNL), FLAC3D (LBNL), GBDEM (KIGAM), FRACOD (DynaFrax), and CASRock (CAS). The general consistency between DFN and MP cases as predicted by all the codes/software demonstrates that major paths can be used to simplify the geometry of DFNs in a wide range of software. Disagreement in results made by some software and potential future improvements are discussed. We show that (1) shearing of one or multiple major fractures can be reduced if there are multiple smaller intersecting fractures in that area, which is a useful basis for understanding and controlling induced seismicity and merits further analysis, and (2) the agreement achieved in the benchmark examples provide confidence that the simplified DFN model is a promising conceptual model that can be used for different types of numerical approaches and software for simplifying the analysis of the shearing of intersecting fractures and faults.

58 GEOSCIENCES↗