Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Benchmarking Software”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Comparison of GSFC and JPL VLBI modeling software: Benchmark

The Very Long Base Interferometry. The VLBI modeling software packages CALC 6.0 (GSFC) and MASTERFIT (JPL) are compared in some detail. Theoretical model delays are calculated for a set of 120 fictitious observations which involve a variety of sources, baselines, and antennas. Discrepancies between the total delays given by the two programs are of the order of 2 cm (RMS). The modeling of antenna offsets appears to account for approximately half of this difference. Relativistic bending contributions to the delay differ by 3 cm (RMS), and the there appears to be some mutual cancellation of errors involving antenna offsets, bending, and the effects of the two different Solar System ephemerides employed by CALC and MASTERFIT. This cancellation has not been completely characterized here.

Sovers, O. J.↗

Evaluation of Ocean Biogeochemistry and Carbon Cycling in CMIP Earth System Models With the International Ocean Model Benchmarking (IOMB) Software System

Abstract The International Ocean Model Benchmarking (IOMB) software package is a new community resource that we use here to evaluate surface and upper ocean biogeochemical variables and integrated anthropogenic carbon uptake from earth system models (ESMs) contributing to the 5th and 6th phases of the Coupled Model Intercomparison Project (CMIP5 and CMIP6). IOMB generates graphics and tables for systematically comparing model predictions against multiple datasets. Our analysis reveals some improvement in the multi‐model mean from CMIP5 to CMIP6 for most of the variables we examined. Compared to data‐constrained estimates of ocean anthropogenic carbon uptake for the 1994–2007 period, negative biases exist for many models between 30 and 50°S. Global model estimates of anthropogenic carbon uptake for the same period do not change significantly from CMIP5 to CMIP6, with the combined ensemble mean estimate of 27.8 ± 0.5 Pg C lower than a data‐constrained estimate of 33.0 ± 4.0 Pg C. At the same time, the change in the natural carbon inventory from CMIP is estimated to be a source of 0.7 ± 0.3 Pg C, which is considerably smaller in magnitude than a data‐constrained estimate of 5.0 ± 3.0 Pg C. With chlorofluorocarbon (CFC) predictions available for several models, we demonstrate that negative anthropogenic dissolved inorganic carbon biases coincide with negative biases in CFC concentration, highlighting the importance of weak exchange between the surface and interior ocean in regulating rates of anthropogenic carbon uptake. To examine the robustness of this attribution across the CMIP models, we calculate the global vertical temperature gradient between 200 and 1,000 m as a metric for global stratification and exchange between the surface and deeper waters. We find a linear relationship between the bias of the vertical temperature gradients and the bias in global anthropogenic carbon uptake, consistent with the hypothesis that model biases in anthropogenic carbon uptake are related to biases in surface‐to‐interior exchange by physical processes.

58 GEOSCIENCES↗

NASA Software Engineering Benchmarking Study

To identify best practices for the improvement of software engineering on projects, NASA's Offices of Chief Engineer (OCE) and Safety and Mission Assurance (OSMA) formed a team led by Heather Rarick and Sally Godfrey to conduct this benchmarking study. The primary goals of the study are to identify best practices that: Improve the management and technical development of software intensive systems; Have a track record of successful deployment by aerospace industries, universities [including research and development (R&D) laboratories], and defense services, as well as NASA's own component Centers; and Identify candidate solutions for NASA's software issues. Beginning in the late fall of 2010, focus topics were chosen and interview questions were developed, based on the NASA top software challenges. Between February 2011 and November 2011, the Benchmark Team interviewed a total of 18 organizations, consisting of five NASA Centers, five industry organizations, four defense services organizations, and four university or university R and D laboratory organizations. A software assurance representative also participated in each of the interviews to focus on assurance and software safety best practices. Interviewees provided a wealth of information on each topic area that included: software policy, software acquisition, software assurance, testing, training, maintaining rigor in small projects, metrics, and use of the Capability Maturity Model Integration (CMMI) framework, as well as a number of special topics that came up in the discussions. NASA's software engineering practices compared favorably with the external organizations in most benchmark areas, but in every topic, there were ways in which NASA could improve its practices. Compared to defense services organizations and some of the industry organizations, one of NASA's notable weaknesses involved communication with contractors regarding its policies and requirements for acquired software. One of NASA's strengths was its software assurance practices, which seemed to rate well in comparison to the other organizational groups and also seemed to include a larger scope of activities. An unexpected benefit of the software benchmarking study was the identification of many opportunities for collaboration in areas including metrics, training, sharing of CMMI experiences and resources such as instructors and CMMI Lead Appraisers, and even sharing of assets such as documented processes. A further unexpected benefit of the study was the feedback on NASA practices that was received from some of the organizations interviewed. From that feedback, other potential areas where NASA could improve were highlighted, such as accuracy of software cost estimation and budgetary practices. The detailed report contains discussion of the practices noted in each of the topic areas, as well as a summary of observations and recommendations from each of the topic areas. The resulting 24 recommendations from the topic areas were then consolidated to eliminate duplication and culled into a set of 14 suggested actionable recommendations. This final set of actionable recommendations, listed below, are items that can be implemented to improve NASA's software engineering practices and to help address many of the items that were listed in the NASA top software engineering issues. 1. Develop and implement standard contract language for software procurements. 2. Advance accurate and trusted software cost estimates for both procured and in-house software and improve the capture of actual cost data to facilitate further improvements. 3. Establish a consistent set of objectives and expectations, specifically types of metrics at the Agency level, so key trends and models can be identified and used to continuously improve software processes and each software development effort. 4. Maintain the CMMI Maturity Level requirement for critical NASA projects and use CMMI to measure organizations developing software for NASA. 5.onsolidate, collect and, if needed, develop common processes principles and other assets across the Agency in order to provide more consistency in software development and acquisition practices and to reduce the overall cost of maintaining or increasing current NASA CMMI maturity levels. 6. Provide additional support for small projects that includes: (a) guidance for appropriate tailoring of requirements for small projects, (b) availability of suitable tools, including support tool set-up and training, and (c) training for small project personnel, assurance personnel and technical authorities on the acceptable options for tailoring requirements and performing assurance on small projects. 7. Develop software training classes for the more experienced software engineers using on-line training, videos, or small separate modules of training that can be accommodated as needed throughout a project. 8. Create guidelines to structure non-classroom training opportunities such as mentoring, peer reviews, lessons learned sessions, and on-the-job training. 9. Develop a set of predictive software defect data and a process for assessing software testing metric data against it. 10. Assess Agency-wide licenses for commonly used software tools. 11. Fill the knowledge gap in common software engineering practices for new hires and co-ops.12. Work through the Science, Technology, Engineering and Mathematics (STEM) program with universities in strengthening education in the use of common software engineering practices and standards. 13. Follow up this benchmark study with a deeper look into what both internal and external organizations perceive as the scope of software assurance, the value they expect to obtain from it, and the shortcomings they experience in the current practice. 14. Continue interactions with external software engineering environment through collaborations, knowledge sharing, and benchmarking.

Rarick, Heather L.↗

DeepSurveySim: Simulation Software and Benchmark Challenges for Astronomical Observation Scheduling

Modern astronomical surveys have multiple competing scientific goals. Optimizing the observation schedule for these goals presents significant computational and theoretical challenges, and state-of-the-art methods rely on expensive human inspection of simulated telescope schedules. Automated methods, such as reinforcement learning, have recently been explored to accelerate scheduling. However, there do not yet exist benchmark data sets or user-friendly software frameworks for testing and comparing these methods. We present DeepSurveySim -- a high-fidelity and flexible simulation tool for use in telescope scheduling. DeepSurveySim provides methods for tracking and approximating sky conditions for a set of observations from a user-supplied telescope configuration. We envision this tool being used to produce benchmark data sets and for evaluating the efficacy of ground-based telescope scheduling algorithms, particularly for machine learning algorithms that would suffer in efficacy if limited to real data for training.We introduce three example survey configurations and related code implementations as benchmark problems that can be simulated with DeepSurveySim.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

DeepSurveySim: Simulation Software and Benchmark Challenges for Astronomical Observation Scheduling

Modern astronomical surveys have multiple competing scientific goals. Optimizing the observation schedule for these goals presents significant computational and theoretical challenges, and state-of-the-art methods rely on expensive human inspection of simulated telescope schedules. Automated methods, such as Reinforcement Learning (RL), have recently been explored to accelerate scheduling. DeepSurveySim provides methods for tracking and approximating sky conditions for a set of observations from a user-supplied telescope configuration.

79 ASTRONOMY AND ASTROPHYSICS↗

NASA Software Engineering Benchmarking Effort

Benchmarking was very interesting and provided a wealth of information (1) We did see potential solutions to some of our "top 10" issues (2) We have an assessment of where NASA stands with relation to other aerospace/defense groups We formed new contacts and potential collaborations (1) Several organizations sent us examples of their templates, processes (2) Many of the organizations were interested in future collaboration: sharing of training, metrics, Capability Maturity Model Integration (CMMI) appraisers, instructors, etc. We received feedback from some of our contractors/ partners (1) Desires to participate in our training; provide feedback on procedures (2) Welcomed opportunity to provide feedback on working with NASA

Godfrey, Sally↗

Computational Fluid Dynamic Modeling of Dry Cask Simulator with Crosswind

The purpose of this study is to create a STAR-CCM+ model of a Belowground Vertical Dry Cask Simulator (BVDCS) at Sandia National Laboratories (SNL) and validate the model with SNL’s experimental results. The BVDCS consists of a single boiling water reactor assembly fitted with electric heaters encompassed by a containment vessel and shell to represent a belowground spent nuclear fuel (SNF) dry storage system. Blowers are located near the inlet and outlet of the BVDCS to simulate crosswind conditions. In addition to the experimental results, the STAR-CCM+ model developed for this study is compared with a previous computational fluid dynamics (CFD) model in a different software program, which is used as a software-to-software benchmark. The experimental results provide a dataset to compare the STAR-CCM+ model results for a variety of different conditions. The main objective is to validate and improve STAR-CCM+ CFD models for spent nuclear fuel storage systems with explicitly modeled external environments and “wind driven” crossflows. These CFD models aide in the study of external particle deposition in spent nuclear fuel storage systems, which is important to predicting the significance of chloride induced stress corrosion cracking (CISCC). In addition to experimental comparison, a sensitivity analysis study is performed using the STAR-CCM+ model. The sensitivity analysis provides a quantitative assessment of the sensitivity of various parameters. This helps provide information on various parameters that are of particular importance to constructing a model representative of real life systems. The STAR-CCM+ model compared well to the experimental results showing similar responses to changes in cross wind flow, and a number of parameters are identified for model improvement.

Jensen, Ben J.↗

NOSS Altimeter Detailed Algorithm specifications

The details of the algorithms and data sets required for satellite radar altimeter data processing are documented in a form suitable for (1) development of the benchmark software and (2) coding the operational software. The algorithms reported in detail are those established for altimeter processing. The algorithms which required some additional development before documenting for production were only scoped. The algorithms are divided into two levels of processing. The first level converts the data to engineering units and applies corrections for instrument variations. The second level provides geophysical measurements derived from altimeter parameters for oceanographic users.

Hancock, D. W.↗

A small evaluation suite for Ada compilers

After completing a small Ada pilot project (OCC simulator) for the Multi Satellite Operations Control Center (MSOCC) at Goddard last year, the use of Ada to develop OCCs was recommended. To help MSOCC transition toward Ada, a suite of about 100 evaluation programs was developed which can be used to assess Ada compilers. These programs compare the overall quality of the compilation system, compare the relative efficiencies of the compilers and the environments in which they work, and compare the size and execution speed of generated machine code. Another goal of the benchmark software was to provide MSOCC system developers with rough timing estimates for the purpose of predicting performance of future systems written in Ada.

Wilke, Randy↗

Quantum Application Specifications and Benchmarks

This software describes computational tasks for quantum computers that are derived from LANL basic science research applications such as the modeling of materials, chemicals and compounds at atomic scales. The computations focus on quantum simulation tasks, such as quantum dynamics, thermal state preparation and ground state estimation. The primary function of the software is to develop estimates of the requirements for large-scale fault-tolerant quantum computers to solve these scientific computations. The secondary focus of the software are codes for assessing the limitations of conducting quantum computations on classical computers.

Coffrin, Carleton↗

Benchmarking hypercube hardware and software

It was long a truism in computer systems design that balanced systems achieve the best performance. Message passing parallel processors are no different. To quantify the balance of a hypercube design, an experimental methodology was developed and the associated suite of benchmarks was applied to several existing hypercubes. The benchmark suite includes tests of both processor speed in the absence of internode communication and message transmission speed as a function of communication patterns.

Grunwald, Dirk C.↗

DOE FAIR Surrogate Benchmarks Supporting AI and Simulation Research (SBI Surrogate Benchmark Initiative) (Final Report)

Computational Science is being revolutionized by integrating AI and simulation and, in particular, by deep learning surrogate models that can replace all or part of traditional large‐scale HPC computations. Such surrogates can achieve remarkable performance improvements, as much as several orders of magnitude, and save both compute time and energy. The Surrogate Benchmark Initiative (SBI) project creates a community repository and FAIR (Findable, Accessible, Interoperable, and Reusable) data ecosystem for HPC application surrogate benchmarks. The SBI team comes from Argonne National Laboratory (ANL), Indiana University (IU), Rutgers University, the University of Tennessee, Knoxville (UTK), and the University of Virginia(UVA). SBI repositories include data, code, and all relevant collateral artifacts, that the science and engineering community needs to use and reuse these data sets and surrogates. SBI repositories generate active research from both participants in SBI and the broader AI and domain science communities. This project develops surrogates that use several different neural nets to learn and quickly infer the results of simulations and data systems and capture them as surrogate benchmarks with a rich set of metadata, covering. Data; Model; Metrics specification; Machine specification; Science, Speed, Power Results, We research FAIR metadata for these benchmarks. We develop application surrogate examples as benchmarks across many fields (ANL, UTK, IU, UVA). We also study non Surrogate benchmarks that have many common features and similar issues regarding FAIRness. We work with MLCommons (UVA, UTK), which is a major machine learning benchmarking activity where we get metadata ontologies, software, and benchmarks, benchmarks have datasets, models, and metadata, and they need a technical framework developed by UTK and Rutgers and deployed by UVA. We study features of Surrogates, including performance, training set size, and uncertainty quantification (Rutgers, UVA and IU).

97 MATHEMATICS AND COMPUTING↗

FAIR Surrogate Benchmarks Supporting AI and Simulation Research (Final Report)

Computational Science is being revolutionized by integrating AI and simulation and, in particular, by deep learning surrogate models that can replace all or part of traditional large‐scale HPC computations. Such surrogates can achieve remarkable performance improvements, as much as several orders of magnitude, and save both compute time and energy. The Surrogate Benchmark Initiative (SBI) project creates a community repository and FAIR (Findable, Accessible, Interoperable, and Reusable) data ecosystem for HPC application surrogate benchmarks. The SBI team comes from Argonne National Laboratory (ANL), Indiana University (IU), Rutgers University, the University of Tennessee, Knoxville (UTK), and the University of Virginia (UVA). SBI repositories include data, code, and all relevant collateral artifacts that the science and engineering community need to use and reuse these data sets and surrogates. SBI repositories generate active research from both the participants in SBI and the broad community of AI and domain scientists. This project develops surrogates that use several different neural nets to learn and quickly infer the results of simulations and data systems and captures them as surrogate benchmarks with a rich set of metadata covering: Data; Model; Metrics specification; Machine specification; and Science, Speed, and Power Results. We research FAIR metadata for these benchmarks. We develop application surrogate examples as benchmarks across many fields (ANL, UTK, IU, UVA). We also study non-Surrogate benchmarks that have many common features and similar issues as regards FAIRness. We work with MLCommons (UVA, UTK), which is a major machine learning benchmarking activity where we get metadata ontologies, software, and benchmarks, Benchmarks have datasets, models, and metadata and they need a technical framework developed by UTK and Rutgers and deployed by UVA. We study features of Surrogates including performance, training set size, and uncertainty quantification (Rutgers, UVA and IU).

97 MATHEMATICS AND COMPUTING↗

Towards Digital and Performance-Based Supervisory HVAC Control Delivery

Upgrading supervisory HVAC control in commercial buildings is one of the most attractive decarbonization tools at our disposal. Modern controls are software programs and can in theory be deployed at scale and with a low up-front carbon "pulse". In practice, however, control delivery is a disjointed and inefficient process, dominated by manual handoffs of imprecise English language documents. A particularly high barrier exists between control implementation and building energy modeling (BEM) which results in control sequences typically not being tested for correctness or performance before implementation. Together with industry partners, DOE and the national labs are developing an ecosystem of tools and standards that can support fully digital performance-based control delivery workflows. This paper describes this ecosystem, which consists of three mutually supportive efforts. Semantic models of buildings and their systems enable automatic configuration and installation of control software. Platform-neutral control descriptions separate control algorithms from control platforms and enable the creation of libraries of reference control implementations. Dynamic whole-building energy-control simulation that can execute physically realistic control sequences makes it possible to test and evaluate the performance of control sequences and then directly compile them for installation and execution in control systems. In addition to digitizing and streamlining project-level control delivery, these standards and related software support benchmarking of control algorithms, both rule-based and optimization-based, and help to both advance the state of the art and to implement ratings and programs that encourage the adoption of high-performance control.

building controls↗

Update to the Microcontroller Benchmark for Radiation Testing

LANL developed a benchmark of software code for radiation testing of microprocessors several years ago, and it was published under an open-source license on GitHub. Publishing the software is necessary for other researchers to adopt and implement this benchmark for radiation testing of other microprocessors to standardize test practices so that test data can be compared across different microprocessors. The original codes have been used several times by other organizations to test a wide range of microcontrollers and microprocessors. After several years of research, LANL is ready to update the benchmark. Changes include: 1. Addition of new codes that allow common software codes to be tested, 2. Addition of new codes that instrument more microprocessor circuitry, 3. Addition of input patterns that allow for a more compressive understanding of how the memory layout affects the sensitivity to radiation-induced faults and better use of automated test pattern generation standards, and 4. Modification of current codes for faster and more resilient detection, reporting and correction of radiation-induced faults. These codes have been tested by LANL researchers over the last few years, which has been published in the open literature. As the code base for the new benchmarks are stable, it is time to release the update to the GitHub repository, where the original codes were released.

Quinn, Heather↗

Uniformly Ordered Binary Decision Algorithm for Benchmark Experiment Correlations in Whisper Validation

When performing a validation exercise for determining the upper subcritical limit of a nuclear criticality safety application, an analyst should select and perform a statistical analysis on a population of benchmark experiments that are neutronically similar to the application. The size of this population should be sufficiently large such that the statistical analysis has a high degree of confidence that the bias plus bias uncertainty (calculational margin) has been accurately quantified. A complication arises because many benchmark experiments share common components, leading to correlations in their measured effective multiplication factors. Correlations between benchmark experiments within the population reduces its predictive power. This motivates the need for methods that consider benchmark experiment correlations and ensure adequate statistical significance of results. The Whisper code is a statistical analysis pack- age that incorporates nuclear data sensitivity coefficients from MCNP to assess benchmark experiment similarity and then performs an extreme-value analysis to estimate the bias plus bias uncertainty. The original methodology in Whisper does not consider the effect of benchmark experiment correlations when making this estimation, and this summary proposes the uniformly ordered binary decision algorithm to address this shortcoming. The original methodology in Whisper computes similarity coefficients ck for an application compared to all benchmark experiments in its library and develops weighting factors for a selected population proportional to the ck values. The methodology can be interpreted as statistically emulating a validation exercise for a particular application where the weighting factors may be viewed as the likelihood that an analyst would include a particular benchmark experiment within the population. The effective sample size of the population is the expected or mean number of benchmark experiments in the population. The uniformly ordered binary decision algorithm identifies clusters of correlated benchmark experiments within the population and then computes adjusted weighting factors based on the magnitude of the correlation coefficients within the cluster to compute a reduced effective sample size accounting for the lower information content because of correlations. Benchmark experiments within the cluster are ordered randomly with equal probability and probabilistic decisions are made as to whether a benchmark. experiment within the cluster should treated as redundant with a previous one; if two redundant benchmark experiments are included, then the conservative worst case bias plus bias uncertainty is used and the pair is counted as a single benchmark experiment in the population. Results are provided for HEU solutions in a research version of the Whisper software using benchmark experiment correlations provided by DICE, the Database for the International Criticality Safety Benchmark Evaluation Project (ICSBEP). These show that there can be a significant increase in the bias plus bias uncertainty because the effective sample size is reduced, and therefore the algorithm, needing to meet sample size requirements, expands the benchmark experiment population by accepting less similar benchmark experiments that would have otherwise not been included.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

A NASA initiative: Software engineering for reliable complex systems

The objective is the development of methods, technology, and skills that will enable NASA to cost-effectively specify, build, and manage reliable software which can evolve and be maintained over an extended period. The need for such software is rooted in the increasing integration of software and computing components into NASA systems. Current NASA Software Engineering expertise was applied toward some of the largest reliable systems including: shuttle launch; ground support; shuttle simulation; minor control; satellite tracking; and scientific data systems. Unfortunately, no theory exists for reliable complex software systems. NASA is seeking to fill this theoretical gap through a number of approaches. One such approach is to conduct research on theoretical foundations for managing complex software systems. It includes: communication models, new and modified paradigms, and life-cycle models. Another approach is research in the theoretical foundations for reliable software development and validation. It focuses upon formal specifications, programming languages, software engineering systems, software reuse, formal verification, and software safety. Further approaches involve benchmarking a NASA software environment, experimentation within the NASA context, evolution of present NASA methodology, and transfer of technology to the space station software support environment.

Holcomb, Lee B.↗

Tools for unbinned unfolding

Machine learning has enabled differential cross section measurements that are not discretized. Going beyond the traditional histogram-based paradigm, these unbinned unfolding methods are rapidly being integrated into experimental workflows. Here, in order to enable widespread adaptation and standardization, we develop methods, benchmarks, and software for unbinned unfolding. For methodology, we demonstrate the utility of boosted decision trees for unfolding with a relatively small number of high-level features. This complements state-of-the-art deep learning models capable of unfolding the full phase space. To benchmark unbinned unfolding methods, we develop an extension of existing dataset to include acceptance effects, a necessary challenge for real measurements. Additionally, we directly compare binned and unbinned methods using discretized inputs for the latter in order to control for the binning itself. Lastly, we have assembled two software packages for the OmniFold unbinned unfolding method that should serve as the starting point for any future analyses using this technique. One package is based on the widely-used RooUnfold framework and the other is a standalone package available through the Python Package Index (PyPI).

47 OTHER INSTRUMENTATION↗