Ookami: Deployment and Initial Experiences
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Ongoing commercial design activities require a thorough verification of the Argonne Reactor Computation codes to be performed. DIF3D is central to this system and substantial work has been done to verify its accuracy on several identified commercial needs. This manuscript details the verification work done on PERSENT which relies upon the DIF3D for its forward and adjoint flux solution.
Tomographic imaging has benefited from advances in X-ray sources, detectors and optics to enable novel observations in science, engineering and medicine. These advances have come with a dramatic increase of input data in the form of faster frame rates, larger fields of view or higher resolution, so high performance solutions are currently widely used for analysis. Tomographic instruments can vary significantly from one to another, including the hardware employed for reconstruction: from single CPU workstations to large scale hybrid CPU/GPU supercomputers. Furthermore, flexibility on the software interfaces and reconstruction engines are also highly valued to allow for easy development and prototyping. This paper presents a novel software framework for tomographic analysis that tackles all aforementioned requirements. The proposed solution capitalizes on the increased performance of sparse matrix-vector multiplication and exploits multi-CPU and GPU reconstruction over MPI. Furthermore, the solution is implemented in Python and relies on CuPy for fast GPU operators and CUDA kernel integration, and on SciPy for CPU sparse matrix computation. As opposed to previous tomography solutions that are tailor-made for specific use cases or hardware, the proposed software is designed to provide flexible, portable and high-performance operators that can be used for continuous integration at different production environments, but also for prototyping new experimental settings or for algorithmic development. The experimental results demonstrate how our implementation can even outperform state-of-the-art software packages used at advanced X-ray sources worldwide.
Environmental screening of gamma radiation consists of detecting weak nuisance and anomaly signal in the presence of strong and highly varying background. In a typical scenario, a mobile detector-spectrometer continuously measures gamma radiation spectra in short, e.g., one-second, signal acquisition intervals. In principle, gamma radiation sources can be detected and identified by their unique spectral lines. However, detecting sources from data measured in a search scenario is difficult due to the highly varying background because of naturally occurring radioactive material (NORM), and low signal-to-noise ratio (S/N) of spectral signal measured during one-second acquisition intervals. In recent prior work, we have developed a Hopfield Neural Network (HNN) in conjunction with an image processing algorithm to detect a weak signal anomaly hidden among the highly fluctuating background spectra. The objective of this work is to explore quantum computing methods to increase the speed of HNN. The approach is based on the Grover’s search algorithm in conjunction with a 3-SAT problem formalism. The Grover’s algorithm is implemented on a quantum computing simulator using Qiskit software. Performance of HNN algorithm is benchmarked using search data from an environmental screening campaign, where the anomaly is a subset of measurements containing a 137 Cs source. Results indicate that using Grover’s algorithm on a quantum simulator reduces runtime of HNN by two orders of magnitude.
Extensive efforts to adaptively manage nutrient pollution rely on Chesapeake Bay Program’s (Phase 6) Watershed Model, called Chesapeake Assessment Scenario Tool (CAST), which helps decision-makers plan and track implementation of Best Management Practices (BMPs). We describe mathematical characteristics of CAST and develop a constrained nonlinear BMP-subset model, software, and visualization framework. This represents the first publicly available optimization framework for exploring least-cost strategies of pollutant load control for the United States’ largest estuary. The optimization identifies implementation options for a BMP subset modeled with load reduction effectiveness factors, and the web interface facilitates interactive exploration of >30,000 solutions organized by objective, nutrient control level, and for ~200 counties. We assess framework performance and demonstrate modeled cost improvements when comparing optimization-suggested proposals with proposals inspired by jurisdiction plans. Stakeholder feedback highlights the framework’s current utility for investigating cost-effective tradeoffs and its usefulness as a foundation for future analysis of restoration strategies.
When appropriately analyzed, thermoluminescent dosimeter glow curve analysis allows for improved quantification of thermoluminescent material behavior while flagging abnormalities. The mathematical separation of a glow curve into contributions from energetically unique trap states, or glow curve analysis, may be used to remove undesired effects of signal fading for complex materials. A generalized glow curve analysis software for the separation of glow curves is presented in this paper. Written in C ++ , the software uses the first-order kinetics model with automatic peak identification. The automatic identification of peaks is achieved through a unique peak-finding algorithm. Here, the program was performance tested using experimental glow curve data from LiF:Mg,Ti, and comparative results are presented.
This paper presents the latest improvements introduced in Version 4 of the UQpy, Uncertainty Quantification with Python, library. In the latest version, the code was restructured to conform with the latest Python coding conventions, refactored to simplify previous tightly coupled features, and improve its extensibility and modularity. To improve the robustness of UQpy, software engineering best practices were adopted. A new software development workflow significantly improved collaboration between team members, and continuous integration and automated testing ensured the robustness and reliability of software performance. Continuous deployment of UQpy allowed its automated packaging and distribution in system agnostic format via multiple channels, while a Docker image enables the use of the toolbox regardless of operating system limitations.
The Boston University component has focused on the algorithmic development of new Multigrid solver for the critical kernel for the Dirac propagators that dominated the both simulation require for lattice ensemble and the analysis of physical correlation functions. Progress on this has meet the above objects, even exceeding them a bit. The result is the beginning if multiscale lattice QCD applicable to future Exascale hardware and the development of the QUDA software for NVIDIA GPUs to give near optimal performance. As we approach exascale hardware and computation at that scale this provides the infrastructure for further advances.
Software containers are a key channel for delivering portable and reproducible scientific software in high performance computing (HPC) environments. HPC environments are different from other types of computing environments primarily due to usage of the message passing interface (MPI) and drivers for specialized hard- ware to enable distributed computing capabilities. This distinction directly impacts how software containers are built for HPC applications and can complicate software quality assurance efforts including portability and performance. This work introduces a strategy for building containers for HPC applications that adopts layering as a mechanism for software quality assurance. The strategy is demonstrated across three different HPC systems, two of them petaflops scale with entirely different interconnect technologies and/or processor chipsets but running the same container. Performance consequences of the containerization strategy are found to be less than 5-14% while still achieving portable and reproducible containers for HPC systems.
Digital real time simulators have the capability to run electromagnetic transient simulations in real time. This capability allows users to leverage the hardware-software combination to evaluate controller performance, protection device performance, and power device performance. This has helped many field deployment projects to be successful and be cost-effective. In this talk, we will present current state-of-art, and future of real time electromagnetic transient simulation and its impacts on field deployment.
Computational Science is being revolutionized by integrating AI and simulation and, in particular, by deep learning surrogate models that can replace all or part of traditional large‐scale HPC computations. Such surrogates can achieve remarkable performance improvements, as much as several orders of magnitude, and save both compute time and energy. The Surrogate Benchmark Initiative (SBI) project creates a community repository and FAIR (Findable, Accessible, Interoperable, and Reusable) data ecosystem for HPC application surrogate benchmarks. The SBI team comes from Argonne National Laboratory (ANL), Indiana University (IU), Rutgers University, the University of Tennessee, Knoxville (UTK), and the University of Virginia(UVA). SBI repositories include data, code, and all relevant collateral artifacts, that the science and engineering community needs to use and reuse these data sets and surrogates. SBI repositories generate active research from both participants in SBI and the broader AI and domain science communities. This project develops surrogates that use several different neural nets to learn and quickly infer the results of simulations and data systems and capture them as surrogate benchmarks with a rich set of metadata, covering. Data; Model; Metrics specification; Machine specification; Science, Speed, Power Results, We research FAIR metadata for these benchmarks. We develop application surrogate examples as benchmarks across many fields (ANL, UTK, IU, UVA). We also study non Surrogate benchmarks that have many common features and similar issues regarding FAIRness. We work with MLCommons (UVA, UTK), which is a major machine learning benchmarking activity where we get metadata ontologies, software, and benchmarks, benchmarks have datasets, models, and metadata, and they need a technical framework developed by UTK and Rutgers and deployed by UVA. We study features of Surrogates, including performance, training set size, and uncertainty quantification (Rutgers, UVA and IU).
The Belle II experiment, which started taking physics data in April 2019, will multiply the volume of data currently stored on its nearly 30 storage elements worldwide by one order of magnitude to reach about 340 PB of data (raw and Monte Carlo simulation data) by the end of operations. To tackle this massive increase and to manage the data even after the end of the data taking, it was decided to move the Distributed Data Management software from a homegrown piece of software to a widely used Data Management solution in HEP and beyond : Rucio. This contribution describes the work done to integrate Rucio with Belle II distributed computing infrastructure as well as the migration strategy that was successfully performed to ensure a smooth transition.
Many electromagnetic measurements detect signals that are the time derivative of the actual quantity of interest. Examples include so-called “B-dot” and “D-dot” detectors that are used for measuring transient magnetic and electric fields. Integration of these signals to recover the quantity of interest is performed either by hardware integrators in the signal line, or by software coding in the analysis programming. Hardware integration is usually done using a classic resistive-capacitive (RC) circuit, or internally in the detector by induction (L/R) or stray capacitance (RC). These hardware integration methods are only approximate, and must be corrected for distortion. This is usually called “droop correction.” In this note, we will examine this correction in some detail. Transient analysis via Laplace transforms will be used to simplify the mathematics. A rudimentary understanding of this method is assumed.
The demand for high performance computing (HPC) resources continues to grow, driven by the increasing complexity of modeling and simulation, artificial intelligence (AI), and machine learning (ML) workloads [Porter]. The growing energy consumption demand of these HPC systems is a significant concern, both in terms of operational costs and environmental impact. AI hardware accelerators are expected to reach 1.5% of the world’s power consumption by 2029 [Shah].
The U.S. Department of Energy (DOE) Exascale Computing Project (ECP) funded the development of new (and the transformation of important existing) applications, libraries, and tools that realized improvement in performance and capabilities of often 100 times or more on emerging exascale computers. This exceptional gain inspired the title of this special issue: Transforming Science through Software: Improving while delivering 100X. The term 100X refers to advancing capabilities in modeling, simulation, and analysis by a factor of 100 or more using some combination of new algorithms, optimization techniques, software libraries, and programming models, coupled with the next generation of hardware for high-performance computing (HPC). The papers in this issue share experiences with the practice and science of scientific software development, with an emphasis on developing a coherent, portable, and sustainable HPC software ecosystem for next-generation computational science. Finally, we hope to foster expanded community efforts related to the fundamental role of sustainable scientific software ecosystems in advancing the computing sciences.
Advanced, high-performance computing at the National Renewable Energy Laboratory (NREL) has enabled access to vast data resources with cutting-edge software techniques to understand, design, plan for, and maintain energy-efficient and resilient infrastructure. We have focused on cities and airports, but the technology we have developed will easily translate to seaports, inland ports, military installations, or other complex and large-scale energy-intensive systems. We can digitally simulate and explore current and future scenarios to make datadriven decisions for optimizing advanced energy systems, transportation and building operations, infrastructure planning and expansion, and battery storage to guide short- and long-term investments, electrification strategies, and integration of new technologies.
As core counts increase in new HPC systems with comparatively little increase in memory bandwidth, the trend is an effective decrease in memory bandwidth per core. Other bandwidth limitations in HPC systems exist between CPU and GPU memory, between system nodes, and between node memory and storage. Compression of floating-point data has the potential to reduce data movement and pressure across these communication channels. Furthermore, it has the potential to reduce the footprint of floating-point arrays stored in memory. ZFP, implemented in software, is gaining traction as an effective method in floating-point compression; however, performance gains are limited to the spare compute cycles available before reaching the bandwidth limitations of the communication channel. A hardware implementation of ZFP has the potential to raise the bar on performance. From the inception of ZFP, it was designed to accommodate a hardware implementation.
This report documents the software testing which has been performed for the PARET/ANL version 7.6 software. The capabilities addressed in this testing were selected based on the use of the software for analysis work under the Research and Test Reactors Department quality assurance program. The verification and validation have been performed and documented to address the steady-state capabilities of the software, as described in Chapter 2, and the transient capabilities, as described in Chapter 3. Testing based on the comparison between PARET/ANL calculations and analytical solutions, hand calculations, or another code calculations of the test cases confirms that all the identified capabilities of the software were implemented correctly.