Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “heterogeneous hardware”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Errant Beam Detection Using the AMD Versal ACAP and Vitis AI

The prevalence of ML and AI-powered solutions along with the slowing of Moore's Law has given rise to novel hardware platforms aimed at accelerating ML and AI. While programming these hardware platforms can be difficult, particularly for non-hardware experts, hardware vendors provide high-level tooling in an effort to address this difficulty. The Versal ACAP is an SoC designed by AMD that combines CPU cores, FPGA fabric, and a tiled, vector architecture called an AI engine all on the same socket. In an effort to more easily program this heterogeneous system, AMD has provided the Vitis AI development stack. In this work, we leverage Vitis AI to program a Versal ACAP to perform errant beam detection in the Spallation Neutron Source at Oak Ridge National Laboratory. Our initial work shows that after quantization and compilation of the model for the Versal ACAP, the classification accuracy, as measured by the AUC metric, is over 95% accurate while achieving this accuracy in 46 microseconds on average.

Cabrera, Anthony↗

Using MLIR Framework for Codesign of ML Architectures Algorithms and Simulation Tools

MLIR (Multi-Level Intermediate Representation), is an extensible compiler framework that supports high-level data structures and operation constructs. These higher-level code representations are particularly applicable to the artificial intelligence and machine learning (AI/ML) domain, allowing developers to more easily support upcoming heterogeneous AI/ML accelerators and develop flexible domain specific compilers/frameworks with higher-level intermediate representations (IRs) and advanced compiler optimizations. The result of using MLIR within the LLVM compiler framework is expected to yield significant improvement in the quality of generated machine code, which in turn will result in improved performance and hardware efficiency

97 MATHEMATICS AND COMPUTING↗

A High-Performance Design for Hierarchical Parallelism in the QMCPACK Monte Carlo code

We introduce a new high-performance design for parallelism within the Quantum Monte Carlo code QMCPACK. We demonstrate that the new design is better able to exploit the hierarchical parallelism of heterogeneous architectures compared to the previous GPU implementation. The new version is able to achieve higher GPU occupancy via the new concept of crowds of Monte Carlo walkers, and by enabling more host CPU threads to effectively offload to the GPU. The higher performance is expected to be achieved independent of the underlying hardware, significantly improving developer productivity and reducing code maintenance costs. Scientific productivity is also improved with full support for fallback to CPU execution when GPU implementations are not available or CPU execution is more optimal.

Luo, Ye↗

COHORT: Coordination of Heterogeneous Thermostatically Controlled Loads for Demand Flexibility

Demand flexibility is increasingly important for power grids. Careful coordination of thermostatically controlled loads (TCLs) can modulate energy demand, decrease operating costs, and increase grid resiliency. We propose a novel distributed control framework for the Coordination Of HeterOgeneous Residential Thermostatically controlled loads (COHORT). COHORT is a practical, scalable, and versatile solution that coordinates a population of TCLs to jointly optimize a grid-level objective, while satisfying each TCL’s end-use requirements and operational constraints. To achieve that, we decompose the grid-scale problem into subproblems and coordi- nate their solutions to find the global optimum using the alternating direction method of multipliers (ADMM). The TCLs’ local problems are distributed to and computed in parallel at each TCL, making COHORT highly scalable and privacy-preserving. While each TCL poses combinatorial and non-convex constraints, we characterize these constraints as a convex set through relaxation, thereby making COHORT computationally viable over long planning horizons. After coordination, each TCL is responsible for its own control and tracks the agreed-upon power trajectory with its preferred strategy. In this work, we translate continuous power back to discrete on/off actuation, using pulse width modulation. COHORT is generalizable to a wide range of grid objectives, which we demonstrate through three distinct use cases: generation following, minimizing ramping, and peak load curtailment. In a notable experiment, we validated our approach through a hardware-in-the-loop simulation, including a real-world air conditioner (AC) controlled via a smart thermostat, and simulated instances of ACs modeled after real-world data traces. During the 15-day experimental period, COHORT reduced daily peak loads by an average of 12.5% and maintained comfortable temperatures.

demand response↗

Computer Science Research Needs for Parallel Discrete Event Simulation (PDES)

Historically, scientific computing efforts have demonstrated the clear need for, and effective use of, supercomputing with traditional time-stepped simulations. Nevertheless, there are several areas in the mission spaces of the U.S. Department of Energy and other agencies waiting to tap advanced computing research using a different, discrete event style of modeling, simulation, and analysis. These span a wide spectrum of applications including energy grid resilience, urban planning and policy, transportation science, building technologies, emergency response and planning, environmental impact analysis, computational epidemiology, Internet communications, cyber security, and cyber-physical systems, to name only a few. Even within traditional scientific applications, the role of discrete event modes of execution is increasing in the form of new event-based mathematical solvers such as quantized state integration methods and discrete-continuous hybrid system solvers. Co-design of advanced supercomputing hardware systems is another area that exploits discrete event simulation at its core for effective analyses. Complex systems, entity behaviors and interconnections play a significant role in all these applications, which are mapped to large-scale models with discrete event formulations. To make advancements in all the aforementioned scientific areas, many technical aspects need to be more thoroughly studied and deeply understood in parallel discrete event simulation (PDES). The unique dynamics inherent in a discrete event modeling approach, by their very nature, intersect and influence the entire stack of the computing system, including (a) the unique nature of the instruction sets exercised in PDES workloads without a predominance of high-precision floating point operations, (b) virtual time-constrained multi-threaded execution of many logical processes per processor, (c) extremely variable and difficult to predict network traffic characteristics, (d) interfaces and inter-dependencies with machine learning and artificial intelligence codes at higher software layers, and (e) highly challenging load balancing needs, especially in effectively accounting for accelerated/extremely heterogeneous computing in current and future high-performance computing systems. Efficient and accurate parallel execution of PDES workloads is also dominated by challenges in dealing with their asynchronous concurrency fundamentally present at the model level. Conservative synchronization, optimistic/speculative synchronization, and their hybrid schemes open new questions in fundamental computer science with respect to reversibility of computation and prediction (lookahead) of behaviors inherent within model codes. On the implementation front, there are relatively few scalable, general-purpose parallel discrete event simulators in the world, and even fewer have been studied on emerging hardware platforms. To enable scientific advances using PDES, the research needs in computer science must also be pursued and met in the intersection of the algorithmic and hardware-aware aspects of scalable PDES engines. This report is aimed at capturing a computer science-oriented view of this important area of research in PDES, presenting a sample of important applications with their inherent discrete event technology elements. Needs are outlined in core areas of parallel discrete event research as well as cross-cutting directions in computer science research that positively impact scientific advancements across several important application areas. A selection of priority research opportunities in advanced computing for PDES is identified to serve as reference for key research topics and their order of importance for scientific advancements.

97 MATHEMATICS AND COMPUTING↗

Exact and Fixed-Point Grover Search with Qudits

Grover's algorithm provides a quadratic speedup for searching unstructured databases and is traditionally implemented with qubits in Hilbert spaces whose dimensions are powers of two. With the advent of quantum platforms utilizing qudits---quantum systems with more than two levels---there is a need to generalize Grover search to these architectures, including heterogeneous systems with qudits of varying dimensions. Here, we present a unified framework for qudit-based Grover search, detailing the construction of oracles and diffusion operators with and without ancilla qubits and generalizing deterministic and fixed-point search variants that ensure exact or bounded success probabilities. We analyze phase-matching techniques and provide explicit circuit decompositions suitable for diverse hardware platforms. We also compare the corresponding trajectories on the Bloch sphere to provide an intuitive visualization of how the different phase choices amplify the target state. These results facilitate flexible, hardware-oriented protocols for implementing Grover search on qudit processors, potentially reducing circuit depth and enhancing success probabilities, thereby offering a practical toolkit for quantum computation and sensing applications leveraging multilevel quantum systems.

Roy, Tanay [Fermilab] (ORCID:000000019442862X)↗

Profiling the BLAST bioinformatics application for load balancing on high-performance computing clusters

Abstract Background The Basic Local Alignment Search Tool (BLAST) is a suite of commonly used algorithms for identifying matches between biological sequences. The user supplies a database file and query file of sequences for BLAST to find identical sequences between the two. The typical millions of database and query sequences make BLAST computationally challenging but also well suited for parallelization on high-performance computing clusters. The efficacy of parallelization depends on the data partitioning, where the optimal data partitioning relies on an accurate performance model. In previous studies, a BLAST job was sped up by 27 times by partitioning the database and query among thousands of processor nodes. However, the optimality of the partitioning method was not studied. Unlike BLAST performance models proposed in the literature that usually have problem size and hardware configuration as the only variables, the execution time of a BLAST job is a function of database size, query size, and hardware capability. In this work, the nucleotide BLAST application BLASTN was profiled using three methods: shell-level profiling with the Unix “time” command, code-level profiling with the built-in “profiler” module, and system-level profiling with the Unix “gprof” program. The runtimes were measured for six node types, using six different database files and 15 query files, on a heterogeneous HPC cluster with 500+ nodes. The empirical measurement data were fitted with quadratic functions to develop performance models that were used to guide the data parallelization for BLASTN jobs. Results Profiling results showed that BLASTN contains more than 34,500 different functions, but a single function, RunMTBySplitDB, takes 99.12% of the total runtime. Among its 53 child functions, five core functions were identified to make up 92.12% of the overall BLASTN runtime. Based on the performance models, static load balancing algorithms can be applied to the BLASTN input data to minimize the runtime of the longest job on an HPC cluster. Four test cases being run on homogeneous and heterogeneous clusters were tested. Experiment results showed that the runtime can be reduced by 81% on a homogeneous cluster and by 20% on a heterogeneous cluster by re-distributing the workload. Discussion Optimal data partitioning can improve BLASTN’s overall runtime 5.4-fold in comparison with dividing the database and query into the same number of fragments. The proposed methodology can be used in the other applications in the BLAST+ suite or any other application as long as source code is available.

59 BASIC BIOLOGICAL SCIENCES↗

Development of an advanced ultrasonic phased array for the characterization of thick, reinforced concrete components (Final Scientific/Technical Report)

There are no nondestructive evaluation (NDE) tools capable of characterizing microscale damage throughout the thickness of concrete components, due to the multiphase, heterogeneous and multiscale nature of concrete. Ultrasound is only scattered by features at the same length scale, or smaller, than the wavelength of a wave’s dominant frequency. Successful imaging of microscale damage using ultrasound requires that the ultrasonic wavelength be on the order of a few millimeters (or smaller), yet the heterogeneous microstructure of concrete, with its fine and coarse aggregates, is on this same (and higher) micrometer/millimeter length scale, causing excessive ultrasonic wave scattering even in “good” concrete. The proposed solution applies non-collinear wave mixing to spatially image microscale damage, while still maintaining penetration through thick concrete components. This microscale imaging is possible by combining nonlinear wave mixing, with advanced phased array hardware and software to develop a breakthrough tool that will bring revolutionary changes in NDE of concrete infrastructure in terms of image resolution and depth of penetration. This work uses non-collinear wave mixing which exploits the physics that material nonlinearities such as microscale damage, cause interactions between two intersecting ultrasonic waves due to cross-mixing, which can lead to the generation of a third wave with a frequency and wave number of the sum or difference of the incident waves. The concrete material volume at this mixing point is characterized/imaged. This project delivered a single-sided nonlinear ultrasonic phased array imaging device, that can image microscale damage (microcracks of 100 micrometers) through a 0.5 m thick concrete component and assessed the commercial feasibility of such arrays for improved crack detection.

42 ENGINEERING↗

Development and performance of a HemeLB GPU code for human-scale blood flow simulation

In recent years, it has become increasingly common for high performance computers (HPC) to possess some level of heterogeneous architecture - typically in the form of GPU accelerators. In some machines these are isolated within a dedicated partition, whilst in others they are integral to all compute nodes - often with multiple GPUs per node - and provide the majority of a machine's compute performance. In light of this trend, it is becoming essential that codes deployed on HPC are updated to execute on accelerator hardware. Here, in this paper, we introduce a GPU implementation of the 3D blood flow simulation code HemeLB that has been developed using CUDA C++. We demonstrate how taking advantage of NVIDIA GPU hardware can achieve significant performance improvements compared to the equivalent CPU only code on which it has been built whilst retaining the excellent strong scaling characteristics that have been repeatedly demonstrated by the CPU version. With HPC positioned on the brink of the exascale era, we use HemeLB as a motivation to provide a discussion on some of the challenges that many users will face when deploying their own applications on upcoming exascale machines.

59 BASIC BIOLOGICAL SCIENCES↗

Low-Cost, Easy-To-Integrate and Reliable Grid Energy Storage System with 2 nd Life Lithium Batteries

Batteries retired from electric vehicles have the potential to extend their service as low-cost stationary energy storage systems. However, disperse battery state of health (SOH) and nonuniform battery characters often lead to compromised battery performance and reliability, which greatly hinder their adoption. A Heterogenous Unifying Battery (HUB) system is proposed to stage 2 nd life battery bricks for a period, and enable them to attain improved SOH uniformity, performance, and reliability before being sold for 2 nd life applications, while simultaneously providing grid services. It may offer a technically and economically advantageous solution for the broad utilization of 2 nd use batteries. The goal of this project was to develop the hardware and software that enables the key functions of the HUB system. The first achievement of the project was the development of a 1kW scale proof-of-concept (POC) system, which comprises (i) a modular plug-n-play DC-DC power converter matrix with isolated series output connections to achieve fully independent control of energy flow to each of the connected battery units at low voltage; (ii) enhanced model based control that drives each batteries’ SOH towards uniformity while collectively providing grid energy storage services; and (iii) comprehensive procedures to perform battery diagnostics and prognostics. The second achievement was the development of a 100kW scale HUB system and demonstrated its performance of re-establishing battery SOH uniformity through a period of battery cycling operation. The final HUB system incorporates six DC-DC power converter matrices paired with six battery bricks. Hot swapping of a single battery brick while maintaining consistent system power was demonstrated and system operation was validated to be capable of implementing the approved grid duty cycle and of balancing and conditioning the battery bricks. Through the course of the project, the team optimized the building-block design, form-factors, and adjusted life balancing control. An up-sized 250kW Scale was developed and deployed in October 2022 with pack-level battery form factors, see photo in Figure 1 The third achievement of the project was to perform a techno-economic analysis in order to better understand the cost and revenue potentials in this new “recondition-then-resell" value proposition. The final TEA quantified the economics of new Li-ion batteries as well as second-life batteries processed via reconditioning and traditional binning. Results showed the reconditioned second-life batteries in this project to be economically favorable and viable in grid energy storage markets. The TEA results were published in the Applied Energy journal. The fourth achievement of the project was to deliver a tech-to-market plan for the HUB system that includes funding, IP, and manufacturing strategies. The final T2M plan outlines a business strategy in which the HUB provides a B2B service to EV companies as an alternative to battery recycling that can prepare batteries for 2nd life applications. A company named Smartville Inc. was founded to carry on the commercialization, funding, and technical IP licensing activities of the OPEN project.

25 ENERGY STORAGE↗

Verification and Validation of Elastodynamic Simulation Software for Aerospace Research

Physics-based simulation of nondestructive evaluation (NDE) inspection can help to advance the inspectability and reliability of mechanical systems. However, NDE simulations applicable to non-idealized mechanical components often require large compute domains and long run times. This has prompted development of custom NDE simulation software tailored to high performance computing (HPC) hardware. Verification and validation (V&V) is an integral part of developing this software to ensure implementations are robust and applicable to inspection problems, producing tools and simulations suitable for computational NDE research. This presentation addresses factors common to V&V of several elastodynamic simulation codes applicable to ultrasonic NDE. Examples are drawn from in-house simulation software at NASA Langley Research Center, ranging from ensuring reliability in a 1D heterogeneous media wave equation solver to the V&V needs of 3D cluster-parallel elastodynamic software. Factors specific to a research environment are addressed, where individual simulation results can be as relevant as the software product itself. Distinct facets of V&V are discussed including testing to establish software reliability, employing systematic approaches for consistency with fundamental conservation laws, establishing the numerical stability of algorithms, and demonstrating concurrence with empirical data. This talk also addresses V&V practices for small groups of researchers. This includes establishing resources (e.g. time and personnel) for V&V during project planning to mitigate and control the risk of setbacks. Similarly, we identify ways for individual researchers to use V&V during simulation software development itself to both speed up the development process and reduce incurred technical debt.

NDE↗

Leveraging Compiler-Based Translation to Evaluate a Diversity of Exascale Platforms

Accelerator-based heterogeneous computing is the de facto standard in current and upcoming exascale machines. These heterogeneous resources empower computational scientists to select a machine or platform well-suited to their domain or applications. However, this diversity of machines also poses challenges related to programming model selection: inconsistent availability of programming models across different exascale systems, lack of performance portability for those programming models that do span several systems, and inconsistent performance between different models on a single platform. We explore these challenges on exascale-similar hardware, including AMD MI100 and NVIDIA A100 GPUs. By extending the sourceto-source compiler OpenARC, we demonstrate the power of automated translation of applications written in a single frontend programming model (OpenACC) into a variety of backend models (OpenMP, OpenCL, CUDA, HIP) that span the upcoming exascale environments. This translation enables us to compare performance within and across devices and to analyze programming model behavior with profiling tools.

Lambert, Jacob↗

ATHENA: Analytical Tool for Heterogeneous Neuromorphic Architectures

The ASC program seeks to use machine learning to improve efficiencies in its stockpile stewardship mission. Moreover, there is a growing market for technologies dedicated to accelerating AI workloads. Many of these emerging architectures promise to provide savings in energy efficiency, area, and latency when compared to traditional CPUs for these types of applications — neuromorphic analog and digital technologies provide both low-power and configurable acceleration of challenging artificial intelligence (AI) algorithms. If designed into a heterogeneous system with other accelerators and conventional compute nodes, these technologies have the potential to augment the capabilities of traditional High Performance Computing (HPC) platforms [5]. This expanded computation space requires not only a new approach to physics simulation, but the ability to evaluate and analyze next-generation architectures specialized for AI/ML workloads in both traditional HPC and embedded ND applications. Developing this capability will enable ASC to understand how this hardware performs in both HPC and ND environments, improve our ability to port our applications, guide the development of computing hardware, and inform vendor interactions, leading them toward solutions that address ASC’s unique requirements.

97 MATHEMATICS AND COMPUTING↗

HIPLZ: Enabling performance portability for exascale systems

While heterogeneous computing has emerged as a dominant trend in current and future High-Performance Computing (HPC) systems, it is also widely recognized that this shift has led to increased software complexity due to a proliferation of programming systems for different heterogeneous processors. One such example is the Heterogeneous-Compute Interface for Portability from AMD (HIP ), which is composed of a C Runtime API and C++ Kernel Language. Many HPC applications will likely use HIP on future exascale systems (e.g., Frontier and El Capitan), but HIP currently only targets AMD and NVIDIA processors. This limitation creates challenges for users who would also like to run their applications on exascale systems based on other architectures (e.g., Aurora, which is based on Intel hardware) that are currently not targeted by HIP . In this paper, we introduce the design and implementation of HIPLZ , a compiler and runtime system that uses the Intel Level Zero API to support HIP on Intel GPU architectures. We discuss the design of HIPLZ , derived from HIPCL (an implementation of HIP on top of OpenCL ), and portability issues that occur from using the Level Zero runtime as a backend. We evaluate our implementation by running several performance benchmarks and mini-apps written in HIP on Intel architectures using HIPLZ . Our results show that this approach provides competitive performance relative to Intel's OpenCL implementations on Intel Gen9 and UHD Graphics 770 GPUs, while providing good coverage of features needed by HPC applications. Overall, this approach is a promising demonstration of enabling performance portability for exascale systems.

97 MATHEMATICS AND COMPUTING↗

XACC: a system-level software infrastructure for heterogeneous quantum–classical computing

Quantum programming techniques and software have advanced significantly over the past five years, with a majority focusing on high-level language frameworks targeting remote REST library APIs. As quantum computing architectures advance and become more widely available, lower-level, system software infrastructures will be needed to enable tighter, co-processor programming and access models. In this work, we present XACC, a system-level software infrastructure for quantum–classical computing that promotes a service-oriented architecture to expose interfaces for core quantum programming, compilation, and execution tasks. Additionally, we detail XACC's interfaces, their interactions, and its implementation as a hardware-agnostic framework for both near-term and future quantum–classical architectures. We provide concrete examples demonstrating the utility of this framework with paradigmatic tasks. Our approach lays the foundation for the development of compilers, associated runtimes, and low-level system tools tightly integrating quantum and classical workflows.

97 MATHEMATICS AND COMPUTING↗

The FLEX/32 multicomputing environment

The FLEX/32 Multicomputer is a generic environment for cooperating multiple processors. The FLEX/32 supports a number of different processors, making it heterogeneous in terms of the instruction sets it supports, and homogeneous in its ability to provide consistent storage and input/output facilities to its differing processors. These facilities are accessed through standard 32-bit VMEbus connections. The FLEX/32 supports the full UNIX System V Operating System and languages associated with it, plus the extended ConCurrent C and Concurrent FORTRAN 77 languages that allow programming of concurrent software at a high level. Direct programming support at all levels is provided by the environment hardware for concurrent software execution and optimization, including hardware support for shared resource access arbitration, conditional critical region arbitration, and interprocessor messages.

Matelan, N.↗

Flexible and Effective Object Tiering for Heterogeneous Memory Systems

Computing platforms that package multiple types of memory, each with their own performance characteristics, are quickly becoming mainstream. To operate efficiently, heterogeneous memory architectures require new data management solutions that are able to match the needs of each application with an appropriate type of memory. As the primary generators of memory usage, applications create a great deal of information that can be useful for guiding memory management, but the community still lacks tools to collect, organize, and leverage this information effectively. To address this gap, this work introduces a novel software framework that collects and analyzes object-level information to guide memory tiering. The framework includes tools to monitor the capacity and usage of individual data objects, routines that aggregate and convert this information into tier recommendations for the host platform, and mechanisms to enforce these recommendations according to user-selected policies. Moreover, the developed tools and techniques are fully automatic, work on standard Linux systems, and do not require modification or recompilation of existing software. Using this framework, this study evaluates and compares the impact of a variety of design choices for memory tiering, including different policies for prioritizing objects for the fast memory tier as well as the frequency and timing of migration events. In conclusion, the results, collected on a modern Intel platform with conventional DDR4 SDRAM as well as Intel Optane NVRAM, show that guiding data tiering with object-level information can enable significant performance and efficiency benefits compared with standard hardware- and software-directed data-tiering strategies for a diverse set of memory-intensive workloads.

97 MATHEMATICS AND COMPUTING↗

Demand Response Optimization and Management System for Real-TIme (DROMS-RT)

To design and demonstrate DROMS-RT, a highly distributed Demand Response Optimization and Management System for Real-Time (DROMS-RT) power flow control to support large scale integration of distributed renewable generation into the grid. AutoGrid developed a novel control and communications platform to allow highly dispatchable demand response (DR) services in time frames suitable for providing ancillary services to the transmission grid. These services will be substantially less expensive and more efficient than other forms of ancillary services options currently available to manage the intermittency associated with large-scale renewable integration. DROMSRT successfully leveraged Automated Demand Response (ADR) by fundamentally re-thinking the architecture of the DR platform from the ground up and by developing innovative new technologies in a number of areas related to DR. DROMS-RT leveraged the low-cost, open, interoperable DR signaling technology, OpenADR, and low-cost, internet-protocol based telemetry solutions to reduce the cost of hardware. This allowed DROMS-RT to provide dynamic price signals to millions of OpenADR clients. Statistically rigorous signal processing techniques were developed to reliably detect even small load reductions in the presence of noisy baseline profiles. Novel forecasting engines based on modern online machine learning algorithms enabled accurate individualized forecasts for customer loads in the presence of dynamic pricing signals, and a real-time decision engines enabled continuous optimization and optimal dispatch of DR resources across a large portfolio of heterogeneous loads that respond at varying time-scales. Moreover, the real-time optimization conducted by the decision engine can utilize grid physics to maximize load reduction at the transmission system in addition to the distribution sites, for more efficient grid operation. Finally, the Software-as-a-Service (SaaS) availability of the DROMS-RT platform has reduced the cost of deployment and enable participation of small commercial and residential customers in DR who otherwise would not be able to do so.

24 POWER TRANSMISSION AND DISTRIBUTION↗