Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “heterogeneous hardware”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Highly-scalable GPU-accelerated compressible reacting flow solver for modeling high-speed flows

Emerging supercomputing systems utilize a combination of central processing units (CPUs) and graphics processing units (GPUs) in an effort to reach exascale capabilities while minimizing the energy footprint of operating such systems. Such heterogeneous machines introduce new challenges for fluids solvers because the hardware architecture and operation of a GPU are fundamentally different from conventional CPUs. In this work, a general approach for efficient implementation of finite-volume based reacting flow solvers on such heterogeneous systems is presented. Three main challenges, namely, data access pattern, thread divergence, and thread safety, are addressed. Since compressible reacting flows require special methods to deal with chemical reactions, hyperbolic and nonlinear convection terms, and the presence of turbulence, specific algorithms that ensure GPU-based efficiency are developed. The approach is demonstrated on the widely available OpenFOAM open source software by modifying core algorithms for GPU accessibility. The scalability of the resulting solver, is demonstrated using practical test cases, including flow through a scramjet engine and the dynamics of a rotating detonation engine. Here, the solver provides near-ideal scaleup on a large number of GPUs (>3000), and extremely efficient use of the GPUs, with throughput nearly a constant even when processing a large number of control volumes.

42 ENGINEERING↗

Evaluating Nonuniform Reduction in HIP and SYCL on GPUs

Motivated by maturing programming models and portability for heterogeneous computing, we describe the challenges posed by hardware architectures and programming models when migrating an optimized implementation of nonuniform reduction from CUDA to HIP and SYCL. We explain the migration experience, evaluate the performance of the reduction on GPU -based computing platforms, and provide feedback on improving portability for the development of the SYCL programming model.

Jin, Zheming↗

IRIS-BLAS: Towards a Performance Portable and Heterogeneous BLAS Library

This paper presents IRIS-BLAS, a novel heterogeneous and performance portable BLAS library. IRIS-BLAS is built on top of the IRIS runtime and multiple vendor and open-source BLAS libraries. It can transparently use all the architectures/devices available in a heterogeneous system, using the appropriate BLAS library based on the task mapping at run time. Thus, IRIS-BLAS is portable across a broad spectrum of architectures and BLAS libraries, alleviating the worry of application developers about modifying the application source code. Even though the emphasis is on portability, IRIS-BLAS provides competitive or even better performance than other state-of-the-art references. Moreover, IRIS-BLAS offers new features such as efficiently using extremely heterogeneous systems composed of multiple GPUs from different hardware vendors.

Miniskar, Narasinga Rao↗

Coupling Noah-Multiparameterization land-surface Model with Energy Research and Forecasting Model

The Energy Research and Forecasting (ERF) model is a high-performance atmospheric model built on the AMReX adaptive mesh refinement (AMR) framework, enabling efficient simulations on heterogeneous computing platforms that combine multicore processors with hardware accelerators. To support land–atmosphere interactions within ERF’s AMR-based environment, a land-surface model must be capable of operating directly on hierarchically refined meshes. In this work, we present a methodology for coupling the Fortran-based Noah-Multiparameterization (Noah-MP) land-surface model with ERF’s C++ codebase. Rather than rewriting Noah-MP, we construct a Fortran–C interoperability layer using CodeScribe, a tool that leverages large language models (LLMs) to automate the generation of interface code. CodeScribe applies structured prompting techniques to generate bindings that support efficient data exchange and function calls between ERF and Noah-MP. The coupling framework also incorporates AMR-aware data handling strategies, allowing NoahMP to operate seamlessly within ERF’s hierarchical mesh structure. This work provides a structured approach for integrating legacy Fortran models into modern C++-based modeling systems using LLM-assisted code generation.

54 ENVIRONMENTAL SCIENCES↗

Design and analysis of CXL performance models for tightly-coupled heterogeneous computing

Truly heterogeneous systems enable partitioned workloads to be mapped to the hardware that nets the best performance. However, current practice requires that inter-device communication between different vendors' hardware use host memory as an intermediary step. To date, there are no widely adopted solutions that allow accelerators to directly transfer data. A new cache-coherent protocol, CXL, aims to facilitate easier, fine-grained sharing between accelerators. In this work we analyze existing methods for designing heterogeneous applications that target GPUs and FPGAs working collaboratively, followed by an exploration to show the benefits of a CXL-enabled system. Specifically, we develop a test application that utilizes both an NVIDIA P100 GPU and a Xilinx U250 FPGA to show current communication limitations. From this application, we capture overall execution time and throughput measurements on the FPGA and GPU. We use these measurements as inputs to novel CXL performance models to show that using CXL caching instead of host memory results in a 1.31X speedup, while a more tightly-coupled pipelined implementation using CXL-enabled hardware would result in a speedup of 1.45X.

Cabrera, Anthony↗

Cross-Feature Transfer Learning for Efficient Tensor Program Generation

Tuning tensor program generation involves navigating a vast search space to find optimal program transformations and measurements for a program on the target hardware. The complexity of this process is further amplified by the exponential combinations of transformations, especially in heterogeneous environments. This research addresses these challenges by introducing a novel approach that learns the joint neural network and hardware features space, facilitating knowledge transfer to new, unseen target hardware. A comprehensive analysis is conducted on the existing state-of-the-art dataset, TenSet, including a thorough examination of test split strategies and the proposal of methodologies for dataset pruning. Leveraging an attention-inspired technique, we tailor the tuning of tensor programs to embed both neural network and hardware-specific features. Notably, our approach substantially reduces the dataset size by up to 53% compared to the baseline without compromising Pairwise Comparison Accuracy (PCA). Furthermore, our proposed methodology demonstrates competitive or improved mean inference times with only 25–40% of the baseline tuning time across various networks and target hardware. The attention-based tuner can effectively utilize schedules learned from previous hardware program measurements to optimize tensor program tuning on previously unseen hardware, achieving a top-5 accuracy exceeding 90%. This research introduces a significant advancement in autotuning tensor program generation, addressing the complexities associated with heterogeneous environments and showcasing promising results regarding efficiency and accuracy.

97 MATHEMATICS AND COMPUTING↗

P38 heterogeneous multi-tiled system with support for message queues (MoSAIC) v0.1

The proposed system is written in the hardware description language (HDL) verilog targeting an FPGA board. It is intended as a testbed to explore architecture tradeoffs in multi-tiled heterogeneous architectures. Although we target FPGAs, the system can be implemented as a monolithic SoC or a package comprised of many chiplets that are interconnected in the same package using a NoC. The proposed NoC is lightweight and follows an axi-lite interface. The endpoints of the NoC are a heterogeneous mix of "tiles" as endpoints that are general purpose processors, fixed function accelerators, and programmable accelerators. We assume that the network interfaces for the NoC endpoints are all addressable in a global name-space in that they represent an address range (for memory addresses) or a range of unique identifiers that are associated with each individual tile. This makes the functionality abstract from the standpoint of the NoC design details. Message queues offer a direct inter-processor interface between peer general purpose cores and diverse accelerators that comprise an SoC. Although they share the same NoC infrastructure for inter-tile communication within an SoC or SiP, the hardware message queues bypass the memory hierarchy and thus do not pollute the memory state or invoke the cache coherence mechanism.

Gonzalez, LouisaPatricia↗

Radiation specification and testing of heterogenous microprocessor SOCs

Modern commercial microprocessor devices include multiple processor architectures, buses, basic peripherals, and application hardware such as Graphics Processing Units (GPUs) and Digital Signal Processors (DSPs) in one device. Developing RHBD versions of similar devices risks sacrificing processing performance for system-wide radiation requirements. The heterogenous structure of modern commercial system on a chip (SOC) devices, in design and performance goals for subsystems, suggests a similar approach to specifying Radiation Hardened by Design (RHBD) requirements.

Ballast, Jon↗

Deep Learning Method for Detecting Precursors to Adverse Events

With the recent advancements in Deep Learning methods, the ability to model large complex heterogeneous data sets are fundamentally changing industry and research. Coupled with hardware improvements, and ease of implementation, a wide variety of deep neural network architectures can quickly be developed to solve a sweeping range of problems such as: object detection in images, automatic healthcare diagnosis using heterogenous data sources, real time language translating and sentence prediction, upscaling low resolution images, and forecasting of multivariate timeseries. Generally, many of these architectures outperform classical machine learning approaches in their respective tasks, however, this typically comes at a cost of interpretability. These black box algorithms generally suffer from lack of transparency in both model complexity as well as the rationale behind the prediction. This lack of comprehension, is driving an emerging area of interest in “Explainable AI”. An algorithm called: “Deep Temporal Multiple Instance Learning”1 was a recently developed to identify precursors to adverse events and has been applied in the aviation domain. The deep learning architecture is designed to capture the evolution of the probability of the outcome over the time preceding the adverse event using a multiple instance learning approach as illustrated in Figure 1. Precursors are defined when the probability of the event has exceeded a threshold at some point in the timeseries, at which point, a sensitivity analysis is performed to determine contributing factors. The contributing factors are used to explain and define the precursor during the periods where the probability score is high. The identified contributing factors are then presented to subject matter experts to provide objective insights into the leading factors associated with the particular adverse event. The algorithm has been tested on flight data from a commercial airline and has the ability to discover precursors to known adverse events that take the form of safety critical operations, such as unstable approach events on final approach. Apart from detecting precursors to adverse events, the converse can also be leveraged to discover corrective actions. These positive actions manifest themselves as periods in the timeseries when the precursor score has been lowered from an elevated state; meaning that if the system had been left uncorrected, it would have eventually reached the adverse event state. Characterizing these state changes can help identify successful interventions that may not have been known before. Policy makers and procedure designers can use this additional knowledge to craft more safety and efficient resilient procedures for future operations and therefore improve the overall performance of the National Airspace.

Matthews, Bryan L.↗

Portable Programming Model Exploration for LArTPC Simulation in a Heterogeneous Computing Environment: OpenMP vs. SYCL

The evolution of the computing landscape has resulted in the proliferation of diverse hardware architectures, with different flavors of GPUs and other compute accelerators becoming more widely available. To facilitate the efficient use of these architectures in a heterogeneous computing environment, several programming models are available to enable portability and performance across different computing systems, such as Kokkos, SYCL, OpenMP and others. As part of the High Energy Physics Center for Computational Excellence (HEP-CCE) project, we investigate if and how these different programming models may be suitable for experimental HEP workflows through a few representative use cases. One of such use cases is the Liquid Argon Time Projection Chamber (LArTPC) simulation which is essential for LArTPC detector design, validation and data analysis. Following up on our previous investigations of using Kokkos to port LArTPC simulation in the Wire-Cell Toolkit (WCT) to GPUs, we have explored OpenMP and SYCL as potential portable programming models for WCT, with the goal to make diverse computing resources accessible to the LArTPC simulations. In this work, we describe how we utilize relevant features of OpenMP and SYCL for the LArTPC simulation module in WCT. We also show performance benchmark results on multi-core CPUs, NVIDIA and AMD GPUs for both the OpenMP and the SYCL implementations. Comparisons with different compilers will also be given where appropriate.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Formally Verified ZTA Requirements for OT/ICS Environments with Isabelle/HOL

The clean energy transformation includes the integration of distributed energy resources with the power grid, which has led to a substantial increase in the complexity of power grids infrastructure and the underlying operational technology environment. Power grids infrastructure represents an operational technology environment that has become a system of systems, integrating heterogeneous devices which are both software-and hardware-intensive; as a result, there are increasing demands to exploit advances in the commodity of software-hardware infrastructures to improve energy systems requirements such as cybersecurity and resilience. In such a setting, system requirements at different levels mix, which leads to vulnerabilities and undesirable outcomes. The use of formal methods to characterize and prove system requirements removes ambiguity, increases automation, and provides high levels of assurance and reliability. In this paper, we contribute a methodology and a framework for the system-level verification of zero trust architecture requirements in operational technology environments. We define a formal specification for the core functionalities of operational technology environments, the corresponding invariants, and security proofs. Of particular note is our modular approach for the formal verification of asynchronous interactions in operational technology environments. The formal specification and the proofs have been mechanized using the interactive theorem proving environment Isabelle/HOL.

formal methods↗

A Case For Intra-rack Resource Disaggregation in HPC

The expected halt of traditional technology scaling is motivating increased heterogeneity in high-performance computing (HPC) systems with the emergence of numerous specialized accelerators. As heterogeneity increases, so does the risk of underutilizing expensive hardware resources if we preserve today’s rigid node configuration and reservation strategies. This has sparked interest in resource disaggregation to enable finer-grain allocation of hardware resources to applications. However, there is currently no data-driven study of what range of disaggregation is appropriate in HPC. To that end, we perform a detailed analysis of key metrics sampled in NERSC’s Cori, a production HPC system that executes a diverse open-science HPC workload. In addition, we profile a variety of deep-learning applications to represent an emerging workload. We show that for a rack (cabinet) configuration and applications similar to Cori, a central processing unit with intra-rack disaggregation has a 99.5% probability to find all resources it requires inside its rack. In addition, ideal intra-rack resource disaggregation in Cori could reduce memory and NIC resources by 5.36% to 69.01% and still satisfy the worst-case average rack utilization.

97 MATHEMATICS AND COMPUTING↗

A memory-driven mapping algorithm for heterogeneous systems

mpibind is a memory-driven algorithm to map parallel hybrid applications to the underlying hardware resources transparently, efficiently, and portably. There are two fundamental aspects of this algorithm. First, unlike existing mappings, its primary design point is the memory system. Compute elements are selected based on the identified memory components and not vice versa. Second, it embodies a global awareness of hybrid programming abstractions as well as heterogeneous devices.

Leon Borja, EdgarA↗

Face Recognition Oak Ridge (FaRO): A Framework for Distributed and Scalable Biometrics Applications

The facial biometrics community has seen a recent abundance of high-accuracy facial analytic models become freely available. Although these models' capabilities in facial detection, landmark detection, attribute analysis, and recognition are ever-increasing, they aren't always straightforward to deploy in a real-world environment. In reality, the use of the field's ever growing collection of models is becoming exceedingly difficult as library dependencies update and deprecate. Researchers often encounter headaches when attempting to utilize multiple models requiring different or conflicting software packages. Face Recognition Oak Ridge (FaRO) is an open-source project designed to provide a highly modular, flexible framework for unifying facial analytic models through a compartmentalized plug-and-play paradigm built on top of the gRPC (Google Remote Procedure Call) protocol. FaRO's server-client architecture and flexible portability allows easy construction of modularized and heterogeneous face analysis pipelines, distributed over many machines with differing hardware and software resources. This paper outlines FaRO's architecture and current capabilities, along with some experiments in model testing and distributed scaling through FaRO.

Bolme, David↗

Adventitious Carbon on Primary Sample Containment Metal Surfaces

Future missions that return astromaterials with trace carbonaceous signatures will require strict protocols for reducing and controlling terrestrial carbon contamination. Adventitious carbon (AC) on primary sample containers and related hardware is an important source of that contamination. AC is a thin film layer or heterogeneously dispersed carbonaceous material that naturally accrues from the environment on the surface of atmospheric exposed metal parts. To test basic cleaning techniques for AC control, metal surfaces commonly used for flight hardware and curating astromaterials at JSC were cleaned using a basic cleaning protocol and characterized for AC residue. Two electropolished stainless steel 316L (SS- 316L) and two Al 6061 (Al-6061) test coupons (2.5 cm diameter by 0.3 cm thick) were subjected to precision cleaning in the JSC Genesis ISO class 4 cleanroom Precision Cleaning Laboratory. Afterwards, the samples were analyzed by X-ray photoelectron spectroscopy (XPS) and Raman spectroscopy.

Calaway, M. J.↗

A Secondary Control Framework for Microgrid Interoperability With Vendor-Agnostic Grid-Forming Units: Design, Implementation, and Demonstration via Large-Scale Hardware Setup

The reliable operation of islanded microgrids increasingly depends on secondary controls that restore voltage and frequency to nominal values and ensure accurate active and reactive power sharing. Centralized secondary control architectures achieve high accuracy through global coordination at the cost of single-point failures and limited scalability compared with decentralized/distributed approaches. But a critical gap remains in addressing the interoperability and vendor-agnostic operation of secondary controls in real-world microgrids where heterogeneous diesel generator(s) and grid-forming (GFM) inverter(s) from multiple manufacturers always coexist. Practical and vendor-agnostic interoperability guidelines for the secondary control architecture of microgrids with multiple GFM units have not yet been developed; therefore, this paper proposes an interoperable and vendor-agnostic secondary control framework that operates seamlessly across GFM units from different vendors without relying on proprietary controls and protocols, hardware, or lock-ins. The framework leverages existing communication infrastructures (e.g., Modbus TCP/IP) to enable cost-effective deployment while addressing practical challenges, such as packet loss and quantization errors. Mitigation strategies-including data averaging, situational event-triggered control, and finite-iteration execution-are introduced to enhance reliability under real-world conditions. A generalized modeling and design framework is also presented, supported by robustness analysis to demonstrate independence from vendor-specific implementations. The proposed framework is validated through a large-scale hardware demonstration using a 3-$\phi$, 480-V, 60-Hz, 713-kVA laboratory hardware microgrid involving a heterogeneous diesel generator and multiple GFM inverters, showcasing its effectiveness in achieving stable voltage and frequency restoration and accurate power sharing under practical constraints. The results highlight the framework's potential as a scalable and practical solution for next-generation microgrids requiring openness, standard framework, and interoperability.

24 POWER TRANSMISSION AND DISTRIBUTION↗

FFTX-IRIS: Towards Performance Portability and Heterogeneity for SPIRAL Generated Code

FFTX-IRIS is a dynamic system to efficiently utilize novel heterogeneous platforms. This system links two next-generation frameworks, FFTX and IRIS, to navigate the complexity of different hardware architectures. FFTX provides a runtime code generation framework for high-performance Fast Fourier Transform kernels. IRIS runtime provides portability and multi-device heterogeneity, allowing computation on any available compute resource. Together, FFTX-IRIS enables code generation, seamless portability, and performance without user involvement. We show the design of the FFTX-IRIS system along with an evaluation of various small FFT benchmarks. We also demonstrate multi-device heterogeneity of FFTX-IRIS with a larger stencil application.

Rao, Sanil↗

Automated Generation of Integrated Digital and Spiking Neuromorphic Machine Learning Accelerators

The growing numbers of application areas for artificial intelligence (AI) methods have led to an explosion of domain-specific accelerators that could support every new machine learning (ML) algorithm advancement, clearly highlighting the need for a capability to quickly and automatically transition from algorithm definition to hardware implementation and explore design space along a variety of SWaP (size, weight and Power). The software defined architectures (SODA) synthesizer implements a compiler-based modular infrastructure for the end-to-end generation of machine learning accelerators from high-level frameworks to hardware description language. At the same time, neuromorphic computing, by mimicking how the brain operates, promises to perform artificial intelligence tasks at efficiencies orders of magnitude higher than the current conventional tensor-processing based accelerators, as demonstrated by a variety of specialized designs leveraging Spiking Neural Networks (SNNs). Nevertheless, the mapping of an artificial neural network (ANN) to solutions supporting SNNs is still a non-trivial and very device-specific task, and completely lack the possibility to design hybrid systems that integrate conventional and spiking neural models. In this paper we discuss the support for such an integrated generation leveraging the SODA Synthesizer framework and its modular structure. In particular, we present a new MLIR dialect (part of the SODA frontend) that allows expressing spiking neural network features (e.g., available resources, spiking sequences, analog signal reading, etc.) and illustrate how it enables mapping to Spiking Neurons and deployment to the related specialized hardware (which, in the digital domain, could be generated through the other existing layers of the SODA Synthesizer). We then discuss the opportunities for even deeper integration afforded by the hardware compilation infrastructure, providing a path towards the generation of complex heterogeneous artificial intelligence systems.

Curzel, Serena↗