Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “heterogeneous memory systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Scalable Heterogeneous Execution of a Coupled-Cluster Model with Perturbative Triples

The CCSD(T) coupled-cluster model with perturbative triples is considered a gold standard for computational modeling of the correlated behavior of electrons in molecular systems. A fundamental constraint is the relatively small global-memory capacity in GPUs compared to the main-memory capacity on host nodes, necessitating relatively smaller tile sizes for high-dimensional tensor contractions in NWChem's GPU-accelerated implementation of the CCSD(T) method. A coordinated redesign is described to address this limitation and associated data movement overheads, including a novel fused GPU kernel for a set of tensor contractions, along with inter-node communication optimization and data caching. The new implementation of GPU-accelerated CCSD(T) improves overall performance by 3.4x. Finally, we discuss the trade-offs in using this fused algorithm on current and future supercomputing platforms.

Kim, Jinsung↗

Dispatch Manager for NEML2 Constitutive Model Calculations Embedded in MOOSE

This report describes the extended capabilities of the NEML2 constitutive modeling library, including a flexible and efficient work dispatching system designed to leverage both CPU and GPU resources. This enhancement addresses one of the primary computational challenges in large-scale simulations: the ability to distribute and execute batches of material model evaluations across heterogeneous computing devices. The new dispatch system introduces a modular set of dispatcher and scheduler classes that coordinate the flow of data and execution between devices. The dispatcher is responsible for efficiently packaging work, managing device-specific memory operations, and synchronizing results. This modularity allows for extensibility, making it straightforward to integrate additional computing backends in the future. From an implementation standpoint, the dispatcher system interfaces seamlessly with NEML2's existing models. They handle device-aware tensor operations, optimize memory transfers, and support asynchronous execution when applicable. This design ensures that batches of material points can be evaluated concurrently, substantially improving throughput compared to previous single-device or serial implementations. These improvements not only enhance the raw performance of NEML2 but also improve its usability in multiscale and high-fidelity simulations, where the simultaneous evaluation of large material point batches is critical. Benchmarks included in the report demonstrate the system’s scalability, highlighting its effectiveness when leveraging modern GPU architectures.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Phase I Final Technical Report on Energy-Efficient Reconfigurable Universal Accelerator Interconnect

It is well known that application specific computing systems, optimally designed and configured for a given workload, offer much higher energy-efficiency and throughput than general purpose systems. In modern computing systems, heterogeneous computing systems have emerged that exploit the energy and performance benefits of combining various different domain-specific processor architectures. Application domains such as high-performance computing and machine learning now process terabyte-sized data sets, requiring enormous processing and memory resources. These applications have very high-power consumption due to bottlenecks in the electrical interconnection between processing units. This project aims to reduce both communication energy and latency by integrating universally available accelerators with silicon photonics. It also aims to increase system throughput by exploiting emerging technologies in silicon photonic reconfigurable interconnects; this will allow the system to balance itself in real time to accommodate changes in workloads and data flows.

2.5D/3D integration↗

Computing the Properties of Matter with Leadership Computing Resources (Closeout Report for DE-SC0018121)

In order to add more capabilities to Halide, we have designed a new framework called Tiramisu and integrated this framework into Halide. Since Tiramisu enables Halide to target heterogeneous architectures, our development efforts have been refocused on Tiramisu. Most high-performance computer systems today are complex and increasingly heterogeneous; they may have CPUs, GPUs and FPGAs. Achieving best performance requires taking full advantage of all these different architectures. To address this issue, we have designed Tiramisu, an optimization framework that enables Halide (and other DSLs) to target heterogeneous architectures. Tiramisu is an optimization framework that takes as input a high level, architecture-independent representation of code and a set of scheduling and data mapping commands that guide code transformation. The input can either be generated by a domain-specific language (DSL) compiler such as Halide or directly written by a programmer. Tiramisu then applies the user-specified code and data-layout transformations and generates an architecture-specific, low-level intermediate representation (IR) that takes advantage of modern architectural features such as multicore parallelism, non-uniform memory (NUMA) hierarchies, clusters, and accelerators like GPUs and FPGAs. We integrated Tiramisu within Halide and implemented a representative set of benchmarks to evaluate this integration. Tiramisu is now open source and is available for public use (http://tiramisu-compiler.org/). A paper about Tiramisu was published, it shows that Tiramisu extends Halide with many new capabilities and that Tiramisu can generate efficient code for multicores, GPUs, FPGAs and distributed heterogeneous systems. The performance of code generated by the Tiramisu backends matches or exceeds hand optimized reference implementations. For example, the multicore backend matches the highly optimized Intel MKL library on many kernels and shows speedups reaching 4x over the original Halide. In addition to making Tiramisu more robust, we have used Tiramisu to implement a set of representative tensor operation for constructing baryon building blocks required for multi baryon contractions in LQCD. In order to implement this code, we needed to generalize Tiramisu in two ways: first we needed to support indirect array accesses, and second, we needed to add support for complex numbers to Tiramisu. The code generated by Tiramisu is 6x faster than the reference code. Our efforts towards an MPI based multi-node version of tiramisu have matured and the resulting code scales well on multiple nodes (tests up to 512 KNL nodes have been undertaken).

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Performance Improvements of the Griffin Solvers in FY24

The Griffin code is a MOOSE-based reactor physics application jointly developed by Idaho National Laboratory and Argonne National Laboratory under the Department of Energy Office of Nuclear Energy Nuclear Energy Advanced Modeling and Simulation Program. This fiscal year, we have made significant efforts to improve the performance of transport solver options and cross-section generation for the efficient use of Griffin in advanced reactor applications. For the HFEM-PN solver, the residual evaluations of HFEM kernels were optimized by utilizing the pre- computed averaged cross sections for individual elements. Numerical integration involving the evaluation of basis functions at quadrature points was bypassed by facilitating precomputed element mass matrices for response matrices. Red-black iterations were improved by introducing a new generalized minimum residual based solver. The memory usage of response matrix storage was significantly reduced by applying basis function rotations on interfaces and calculating volumetric odd-parity moments on the fly. Additionally, the adjoint flux and transient calculation capabilities of the HFEM-PN solver were successfully implemented and verified using the TWIGL benchmark problem. For the DFEM-SN solver, memory footprint and computation time were significantly reduced by not treating angular flux vectors as the MOOSE nonlinear system vectors. Specifically for IQS, scalar adjoint weighting was introduced to further eliminate angular adjoint flux storage in the MOOSE auxiliary system. It was demonstrated through the three-dimensional Advanced Burner Test Reactor core problem that the memory usage for transient calculations with the IQS method was reduced by over 7.5× compared to before the optimizations. For the self-shielding application programming interface, a new double-heterogeneity treatment method, named the Bell Function-Based Analytic Two-Region Slowing Down Method, was developed to efficiently flux-volume homogenize TRISO particles with the matrix. Additionally, optimizations were made to hyper- fine group (HFG) slowing down calculations by pretabulating collision probability coefficients and grouping isotopes, significantly reducing the computational time for calculating scattering sources per HFG. Lastly, the pin power reconstruction module was extended to account for temporal behavior in a microreactor analysis problem, specifically for a control drum transient. Verification tests for each of these improvements demonstrated significant performance enhancements and memory reduction.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Development of Short-Term Forecasting Models Using Plant Asset Data and Feature Selection

Nuclear power plants collect and store large volumes of heterogeneous data from various components and systems. With recent advances in machine learning (ML) techniques, these data can be leveraged to develop diagnostic and short-term forecasting models to better predict future equipment condition. Maintenance operations can then be planned in advance whenever degraded performance is predicted, thus resulting in fewer unplanned outages and the optimization of maintenance activities. This enables lower maintenance costs and improves the overall economics of nuclear power. This paper focuses on developing a short-term forecasting process that leverages a feature selection process to distill large volumes of heterogeneous data and predict specific equipment parameters. A variety of feature selection methods, including Shapley Additive Explanations (SHAP) and variance inflation factor (VIF), were used to select the optimal features as inputs for three ML methods: long short-term memory (LSTM) networks, support vector regression (SVR), and random forest (RF). Each combination of model and input features was used to predict a pump bearing temperature both 1 and 24 hours in advance, based on actual plant system data. The optimal inputs for the LSTM and SVR were selected using the SHAP values, while the optimal input for the RF consisted solely of the response variable itself. Each model produced similar 1-hour-ahead predictions, with root mean square errors (RMSEs) of roughly 0.006. For the 24-hour-ahead predictions, differences could be seen between LSTM, SVR, and RF, as reflected by model performances of 0.036 +- 0.014, 0.0026 +- 0, and 0.063 +- 0.004 RMSE, respectively. As big data and continuous online monitoring become more widely available, the proposed feature selection process can be used for many applications beyond the prediction of process parameters within nuclear infrastructure.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Analytics-at-scale of Sensor Data for Digital Monitoring in Nuclear Plants (3 rd Annual Report)

Nuclear power plants collect and store large volumes of heterogeneous data from various components and systems. With recent advances in machine learning (ML) techniques, these data can be leveraged to develop diagnostic and short-term forecasting models to better predict future equipment condition. Maintenance operations can then be planned in advance whenever degraded performance is predicted, thus resulting in fewer unplanned outages and the optimization of maintenance activities. This enables lower maintenance costs and improves the overall economics of nuclear power. This report primarily focuses on developing a short-term forecasting process that leverages a feature selection process to distill large volumes of heterogeneous data and predict specific equipment parameters. A variety of feature selection methods, including Shapley Additive Explanations (SHAP) and variance inflation factor (VIF), were used to select the optimal features as inputs for three ML methods: long short-term memory (LSTM) networks, support vector regression (SVR), and random forest (RF). Each combination of model and input features was used to predict a pump bearing temperature both 1 and 24 hours in advance, based on actual plant system data. The optimal inputs for the LSTM and SVR were selected using the SHAP values, while the optimal input for the RF consisted solely of the response variable itself. Each model produced similar 1-hour-ahead predictions, with root mean square errors (RMSEs) of roughly 0.006. For the 24-hour-ahead predictions, differences could be seen between LSTM, SVR, and RF, as reflected by model performances of 0.036 ± 0.014, 0.0026 ± 0, and 0.063 ± 0.004 RMSE, respectively. As big data and continuous online monitoring become more widely available, the proposed feature selection process can be used for many applications beyond the prediction of process parameters within nuclear infrastructure. This report summarizes the Fiscal Year 2021 research progress encompassing the (1) data cleaning and feature selection necessary for ML applications; (2) development of short-term forecasting models to predict future plant process parameters for both single and multiple time steps ahead; and (3) validation of the feature selection methods and short-term forecasting models given new data from different systems.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Therapeutic targeting of membrane-associated proteins in central nervous system tumors

The activity of the most complex system, the central nervous system (CNS) is profoundly regulated by a huge number of membrane-associated proteins (MAP). A minor change stimulates immense chemical changes and the elicited response is organized by MAP, which acts as a receptor of that chemical or channel enabling the flow of ions. Slight changes in the activity or expression of these MAPs lead to severe consequences such as cognitive disorders, memory loss, or cancer. CNS tumors are heterogeneous in nature and hard-to-treat due to random mutations in MAPs; like as overexpression of EGFRvIII/TGFβR/VEGFR, change in adhesion molecules α5β3 integrin/SEMA3A, imbalance in ion channel proteins, etc. Extensive research is under process for developing new therapeutic approaches using these proteins such as targeted cytotoxic radiotherapy, drug-delivery, and prodrug activation, blocking of receptors like GluA1, developing viral vector against cell surface receptor. The combinatorial approach of these strategies along with the conventional one might be more potential. Henceforth, our review focuses on in-depth analysis regarding MAPs aiming for a better understanding for developing an efficient therapeutic approach for targeting CNS tumors.

60 APPLIED LIFE SCIENCES↗

Long short-term memory embedded nudging schemes for nonlinear data assimilation of geophysical flows

Reduced rank nonlinear filters are increasingly utilized in data assimilation of geophysical flows, but often require a set of ensemble forward simulations to estimate forecast covariance. On the other hand, predictor-corrector type nudging approaches are still attractive due to their simplicity of implementation when more complex methods need to be avoided. However, optimal estimate of nudging gain matrix might be cumbersome. In this paper, we put forth a fully nonintrusive recurrent neural network approach based on a long short-term memory (LSTM) embedding architecture to estimate the nudging term, which plays a role not only to force the state trajectories to the observations but also acts as a stabilizer. Furthermore, our approach relies on the power of archival data and the trained model can be retrained effectively due to power of transfer learning in any neural network applications. In order to verify the feasibility of the proposed approach, we perform twin experiments using Lorenz 96 system. Our results demonstrate that the proposed LSTM nudging approach yields more accurate estimates than both extended Kalman filter (EKF) and ensemble Kalman filter (EnKF) when only sparse observations are available. With the availability of emerging AI-friendly and modular hardware technologies and heterogeneous computing platforms, we articulate that our simplistic nudging framework turns out to be computationally more efficient than either the EKF or EnKF approaches.

42 ENGINEERING↗

Scalable Quantum Monte Carlo Method for Polariton Chemistry via Mixed Block Sparsity and Tensor Hypercontraction Method

We present a reduced-scaling auxiliary-field quantum Monte Carlo (AFQMC) framework designed for large molecular systems and ensembles, with or without coupling to optical cavities. Our approach leverages the natural block sparsity of the Cholesky decomposition (CD) of electron repulsion integrals in molecular ensembles and employs tensor hypercontraction (THC) to efficiently compress low-rank Cholesky blocks. By representing the Cholesky vectors in a mixed format, keeping high-rank blocks in block-sparse form and compressing low-rank blocks with THC, we reduce the scaling of exchange-energy evaluation from quartic to robust cubic in the number of molecular orbitals N, while lowering memory from cubic toward quadratic. Benchmark analyses on one-, two-, and three-dimensional molecular ensembles (up to ∼1,200 orbitals) show that (a) the number of nonzeros in Cholesky tensors grows linearly with system size across dimensions; (b) the average numerical rank increases sublinearly and does not saturate at these sizes; and (c) rank heterogeneity─some blocks nearly full rank and many low rank, naturally motivates the proposed mixed block sparsity and THC scheme for efficient calculation of exchange energy. In conclusion, we demonstrate that the mixed scheme yields cubic wall-time scaling with favorable prefactors and preserves AFQMC accuracy.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Large-scale integration of artificial atoms in hybrid photonic circuits

A central challenge in developing quantum computers and long-range quantum networks is the distribution of entanglement across many individually controllable qubits. Colour centres in diamond have emerged as leading solid-state ‘artificial atom’ qubits because they enable on-demand remote entanglement, coherent control of over ten ancillae qubits with minute-long coherence times and memory-enhanced quantum communication. A critical next step is to integrate large numbers of artificial atoms with photonic architectures to enable large-scale quantum information processing systems. So far, these efforts have been stymied by qubit inhomogeneities, low device yield and complex device requirements. Here we introduce a process for the high-yield heterogeneous integration of ‘quantum microchiplets’—diamond waveguide arrays containing highly coherent colour centres—on a photonic integrated circuit (PIC). We use this process to realize a 128-channel, defect-free array of germanium-vacancy and silicon-vacancy colour centres in an aluminium nitride PIC. Photoluminescence spectroscopy reveals long-term, stable and narrow average optical linewidths of 54 megahertz (146 megahertz) for germanium-vacancy (silicon-vacancy) emitters, close to the lifetime-limited linewidth of 32 megahertz (93 megahertz). We show that inhomogeneities of individual colour centre optical transitions can be compensated in situ by integrated tuning over 50 gigahertz without linewidth degradation. The ability to assemble large numbers of nearly indistinguishable and tunable artificial atoms into phase-stable PICs marks a key step towards multiplexed quantum repeaters and general-purpose quantum processors

36 MATERIALS SCIENCE↗

IRIS-GNN: Leveraging Graph Neural Networks for Scheduling on Truly Heterogeneous Runtime Systems

The diversity of accelerators in computer systems poses significant challenges for software developers, such as managing vendor-specific compiler toolchains, code fragmentation requiring different kernel implementations, and performance portability issues. To address these, the Intelligent Runtime System (IRIS) was developed. IRIS works across various systems, from smartphones to supercomputers, enabling automatic performance scaling based on available accelerators. It introduces abstract tasks for seamless execution transitions between accelerators while ensuring memory consistency and task dependencies. Although IRIS simplifies system details, optimal dynamic scheduling still requires user input to understand workload structures. To address this, we introduce a new scheduling policy for IRIS, termed IRIS-GNN, which is the first IRIS hybrid policy that operates in conjunction with the dynamic policies. This policy employs a Graph-Neural Network (GNN) to conduct Graph Classification of any task graphs submitted to IRIS. This GNN analyzes the structure and attributes of the task graph, categorizing it as either locality, concurrency, or mixed. This classification subsequently guides the selection of the dynamic policy used by IRIS. We provide a comparison of the performance of IRIS-GNN against the complete spectrum of IRIS’s dynamic policies, assess the overhead introduced by the GNN within this scheduling framework, and ultimately explore its practical application in real-world scenarios.

Johnston, Beau↗

Asynchronous domain decomposition methods for nonlinear PDEs

One- and two-level parallel asynchronous methods for the numerical solution of nonlinear systems of equations, especially those arising from (nonlinear) partial differential equations, are studied. The proposed methods are based on domain decomposition techniques. Local convergence theorems are presented in several cases, with appropriate hypotheses. Computational results on a shared memory multiprocessor machine for various problems exhibiting nonlinearities are reported, illustrating the potential of these asynchronous methods, especially for heterogeneous clusters.

97 MATHEMATICS AND COMPUTING↗

BULKI-Store v0.3.2

BULKI-Store is a distributed object storage system optimized for high-performance computing environments. Built with a Rust core and Python bindings, it efficiently manages scientific and machine learning datasets across HPC clusters. The system employs a client-server architecture with MPI integration, enabling seamless scaling on supercomputers like Perlmutter. BULKI-Store's object-oriented approach provides intuitive data organization with rich metadata support, contrasting with traditional file-based solutions. Key optimizations include selective checkpoint loading, unified checkpoint files, and object chunking for large data transfers. For machine learning workloads, BULKI-Store offers advantages through fine-grained access patterns, dynamic data sharing between training instances, and reduced memory pressure. Memory management features include strategic Python GC calls, minimized data copies, and batch processing capabilities. The system leverages Rayon's thread pool for asynchronous data prefetching and supports multiple CPU architectures (ARM64, x86, AMD, RISC-V). By combining performance optimizations with developer-friendly APIs, BULKI-Store addresses the complex data management challenges of modern HPC applications while maintaining compatibility across heterogeneous computing environments.

Zhang, Wei [Lawrence Berkeley National Laboratory ↗

Technical note: Using long short-term memory models to fill data gaps in hydrological monitoring networks

Abstract. Quantifying the spatiotemporal dynamics in subsurface hydrological flows over a long time window usually employs a network of monitoring wells. However, such observations are often spatially sparse with potential temporal gaps due to poor quality or instrument failure. In this study, we explore the ability of recurrent neural networks to fill gaps in a spatially distributed time-series dataset. We use a well network that monitors the dynamic and heterogeneous hydrologic exchanges between the Columbia River and its adjacent groundwater aquifer at the U.S. Department of Energy's Hanford site. This 10-year-long dataset contains hourly temperature, specific conductance, and groundwater table elevation measurements from 42 wells with gaps of various lengths. We employ a long short-term memory (LSTM) model to capture the temporal variations in the observed system behaviors needed for gap filling. The performance of the LSTM-based gap-filling method was evaluated against a traditional autoregressive integrated moving average (ARIMA) method in terms of error statistics and accuracy in capturing the temporal patterns of river corridor wells with various dynamics signatures. Our study demonstrates that the ARIMA models yield better average error statistics, although they tend to have larger errors during time windows with abrupt changes or high-frequency (daily and subdaily) variations. The LSTM-based models excel in capturing both high-frequency and low-frequency (monthly and seasonal) dynamics. However, the inclusion of high-frequency fluctuations may also lead to overly dynamic predictions in time windows that lack such fluctuations. The LSTM can take advantage of the spatial information from neighboring wells to improve the gap-filling accuracy, especially for long gaps in system states that vary at subdaily scales. While LSTM models require substantial training data and have limited extrapolation power beyond the conditions represented in the training data, they afford great flexibility to account for the spatial correlations, temporal correlations, and nonlinearity in data without a priori assumptions. Thus, LSTMs provide effective alternatives to fill in data gaps in spatially distributed time-series observations characterized by multiple dominant frequencies of variability, which are essential for advancing our understanding of dynamic complex systems.

54 ENVIRONMENTAL SCIENCES↗

Accurate assessment of land–atmosphere coupling in climate models requires high-frequency data output

Land–atmosphere (L–A) interactions are important for understanding convective processes, climate feedbacks, the development and perpetuation of droughts, heatwaves, pluvials, and other land-centered climate anomalies. Local L–A coupling (LoCo) metrics capture relevant L–A processes, highlighting the impact of soil and vegetation states on surface flux partitioning and the impact of surface fluxes on boundary layer (BL) growth and development and the entrainment of air above the BL. A primary goal of the Climate Process Team in the Coupling Land and Atmospheric Subgrid Parameterizations (CLASP) project is parameterizing and characterizing the impact of subgrid heterogeneity in global and regional Earth system models (ESMs) to improve the connection between land and atmospheric states and processes. A critical step in achieving that aim is the incorporation of L–A metrics, especially LoCo metrics, into climate model diagnostic process streams. However, because land–atmosphere interactions span timescales of minutes (e.g., turbulent fluxes), hours (e.g., BL growth and decay), days (e.g., soil moisture memory), and seasons (e.g., variability in behavioral regimes between soil moisture and latent heat flux), with multiple processes of interest happening in different geographic regions at different times of year, there is not a single metric that captures all the modes, means, and methods of interaction between the land and the atmosphere. And while monthly means of most of the LoCo-relevant variables are routinely saved from ESM simulations, data storage constraints typically preclude routine archival of the hourly data that would enable the calculation of all LoCo metrics. Here, we outline a reasonable data request that would allow for adequate characterization of sub-daily coupling processes between the land and the atmosphere, preserving enough sub-daily output to describe, analyze, and better understand L–A coupling in modern climate models. A secondary request involves embedding calculations within the models to determine mean properties in and above the BL to further improve characterization of model behavior. Higher-frequency model output will (i) allow for more direct comparison with observational field campaigns on process-relevant timescales, (ii) enable demonstration of inter-model spread in L–A coupling processes, and (iii) aid in targeted identification of sources of deficiencies and opportunities for improvement of the models.

54 ENVIRONMENTAL SCIENCES↗

IRIS: A Performance-Portable Framework for Cross-Platform Heterogeneous Computing

From edge to exascale, computer architectures are becoming more heterogeneous and complex. The systems typically have fat nodes, with multicore CPUs and multiple hardware accelerators such as GPUs, FPGAs, and DSPs. This complexity is causing a crisis in programming systems and performance portability. Several programming systems are working to address these challenges, but the increasing architectural diversity is forcing software stacks and applications to be specialized for each architecture. As we show, all of these approaches critically depend on their software framework for discovery, execution, scheduling, and data orchestration. To address this challenge, we believe that a more agile and proactive software framework is essential to increase performance portability and improve user productivity. To this end, we have designed and implemented IRIS: a performance-portable framework for cross-platform heterogeneous computing. IRIS can discover available resources, manage multiple diverse programming platforms (e.g., CUDA, Hexagon, HIP, Level Zero, OpenCL, OpenMP) simultaneously in the same execution, respect data dependencies, orchestrate data movement proactively, and provide for user-configurable scheduling. To simplify data movement, IRIS introduces a shared virtual device memory with relaxed consistency among different heterogeneous devices. IRIS also adds an automatic kernel workload partitioning technique using the polyhedral model so that it can resize kernels for a wide range of devices. Our evaluation on three architectures, ranging from Qualcomm Snapdragon to a Summit supercomputer node, shows that IRIS improves portability across a wide range of diverse heterogeneous architectures with negligible overhead.

97 MATHEMATICS AND COMPUTING↗

Proxy Applications for Converged Workloads: DMC LDRD Initiative

Modern scientific applications are complicated and require coordination of several components. Proxy application driven software-hardware co-design plays a vital role in driving innovation among the developments of applications, software infrastructure and hardware architecture. Proxy applications are self-contained and simplified codes that are intended to model the performance-critical computations within applications. Applications executing on modern High Performance Computing (HPC) systems are susceptible to network congestion, insufficient memory bandwidth within and across compute nodes, and inadvertent loss of performance due to bugs and unoptimized programming models. Modern numerical simulations and machine learning models play a critical role in studying physical phenomenon under myriad uncertainties. Such applications often exhibit irregular computation and memory accesses at specific regions of the application code, which can contribute to various performance bottlenecks at scale. To mitigate such issues and prepare the next generation hardware for a variety of computation and data movement contingencies, a well-known practice is to consider "proxy" applications as representative motifs for various classes of scientific applications. While there is disagreement in the HPC community on the mechanisms of construction of the proxy applications, there is a strong consensus on their positive impact in co-design. Proxy Applications for Converged Workloads (PACER) is about facilitating software-hardware co-design through proxy applications with the goal of improving the performance of converged science workflows on heterogeneous systems.

97 MATHEMATICS AND COMPUTING↗