Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “heterogeneous computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

From Reproducible Edge–Cloud Experimentation to Real-World Practice: The E2Clab Experience

Reproducibility is already difficult in distributed systems; on the computing continuum, it becomes substantially harder. Applications that span sensing devices, edge and fog resources, and cloud platforms must be evaluated across heterogeneous hardware, variable network conditions, cross-layer orchestration decisions, and long-running workflow lifecycles. We use E2Clab as a case study to examine these challenges and their implications for experimental methodology. We explain why reproducible experimentation is harder on the continuum, then revisit E2Clab as an initial response based on explicit modeling of infrastructure, workflow lifecycle, and artifacts. Lastly, we discuss how its evolution toward more realistic application settings can be understood through the lens of Translational Computer Science. We argue that reproducible continuum experimentation requires methods that are rigorous enough for research while remaining adaptable to real-world practice.

42 ENGINEERING↗

Portable, heterogeneous ensemble workflows at scale using libEnsemble

libEnsemble is a Python-based toolkit for running dynamic ensembles, developed as part of the DOE Exascale Computing Project. The toolkit utilizes a unique generator–simulator–allocator paradigm, where generators produce input for simulators, simulators evaluate those inputs, and allocators decide whether and when a simulator or generator should be called. The generator steers the ensemble based on simulation results. Generators may, for example, apply methods for numerical optimization, machine learning, or statistical calibration. libEnsemble communicates between a manager and workers. Flexibility is provided through multiple manager–worker communication substrates each of which has different benefits. These include Python’s multiprocessing, mpi4py, and TCP. Multisite ensembles are supported using Balsam or Globus Compute. We overview the unique characteristics of libEnsemble as well as current and potential interoperability with other packages in the workflow ecosystem. We highlight libEnsemble’s dynamic resource features: libEnsemble can detect system resources, such as available nodes, cores, and GPUs, and assign these in a portable way. These features allow users to specify the number of processors and GPUs required for each simulation; and resources will be automatically assigned on a wide range of systems, including Frontier, Aurora, and Perlmutter. Such ensembles can include multiple simulation types, some using GPUs and others using only CPUs, sharing nodes for maximum efficiency. We also describe the benefits of libEnsemble’s generator–simulator coupling, which easily exposes to the user the ability to cancel, and portably kill, running simulations based on models that are updated with intermediate simulation output. We demonstrate libEnsemble’s capabilities, scalability, and scientific impact via a Gaussian process surrogate training problem for the longitudinal density profile at the exit of a plasma accelerator stage. In conclusion, the study uses gpCAM for the surrogate model and employs either Wake-T or WarpX simulations, highlighting efficient use of resources that can easily extend to exascale.

Dynamic ensembles↗

Atomic-scale identification of active sites of oxygen reduction nanocatalysts

Heterogeneous nanocatalysts play a crucial role in both the chemical and energy industries. Despite substantial advancements in theoretical, computational and experimental studies, identifying their active sites remains a major challenge. Here we utilize atomic electron tomography to determine the three-dimensional atomic structure of PtNi and Mo-doped PtNi nanocatalysts for the electrochemical oxygen reduction reaction. We then employ the experimental atomic structures as input to first-principles-trained machine learning to identify the active sites of the nanocatalysts. Through the analysis of the structure–activity relationships, we formulate an equation termed the local environment descriptor, which balances the strain and ligand effects to provide physical and chemical insights into active sites in the oxygen reduction reaction. The ability to determine the three-dimensional atomic structure and chemical composition of realistic nanoparticles, combined with machine learning, could transform our fundamental understanding of the active sites of catalysts and guide the rational design of optimal nanocatalysts.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Performance Results on CPU/GPU Exascale Architectures for OMEGA: The Ocean Model for E3SM Global Applications

The US Department of Energy (DOE) conducts climate simulations on some of the world’s largest supercomputers. These exascale machines use heterogeneous architectures with both CPUs and GPUs, and scientific codes must adapt to make full use of this computing power. Los Alamos National Lab is developing Omega: The Ocean Model for E3SM Global Applications, which is specifically designed for modern exascale computers. It uses external libraries that have been optimized for a variety of architectures to run on different supercomputers. Omega is an unstructured-mesh ocean model based on TRiSK numerical methods. It will be the new ocean component of the DOE’s Energy Exascale Earth System Model (E3SM). The algorithms in Omega follow those of the current ocean component, MPAS-Ocean, but it will be written in C++ rather than Fortran to take advantage of the Kokkos performance portability library. Omega spatial operators are written as Kokkos kernels to run efficiently on both CPUs and GPUs. Work on Omega began in 2023 with a new C++ framework for unstructured mesh partitioning, halo exchanges, parallel IO, and Kokkos interfaces. The current version, Omega-0, is being developed to solve the shallow water equations and at present includes all of the tendency terms but not time stepping. Here we share the results of Omega-0 verification and performance testing. Verification includes unit tests implemented with CTest as well as convergence tests in Polaris, an in-house python package with a large suite of test problems. Performance tests compare simulations conducted on CPUs versus GPUs and across different architectures: tests are run on Frontier, which has AMD “Optimized 3rd Gen EPYC” CPUs and AMD MI250X GPUs, as well as Perlmutter, which is composed of AMD EPYC 7763 CPUs and NVIDIA A100 GPUs.

58 GEOSCIENCES↗

The Heterogeneous Integration of Electronic Components

Heterogeneous integration (HI) of electronics components is broadly recognized as a powerful and crucial enabler for the continued growth of computing and communication. From 2010 onwards, the value of HI is increasingly visible in the advanced packaging used in artificial intelligence, high-performance computing, smartphones and communications product implementations. In this Perspective, we argue that HI is crucial to semiconductors and more broadly to the continued evolution of computing and communications. We use leading-edge advanced packaging examples to represent the value, advancements and opportunities for HI. To succeed, it is critical to develop comprehensive HI roadmaps that inform collaborations across the design, manufacturing and reliability spectrum between systems architects, packaging and semiconductor technologists to common goals. Although this article does not provide a full roadmap, we instead detail additional parameters for artificial intelligence, smartphone and other cellular communication devices, and their constituent building blocks including interconnects, power electronics, photonics, thermal management, reliability, modelling and co-design, to foster greater collaboration opportunities among academia, research laboratories and industry.

42 ENGINEERING↗

The impact of capillary heterogeneity on CO 2 flow and trapping across scales

Capillary heterogeneity has been identified over the last decade as a key control on subsurface CO 2 flow behavior during geological CO 2 sequestration. These heterogeneities can be formed in all sedimentary rocks, ranging from slight variations in the sand grain sizes to extensive sequences of interbedded sands, shales, and limestones. Capillary heterogeneity has been largely, although not entirely, overlooked in subsurface flow modeling because it is assumed to only directly influence fluid redistribution over scales of centimeters to meters. However, even small-scale fluid movements can result in dramatic impacts on the mobility and trapping of the CO 2 over kilometers. Therefore, neglecting capillary heterogeneity at multiple scales could potentially lead to errors in modeling and predicting field-scale plume migration. In this review paper, we aim to provide a consistent overview to (1) establish that capillary heterogeneity can have a major impact on CO 2 plume migration, (2) establish the respective length scales at which capillary heterogeneity matters, and (3) provide guidance for numerical modeling. This review covers pertinent literature and extracts key observations from the core to the field scales. Experimental studies have shown that millimeter-decimeter scale capillary heterogeneity can cause the so-called capillary heterogeneity trapping in addition to pore-scale residual trapping. Even at such a small scale, capillary heterogeneity can already lead to complex upscaled constitutive relationships, such as flow-rate dependent and anisotropic relative permeability, which affects field-scale CO 2 migration even when field-scale heterogeneities are present. Under gravity-dominated flow regimes, centimeter-meter scale capillary heterogeneity can entrap a significant amount of CO 2 at field scale, not just after imbibition but also during drainage. In certain cases, the presence of capillary heterogeneity can even completely stop the vertical movement of the CO 2 plume, hence greatly reducing leakage risks. At meter-kilometer scale, the influence of capillary heterogeneity is more pronounced and can hinder or redirect CO 2 migration in both lateral and vertical directions. The impact of capillary heterogeneity across multiple spatial scales poses a great challenge in modeling CO 2 migration at field scale, because it is practically impossible to build a field-scale earth model with grid blocks at millimeter scale. We recommend a hierarchical modeling approach to address this challenge. At field scale, earth models are built to capture geological features and heterogeneities in high but still practical grid resolutions. For each facies or rock type of the field-scale model, high- resolution meter-scale “conceptual” models are built with millimeter-scale grid blocks to capture representative fine-scale bedding geometries and heterogeneities in various environments of deposition, bridging the gap from subcore scale to the size of a field-scale simulation grid block. Upscaling is then used to preserve the smaller-scale flow dynamics of various rock types in field-scale simulations. Here, future work is needed to (1) refine, improve, and validate the hierarchical modeling approach; (2) build libraries of fine-scale bedding models for facies in various environments of deposition; (3) quantify multiscale capillary heterogeneity effects under subsurface uncertainties; (4) gain learning from different storage formations; and (5) establish best practices that balance accuracy and computational speed.

Capillary heterogeneity↗

Physics-informed neural networks for heterogeneous poroelastic media

This study presents a novel physics-informed neural network (PINN) framework for modeling poroelasticity in heterogeneous media with material interfaces. The approach introduces a composite neural network (CoNN) where separate neural networks predict displacement and pressure variables for each material. While sharing identical activation functions, these networks are independently trained for all other parameters. To address challenges posed by heterogeneous material interfaces, the CoNN is integrated with the Interface-PINNs (I-PINNs) framework (Sarma et al., Comput. Methods Appl. Mech. Eng. 429: 117135, 2024), allowing different activation functions across material interfaces. Further, this ensures accurate approximation of discontinuous solution fields and gradients. Performance and accuracy of this combined architecture were evaluated against the conventional PINNs approach, a single neural network (SNN) architecture, and the eXtended PINNs (XPINNs) framework through two one-dimensional benchmark examples with discontinuous material properties. The results show that the proposed CoNN with I-PINNs architecture achieves an RMSE that is two orders of magnitude better than the conventional PINNs approach and is at least 40 times faster than the SNN framework. Compared to XPINNs, the proposed method achieves an RMSE at least one order of magnitude better and is 40% faster.

42 ENGINEERING↗

Applying Corrective Machine Learning in the E3SM Atmosphere Model in C++ (EAMxx)

The Simplified Cloud-Resolving E3SM Atmosphere Model (SCREAM) is the newest addition to the family of Earth System Models capable of explicitly resolving convective systems. SCREAM is a kilometer-scale configuration of the advanced E3SM Atmosphere Model (EAMxx), designed for heterogeneous systems. While the enhanced accuracy of kilometer-scale modeling offers significant benefits, it comes with a substantial computational cost, limiting feasible simulation durations to only a few years, even on the fastest supercomputers. Machine learning presents an opportunity for scientists to achieve the high accuracy of storm-resolving models at a significantly reduced cost. Building on the previous success of applying corrective machine learning (ML) to the FV3 model, this study explores the effects of implementing corrective ML in EAMxx-SCREAM. We also address the computational challenges of integrating the corrective ML, which is written in Python, with the C++/Kokkos EAMxx driver, as well as the potential pitfalls of generalizing an approach that was effective with one atmosphere model to another.

54 ENVIRONMENTAL SCIENCES↗

Energy–Performance Trade-offs in Privacy-Preserving Federated Learning on SmartNIC-Enabled HPC Systems

Federated learning (FL) is increasingly deployed on accelerator-rich high-performance computing (HPC) systems, yet the system-level energy cost of privacy-aware FL remains poorly understood, particularly across heterogeneous networking and server-placement options. We present a measurement-driven study of energy–performance trade-offs for FL on GH200-class nodes across three deployment configurations: CPU-Ethernet, CPU-InfiniBand (RDMA-capable), and a DPU-hosted FL server over InfiniBand using a BlueField-3 SmartNIC/DPU. Using NVIDIA FLARE (NVFLARE), we align node-level power telemetry with per-round timing extracted from NVFLARE logs to quantify time-to-solution (TTS), energy-to-solution (ETS), energy-delay product (EDP), and synchronization behavior for three transformer models (ALBERT, DistilBERT, BERT), trained with and without differential privacy (DP). We find that interconnect choice is the dominant driver of runtime and energy: host-managed InfiniBand consistently reduces communication overhead versus Ethernet, yielding lower TTS/ETS/EDP. In contrast, in our NVFLARE deployment, placing the FL server on the DPU does not consistently match CPU-InfiniBand performance and can be slower—especially for larger models—highlighting that server placement alone is not sufficient to guarantee end-to-end gains. Finally, under our fixed-round protocol, DP increases per-round cost and runtime variance; ETS increases largely in proportion to TTS because average node power remains relatively stable across configurations.

Kotevska, Olivera [ORNL] (ORCID:0000000316772243)↗

Fair Concurrent Training of Multiple Models in Federated Learning

Federated learning (FL) enables collaborative learning across multiple clients. In most FL work, all clients train a single learning task. However, the recent proliferation of FL applications may increasingly require multiple FL tasks to be trained simultaneously, sharing clients’ computing resources, which we call Multiple-Model Federated Learning (MMFL). Current MMFL algorithms use naïve average-based client-task allocation schemes that often lead to unfair performance when FL tasks have heterogeneous difficulty levels, as the more difficult tasks may need more client participation to train effectively. Furthermore, in the MMFL setting, we face a further challenge that some clients may prefer training specific tasks to others, and may not even be willing to train other tasks, e.g., due to high computational costs, which may exacerbate unfairness in training outcomes across tasks. We address both challenges by firstly designing FedFairMMFL, a difficulty-aware algorithm that dynamically allocates clients to tasks in each training round, based on the tasks’ current performance levels. We provide guarantees on the resulting task fairness and FedFairMMFL’s convergence rate. We then propose novel auction designs that incentivizes clients to train multiple tasks, so as to fairly distribute clients’ training efforts across the tasks, and extend our convergence guarantees to this setting. Here, we finally evaluate our algorithm with multiple sets of learning tasks on real world datasets, showing that our algorithm improves fairness by improving the final model accuracy and convergence speed of the worst performing tasks, while maintaining the average accuracy across tasks.

Federated learning↗

Polysulfide-incompatible additive suppresses spatial reaction heterogeneity of Li-S batteries

Rational electrolyte engineering for practical pouch cells remains elusive because the correlation between the cathode/solid-electrolyte interphase layer and cell-level reaction behavior is poorly understood. Here, by combining multiscale characterization and computational modeling, we show that—counter to the conventional perception of polysulfide-incompatible additives—the spontaneous reaction of sparingly solvated polysulfides with Lewis acid additives (LAAs) can induce in situ formation of a homogeneous interphase on thick and tortuous S cathode. Multiscale synchrotron X-ray characterization consistently affirms that such interface design could effectively eliminate the notorious problems of polysulfide shuttle and lithium corrosion and, more importantly, provide an interconnected “ion transport highway” to alleviate the uneven ion transport within the tortuous S cathode. Hence, this design dramatically reduces the reaction heterogeneity of lithium-sulfur (Li-S) pouch cells under lean electrolyte conditions. Further, this work resolves controversy around the role of polysulfide-incompatible additives in high-energy Li-S pouch cells and highlights the importance of suppressing reaction heterogeneity for practical batteries.

36 MATERIALS SCIENCE↗

Stress and Strain Heterogeneity and Persistence in Uniaxially‐ and Triaxially‐Loaded Sandstone

Two critical questions in brittle rock mechanics are how rocks developed localized strains and to what extent internal stress heterogeneity controls this localization and subsequent macroscopic failure. Definitive answers have not yet been found, but would provide insight into rock fracture mechanics as relevant to hydrocarbon extraction and sequestration. Here, we use synchrotron X‐ray tomography (XRT) and 3D X‐ray diffraction (3DXRD) during uniaxial and triaxial tests on Nugget and Bentheimer sandstones to examine strain and stress localization prior to mechanical failure. 3DXRD was used to measure intra‐granular lattice strains which were used to compute elastic stress tensors of each grain. Digital volume correlation (DVC) was applied to XRT images to determine the strain field in the sample. Both samples featured marked spatial heterogeneity, localization, and temporal persistence of elevated stresses and strains during their mechanical deformation toward failure. Both samples featured a majority of grains with at least one principal stress component that was tensile, a signature of the influence of heterogeneity on stress transmission. Measurements further revealed that compressive stress orientations and statistics evolved in a similar manner to those of inter‐particle forces in loose granular materials, with triaxially‐compressed rock exhibiting enhanced grain stress heterogeneity compared to uniaxially‐compressed rock. Our results complement recent work by others who employed XRT and scanning 3DXRD to study triaxially‐compressed sandstone, but extend those results to uniaxial compression, sandstones of varied porosity, and grain stress measurements throughout the 3D full extent of the samples rather than in a single layer examined with scanning 3DXRD.

58 GEOSCIENCES↗

Ethical considerations in infectious disease modelling for public health policy: the case of school closures

Mathematical models of infectious diseases are frequently used as a tool to support public health policy and decisions around the implementation of interventions such as school closures. However, most publications on policy-relevant modelling lack an ethical framework and do not explicitly consider the ethical implications of the work. This creates a risk that the unintended consequences of interventions are overlooked or that models are used to justify decisions that are inconsistent with public health ethics. In this article, we focus on the case study of school closures as a commonly modelled intervention against pandemic influenza, COVID-19 and other infectious disease threats. We briefly review some of the key concepts in public health ethics and describe approaches to modelling the effects of school closures. We then identify a series of ethical considerations involved in modelling school closures. These include accounting for population heterogeneity and inequalities; including a diversity of viewpoints and expertise in model design; considering the distribution of benefits and harms; and model transparency and contextualization. Furthermore, we conclude with some recommendations to ensure that policy-relevant modelling is consistent with some key ethics values.

97 MATHEMATICS AND COMPUTING↗

From RNNs to Foundation Models: An Empirical Study on Commercial Building Energy Consumption

Accurate short-term energy consumption forecasting for commercial buildings is crucial for smart grid operations. While smart meters and deep learning models enable forecasting using past data from multiple buildings, data heterogeneity from diverse buildings can reduce model performance. The impact of increasing dataset heterogeneity in time series forecasting, while keeping size and model constant, is understudied. We tackle this issue using the ComStock dataset, which provides synthetic energy consumption data for U.S. commercial buildings. Two curated subsets, identical in size and region but differing in building type diversity, are used to assess the performance of various time series forecasting models, including finetuned open-source foundation models (FMs). The results show that dataset heterogeneity and model architecture have a greater impact on post-training forecasting performance than the parameter count. Moreover, despite the higher computational cost, finetuned FMs demonstrate competitive performance compared to base models trained from scratch.

commercial buildings↗

Site heterogeneity and broad surface-binding isotherms in modern catalysis: Building intuition beyond the Sabatier principle

Learning the science of heterogeneous catalysis and electrocatalysis always starts with the simple case of a flat, uniform surface with an ideal adsorbate. It has of course been recognized for a century that real catalysts are more complicated. For the increasingly complex catalysts of the 21st century, this Perspective argues that surface heterogeneity and non-ideal binding isotherms are central features, and their implications need to be incorporated in current thinking. A variety of systems are described herein where catalyst complexity leads to broad, non-Langmuirian surface isotherms for the binding of hydrogen atoms – and this occurs even for ideal, flat Pt(111) surfaces. Modern catalysis employs nanoscale materials whose surfaces have substantial step, edge, corner, impurity, and other defect sites, and they increasingly have both metallic and non-metallic elements M n X m , including metal oxides, chalcogenides, pnictides, carbides, doped carbons, etc. The surfaces of such catalysts are often not crystal facets of the bulk phase underneath, and they typically have a variety of potential active sites. Catalytic surfaces in operando are often non-stoichiometric, amorphous, dynamic, and impure, and often vary from one part of the surface to another. Understanding of the issues that arise at such nanoscale, multi-element catalysts is just beginning to emerge. Yet these catalysts are widely discussed using Brønsted/Bell-Evans-Polanyi (BEP) relations, volcano plots, Tafel slopes, the Butler-Volmer equation, and other linear free energy relations (LFERs), which all depend on the implicit assumption that the active sites are “similar” and that surface adsorption is close to ideal. These assumptions underly the ubiquitous intuition based on the Sabatier Principle, that the fastest catalysis will occur when key intermediates have free energies of adsorption that are not too strong nor too weak. Current catalysis research often aims to minimize the complexity of non-ideal isotherms through experimental and computational design (e.g., the use of single crystal surfaces), and these studies are the foundation of the field. In contrast, this Perspective argues that the heterogeneity of binding sites and binding energies is an inherent strength of these catalysts. Here, this diversity makes many nanoscale catalysts inherently a high-throughput screen wrapped in a tiny package. Only by making the heterogeneity part of the foundation of catalysis models, sorting the types of active sites and dissecting non-ideal binding isotherms, will modern catalysis learn to harness the inherent diversity of real catalysts. Controlling and exploiting diversity rather than avoiding it will help to optimize complex modern catalysts and catalytic conditions.

Mayer, James M.↗

Shaping the Future of Self-Driving Autonomous Laboratories Workshop

The "Shaping the Future of Self-Driving Autonomous Laboratories" workshop, held in Denver on November 7-8, 2024, brought together leading experts from materials science and computing to address the growing need to revolutionize scientific research through AI-driven autonomous laboratories. The workshop identified critical challenges, including the integration of heterogeneous data, development of AI systems that understand fundamental physical principles, and comprehensive safety protocols. Key recommendations emerged around developing universal laboratory equipment interfaces, implementing automated metadata collection systems, and creating hybrid AI approaches that combine data-driven learning with scientific principles. The workshop emphasized maintaining human oversight while leveraging automation, transforming scientific education to prepare the next generation of researchers, and establishing a national consortium leveraging DOE facilities as anchors for broader collaboration with academia and industry. Participants stressed the urgency of addressing the growing disconnect between human decision-making timescales and modern instrumentation capabilities, highlighting the need for strategic automation while preserving essential human insight and oversight in the research process.

36 MATERIALS SCIENCE↗

Privacy-Preserving Federated Learning for Science: Challenges and Research Directions

This paper discusses the key challenges and future research directions for privacy-preserving federated learning (PPFL), with a focus on its application to large-scale scientific AI models, in particular, foundation models~(FMs). PPFL enables collaborative model training across distributed datasets while preserving privacy-- an important collaborative approach for science. We discuss the need for efficient and scalable algorithms to address the increasing complexity of FMs, particularly when dealing with heterogeneous clients. In addition, we underscore the need for developing advance privacy-preserving techniques, such as differential privacy, to balance privacy and utility in large FMs emphasizing fairness and incentive mechanisms to ensure equitable participation among heterogeneous clients. Finally, we emphasize the need for a robust software stack supporting scalable and secure PPFL deployments across multiple high-performance computing facilities. We envision that PPFL would play a crucial role to advance scientific discovery and enable large-scale, privacy-aware collaborations across science domains.

Kim, Kibaek [Argonne National Laboratory (ANL)]↗