Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “concurrent computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Concurrent Cooling Effects of Dynamic Line Ratings on Wind Plant Gen-Tie Lines

This report was prepared for the Wind Energy Technology Office for the FY 2022, quarter 2 deliverable. This details the use of dynamic line rating technology to rate a series of gen tie lines connecting wind plants to the regional transmission lines. This consists of two primary study regions, the first on the desert west of Idaho Falls, and the second in the region east of the Cascades along the Columbia River Gorge. The TREAD program that was developed at INL was used to create generated gen-tie lines based on nearby regional transmission line connections. The capacity availability of the gen tie lines are compared with the power production of the wind farms. Overall, the INL Site location shows a much greater capacity for the gen-tie lines due to the higher wind speeds. Across both locations, the HRRR data shows higher wind speeds than observed at the observational weather stations. For both the Columbia Gorge and Idaho areas, the sites show that a statically rate gen-tie could carry additional capacity far above the rated during periods of high wind due to the concurrent cooling effects.

17 WIND ENERGY↗

AMRIC: A Novel In Situ Lossy Compression Framework for Efficient I/O in Adaptive Mesh Refinement Applications

As supercomputers advance towards exascale capabilities, computational intensity increases significantly, and the volume of data requiring storage and transmission experiences exponential growth. Adaptive Mesh Refinement (AMR) has emerged as an effective solution to address these two challenges. Concurrently, error-bounded lossy compression is recognized as one of the most efficient approaches to tackle the latter issue. Despite their respective advantages, few attempts have been made to investigate how AMR and error-bounded lossy compression can function together. To this end, this study presents a novel in-situ lossy compression framework that employs the HDF5 filter to improve both I/O costs and boost compression quality for AMR applications. We implement our solution into the AMReX framework and evaluate on two real-world AMR applications, Nyx and WarpX, on the Summit supercomputer. Experiments with 512 cores demonstrate that AMRIC improves the compression ratio by 81X and the I/O performance by 39X over AMReX's original compression solution.

Wang, Daoce↗

Bringing heterogeneity to the CMS software framework

The advent of computing resources with co-processors, for example Graphics Processing Units (GPU) or Field-Programmable Gate Arrays (FPGA), for use cases like the CMS High-Level Trigger (HLT) or data processing at leadership-class supercomputers imposes challenges for the current data processing frameworks. These challenges include developing a model for algorithms to offload their computations on the co-processors as well as keeping the traditional CPU busy doing other work. The CMS data processing framework, CMSSW, implements multithreading using the Intel Threading Building Blocks (TBB) library, that utilizes tasks as concurrent units of work. In this paper we will discuss a generic mechanism to interact effectively with non-CPU resources that has been implemented in CMSSW. In addition, configuring such a heterogeneous system is challenging. In CMSSW an application is configured with a configuration file written in the Python language. The algorithm types are part of the configuration. The challenge therefore is to unify the CPU and co-processor settings while allowing their implementations to be separate. We will explain how we solved these challenges while minimizing the necessary changes to the CMSSW framework. We will also discuss on a concrete example how algorithms would offload work to NVIDIA GPUs using directly the CUDA API.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

eCounter: Inline Per-IP Network Monitoring at Millisecond Resolution via eBPF

Scientific data acquisition (SciDAQ) systems are shifting from archive-based workflows to streaming paradigms, where real-time, fine-grained network monitoring becomes essential. While P4-enabled devices offer per-packet in-band observability, they require specialized switches and routers. Host-side tools like Prometheus exporters lack sufficient temporal granularity. To bridge this gap, we present eCounter, a lightweight, hardware-agnostic, inline telemetry agent built on extended Berkeley Packet Filter (eBPF). eCounter captures per-interface ingress and egress traffic, categorized by IP address and protocol, at millisecond to sub-millisecond resolution. In a 100 Gbps environment, it continuously exports up to 3,257 time-series bins per second with only 4% CPU utilization at a 35¿KiB/s data rate. We evaluate eCounter across diverse NIC MTU settings, hook types, CPU architectures and operating systems, and observed negligible impact on concurrent high-throughput streaming applications. Complexity analysis confirms that it can be readily scaled to distributed SciDAQ deployments.

Mei, Xinxin [Computational Sciences and Technology↗

Toward Resilient Heterogeneous Computing Workflow through Kokkos-DataSpaces Integration

With the growing number of applications designed for heterogeneous HPC devices, application programmers and users are finding it challenging to compose scalable workflows as ensembles of these applications, that are portable, performant and resilient. The Kokkos C++ library has been designed to simplify this cumbersome procedure by providing an intra-application uniform programming model and portable performance. However, assembling multiple Kokkos-enabled applications into a complex workflow is still a challenge. Although Kokkos enables a uniform programming model, the inter-application data exchange still remains a challenge from both performance and software development cost perspectives. In order to address this issue, we propose a Kokkos-DataSpaces Integration, with the goal of providing a virtual shared-space abstraction that can be accessed concurrently by all applications in an Kokkos workflow, thus extending Kokkos to support inter-application data exchange.

97 MATHEMATICS AND COMPUTING↗

Parallel computing for power system climate resiliency: Solving a large-scale stochastic capacity expansion problem with mpi-sppy

Here we propose a nodal stochastic generation and transmission expansion planning model that incorporates the output from high-resolution global climate models through load and generation availability scenarios. We implement our model in Pyomo and perform computational studies on a realistically-sized test case of the California electric grid in a high performance computing environment. We propose model reformulations and algorithm tuning to efficiently solve this large problem using a variant of the Progressive Hedging Algorithm. We utilize the parallelization capabilities and overall versatility of mpi-sppy, exploiting its hub-and-spoke architecture to concurrently obtain inner and outer bounds on an optimal expansion plan. Initial results show that instances with 360 representative days on a system with over 8,000 buses can be solved to within 5% of optimality in under 4 h of wall clock time, a first step towards solving a large-scale power system expansion planning problem across a wide range of climate-informed operational scenarios.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Fair Concurrent Training of Multiple Models in Federated Learning

Federated learning (FL) enables collaborative learning across multiple clients. In most FL work, all clients train a single learning task. However, the recent proliferation of FL applications may increasingly require multiple FL tasks to be trained simultaneously, sharing clients’ computing resources, which we call Multiple-Model Federated Learning (MMFL). Current MMFL algorithms use naïve average-based client-task allocation schemes that often lead to unfair performance when FL tasks have heterogeneous difficulty levels, as the more difficult tasks may need more client participation to train effectively. Furthermore, in the MMFL setting, we face a further challenge that some clients may prefer training specific tasks to others, and may not even be willing to train other tasks, e.g., due to high computational costs, which may exacerbate unfairness in training outcomes across tasks. We address both challenges by firstly designing FedFairMMFL, a difficulty-aware algorithm that dynamically allocates clients to tasks in each training round, based on the tasks’ current performance levels. We provide guarantees on the resulting task fairness and FedFairMMFL’s convergence rate. We then propose novel auction designs that incentivizes clients to train multiple tasks, so as to fairly distribute clients’ training efforts across the tasks, and extend our convergence guarantees to this setting. Here, we finally evaluate our algorithm with multiple sets of learning tasks on real world datasets, showing that our algorithm improves fairness by improving the final model accuracy and convergence speed of the worst performing tasks, while maintaining the average accuracy across tasks.

Federated learning↗

Data-flow parallelism for high-energy and nuclear physics computing frameworks

The processing tasks of a scientific workflow in high-energy and nuclear physics (HENP) can typically be represented as a directed acyclic graph formed according to the data flow—i.e. the data dependencies among algorithms executed as part of the workflow. With this representation, an HENP computing framework can optimally execute a workflow, exploiting the parallelism inherent among independent tasks. Despite such a natural description of a workflow, most HENP frameworks do not make use of technologies that provide concurrent execution of graph-based tasking structures. In this session, we describe Fermilab efforts to adopt a graph-based technology (specifically Intel’s oneTBB flow graph) for meeting the framework needs of its experiments, notably DUNE. After introducing the physics DUNE intends to explore, we will show that all common processing idioms supported by current HENP frameworks can naturally be supported by oneTBB’s data-flow technology, optimally leveraging the concurrent capabilities of the machine. In addition, we discuss collaborative efforts between Fermilab and the Intel oneTBB development team, who is considering improvements to the flow-graph technology to better support HENP use cases.

43 PARTICLE ACCELERATORS↗

Optimizing High Performance Markov Clustering for Pre-Exascale Architectures

HipMCL is a high-performance distributed memory implementation of the popular Markov Cluster Algorithm (MCL) and can cluster large-scale networks within hours using a few thousand CPU-equipped nodes. It relies on sparse matrix computations and heavily makes use of the sparse matrix-sparse matrix multiplication kernel (SpGEMM). The existing parallel algorithms in HipMCL are not scalable to Exascale architectures, both due to their communication costs dominating the runtime at large concurrencies and also due to their inability to take advantage of accelerators that are increasingly popular. In this work, we systematically remove scalability and performance bottlenecks of HipMCL. We enable GPUs by performing the expensive expansion phase of the MCL algorithm on GPU. Additionally, we propose a CPU-GPU joint distributed SpGEMM algorithm called pipelined Sparse SUMMA and integrate a probabilistic memory requirement estimator that is fast and accurate. Furthermore, we develop a new merging algorithm for the incremental processing of partial results produced by the GPUs, which improves the overlap efficiency and the peak memory usage. We also integrate a recent and faster algorithm for performing SpGEMM on CPUs. We validate our new algorithms and optimizations with extensive evaluations. With the enabling of the GPUs and integration of new algorithms, HipMCL is up to 12.4x faster, being able to cluster a network with 70 million proteins and 68 billion connections just under 15 minutes using 1024 nodes of ORNL's Summit supercomputer.

97 MATHEMATICS AND COMPUTING↗

Disorderly Conduct of Benzamide IV: Crystallographic and Computational Analysis of High Entropy Polymorphs of Small Molecules

Benzamide, a simple derivative of benzoic acid and a common intermediate of pharmaceutical compounds, was reported to form two polymorphs in 1832, but the single crystal structure of the more stable form was not solved until 1959. Nearly 50 years later, the second form was characterized by powder diffraction, followed shortly thereafter by characterization of a third form, a polytype of the most thermodynamically stable Form I. These two new forms, Forms II and III, are metastable. Herein, we describe a fourth polymorph, Form IV, discovered by melt crystallization concurrently with its crystallization under confinement at small length scales (<10 nm), where it is stable indefinitely. Form III exists under confinement in larger pores, and melting point data for different pore sizes corroborate the existence of Form IV below 10 nm. Form IV is highly disordered, precluding indexing of powder diffraction data other than hk0 reflections. Nonetheless, a combination of powder X-ray diffraction and computational crystal structure prediction reveals that Form IV contains a 2D motif resembling that of Form II, but with longrange order in the third dimension masked by ubiquitous stacking faults. This approach relies on distilling a large number of candidate structures to a few possible disorder models based on benzamide tetrads that organize in 2D parquet-like tiles, with organization along the third dimension, that can be modeled with various stacking fault configurations having distinct intermolecular interactions and translations in the dimension orthogonal to the tiling planes. These observations reveal a bewildering crystallographic complexity for such a simple molecule. Nonetheless, the approach described herein demonstrates that challenging structures that may be abandoned prematurely because of poor crystallinity, twinning, or disorder can be solved.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Dispatch Manager for NEML2 Constitutive Model Calculations Embedded in MOOSE

This report describes the extended capabilities of the NEML2 constitutive modeling library, including a flexible and efficient work dispatching system designed to leverage both CPU and GPU resources. This enhancement addresses one of the primary computational challenges in large-scale simulations: the ability to distribute and execute batches of material model evaluations across heterogeneous computing devices. The new dispatch system introduces a modular set of dispatcher and scheduler classes that coordinate the flow of data and execution between devices. The dispatcher is responsible for efficiently packaging work, managing device-specific memory operations, and synchronizing results. This modularity allows for extensibility, making it straightforward to integrate additional computing backends in the future. From an implementation standpoint, the dispatcher system interfaces seamlessly with NEML2's existing models. They handle device-aware tensor operations, optimize memory transfers, and support asynchronous execution when applicable. This design ensures that batches of material points can be evaluated concurrently, substantially improving throughput compared to previous single-device or serial implementations. These improvements not only enhance the raw performance of NEML2 but also improve its usability in multiscale and high-fidelity simulations, where the simultaneous evaluation of large material point batches is critical. Benchmarks included in the report demonstrate the system’s scalability, highlighting its effectiveness when leveraging modern GPU architectures.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Enabling Fortran Standard Parallelism in GAMESS for Accelerated Quantum Chemistry Calculations

The performance of Fortran 2008 DO CONCURRENT (DC) relative to OpenACC and OpenMP target offloading (OTO) with different compilers is studied for the GAMESS quantum chemistry application. Specifically, DC and OTO are used to offload the Fock build, which is a computational bottleneck in most quantum chemistry codes, to GPUs. The DC Fock build performance is studied on NVIDIA A100 and V100 accelerators and compared with the OTO versions compiled by the NVIDIA HPC, IBM XL, and Cray Fortran compilers. The results show that DC can speed up the Fock build by 3.0× compared with that of the OTO model. Finally, with similar offloading efforts, DC is a compelling programming model for offloading Fortran applications to GPUs.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Hardware-Accelerated Ray Tracing of CAD-Based Geometry for Monte Carlo Radiation Transport

Monte Carlo radiation transport (MCRT) methods have been used to simulate radiation environments for many decades by tracking individual particles through a model to accumulate statistical information. MCRT geometry is historically formed using the constructive solid geometry (CSG). Recently, significant work has been performed to support simulations using computer-aided design (CAD)-based tessellated surfaces to support highly complex geometries. Ray tracing acceleration data structures from the rendering and visualization community are applied to accelerate particle tracking in CAD-based models. Despite these efforts, CSG representations provide the superior performance in surface intersection operations during particle flight. Concurrently, pseudo Monte Carlo methods have become prevalent in rendering applications to support more realistic models for scattering media, motivating innovations that are advantageous for MCRT simulations. Finally, the authors’ work extends these innovations by employing Intel’s Embree ray tracing kernel within a geometry toolkit for Monte Carlo to improve the simulation performance using CAD-based models by factors of 1.5 to 2.

42 ENGINEERING↗

Preparing MPICH for exascale

The advent of exascale supercomputers heralds a new era of scientific discovery, yet it introduces significant architectural challenges that must be overcome for MPI applications to fully exploit its potential. Among these challenges is the adoption of heterogeneous architectures, particularly the integration of GPUs to accelerate computation. Additionally, the complexity of multithreaded programming models has also become a critical factor in achieving performance at scale. The efficient utilization of hardware acceleration for communication, provided by modern NICs, is also essential for achieving low latency and high throughput communication in such complex systems. In response to these challenges, the MPICH library, a high-performance and widely used Message Passing Interface (MPI) implementation, has undergone significant enhancements. Here, this paper presents four major contributions that prepare MPICH for the exascale transition. First, we describe a lightweight communication stack that leverages the advanced features of modern NICs to maximize hardware acceleration. Second, our work showcases a highly scalable multithreaded communication model that addresses the complexities of concurrent environments. Third, we introduce GPU-aware communication capabilities that optimize data movement in GPU-integrated systems. Finally, we present a new datatype engine aimed at accelerating the use of MPI derived datatypes on GPUs. These improvements in the MPICH library not only address the immediate needs of exascale computing architectures but also set a foundation for exploiting future innovations in high-performance computing. By embracing these new designs and approaches, MPICH-derived libraries from HPE Cray and Intel were able to achieve real exascale performance on OLCF Frontier and ALCF Aurora respectively.

Guo, Yanfei [Argonne National Laboratory (ANL), Ar↗

Drugsniffer: An Open Source Workflow for Virtually Screening Billions of Molecules for Binding Affinity to Protein Targets

The SARS-CoV2 pandemic has highlighted the importance of efficient and effective methods for identification of therapeutic drugs, and in particular has laid bare the need for methods that allow exploration of the full diversity of synthesizable small molecules. While classical high-throughput screening methods may consider up to millions of molecules, virtual screening methods hold the promise of enabling appraisal of billions of candidate molecules, thus expanding the search space while concurrently reducing costs and speeding discovery. Here, we describe a new screening pipeline, called drugsniffer, that is capable of rapidly exploring drug candidates from a library of billions of molecules, and is designed to support distributed computation on cluster and cloud resources. As an example of performance, our pipeline required ~40,000 total compute hours to screen for potential drugs targeting three SARS-CoV2 proteins among a library of ~3.7 billion candidate molecules.

59 BASIC BIOLOGICAL SCIENCES↗

Qudit Gate Decomposition Dependence for Lattice Gauge Theories

In this work, we investigate the effect of decomposition basis on primitive qudit gates on superconducting radio-frequency cavity-based quantum computers with applications to lattice gauge theory. Three approaches are tested: SNAP & Displacement gates, ECD & single-qubit rotations $R(\theta,\phi)$, and optimal pulse control. For all three decompositions, implementing the necessary sequence of rotations concurrently rather then sequentially can reduce the primitive gate run time. The number of blocks required for the faster ECD &$R_p(\theta)$ is found to scale $\mathcal{O}(d^2)$, while slower SNAP & Displacement set scales at worst $\mathcal{O}(d)$. For qudits with $d<10$, the resulting gate times for the decompositions is similar, but strongly-dependent on experimental design choices. Optimal control can outperforms both decompositions for small $d$ by a factor of 2-12 at the cost of higher classical resources. Lastly, we find that SNAP & Displacement are slightly more robust to a simplified noise model.

Kürkçüoglu, Doga Murat↗

Importance learning estimator for the site-averaged turnover frequency of a disordered solid catalyst

For disordered catalysts such as atomically dispersed “single-atom” metals on amorphous silica, the active sites inherit different properties from their quenched-disordered local environments. The observed kinetics are site-averages, typically dominated by a small fraction of highly active sites. Standard sampling methods require expensive ab initio calculations at an intractable number of sites to converge on the siteaveraged kinetics. We present a new method that efficiently estimates the site-averaged turnover frequency (TOF). The new estimator uses the same importance learning algorithm [Vandervelden et al., React. Chem. Eng. 5, 77 (2020)] that we previously used to compute the siteaveraged activation energy. We demonstrate the method by computing the site-averaged TOF for a simple disordered lattice model of an amorphous catalyst. The results show that with the importance learning algorithm, the site-averaged TOF and activation energy can now be obtained concurrently with orders of magnitude reduction in required ab initio calculations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗