Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “concurrent computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Physics-inspired spatiotemporal-graph AI ensemble for the detection of higher order wave mode signals of spinning binary black hole mergers

We present a new class of AI models for the detection of quasi-circular, spinning, non-precessing binary black hole mergers whose waveforms include the higher order gravitational wave modes ($\ell$, |m|) = {(2,2), (2,1), (3,3), (3,2), (4,4)}, and mode mixing effects in the $\ell$ = 3, |m| = 2 harmonics. These AI models combine hybrid dilated convolution neural networks to accurately model both short- and long-range temporal sequential information of gravitational waves; and graph neural networks to capture spatial correlations among gravitational wave observatories to consistently describe and identify the presence of a signal in a three detector network encompassing the Advanced LIGO and Virgo detectors. We first trained these spatiotemporal-graph AI models using synthetic noise, using 1.2 million modeled waveforms to densely sample this signal manifold, within 1.7 h using 256 NVIDIA A100 GPUs in the Polaris supercomputer at the Argonne Leadership Computing Facility. This distributed training approach exhibited optimal classification performance, and strong scaling up to 512 NVIDIA A100 GPUs. With these AI ensembles we processed data from a three detector network, and found that an ensemble of 4 AI models achieves state-of-the-art performance for signal detection, and reports two misclassifications for every decade of searched data. We distributed AI inference over 128 GPUs in the Polaris supercomputer and 128 nodes in the Theta supercomputer, and completed the processing of a decade of gravitational wave data from a three detector network within 3.5 h. Finally, we fine-tuned these AI ensembles to process the entire month of February 2020, which is part of the O3b LIGO/Virgo observation run, and found 6 gravitational waves, concurrently identified in Advanced LIGO and Advanced Virgo data, and zero false positives. This analysis was completed in one hour using one NVIDIA A100 GPU.

79 ASTRONOMY AND ASTROPHYSICS↗

FitCache: A Transparent Drop-In Framework for Multi-Tier Caching to Accelerate Distributed Deep Learning Workloads

Training in Deep learning (DL) remains highly compute- and data-intensive, with I/O becoming a critical bottleneck as models and datasets scale. Recent studies report that data loading can dominate training time, especially on large-scale HPC systems with shared parallel file systems (PFS). Existing caching approaches either rely on single-tier designs or require intrusive modifications to training pipelines, limiting their portability and effectiveness. In this work, we present FitCache, a transparent drop-in framework for multi-tier caching to accelerate distributed DL training by coordinating fast local memory (e.g., DRAM, Persistent Memory (PMem)) and NVMe as hierarchical caches atop PFS. Our design adapts to hardware diversity, i.e., if NVMe is missing, memory transparently acts as a caching tier, ensuring stable performance. FitCache transparently intercepts I/O requests and issues concurrent fetches across all tiers, returning data from the fastest responder without centralized metadata or static redirection paths. FitCache adapts to dynamic workloads and heterogeneous clusters while maintaining POSIX compatibility. Experiments on Frontier (2048 GPUs) and smaller research clusters show that FitCache reduces training time by up to 40% and per-batch I/O latency by up to 71.6% compared to Lustre Orion PFS, offering a drop-in solution for scalable DL training.

Hu, Guangxing [ORNL] (ORCID:0009000283203614)↗

Concurrent Shape and Topology Optimization

The typical topology optimization workflow uses a design domain that does not change during the optimization process. Consequently, features of the design domain, such as the location of loads and constraints, must be determined in advance and are not optimizable. A method is proposed herein that allows the design domain to be optimized along with the topology. This approach uses topology and shape derivatives to guide nested optimizers to the optimal topology and design domain. The details of the method are discussed, and examples are provided that demonstrate the utility of this approach.

42 ENGINEERING↗

Scalable balanced training of conditional generative adversarial neural networks on image data

Here, we propose a distributed approach to train deep convolutional generative adversarial neural network (DC-CGANs) models. Our method reduces the imbalance between generator and discriminator by partitioning the training data according to data labels, and enhances scalability by performing a parallel training where multiple generators are concurrently trained, each one of them focusing on a single data label. Performance is assessed in terms of inception score, Fréchet inception distance, and image quality on MNIST, CIFAR10, CIFAR100, and ImageNet1k datasets, showing a significant improvement in comparison to state-of-the-art techniques to training DC-CGANs. Weak scaling is attained on all the four datasets using up to 1000 processes and 2000 NVIDIA V100 GPUs on the OLCF supercomputer Summit.

97 MATHEMATICS AND COMPUTING↗

Performance assessment of ensembles of in situ workflows under resource constraints

Summary Scientific breakthroughs in biomolecular methods and improvements in hardware technology have shifted from a long‐running simulation to a large set of shorter simulations running simultaneously, called an ensemble. In an ensemble, simulations are usually coupled with analyses of data produced by the simulations. In situ methods can be used to analyze large volumes of data generated by scientific simulations at runtime (i.e., simulations and analyses are performed concurrently). In this work, we study the execution of ensemble‐based simulations paired with in situ analyses using in‐memory staging methods. Using an ensemble of molecular dynamics in situ workflows with multiple simulations and analyses, we first show that collecting traditional metrics such as makespan, instructions per cycle, memory usage, or cache miss ratio is not sufficient to characterize complex behaviors of ensembles. We propose a method to evaluate the performance of ensembles of workflows that captures multiple resource usage aspects: resource efficiency, resource allocation, and resource provisioning. Experimental results demonstrate that the proposed method can effectively distinguish the performance of different component placements in an ensemble with up to 32 ensemble members. By evaluating different co‐location scenarios, our proposed performance indicators demonstrate benefits of co‐locating simulation and coupled analyses within a compute node.

Do, Tu Mai Anh↗

Photoinduced Electron and Energy Transfer Pathways and Photocatalytic Mechanisms in Hybrid Plasmonic Photocatalysis

Hybrid plasmonic nanostructures are built on plasmonic metalnanostructures surrounded by catalytic metals or metal oxides. We report that recent studies have shown that hybrid plasmonic nanocatalysts can concurrently utilize thermal energy and photon stimuli and exhibit high catalytic activity, selectivity, and stability that are not attainable in conventional purely thermally activated catalytic processes. The hybrid plasmonic photocatalytic approach has recently emerged as an attractive concept for the conversion of solar energy into chemical energy, the distributed synthesis of valuable chemicals such as ammonia with little to no requirement of external heating, and the development of coke-resistant and selective catalytic processes. The field of hybrid plasmonic photocatalysis has grown tremendously in the last decade. In this review article, the advantages of visible-light-augmented hybrid plasmonic photocatalysis over conventional pure thermally activated heterogeneous catalysis are discussed. Fundamental insights are provided into photocatalytic mechanisms by which the photoexcited charge carriers (electrons and holes) are formed and transferred to adsorbates triggering chemical transformations on the surface of hybrid plasmonic nanocatalysts. Computational modeling used for predicting and understanding the photocatalytic activity and selectivity on hybrid plasmonic nanostructures is also reviewed. The review closes with a discussion of the current challenges, new opportunities, and future outlook for hybrid plasmonic photocatalysis.

36 MATERIALS SCIENCE↗

Predicting temperature-dependent ultimate strengths of body-centered-cubic (BCC) high-entropy alloys

This paper presents a bilinear log model, for predicting temperature-dependent ultimate strength of high-entropy alloys (HEAs) based on 21 HEA compositions. We consider the break temperature, T break , introduced in the model, an important parameter for design of materials with attractive high-temperature properties, one warranting inclusion in alloy specifications. For reliable operation, the operating temperature of alloys may need to stay below T break . We introduce a technique of global optimization, one enabling concurrent optimization of model parameters over low-temperature and high-temperature regimes. Furthermore, we suggest a general framework for joint optimization of alloy properties, capable of accounting for physics-based dependencies, and show how a special case can be formulated to address the identification of HEAs offering attractive ultimate strength. We advocate for the selection of an optimization technique suitable for the problem at hand and the data available, and for properly accounting for the underlying sources of variations.

36 MATERIALS SCIENCE↗

Three-dimensional lattice modulations in the charge density wave system Lu 2 Ir 3 Si 5

Using total and resonant x-ray scattering coupled to large-scale computer modeling, we study the lattice modulations in the complex charge density wave (CDW) material Lu 2 ⁢Ir 3 ⁢Si 5 . Here, we find that it is a unique quantum system where periodic lattice modulations related to emergent CDW order occur in three orthogonal atomic planes of the crystal lattice, leading to the emergence of an unusual three-dimensional (3D) pattern of short and long Ir-Ir and Lu-Lu bonds. The 3D character of observed lattice modulations explains the largely isotropic character of the changes in the electronic properties occurring when the CDW order sets in, demonstrating the strong electron-lattice coupling in Lu 2 ⁢Ir 3 ⁢Si 5 . The result is supported by DFT calculations based on the experimental structure data. Altogether, our work provides strong evidence for the presence of a relationship between the dimensionality of emergent lattice distortions and that of concurrent changes in the electronic properties of CDW materials. The relationship may need to be accounted for when these materials are explored for practical applications.

36 MATERIALS SCIENCE↗

A Full-Stack Exploration of Language-Based Parallelism in Fortran 2023

This poster explores native parallel features in Fortran 2023 through the lens of supporting applications with libraries, compilers, and parallel runtimes. The language revision informally named Fortran 2008 introduced parallelism in the form of Single Program Multiple Data (SPMD) execution with two broad feature sets: (1) loop-level parallelism via do concurrent and (2) a Partitioned Global Address Space (PGAS) comprised of distributed “coarray” data structures. Fortran’s native parallelism has demonstrated high performance [1] and reduced the burden of inserting what sometimes amounts to more directives than code. Several compilers support both feature sets, typically by translating do concurrent into serial do loops annotated by parallel directives and by translating SPMD/PGAS features into direct calls to a communication library. Our research focuses primarily on two questions: (1) can the compiler’s parallel runtime library be developed in the language being compiled (Fortran) and (2) can we define an interface to the runtime that liberates compilers from being hardwired to one runtime and vice versa. We are answering these questions by developing the Parallel Runtime Interface for Fortran (PRIF) [2] and the Co-Array Fortran Framework of Efficient Interfaces to Network Environments (Caffeine) [3]. Caffeine is initially targeting adoption by LLVM Flang, a new open-source Fortran compiler developed by a broad community in industry, academia, and government labs. We are also exploring the use of these features in Inference-Engine, a deep learning library designed to facilitate neural network training and inference for high-performance computing applications written in modern Fortran.

Rasmussen, Katherine↗

IRIS-GNN: Leveraging Graph Neural Networks for Scheduling on Truly Heterogeneous Runtime Systems

The diversity of accelerators in computer systems poses significant challenges for software developers, such as managing vendor-specific compiler toolchains, code fragmentation requiring different kernel implementations, and performance portability issues. To address these, the Intelligent Runtime System (IRIS) was developed. IRIS works across various systems, from smartphones to supercomputers, enabling automatic performance scaling based on available accelerators. It introduces abstract tasks for seamless execution transitions between accelerators while ensuring memory consistency and task dependencies. Although IRIS simplifies system details, optimal dynamic scheduling still requires user input to understand workload structures. To address this, we introduce a new scheduling policy for IRIS, termed IRIS-GNN, which is the first IRIS hybrid policy that operates in conjunction with the dynamic policies. This policy employs a Graph-Neural Network (GNN) to conduct Graph Classification of any task graphs submitted to IRIS. This GNN analyzes the structure and attributes of the task graph, categorizing it as either locality, concurrency, or mixed. This classification subsequently guides the selection of the dynamic policy used by IRIS. We provide a comparison of the performance of IRIS-GNN against the complete spectrum of IRIS’s dynamic policies, assess the overhead introduced by the GNN within this scheduling framework, and ultimately explore its practical application in real-world scenarios.

Johnston, Beau↗

Bayesian Entropy Neural Networks for physics-aware prediction

This article addresses the need for deep learning models to integrate well-defined constraints into their outputs, driven by their application in surrogate models, learning with limited data and partial information, and scenarios requiring flexible model behavior to incorporate non-data sample information. We introduce Bayesian Entropy Neural Networks (BENN), a framework grounded in Maximum Entropy (MaxEnt) principles, designed to impose constraints on Bayesian Neural Network (BNN) predictions. BENN is capable of constraining not only the predicted values but also their derivatives and variances, ensuring a more robust and reliable model output. To achieve simultaneous uncertainty quantification and constraint satisfaction, we employ the method of multipliers approach. This allows for the concurrent estimation of neural network parameters and the Lagrangian multipliers associated with the constraints. Our experiments, spanning diverse applications such as beam deflection modeling and microstructure generation, demonstrate the effectiveness of BENN. The results highlight significant improvements over traditional BNNs and showcase competitive performance relative to contemporary constrained deep learning methods.

14 SOLAR ENERGY↗

Models and Strategies for Optimal Demand Side Management in the Chemical Industries

Deregulation and the increase of renewable electricity generation from wind and solar photovoltaics have transformed the U.S. electricity market. Economic and environmental benefits notwithstanding, the presence of renewables has increased variability and uncertainty on the supply side of the grid. Managing demand, rather than generation – a strategy referred to as “demand response (DR)” – is an attractive approach for mitigating this imbalance. DR efforts aim to reduce electricity usage during peak demand times, lessening stress on the grid. Industrial users are particularly attractive entities for DR participation since they present large, localized loads that can provide significant relief on grid demand and –unlike other large loads, such as buildings – are minimally dependent on human needs and preferences. In this project, we accomplished three main objectives. (1) We developed data-driven low-order DR scheduling-relevant dynamic models of chemical processes. Concurrently, we studied the formulation and solution of the associated optimal DR production scheduling problems. (a) A prototype air separation unit (ASU) model was used to generate simulated operating data for initial modeling efforts, which enabled the later use of industrial data for data-driven modeling. (b) We utilized Hammerstein-Wiener (HW) and Finite Step Response (FSR) models to represent nonlinear plant dynamics. (c) The HW models were linearized using exact linearization so they could potentially be embedded in power system models, which are formulated as mixed integer linear programs (MILPs). (d) We solved DR optimization problems under uncertainty and found that even naïve predictions of electricity price and product demand led to significant cost savings benefits. (2) Our DR scheduling optimization problem formulations are amenable to real-time solution. (a) We utilized Lagrangian Relaxation (LR) to efficiently solve the optimization problem by decoupling subproblems linked by complicating constraints. (b) We have achieved computation times for the 3-day DR scheduling problem of an ASU as low as 1.88 minutes. (3) Our representations of the DR behavior of chemical process as grid-level batteries were embedded in power system models. (a) For a small-scale grid, we found that incorporating the dynamics of the chemical plant in the optimal power flow calculations resulted in better resource management leading to up to 15% and 46% cost reduction for the grid and chemical plant operations, respectively, during periods of power line congestion. We have published several works dedicated to modeling and solving DR optimization problems from the user side. These were published in top peer-reviewed journals and are summarized in this report. The most recent work (and papers in preparation) considers DR scheduling from the grid side. Future efforts will consider networked plants (e.g., air separation units operating on a common pipeline) for DR participation, which is expected to amplify the capabilities of industrial DR participants to perform load-shifting. Our consideration of uncertainty in DR has inspired future directions in this area as well: we plan to develop multistage methods to fully account for the effects of uncertainty in DR scheduling.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Predicting resistive wall mode stability in NSTX through balanced random forests and counterfactual explanations

Abstract Recent progress in the disruption event characterization and forecasting framework has shown that machine learning guided by physics theory can be easily implemented as a supporting tool for fast computations of ideal stability properties of spherical tokamak plasmas. In order to extend that idea, a customized random forest (RF) classifier that takes into account imbalances in the training data is hereby employed to predict resistive wall mode (RWM) stability for a set of high beta discharges from the NSTX spherical tokamak. More specifically, with this approach each tree in the forest is trained on samples that are balanced via a user-defined over/under-sampler. The proposed approach outperforms classical cost-sensitive methods for the problem at hand, in particular when used in conjunction with a random under-sampler, while also resulting in a threefold reduction in the training time. In order to further understand the model’s decisions, a diverse set of counterfactual explanations based on determinantal point processes (DPP) is generated and evaluated. Via the use of DPP, the underlying RF model infers that the presence of hypothetical magnetohydrodynamic activity would have prevented the RWM from concurrently going unstable, which is a counterfactual that is indeed expected by prior physics knowledge. Given that this result emerges from the data-driven RF classifier and the use of counterfactuals without hand-crafted embedding of prior physics intuition, it motivates the usage of counterfactuals to simulate real-time control by generating the β N levels that would have kept the RWM stable for a set of unstable discharges.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Custom surface reflectance, shade mask, and equivalent water thickness maps for the Colorado Headwaters Ecological Spectroscopy Study (2025)

This dataset contains land surface reflectance estimates and additional derived products generated from NEON Imaging Spectrometer (NIS) data collected in the Upper Gunnison river basin during June and July of 2025. Data was collected over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). These products were derived from radiance and LiDAR data collected by the NEON Airborne Observation Platform (AOP) campaign funded by the Colorado Headwaters Ecological Spectroscopy Study (CHESS) (doi:10.15485/3017965). Products include per-pixel surface reflectance (rfl) and reflectance uncertainty (rfl_unc), observational data (obs), canopy equivalent water thickness (ewt), and shade masks. Atmospheric correction was performed per flightline using the ISOFIT (Imaging Spectrometer Optimal FITting) optimal estimation framework to estimate surface reflectance and the associated per-band reflectance uncertainty. Reflectance retrievals achieved a mean absolute error of 1.5% across diverse validation surfaces (see validation report.pdf). Equivalent water thickness was calculated from surface reflectance using the Beer–Lambert absorption of liquid water. Shade masks were generated based on the geometry between the sun angle, ground surface, and sensor at the time of flight. Data products are provided per-flightline and as mosaics for each domain. Flightline data products are provided as ENVI-formatted binary files (rfl, rfl_unc, ewt) and GeoTIFFs (shade). Reflectance and uncertainty mosaics are provided as tiled NetCDFs, while all other mosaicked products are provided as cloud-optimized GeoTIFFs. These formats are supported by common geospatial software (e.g., QGIS, ArcGIS, ENVI) and programmatic libraries in Python (e.g., rasterio, xarray, spectral, netCDF4) and R (e.g., terra, ncdf4). Processing workflows were designed to be equivalent to those used to generate the 2018 CHESS campaign airborne imaging spectroscopy data products (doi:10.15485/3013527). All outputs were co-registered to a common spatial grid to support time series analyses. CHESS Project Description: The Colorado Headwaters Ecological Spectroscopy Study (CHESS) comprised a multi-week airborne remote sensing and field observation campaign in the Upper Gunnison Basin, Colorado, conducted in June and July of 2025. Airborne remote sensing was conducted by the National Ecological Observatory Network Airborne Observation Platform (NEON AOP), concurrent with a field campaign run by the Rocky Mountain Biological Laboratory (RMBL), the Lawrence Berkeley National Laboratory (LBNL) and SLAC National Accelerator Laboratory Watershed Function Science Focus Area (SFA), and NASA-JPL (Jet Propulsion Laboratory) Earth Surface Mineral Dust Source Investigation (EMIT) program. Between June 10 and July 18, 2025, the NEON AOP flight team collected high-resolution aerial imaging spectroscopy and Light Detection and Ranging (LiDAR) data over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). In coordination with the flights, a field campaign acquired ground-truth observations, including observations of vegetation composition, foliar traits, forest demography, and subsurface properties in 18 core sampling areas within the domains. Additional surface water observations were taken at over 380 point locations. All CHESS campaign datasets can be found within the CHESS ESS-DIVE data portal: https://data.ess-dive.lbl.gov/portals/chess. Funding Acknowledgment: Data acquisition was performed under a grant from the National Aeronautics and Space Administration (80NSSC24K1005). Computational research was carried out at the Jet Propulsion Laboratory, California Institute of Technology, under a contract with the National Aeronautics and Space Administration (80NM0018D0004) and was funded by EMIT Extended Mission Phase E Science.

2018 NEON and 2025 CHESS Campaigns↗

Sodium Ion Expansion Power Block for Distributed CSP

The Sodium Ion Expansion Power Block for Distributed CSP was a three-plus-one-year effort under the Concentrating Solar Power: Advanced Projects Offering Low LCOE Opportunities (CSP: APOLLO) funding program within the U.S. Department of Energy Solar Energy Technologies Office. The primary objective of this project is to develop a dual-stage modular sodium thermal electrochemical converter (Na-TEC) heat engine power block, which can be potentially integrated with either a small-scale dish solar or large-scale heliostats and parabolic trough CSP. Na-TEC is a heat engine that generates electricity through the isothermal expansion of sodium ions. The Na-TEC is a closed system that can theoretically achieve conversion efficiencies above 45% when operating between thermal reservoirs at 1150 K and 550 K. However, thermal designs have confined previous single-stage devices to thermal efficiencies below 20%. To mitigate some of these limitations, we consider dividing the isothermal expansion into two stages; one at the evaporator temperature (1150 K) and another at an intermediate temperature (650 K –1050 K). This dual-stage Na-TEC takes advantage of regeneration and reheating, and could be amenable to better thermal management. In light of this, we first designed and developed a thermo-electrochemical model, and thermodynamically demonstrated how the dual-stage device can improve the efficiency by up to 8% points over the best performing single-stage device. We also established an application regime map for the single- and dual-stage Na-TEC in terms of the power density and the total thermal parasitic loss. Moreover, a thermal design of an axisymmetric dual-stage Na-TEC is developed to guide the scale-up and fabrication of sub-components of prototype module. A reduced-order finite-element model is used in conjunction with a Na-TEC thermodynamic model that was developed to determine the total parasitic heat loss of this dual-stage design. A number of simplifications are applied in the reduced-order model to decrease the computational time while maintaining acceptable accuracy. According to this analysis, a maximum efficiency of 29% and a maximum power output of 125 W can be achieved. Ultimately, we were able to demonstrate thermal efficiency improvements of the Na-TEC heat engine from 19% up to 40.3%, in a dual-stage (non-optimized) prototype module that we designed, fabricated, and tested with high temperature stage at 923 K. Furthermore, a cost-performance analysis for this improved dual-stage design was carried out for distributed-CSP systems. A high-level techno-economic analysis (TEA) explores four scenarios where a Na-TEC is used as the heat engine for a distributed-CSP system. Overnight capital cost and levelized cost of electricity (LCOE) are estimated for a system lifetime of 30 years, revealing that overnight capital costs in a range from $3.57 to $17.71 per We are feasible, which equate to LCOEs from 6.9 to 17.2 cents/kWh e -1 . This analysis makes a significant contribution by concurrently quantifying the efficiency and unit costs for a range of multistage configurations, and demonstrating that a Na-TEC may be a promising alternative to Stirling engines for distributed-CSP systems at residential scale of 1–5 kW e .

14 SOLAR ENERGY↗

Reducing Communication in Graph Neural Network Training

Graph Neural Networks (GNNs) are powerful and flexible neural networks that use the naturally sparse connectivity information of the data. GNNs represent this connectivity as sparse matrices, which have lower arithmetic intensity and thus higher communication costs compared to dense matrices, making GNNs harder to scale to high concurrencies than convolutional or fully-connected neural networks. Here, we introduce a family of parallel algorithms for training GNNs and show that they can asymptotically reduce communication compared to previous parallel GNN training methods. We implement these algorithms, which are based on 1D, 1. 5D, 2D, and 3D sparse-dense matrix multiplication, using torch.distributed on GPU-equipped clusters. Our algorithms optimize communication across the full GNN training pipeline. We train GNNs on over a hundred GPUs on multiple datasets, including a protein network with over a billion edges.

97 MATHEMATICS AND COMPUTING↗

Code modernization strategies for short-range non-bonded molecular dynamics simulations

Modern HPC systems are increasingly relying on greater core counts and wider vector registers. Thus, applications need to be adapted to fully utilize these hardware capabilities. One class of applications that can benefit from this increase in parallelism are molecular dynamics simulations. In this paper, we describe our efforts at modernizing the ESPResSo++ simulation package for molecular dynamics by restructuring its particle data layout for efficient memory accesses and applying vectorization techniques to benefit the calculation of short-range non-bonded forces, which results in an overall three times speedup and serves as a baseline for further optimizations. We also implement fine-grained parallelism for multi-core CPUs through HPX, a C++ runtime system which uses lightweight threads and an asynchronous many-task approach to maximize concurrency. Our goal is to evaluate the performance of an HPX-based approach compared to the bulk-synchronous MPI-based implementation. This requires the introduction of an additional layer to the domain decomposition scheme that defines the task granularity. On spatially inhomogeneous systems, which impose a corresponding load-imbalance in traditional MPI-based approaches, we demonstrate that by choosing an optimal task size, the efficient work-stealing mechanisms of HPX can overcome the overhead of communication resulting in an overall 1.4 times speedup compared to the baseline MPI version.

97 MATHEMATICS AND COMPUTING↗

Evaluation of Apex Alpha with LabWare-LIMS Data Reduction for Alpha Analyses

Savannah River Site's Environmental Bioassay Laboratory is migrating its laboratory information management system (LIMS) from SQL-LIMS to Oracle Labware-LIMS (LW-LIMS) systems to align with the Department of Energy's cyber security policies. Concurrently with the LIMS upgrade, the current VMS based Alpha Measurement System (AMS) software is being replaced with Apex Alpha software to reduce other cyber vulnerabilities, streamline procedural production aspects of data handling, and improve outdated reduction processes to align the program with requirements outlined in ANSI N13.30. The primary technical changes which are being implemented are: - Reagent contributions are added in data reduction, - Moving average background equivalent activity are used for gross signal corrections, - Measured uncertainty of both are included in the decision level and propagated uncertainty, - Tracer levels are increased for precision and accuracy improvements, and - Reports are modified for compliance.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗