Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “high performance analytics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Automated pipeline framework for processing of large-scale building energy time series data

Commercial buildings account for one third of the total electricity consumption in the United States and a significant amount of this energy is wasted. Therefore, there is a need for “virtual” energy audits, to identify energy inefficiencies and their associated savings opportunities using methods that can be non-intrusive and automated for application to large populations of buildings. Here we demonstrate virtual energy audits applied to large populations of buildings’ time-series smart-meter data using a systematic approach and a fully automated Building Energy Analytics (BEA) Pipeline that unifies, cleans, stores and analyzes building energy datasets in a non-relational data warehouse for efficient insights and results. This BEA pipeline is based on a custom compute job scheduler for a high performance computing cluster to enable parallel processing of Slurm jobs. Within the analytics pipeline, we introduced a data qualification tool that enhances data quality by fixing common errors, while also detecting abnormalities in a building’s daily operation using hierarchical clustering. We analyze the HVAC scheduling of a population of 816 buildings, using this analytics pipeline, as part of a cross-sectional study. With our approach, this sample of 816 buildings is improved in data quality and is efficiently analyzed in 34 minutes, which is 85 times faster than the time taken by a sequential processing. The analytical results for the HVAC operational hours of these buildings show that among 10 building use types, food sales buildings with 17.75 hours of daily HVAC cooling operation are decent targets for HVAC savings. Overall, this analytics pipeline enables the identification of statistically significant results from population based studies of large numbers of building energy time-series datasets with robust results. These types of BEA studies can explore numerous factors impacting building energy efficiency and virtual building energy audits. This approach enables a new generation of data-driven buildings energy analysis at scale.

36 MATERIALS SCIENCE↗

Validation of SPH code Spheral to model interacting solid bodies in a supersonic flow

Contemporary discussions of planetary defense involve analyzing the risks posed by smaller sized, 20 to 200 m diameter, asteroids which are capable of breaking up in the atmosphere and generating a blast wave. Consequence assessments for this size class of asteroids are performed through fast-running analytic or semi-analytic models which are informed by high-fidelity hydrocode simulations of asteroid entry and breakup. However, insufficient historical data necessitates validating the independent physical processes which dominate airburst events. Here, the Fluid Solid Interface Smoothed Particle Hydrodynamics solver was previously used by Pearl et al. in 2023 to model the Chelyabinsk airburst and is used here to perform a series of validation simulations. The first effort involves modeling a cylinder in a hypersonic flow and comparing the bow shock geometry to that predicted by analytic theory. The second effort involves modeling the separation of two spherical bodies in supersonic flow and validating against experimental footage. Combined, these exercises demonstrate the ability of the code to model the flight-path of interacting solid bodies in a hypersonic flow.

Airburst↗

High-performance strategies for the recent MRSF-TDDFT in GAMESS

Multiple ERI (Electron Repulsion Integral) tensor contractions (METC) with several matrices are ubiquitous in quantum chemistry. In response theories, the contraction operation, rather than ERI computations, can be the major bottleneck, as its computational demands are proportional to the multiplicatively combined contributions of the number of excited states and the kernel pre-factors. Here, this paper presents several high-performance strategies for METC. Optimal approaches involve either the data layout reformations of interim density and Fock matrices, the introduction of intermediate ERI quartet buffer, and loop-reordering optimization for a higher cache hit rate. The combined strategies remarkably improve the performance of the MRSF (mixed reference spin flip)-TDDFT (time-dependent density functional theory) by nearly 300%. The results of this study are not limited to the MRSF-TDDFT method and can be applied to other METC scenarios.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Medium-induced radiative kernel with the Improved Opacity Expansion

We calculate the fully differential medium-induced radiative spectrum at next-to-leading order (NLO) accuracy within the Improved Opacity Expansion (IOE) framework. This scheme allows us to gain analytical control of the radiative spectrum at low and high gluon frequencies simultaneously. The high frequency regime can be obtained in the standard opacity expansion framework in which the resulting power series diverges at the characteristic frequency ω c ~ q^L 2 . In the IOE, all orders in opacity are resumed systematically below ωc yielding an asymptotic series controlled by logarithmically suppressed remainders down to the thermal scale T « ω c , while matching the opacity expansion at high frequency. Furthermore, we demonstrate that the IOE at NLO accuracy reproduces the characteristic Coulomb tail of the single hard scattering contribution as well as the Gaussian distribution resulting from multiple soft momentum exchanges. Finally, we compare our analytic scheme with a recent numerical solution, that includes a full resummation of multiple scatterings, for LHC-inspired medium parameters. Furthermore, we find a very good agreement both at low and high frequencies showcasing the performance of the IOE which provides for the first time accurate analytic formulas for radiative energy loss in the relevant perturbative kinematic regimes for dense media.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Strategies for Integrating Deep Learning Surrogate Models with HPC Simulation Applications

The emerging trend of the convergence of high performance computing (HPC), machine learning/deep learning (ML/DL), and big data analytics presents a host of challenges for large-scale computing campaigns that seek best practices to interleave traditional scientific simulation-based workloads with ML/DL models. A portfolio of systematic approaches to incorporate deep learning into modeling and simulation serves a vital need when we support AI for science at a computing facility. In this paper, we evaluate several strategies for deploying deep learning surrogate models in a representative physics application on supercomputers at the Oak Ridge Leadership Computing Facility (OLCF). We discuss a set of recommended deployment architectures and implementation approaches. We analyze and evaluate these alternatives and show their performance and scalability up to 1000 GPUs on two mainstream platforms equipped with different deep learning hardware and software stacks.

Yin, Junqi↗

Transport coefficient sensitivities in a semi-analytic model for magnetized liner inertial fusion

Performance of magnetized liner inertial fusion (MagLIF) experiments is highly dependent on transport processes including magnetized heat flows and magnetic flux losses. Magnetohydrodynamic simulations used to model these experiments require a choice of model for the transport coefficients, which are the constants of proportionality relating driving terms, such as temperature gradients and currents, to the associated heat and magnetic field transport. The coefficients have been the subject of repeated recalculation using various methods throughout the years. Using a semi-analytic MagLIF model, we compare models for the transport coefficients. The choice of model modifies magnetic-flux losses caused by the Nernst thermoelectric effect and thermal conduction losses. We present simulated results from parameter scans conducted in order to compare the effects of the different models on parameters of interest in MagLIF. In some regions of parameter space, discrepancies of up to 38% are found in integrated quantities like the fusion yield. These results may serve as a guide for experimental validation of the various models, particularly as laser preheat energies and initial axial field strengths are increased on MagLIF experiments.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Microstructural Characterization of the Second High Fluence Baffle-Former Bolt Retrieved from a Westinghouse Two-loop Downflow Type PWR

As one of the pressurized water reactor (PWR) internal components, baffle-former bolts (BFBs) are subjected to significant mechanical stress and neutron irradiation from the reactor core during the plant operation. Over the long operation period, these conditions lead to potential degradation and reduced load-carrying capacity of the bolts. In support of evaluating long-term operational performance of materials used in core internal components, the Oak Ridge National Laboratory (ORNL), through the Department of Energy (DOE), Light Water Reactor Sustainability (LWRS) Program, Materials Research Pathway (MRP) has harvested two high fluence BFBs from a commercial Westinghouse two-loop downflow type PWR. The two bolts of interest, i.e. bolts # 4412 and 4416, were withdrawn from service in 2011 as part of a preventative replacement plan. No identification of cracking or potential damage was found for these bolts during their removal in 2011. However, the bolts required a lower torque for removal from the baffle structure than the original torque specified during installation. Irradiation displacement damage levels in the bolts range from 15 to 41 displacements per atom. The goal of this project is to perform detailed microstructural and mechanical property characterization of BFBs following in-service exposures. The information from these bolts will be integral to the LWRS program initiatives in evaluating end of life microstructure and properties. Furthermore, valuable data will be obtained that can be incorporated into model predictions of long-term irradiation behavior and compared to results obtained in high flux experimental reactor conditions. In this report, we present our latest study in FY22 on microstructural characterizations of the second high fluence baffle-former bolt, i.e., bolt # 4412. Analytical electron microscopy and atom probe tomography characterization were performed. The radiation-induced defects in the material add to the large wealth of knowledge for neutron-induced defects in 304/316 grades of stainless steels, specifically for radiation-induced precipitation after high fluence commercial PWR irradiation. The main findings are summarized as follows: 1) The cavity size was considerably larger in the bolt thread section than in the bolt head, with the bolt thread section having a bimodal distribution of cavities greater than ~6 nm in diameter and less than ~3 nm in diameter. The bolt head only had the small-sized cavities. In addition, there was a denuded zone of large cavities near grain boundaries in the thread section of the bolt. 2) Radiation-induced precipitation in the BFB #4412 was highly complex, with the volume fraction, size, and number density of Ni/Si and Cu-rich precipitates depending strongly on the radiation temperature/dose. In many cases, co-precipitates of adjoined clusters were found with Ni/Si-rich precipitates sandwiched between Cu-rich clusters and Mo/Cr/P-rich clusters. 3) Solute segregation out of solution was highest for most solutes in the thread section of the bolt #4412 with the exception of Cu, which experienced more separation out of solution into Curich clusters in the bolt head section. This highlights the difference in the mechanisms for precipitation of Ni/Si clusters, which have the Ni 3 Si phase composition, and precipitation of Cu-rich clusters. 4) There appear to be multiple simultaneous influences that affect the microstructural variation along the length of the bolt that overcomes the ~2X difference in irradiation dose between the bolt head and the bolt thread. The irradiation temperature, thermal/fast neutron ratio variation, potential strain gradient, and exposure to PWR coolant water that each section of the bolt sees may have more influence on the microstructural evolution than the total irradiation dose. The bolt thread and shank, with higher temperature, higher relative fast neutron flux, higher strain, and exposure to coolant but lower dose, underwent more enhanced cavity formation, precipitation, and solute segregation than the bolt head section.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Applications of analytical electron microscopy to guide the design of boron carbide

Compositional analysis of boron carbide on nanometer length scales to examine or interpret atomic mechanisms, for example, solid-state amorphization or grain-boundary segregation, is challenging. This work reviews advancements in high-resolution microanalysis to characterize multiple generations of boron carbide. First, ζ-factor microanalysis will be introduced as a powerful (scanning) transmission electron microscopy ((S)TEM) analytical framework to accurately characterize boron carbide. Three case studies involving the application of ζ-factor microanalysis will then be presented: (1) accurate stoichiometry determination of B-doped boron carbide using ζ-factor microanalysis and electron energy loss spectroscopy, (2) normalized quantification of silicon grain-boundary segregation in Si-doped boron carbide, and (3) calibration of a scanning electron microscope X-ray energy-dispersive spectroscopy (XEDS) system to measure compositional homogeneity differences of B/Si-doped arc-melted boron carbides in the as-melted and annealed conditions. Overall, the improvement and application of advanced analytical tools have helped better understand processing–microstructure–property relationships and successfully manufacture high-performance ceramics.

36 MATERIALS SCIENCE↗

Nanoporous Materials Genome Center Final Technical Report

Nanoporous materials (NPMs), including zeolites/zeotypes, metal-organic frameworks (MOFs), covalent organic frameworks, polymers with intrinsic microporosity, and molecular cages, possess enormous potential in diverse areas relevant to the DOE Office of Science Basic Energy Sciences (BES) mission and objectives. The Nanoporous Materials Genome Center (NMGC) has developed exascale-ready software, computational/theoretical chemistry methods, and data-driven science approaches that enable (i) the de-novo design of functional NPMs for chemical separation and catalysis tasks of increasing complexity, (ii) the discovery of the most promising functional NPMs from databases of synthesized and hypothetical adsorbent structures and the optimization of process conditions for specific applications, and (iii) the microscopic-level understanding of the fundamental interactions underlying the function of NPMs including hierarchical architectures, composite materials, responsive frameworks that may undergo phase transitions or post-synthetic modifications, and materials containing defects, partial disorder, or interfaces. A pivotal part of the NMGC project has been a tight collaboration between leading experimental groups for synthesis and characterization of NPMs and of computational groups that allowed for iterative feedback. The NMGC project has resulted in the publication of more than 290 research and review articles including more than 60 publications in high-impact journals and more than 15 journal covers. NMGC publications have already received more than 20,000 citations (with more than 3,000 citations per year in 2021, 2022, and 2023) and contribute to an h-index of more than 72. The NMGC award has supported collaborative research involving 28 research groups and contributed to the training of more than 40 postdocs, more than 60 graduate students, and more than 20 undergraduate students with broad expertise in data-driven science approaches, computational chemistry methods, and high-performance computing, in addition to the skills to thrive in an integrated experimental and computational research environment.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

HP-MDR: High-performance and Portable Data Refactoring and Progressive Retrieval with Advanced GPUs

Scientific applications produce vast amounts of data, posing grand challenges in the underlying data management and analytic tasks. Progressive compression is a promising way to address this problem, as it allows for on-demand data retrieval with significantly reduced data movement cost. However, most existing progressive methods are designed for CPUs, leaving a gap for them to unleash the power of today’s heterogeneous computing systems with GPUs.In this work, we propose HP-MDR, a high-performance and portable data refactoring and progressive retrieval framework for GPUs. Our contributions are four-fold: (1) We carefully optimize the bitplane encoding and lossless encoding, two key stages in progressive methods, to achieve high performance on GPUs; (2) We propose pipeline optimization and incorporate it with data refactoring and progressive retrieval workflows to further enhance the performance for large data process; (3) We leverage our framework to enable high-performance data retrieval with guaranteed error control for common Quantities of Interest; (4) We evaluate HP-MDR and compare it with state of the arts using five real-world datasets. Experimental results demonstrate that HP-MDR delivers an average 13.68 × and 6.31 × throughput in data refactoring and progressive retrieval tasks, respectively. It also leads to 11.22 × throughput for recomposing required data representations under Quantity-of-Interest error control and 6.04 × performance for the corresponding end-to-end data retrieval, when compared with state-of-the-art solutions.

Li, Yanliang [University of Oregon]↗

Higher-order Van Hove singularities in kagome topological bands

Motivated by the growing interest in band structures featuring higher-order Van Hove singularities (HOVHS), we investigate a spinless fermion kagome system characterized by nearest-neighbor (NN) and next-nearest-neighbor (NNN) hopping amplitudes. While NN hopping preserves time-reversal symmetry, NNN hopping, akin to chiral hopping on the Haldane lattice, breaks time-reversal symmetry and leads to the formation of topological bands with Chern numbers ranging from 𝐶 = ±1 to ±4. We perform analytical and numerical analysis of the energy bands near the high-symmetry points Γ, ±𝐊, and 𝐌 𝑖 (𝑖 = 1, 2, and 3), which uncover a rich and complex landscape of HOVHS, controlled by the magnitude and phase of the NNN hopping. We observe power-law divergences in the density of states (DOS), 𝜌⁡(𝜀)∼|𝜀| −𝜈 , with exponents 𝜈 = 1/2, 1/3, 1/4, which can significantly affect the anomalous Hall response at low temperatures when the Fermi level crosses the HOVHS. Additionally, the NNN hopping induces the formation of higher Chern number bands 𝐶 = ±2, ±4 in the middle of the spectrum obeying a sublattice interference whereupon electronic states are maximally localized in each of the sublattices when the momentum approaches the three high-symmetry points 𝐌 𝑖 (𝑖 = 1, 2, and 3) on the Brillouin zone boundary. Finally, this classification of HOVHS in kagome systems provides a platform to explore unconventional electronic orders induced by electronic correlations.

Chern insulators↗

ChemComp: Compiling and Computing with Chemical Reaction Networks

The exponential growth in computing demands driven by scientific computing, data analytics, and artificial intelligence is pushing conventional CMOS-based high-performance computing systems to their physical and energy efficiency limits. As we approach the era of post-exascale computing, disruptive approaches are necessary to overcome these barriers and achieve substantial gains in energy efficiency. Analog and hybrid digital-analog computing systems have emerged as promising alternatives, offering the potential for orders-of-magnitude improvements in efficiency. Among these, biochemical computing stands out as a novel paradigm capable of leveraging the natural efficiency of chemical reactions, which have shown promise in solving optimization problems by converging to steady states. By scaling up reaction networks or reaction vessel sizes, biochemical systems present an opportunity to meet the high-performance demands of modern computing tasks. Despite their promise, significant theoretical and practical challenges remain, particularly in formulating and mapping computational problems to chemical reaction networks (CRNs) and designing viable biochemical computing devices. This paper addresses these challenges by introducing new ideas to ChemComp, a compilation and emulation framework for chemical computation. This work describes the mechanisms through which solutions to ordinary differential equations (ODEs) that can be represented as CRN systems can be achieved. Furthermore, we explain the design principles of an ODE dialect implemented as a multi-level intermediate representation (MLIR) compiler extension that will be coupled with existing infrastructure. We demonstrate the potential of our framework through a case study emulating a simplified chemical reservoir computing device. This work establishes foundational tools and methodologies necessary to harness the computational power of chemistry, paving the way for the development of energy-efficient, high-performance computing systems tailored to contemporary and future computational needs.

Bohm Agostini, Nicolas↗

Rapid approach for structural design of the tower and monopile for a series of 25 MW offshore turbines

The goal of further reducing the Levelized Cost of Energy (LCOE) has driven the investigation of large-scale wind turbines. This work presents a simple, rapid and detailed approach for the structural design of the tower and monopile without a controller, but with frequency and high fidelity structural verification. The approach uses an optimization to reduce the mass of the structures while meeting strength, buckling and geometric constraints by using analytical equations. A verification of frequency constraints is performed with BModes, and ANSYS Mechanical APDL is used for high fidelity verification of stress and buckling. The approach is applied to study the design space of three 25 MW offshore wind turbines with different rotor diameters and cone angles, and to evaluate the nacelle center of mass fore-aft location effect. Results obtained show that the tower and monopile are more susceptible to changes in the rotor thrust than the overturning moment even for designs with high pre-cone angle and large distance of the nacelle center of mass from the tower axis. But it is possible to obtain structurally feasible tower and monopile designs for the three 25 MW turbines studied while not exceeding diameter and wall thickness limits. However, mass penalties can be decreased by 0.8-14%, to further reduce the cost of energy, by increasing the diameter limit which may require manufacturing technology development. The approach applied and studies serve to understand the design space of the tower and monopile for a 25 MW turbine, and provide baseline designs that can be used in the development of a controller and evaluation of a full suite of design load cases.

17 WIND ENERGY↗

Comparing the extraction performance of cyclodextrin-containing supramolecular deep eutectic solvents versus conventional deep eutectic solvents by headspace single drop microextraction

A headspace single drop microextraction (HS-SDME) method coupled with high performance liquid chromatography was developed to compare the extraction of eighteen aromatic organic pollutants from aqueous solutions using cyclodextrin-based supramolecular deep eutectic solvents (SUPRADESs) and alkylammonium halide-based conventional deep eutectic solvents (DESs). Different derivatives of beta-cyclodextrin (β-CD) were employed as hydrogen bond acceptors (HBA) in SUPRADESs and the extraction performance investigated. SUPRADES comprised of the 20 wt% native β-CD HBA provided the highest enrichment factors of analytes compared to SUPRADESs comprised of other derivatives of β-CD (random methylated β-cyclodextrin, heptakis(2,3,6-tri-O-methyl)-β-cyclodextrin, and 2-hydroxypropyl β-cyclodextrin). In addition, native β-CD and its derivatives were dissolved in the neat DESs and their effect on the extraction of analytes examined. Dissolution of 20 wt% native β-CD in the choline chloride ([Ch + ][Cl - ]):2Urea DES resulted in a significant increase in the extraction efficiencies of target analytes compared to the neat [Ch + ][Cl - ]:2Urea DES. Under optimum conditions, the extraction method required a solvent microdroplet of 6.5 μL, 1000 rpm stir rate, 30% (w/v) salt concentration, and a temperature of 40 °C. The tetrabutylammonium chloride: 2 lactic acid DES resulted in the highest enrichment factors while the [Ch + ][Cl - ]:2Urea DES had the lowest for most of the analytes among the evaluated solvents. The method provided limits of detection (LODs) down to 35 μg L -1 . Finally, the developed method was applied for the analysis of spiked tap and lake water, where relative recoveries ranging from 83.7% -119.7% and relative standard deviations lower than 19.2% were achieved.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Analysis of persistent contaminants and personal care products by dispersive liquid-liquid microextraction using hydrophobic magnetic deep eutectic solvents

Here, in this work, hydrophobic magnetic deep eutectic solvents (HMDESs) were used in the development of a simple and rapid dispersive liquid-liquid microextraction (DLLME) approach coupled to high performance liquid chromatography with UV detection (HPLC-UV) for the determination of ten organic contaminants including five polycyclic aromatic hydrocarbons, four UV filters, and a pesticide from water at trace levels. The HMDESs were prepared by mixing a hydrogen bond acceptor, metal halide salt, and hydrogen bond donor in suitable molar ratios. Two HMDESs, 2 tetraoctylammonium bromide ([N 8888 + ][Br - ]): cobalt chloride (CoCl 2 ): 4 octanoic acid (OA) and 3 trioctylphosphine oxide (TOPO): neodymium chloride (NdCl 3 ): 3 OA, offered the highest analyte extraction efficiency overall and were chosen as suitable solvents for validation of the microextraction method. Under optimized extraction conditions, the method required 30 µL of HMDES as extraction solvent, acetone (87.5 µL) as disperser solvent, a NaCl concentration of 30% (w/v), and an extraction time of 120 s at 20°C. Enrichment factors of the analytes ranged from 44.6 for 3-(4-methylbenzylindene) camphor to 66.0 for 2-ethylhexyl-4-(dimethyl)aminobenzoate. The method provided low limits of detection (LODs) ranging from 0.5 to 4.5 µg L -1 , and acceptable precision, with RSD values lower than 9.6%. Furthermore, the validated method was successfully applied for tap and lake water analysis, resulting in relative recoveries of spiked samples ranging between 94.7 and 119.2%.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

An Integrated Platform for Collaborative Data Analytics

While collaboration among data scientists is a key to organizational productivity, data analysts face significant barriers to achieving this end, including data sharing, accessing and configuring the required computational environment, and a unified method of sharing knowledge. Each of these barriers to collaboration is related to the fundamental question of knowledge management “how can organizations use knowledge more effectively?”. In this paper, we consider the problem of knowledge management in collaborative data analytics and present ShareAL, an integrated knowledge management platform, as a solution to that problem. The ShareAL platform consists of three core components: a full stack web application, a dashboard for analyzing streaming data and a High Performance Computing (HPC) cluster for performing real time analysis. Prior research has not applied knowledge management to collaborative analytics or developed a platform with the same capabilities as ShareAL. ShareAL overcomes the barriers data scientists face to collaboration by providing intuitive sharing of data and analytics via the web application, a shared computing environment via the HPC cluster and knowledge sharing and collaboration via a real time messaging application.

Oesch, T↗

Developing And Scaling an OpenFOAM Model to Study Turbulent Flow in a HFIR Coolant Channel

Improving the understanding of how computational fluid dynamics (CFD) direct numerical simulations (DNS) of flows in the High Flux Isotope Reactor (HFIR) perform when run in parallel using the high performance computing (HPC) platform Summit at the Oak Ridge Leadership Computing Facility (OLCF) is of particular importance to boost the computational tools used to support HFIR conversion to low enriched fuel (LEU). Evaluation of scaling performance was driven by the increasing importance of graphics processing unit (GPU) usage in HPC, which is becoming the standard for modern supercomputers such as Summit. The desired results are to obtain a strong positive correlation between the computational resources dedicated to a problem and the relative speed-up of the simulation in comparison to a benchmark. This capability will allow substantially improvement in HFIR flow analytical capabilities, specifically when predicting turbulence properties at high Reynolds numbers. The study leverages previous simulation results performed with code PHASTA (finite element) on HPC platforms Cori (NERSC) and Theta (ALCF) [1] with computing options provided in the computing platform OpenFOAM (finite volume) at OLCF. Transitioning from PHASTA to OpenFOAM will (1) eliminate dependence on third-party software for mesh generation and manipulation, (2) reduce resource needs by employing modern architectures, and (3) build expertise for future modeling of HFIR-specific problems like heat transfer in involute geometry, entrance effects, flow structure in channel corners, and so on—all important issues when defining the available thermal margins in the transition to LEU. CPUs and GPUs differ significantly in their architecture and utilization, as discussed in the literature [2]. The most important differences are in the approach to computations and their memory. A single GPU contains a large quantity of cores, enabling it to perform with a much higher throughput than a CPU, but execution requires a different approach. GPU codes execute instructions using the Single-Instruction Multiple-Thread (SIMT) approach in which a single instruction is used for groups of threads called warps. A warp typically consists of 32 threads which must execute the same set of instructions, although on separate threads. Alternately, a CPU has far fewer cores that are much more flexible in their operation, excelling at quickly performing more complex serial computations. This is why GPUs have greater throughput when properly utilized. The second important difference is seen when comparing their memory spaces. Limited memory allocations and CPU–GPU communications cause a significant bottleneck in GPU-accelerated programs. Further study was required to properly take advantage of GPU resources. A comprehensive analysis of code performance and the model-specific features of turbulence constitutes the core of this work. In this study, a DNS simulation of HFIR channel turbulence was performed with the finite volume CFD code OpenFOAM v2112 and CUDA v11.0 on Red Hat Enterprise Linux v8.2. The OpenFOAM installation had AMGx integrated to enable GPU acceleration and utilizes the PETSc4FOAM library. The computational resources and the problem size were scaled on CPU and CPU + GPU architectures to gain a better understanding of the performance of a DNS problem on modern computing hardware. The study aimed to analyze the scaling of the code exclusively on CPUs and then to examine the scaling of the codes with GPU acceleration enabled. Scaling studies included CPU and GPU acceleration on a mesh of varying resolution to analyze the impact of problem size relative to computational resources. In the course of preparing the GPU configuration on Summit, mainly using the AMGX solvers, difficulties were encountered stemming from constant changes resulting from extensive ongoing development activities and the changing environment. This resulted in the inability to complete the GPU portion of the work. The code was compiled and tested, but production runs to assess acceleration were not performed because the used discretional compute time allocation expired as year-end approached. The Summit HPC platform is scheduled for decommissioning in 2024, making it unattractive for future use with Nvidia-based GPUs. Therefore, the work will be moved onto NERSC machines in FY24. An application was prepared and submitted, and sufficient node-hours were awarded to continue the research in the next calendar year. This report summarizes work performed thus far, which mostly focused on CPU OpenFOAM computing.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗