Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Two level solver”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Orbital Conflict: Cutting Planes for Symmetric Integer Programs

Cutting planes have been an important factor in the impressive progress made by integer programming (IP) solvers in the past two decades. However, cutting planes have had little impact on improving performance for symmetric IPs. Rather, the main breakthroughs for solving symmetric IPs have been achieved by cleverly exploiting symmetry in the enumeration phase of branch and bound. In this work, we introduce a hierarchy of cutting planes that arise from a reinterpretation of symmetry-exploiting branching methods. There are too many inequalities in the hierarchy to be used efficiently in a direct manner. However, the lowest levels of this cutting-plane hierarchy can be implicitly exploited by enhancing the conflict graph of the integer programming instance and by generating inequalities such as clique cuts valid for the stable set relaxation of the instance. We provide computational evidence that the resulting symmetry-powered clique cuts can improve state-of-the-art symmetry-exploiting methods. Furthermore, the inequalities are then employed in a two-phase approach with high-throughput computations to solve heretofore unsolved symmetric integer programs arising from covering designs, establishing for the first time the covering radii of two binary-ternary codes.

97 MATHEMATICS AND COMPUTING↗

Radial electric field and density fluctuations measured by Doppler reflectometry during the post-pellet enhanced confinement phase in W7-X

Radial profiles of density fluctuations and the radial electric field, Er, have been measured using Doppler reflectometry during the post-pellet enhanced confinement phase achieved, under different heating power levels and magnetic configurations, during the 2018 W7-X experimental campaign. A pronounced Er-well is measured with local values as high as -40 kV m -1 in the radial range ρ ~ 0.7–0.8 during the post-pellet enhanced confinement phase. The maximum Er intensity scales with both the plasma density and electron cyclotron heating power level, following a similar trend to the plasma energy content. A good agreement is found when the experimental Er profiles are compared to simulations carried out using the neoclassical codes, the drift kinetic equation solver (DKES) and kinetic orbit-averaging solver for stellarators (KNOSOS). The density fluctuation level decreases from the plasma edge toward the plasma core and the drop is more pronounced in the post-pellet enhanced confinement phase than in reference gas-fuelled plasmas. Besides, in the post-pellet phase, the density fluctuation level is lower in the high iota magnetic configuration than in the standard one. Therefore, to determine whether this difference is related to the differences in the plasma profiles or to the stability properties of the two configurations, gyrokinetic simulations have been carried out using the codes stella and EUTERPE. The simulation results point to the plasma profile evolution after the pellet injection and the stabilization effect of the radial electric field profile as the dominant players in the stabilization of the plasma turbulence.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

The Effective Fragment Molecular Orbital Method: Achieving High Scalability and Accuracy for Large Systems

The effective fragment molecular orbital (EFMO) method has been developed to predict the total energy of a very large molecular system accurately (with respect to the underlying quantum mechanical method) and efficiently by taking advantage of the locality of strong chemical interactions and employing a two-level hierarchical parallelism. The accuracy of the EFMO method is partly attributed to the accurate and robust intermolecular interaction prediction between distant fragments, in particular, the many-body polarization and dispersion effects, which require the generation of static and dynamic polarizability tensors by solving the coupled perturbed Hartree–Fock (CPHF) and time-dependent HF (TDHF) equations, respectively. Solving the CPHF and TDHF equations is the main EFMO computational bottleneck due to the inefficient (serial) and I/O-intensive implementation of the CPHF and TDHF solvers. In this work, the efficiency and scalability of the EFMO method are significantly improved with a new CPU memory-based implementation for solving the CPHF and TDHF equations that are parallelized by either message passing interface (MPI) or hybrid MPI/OpenMP. Here, the accuracy of the EFMO method is demonstrated for both covalently bonded systems and noncovalently bound molecular clusters by systematically examining the effects of basis sets and a key distance-related cutoff parameter, R cut . R cut determines whether a fragment pair (dimer) is treated by the chosen ab initio method or calculated using the effective fragment potential (EFP) method (separated dimers). Decreasing the value of Rcut increases the number of separated (EFP) dimers, thereby decreasing the computational effort. It is demonstrated that excellent accuracy (<1 kcal/mol error per fragment) can be achieved when using a sufficiently large basis set with diffuse functions coupled with a small R cut value. With the new parallel implementation, the total EFMO wall time is substantially reduced, especially with a high number of MPI ranks. Given a sufficient workload, nearly ideal strong scaling is achieved for the CPHF and TDHF parts of the calculation. For the first time, EFMO calculations with the inclusion of long-range polarization and dispersion interactions on a hydrated mesoporous silica nanoparticle with explicit water solvent molecules (more than 15k atoms) are achieved on a massively parallel supercomputer using nearly 1000 physical nodes. In addition, EFMO calculations on the carbinolamine formation step of an amine-catalyzed aldol reaction at the nanoscale with explicit solvent effects are presented.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Advanced System Thermal Fluids Solver Development for SAM

This work summarizes a feasibility study on testing numerical algorithms that are suitable and efficient for advanced system analysis code development under the mutli-physics framework, MOOSE. The key is the implementation of a high-order one-dimensional staggered-grid finite volume method (SG-FVM), and its direct interaction with the linear/nonlinear solver, PETSc. Leveraging the existing capabilities of the SAM code, significant code coverages were established in the finite volume method code. This in turn allows for a suite of test problems with different problem sizes and levels of complexity to be used to quantify the performance improvement of the finite volume method code. As evidently shown in this study, the implemented SG-FVM demonstrated superior performance improvement against a direct finite element method implementation through MOOSE for the wide range of selected problems. On two computer systems, the speedup was observed to be significant, with at least one order of magnitude of solving time reduction. In addition, for a complex reactor model, transient simulation was performed using the finite volume method code, the results of which agree very well with the reference results from the finite element method code. Overall, this study demonstrates a successful feasibility study on the proposed numerical algorithms and software structure to support advanced system analysis tool development. In this work, short-term priority development and testing items were identified, and long-term code adoption and integration plans were made for the eventual deployment of the finite volume method in the SAM code.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Multilevel Graph Partitioning for Three-Dimensional Discrete Fracture Network Flow Simulations

We present a topology-based method for mesh-partitioning in three-dimensional discrete fracture network (DFN) simulations that takes advantage of the intrinsic multi-level nature of a DFN. DFN models are used to simulate flow and transport through low-permeability fractured media in the subsurface by explicitly representing fractures as discrete entities. The governing equations for flow and transport are numerically integrated on computational meshes generated on the interconnected fracture networks. Modern high-fidelity DFN simulations require high-performance computing on multiple processors where performance and scalability depends partially on obtaining a high-quality partition of the mesh to balance work-loads and minimize communication across all processors. The discrete structure of a DFN naturally lends itself to various graph representations, which can be thought of as coarse-scale representations of the computational mesh. Using this concept, we develop two applications of the multilevel graph partitioning algorithm to partition the mesh of a DFN. In the first, we project a partition of the graph based on the DFN topology onto the mesh of the DFN and in the second, this DFN-based projection is used as the initial condition for further partitioning refinement of the mesh. We compare the performance of these methods with standard multi-level graph partitioning using graph-based metrics (cut, imbalance, partitioning time), computational-based metrics (FLOPS, iterations, solver time), and total run time. The DFN-based and the mesh-based partitioning methods are comparable in terms of the graph-based metrics, but the time required to obtain the partition is several orders of magnitude faster using the DFN-based partitions. The computation-based metrics show comparable performance between both methods so, in combination, the DFN-based partitions are several orders of magnitude faster than the mesh-based partition. Furthermore, the method which uses the DFN-partition solution as the initial condition of the mesh partition provided cut and imbalance values that were close to the mesh-based partition but in a fraction of the time. In turn, this hybrid method outperformed both of the other methods in terms of the total run time.

58 GEOSCIENCES↗

Impact of Radial Reflector Fidelity on Neutronics and Vessel Fluence Simulations

The Consortium for Advanced Simulation of Light Water Reactors is developing the Virtual Environment for Reactor Applications (VERA), and the MPACT code, which is the primary deterministic neutron transport solver in VERA, provides sub-pin level flux and power distributions as part of full-scale cycle depletion and analysis. In such calculations, an important aspect is the radial reflector treatment. To improve the fidelity of the radial reflector treatment, MPACT was extended to approximate the modeling of the reactor’s structural components such as the core shroud, barrel, neutron pads, and vessel. This work explores several modeling configurations with varying levels of fidelity and computational burden and assesses the importance of modeling fidelity on the eigenvalue and pin power distribution. Two two-dimensional (2-D) problems were analyzed to assess the impact on eigenvalue and pin power distributions with low-fidelity, coarse square cell reflector representations: (1) a Watts Bar Nuclear Plant Unit 1 (WBN1) quarter-core slice with depletion and (2) an AP1000 quarter-core slice. In this work, the analyses showed that the effect on eigenvalue is fairly small, but the effect on pin power is more pronounced, especially locally in the assemblies closest to the periphery, where the maximum pin power difference is nearly 3.5% in the AP1000 case. Two additional 2-D problems were used to assess the comparison between the low-fidelity coarse square cell treatment and a high-fidelity geometric representation that uses subpin material specification: (1) the same WBN1 quarter-core slice and (2) a representative model of the NuScale small modular reactor (SMR), which features a solid reflector design with moderator holes. These results demonstrate that even a coarse, low-fidelity representation adequately captures the necessary simulation characteristics. Last, these capabilities were applied to the 2-D WBN1 quarter-core depletion to assess the impact on vessel fluence using VeraShift. From adjoint calculations, pins along the periphery were observed to be of highest importance for fluence calculation, so the impact of the reflector representation in MPACT could theoretically substantially affect the predicted result. However, it was observed that the change in pin powers along the periphery minimally impacts the maximum vessel fluence with a difference within the statistical uncertainty but provides terrific insight on the sensitivity of the peripheral pins.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

SNICAR-ADv4: a physically based radiative transfer model to represent the spectral albedo of glacier ice

Abstract. Accurate modeling of cryospheric surface albedo is essential for our understanding of climate change as snow and ice surfaces regulate the global radiative budget and sea-level through their albedo and mass balance. Although significant progress has been made using physical principles to represent the dynamic albedo of snow, models of glacier ice albedo tend to be heavily parameterized and not explicitly connected with physical properties that govern albedo, such as the number and size of air bubbles, specific surface area (SSA), presence of abiotic and biotic light absorbing constituents (LACs), and characteristics of any overlying snow. Here, we introduce SNICAR-ADv4, an extension of the multi-layer two-stream delta-Eddington radiative transfer model with the adding–doubling solver that has been previously applied to represent snow and sea-ice spectral albedo. SNICAR-ADv4 treats spectrally resolved Fresnel reflectance and transmittance between overlying snow and higher-density glacier ice, scattering by air bubbles of varying sizes, and numerous types of LACs. SNICAR-ADv4 simulates a wide range of clean snow and ice broadband albedo (BBA), ranging from 0.88 for (30 µm) fine-grain snow to 0.03 for bare and bubble-free ice under direct light. Our results indicate that representing ice with a density of 650 kg m−3 as snow with no refractive Fresnel layer, as done previously, generally overestimates the BBA by an average of 0.058. However, because most naturally occurring ice surfaces are roughened “white ice”, we recommend modeling a thin snow layer over bare ice simulations. We find optimal agreement with measurements by representing cryospheric media with densities less than 650 kg m−3 as snow and larger-density media as bubbly ice with a Fresnel layer. SNICAR-ADv4 also simulates the non-linear albedo impacts from LACs with changing ice SSA, with peak impact per unit mass of LACs near SSAs of 0.1–0.01 m2 kg−1. For bare, bubble-free ice, LACs actually increase the albedo. SNICAR-ADv4 represents smooth transitions between snow, firn, and ice surfaces and accurately reproduces measured spectral albedos of a variety of glacier surfaces. This work paves the way for adapting SNICAR-ADv4 to be used in land ice model components of Earth system models.

54 ENVIRONMENTAL SCIENCES↗

Accelerated impurity solver for DMFT and its diagrammatic extensions

Here, we present ComCTQMC, a GPU accelerated quantum impurity solver. It uses the continuous-time quantum Monte Carlo (CTQMC) algorithm wherein the partition function is expanded in terms of the hybridisation function (CT-HYB). ComCTQMC supports both partition and worm-space measurements, and it uses improved estimators and the reduced density matrix to improve observable measurements whenever possible. ComCTQMC efficiently measures all one and two-particle Green's functions, all static observables which commute with the local Hamiltonian, and the occupation of each impurity orbital. ComCTQMC can solve complex-valued impurities with crystal fields that are hybridized to both fermionic and bosonic baths. Most importantly, ComCTQMC utilizes graphical processing units (GPUs), if available, to dramatically accelerate the CTQMC algorithm when the Hilbert space is sufficiently large. We demonstrate acceleration by a factor of over 600 (100) in a simulation of δ-Pu at 600 K with (without) crystal fields. In easier problems, the GPU offers less impressive acceleration or even decelerates the CTQMC. Here we describe the theory, algorithms, and structure used by ComCTQMC in order to achieve this set of features and level of acceleration.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

The Random Ray Method Versus Multigroup Monte Carlo: The Method of Characteristics in OpenMC and SCONE

The Random Ray Method (TRRM) is a recently developed approach to solving neutral particle transport problems based on the Method of Characteristics. While the method previously has been implemented only in closed-source or limited-functionality codes, this work describes its implementation in two open-source Monte Carlo codes: OpenMC and SCONE. The random ray implementations required small modifications to the existing Multigroup Monte Carlo (MGMC) solvers, offering a rare venue for redundant, fine-grained, "apples-to-apples" speed and accuracy comparisons between transport methods. To this end, TRRM and MGMC solvers are evaluated against each other using each code's native capabilities on reactor eigenvalue problems with different degrees of energy discretization. On the C5G7 benchmark (featuring only seven energy groups), TRRM achieves a maximum pin power error comparable to or lower than that of MGMC for a given run time. On a problem with 69 energy groups, MGMC is found to scale more efficiently, obtaining a lower pin power error for a given run time. However, the defining difference between the two transport methods is found to be their vastly different uncertainty distributions. Specifically, TRRM is found to maintain similar levels of accuracy and uncertainty throughout the simulation domain whereas MGMC can exhibit orders-of-magnitude greater errors in areas of the problem that feature low neutron flux. For instance, TRRM provided an up to 373 times speed advantage compared with MGMC for computing the flux in low-flux regions in the moderator surrounding the C5G7 core.

42 ENGINEERING↗

Mutual Inductance Level Sensor for Use in Liquid Metals

Electromagnetic instrumentation capable of measuring the level of liquid metal has been developed at Argonne National Laboratory’s (ANL) Mechanisms Engineering Test Loop (METL). The mutual inductance level sensor (MILS) utilizes the electromagnetic coupling between two coil conductors and a surrounding liquid metal to determine the level of the liquid metal. Mineral insulated cables wrapped on a stainless steel core provide a durable sensor construction. The use of a sealed stainless steel thimble isolates the sensor from the high-temperature liquid metal, allowing for easy sensor maintenance. Modern digital electronics allow for reliable and accurate operation of the sensor. Two sensor variations have been designed and fabricated, and the first variation has seen extensive testing in a non-sodium testing stand as well as in the high-temperature sodium environment of METL. Electromagnetic finite-element analysis studies have been performed using the COMSOL Magnetic Fields Solver. Regular operation of the MILS will commence following in-situ calibration of the sensors in the METL environment.

36 MATERIALS SCIENCE↗

Application of linear prolongation to coarse mesh finite difference acceleration in CASMO5

The Coarse-Mesh Finite Difference (CMFD) method has been used for over a decade to accelerate the convergence of the Method of Characteristics (MOC) solution to the two- dimensional particle transport equation in CASMO5. Numerical testing, along with widespread use in production-level calculations, have shown that the current CMFD implementation provides stability and robustness for a wide range of realistic reactor physics problems. However, the recent development of linear prolongation has attracted attention from the community as a way to further improve the performance and stability of CMFD. Two interpolation methods for linear prolongation are presented in this work and implemented into a test version of CASMO5. The performance of the proposed interpolations, referred to as the bilinear and linear directional schemes, is evaluated in terms of runtime relative to the default constant or uniform scaling update. Numerical results indicate that the use of linear prolongation can reduce the transport solver runtime on average by approximately 10% when tested with two hundred randomly selected cases. The new directional linear interpolation, combined with default constant boundary updates, is found to provide the highest reduction in runtime for the cases analyzed. (authors)

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Initial demonstration of automated fuel performance modeling with 1977 EBR-II metallic fuel pins using BISON code with FIPD and IMIS databases

Using the BISON fuel performance code, simulations were conducted using an automated process to read initial and operating conditions from the Fuels Irradiation and Physics Database (FIPD) and Integral Fast Reactor materials information system (IMIS) database, which contains metallic fuel data from the Experimental Breeder Reactor-II (EBR-II). This work demonstrates use of an integrated framework to access the vast majority of EBR-II experimental fuel pin data to support rapid development of fuel performance models for next-generation metallic fuel systems. With this capability, validation for fuel qualification can be performed rapidly. Between IMIS and FIPD, there is enough information to conduct 1977 unique EBR-II metallic fuel pin histories from 24 different experiments, at varying levels of detail between the two databases. Each of these histories includes a high-resolution power history, flux history, coolant channel flow rates, and coolant channel temperatures. Fission gas release (FGR), cumulative damage fraction (CDF), fuel axial swelling, cladding profilometry, and burnup were all simulated in BISON. The results were compared to post-irradiation examination (PIE) results for the initial demonstration of automated BISON modeling. BISON simulations conducted with IMIS and FIPD were in rough agreement with PIE measurements and calculations. Cladding profilometry, FGR, and fuel axial swelling were found to be in rough agreement with PIE measurements, depending on the physics used within the BISON input files. Here, the mechanical contact solver chosen was found to significantly impact axial fuel swelling and cladding strain predictions. CDF values were assessed to see whether pin failure may have been predicted (CDF ≥ 1). This work suggests that continued development of an automated tool for BISON should focus on inclusion of the Fast Flux Test Facility (FFTF) experimental data for a larger database for metallic fuel, improved physical models to better capture fuel performance, such as fuel-cladding interactions, and a more detailed comparison with available PIE data to further the BISON model development.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

An open source fast fluid dynamics model for data center thermal management

Although computational fluid dynamics (CFD) has been widely adopted to improve data center thermal management, the high computational demand limits its applications, such as multivariate optimal design and operation. Fast fluid dynamics (FFD), which has been applied for fast airflow simulation, shows great potential. However, few research applied FFD for optimal design and operation of data center thermal management. This research improves the FFD model for data centers and conducts a comprehensive evaluation and demonstration. First, the FFD model is improved by solving the advection and diffusion equations together using an upwind scheme instead of a semi-Lagrangian advection solver in the conventional FFD model. Second, new features for data centers are added, such as a pressure correction method to simulate plenum airflow and dynamic boundary conditions for IT racks. The new FFD model is first validated with two indoor environment cases and the results show that the new FFD model has slightly better overall prediction accuracy and faster speed compared to the conventional FFD model. It is also observed that both FFD models achieve acceptable accuracy, except for a few localized disparities with experimental data, which might be due to simplified handling of turbulence viscosity near the boundaries. Furthermore, validation with a real data center shows that the FFD model achieves a similar level of accuracy as CFD when compared to the experimental measurements with some level of uncertainties. It is then demonstrated for data center optimal design and operation, which saves 53.4–58.8% of annual energy while still meeting the thermal requirements. In conclusion, with a much faster speed and comparable accuracy compared to CFD, the FFD model parallelized on a graphics processing unit is promising for practical model-based data center early design and operation.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Simulating self-powered neutron detector responses to infer burnup-induced power distribution perturbations in next-generation light water reactors

Understanding how 3D power distribution will be monitored throughout reactor core volumetric space in next-generation nuclear power reactors is crucial to the design, deployment, and licensing of these reactors. Although numerous techniques exist for 3D power distribution monitoring based on the response of both in situ and ex situ sensors currently implemented or proposed for use in the US reactor fleet, crucial details about these techniques are often unclear. The publicly available documentation does not include information such as how well these techniques are characterized and optimized in their implementations and the levels of uncertainty in the inferred 3D power distribution. The work described herein investigated a recently developed 3D power distribution inferencing method as applied to two next-generation reactor simulations: (1) the NuScale small modular reactor design and (2) the Westinghouse AP1000 design, both of which contain in-core strings of vanadium self-powered neutron detectors (SPNDs). This investigation considered a range of SPND string sensor densities, as well as a range of 3D power distribution axial segment sizes. In this work, SPND response simulation is informed by neutron flux calculations in representative homogenized cores. For the different sensor densities and power distribution axial segment sizes in these simulations, the average solution error, solver iterations, and run time were tracked to parameterize the sensor-core configuration.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

NREL Marine Energy Desalination Research and Development Portfolio: Preprint

The National Renewable Energy Laboratory (NREL), in collaboration with the U.S. Department of Energy's (DOE's) Water Power Technologies Office (WPTO), has developed a unique R&D approach to advance marine energy desalination. Desalination is a foundational investment within WPTO's Powering the Blue Economy portfolio, and was the first investment within this portfolio. NREL's marine energy desalination spans techno-economic feasibility studies, numerical modeling, and laboratory testing at the component and subsystem level, as well as development of the Hydraulic and Electric Reverse Osmosis Wave Energy Converter (HERO WEC). This multilayered approach enables an innovative feedback loop where the data and lessons learned from laboratory and field experiments are used to refine modeling tools and analysis techniques, prioritize out-year activities, and refine strategic directions within NREL and across the WPTO portfolio. The primary objective of the NREL-led research is to identify key barriers associated with the commercialization of wave-powered desalination and develop solutions that can be adopted by the marine energy industry. In parallel, these R&D activities can help inform technical assistance and support of industry and academic technologies. These two tracks help build a common solver community approach, while also identifying key stakeholders, government agencies, and other organizations outside of the marine sector that will be necessary for developing a robust industry. For the Spanish version of this report, see NREL/CP-5700-88483 (https://www.nrel.gov/docs/fy24osti/88483.pdf).

analysis↗

Cartera de Investigacion y Desarrollo Sobre Desalinizacion de Energia Marina del NREL: Preprint (Spanish Translation)

The National Renewable Energy Laboratory (NREL), in collaboration with the U.S. Department of Energy's (DOE's) Water Power Technologies Office (WPTO), has developed a unique R&D approach to advance marine energy desalination. Desalination is a foundational investment within WPTO's Powering the Blue EconomyTM portfolio [1], and was the first investment within this portfolio. NREL's marine energy desalination spans techno-economic feasibility studies, numerical modeling, and laboratory testing at the component and subsystem level, as well as development of the Hydraulic and Electric Reverse Osmosis Wave Energy Converter (HERO WEC). This multilayered approach enables an innovative feedback loop where the data and lessons learned from laboratory and field experiments are used to refine modeling tools and analysis techniques, prioritize out-year activities, and refine strategic directions within NREL and across the WPTO portfolio. The primary objective of the NREL-led research is to identify key barriers associated with the commercialization of wave-powered desalination and develop solutions that can be adopted by the marine energy industry. In parallel, these R&D activities can help inform technical assistance and support of industry and academic technologies. These two tracks help build a common solver community approach, while also identifying key stakeholders, government agencies, and other organizations outside of the marine sector that will be necessary for developing a robust industry. For the English version of this report, see NREL/CP-5700-86724 (https://www.nrel.gov/docs/fy24osti/86724.pdf).

analysis↗

A physics-based model for frost buildup under turbulent flow using direct numerical simulations

We present a new model for frost buildup under turbulent (and laminar) flow using direct numerical simulations. The physical model consists of two layers, the air and the frost. The air layer is fully resolved and consists of solving for the velocity, temperature, and vapor mass fraction fields. The frost layer thickness is resolved using conservation of mass and energy. Both phases are dynamically coupled using the immersed boundary method. Three-dimensional simulations are conducted in an open-channel configuration. A number of challenges need to be overcome to make these simulations feasible. First, to enforce far-field conditions of zero gradient and prescribed mean temperature and humidity, a source term is added to the energy and transport equations in the flow solver. Second, the mean frost thickness is subtracted after each time step to ensure a constant mean flow thickness and level of turbulence in the numerical domain. Third, a slow-time acceleration approach, which accelerates the frost buildup by a predetermined factor, is employed to bridge the gap between the fast turbulent and slow frost buildup time scales. Finally, a frost densification scheme is used to overcome the difficulties of vertically varying frost properties. The model is validated by comparing the frost thickness and frost thickness buildup rate over a period of one hour from a cooled flat plate experiment. As a result, both quantities compare favorably with experiments.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Characterizing GPU Energy Usage in Exascale-Ready Portable Science Applications

We characterize the GPU energy usage of two widely adopted exascale-ready applications representing two classes of particle and mesh solvers: (i) QMCPACK, a quantum Monte Carlo package, and (ii) AMReX-Castro, an adaptive mesh astrophysical code. We analyze power, temperature, utilization, and energy traces from double-/single (mixed)-precision benchmarks on NVIDIA’s A100 and H100 and AMD’s MI250X GPUs using queries in NVML and rocm_smi_lib, respectively. We explore application-specific metrics to provide insights on energy vs. performance trade-offs. Our results suggest that mixed-precision energy savings range between 6–25% on QMCPACK and 45% on AMReX-Castro. Also, we found gaps in the AMD tooling used on Frontier GPUs that need to be understood, while query resolutions on NVML have little variability between 1 ms-1 s. Overall, application level knowledge is crucial to define energy-cost/science-benefit opportunities for the codesign of future supercomputer architectures in the post-Moore era.

Godoy, William [ORNL] (ORCID:0000000225905178)↗