Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “UPS”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

High Performance Computing Peak Shaving for Microreactor Operation

There are multiple nuclear microreactors currently under development that are designed to provide autonomous power for as many as ten or more years without refueling and are designed to power high performance computing (HPC) datacenters. But the load-follow speeds for a nuclear microreactor will be much slower than grid power and slower than the power variance typical of a HPC system. HPC datacenters experience peak power load variance driven by several factors ranging from the operation of cooling systems to remove heat from the servers to supporting a wide range of user application workflows and architectures each with different power signatures. One mechanism to support the limited load-follow of a microreactor is peak shaving where an energy storage mechanism is used to shed peak load and reduce significant power variance. This work explores peak electrical load shaving using uninterruptible power supply (UPS) systems designed for HPC support in the context of peak shaving when operating using a nuclear microreactor with a load-follow limited to 10% of load per minute. Using a self contained HPC datacenter complete with stand-alone cooling system and provisioned with an x86 cluster, an ARM cluster, and a graphics processing unit (GPU) cluster, peak shaving for microreactor operation using the UPS battery backup is explored while running two classes of typical HPC user applications. HPC architecture suitability for microreactor operation under this type of peak shaving is examined.

97 MATHEMATICS AND COMPUTING↗

KRT18 as a Novel Biomarker of Urothelial Papilloma while Evaluating Low-Grade Papillary Urothelial Neoplasms: Bi-Center Analysis

Introduction: Although urothelial papilloma (UP) is an indolent papillary neoplasm that can mimic the morphology of low-grade papillary urothelial carcinoma (PUC), there is no immunomarker to differentiate reliably these two entities. In addition, the molecular characteristics of UP are not fully understood. Methods: We conducted an in-depth proteomic analysis of papillary urothelial lesions (n = 31), including UP and PUC along with normal urothelium. Protein markers distinguishing UP and PUC were selected with machine learning analysis, followed by internal and external validation using immunohistochemistry. Results: In the proteomic analysis, UP and PUC showed overlapping proteomic profiles. Here, we identified EHD4 and KRT18 as candidate diagnostic biomarkers of UP. Through immunohistochemical validation in two independent cohorts (n = 120), KRT18 was suggested as a novel UP diagnostic marker, able to differentiate UP from low-grade PUC. We also found that 3.5% of patients with UP developed urothelial carcinoma in subsequent resections, supporting the malignant potential of UP. KRT18 downregulation was significantly associated with UPs subsequently progressing to urothelial carcinoma, following their initial diagnosis. Conclusion: This is the first study that successfully revealed UPs comprehensive proteomic landscape, while it also identified KRT18 as a potential diagnostic biomarker of UP.

Biomarkers↗

Latent Space Simulation for Carbon Capture Design Optimization

The CO2 capture efficiency in solvent-based carbon capture systems (CCSs) critically depends on the gas-solvent interfacial area (IA), making maximization of IA a foundational challenge in CCS design. While the IA associated with a particular CCS design can be estimated via a computational fluid dynamics (CFD) simulation, using CFD to derive the IAs associated with numerous CCS designs is prohibitively costly. However, previous works such as Deep Fluids (Kim et al., 2019) show that large simulation speed-ups and low error are achievable by replacing CFD simulators with neural-network (NN) surrogates that mimic the CFD simulation process. This raises the possibility for a fast, accurate replacement for a CFD simulator, and thus computationally feasible IA-based CCS-design optimization. As such, here, we explore whether an existing NN-surrogate approach (and variants we develop) can successfully be applied to our complex carbon-capture CFD simulations, with the ultimate goal of obtaining a fast and accurate simulator for our CCS-design application. Our experiments build on the Deep Fluids approach and find that resulting surrogates can produce large speed ups (4000x) while maintaining IA relative errors as low as 4% on unseen CCS configurations (interpolating between configurations seen during training). Thus, despite less faithfulness to the underlying physics of the problem, NN surrogates may be a promising tool for our CCS design optimization problem. Notably, though, the Deep Fluids approach has limitations for our application (such as model non-transferability to CCS-packing changes), which we discuss. We conclude with potential directions for future work that may, like innovations we introduced here (e.g., transformer-based dynamics prediction), improve performance on our complicated dataset of CCS CFD simulations.

AI Surrogate Simulation, Carbon Capture, design op↗

Thermal simulation of the HSR arc BPM Module for EIC

The Electron Ion Collider (EIC) Hadron Storage Ring (HSR) will reuse most of the existing superconducting magnets from the RHIC storage ring. However, the existing stripline beam position monitors (BPM) used for RHIC will not be compatible with the planned EIC hadron beam parameters that include higher intensity, shorter bunches, and some operational scenarios with large radial offsets of the beam in the vacuum chamber. To address these challenges, the existing RHIC stripline BPMs will be shielded, and a new BPM design using button pick-ups and integrated in a new vacuum interconnect/bellows assembly that will be installed adjacent to the existing BPMs. A thermal analysis of the new arc BPM housing and button pick-up design has been conducted to assess the effects caused by beam induced resistive wall heating and signal propagation through the button pick-up cables for several operational scenarios. This report will describe the analysis results to quantify the heat transfer and temperature distribution that can be expected on the new HSR cryogenic arc BPM housings, button pick-ups, and cryo-signal cables.

43 PARTICLE ACCELERATORS↗

The Port of Los Angeles Zero- and Near-Zero-Emission Freight Facilities "Shore to Store" Project (Final Project Report)

The City of Los Angeles Harbor Department (Harbor Department, POLA) partnered with Equilon Enterprises LLC (d/b/a Shell Oil Products US) (Shell), Toyota Motor North America (Toyota) and Kenworth Truck Company (Kenworth) partnered with the Port of Hueneme (POH), United Parcel Service (UPS), Total Transportation Services Inc. (TTSI), Southern Counties Express (SCE), Toyota Logistics Services (TLS), Air Liquide, National Renewable Energy Laboratory (NREL), Coalition For A Safe Environment, and the South Coast Air Quality Management District (South Coast AQMD) to introduce hydrogen (H 2 ) fuel into the Southern California drayage truck market by demonstrating near-commercial heavy-duty H 2 fuel cell electric trucks at and between freight facilities throughout the region, while continuing to lay the groundwork for battery-electric operations. The "Shore to Store" (S2S) project built on project team experience to help realize our vision of zero-emission freight operations in the future. Ten Kenworth zero-emission Class 8 fuel cell electric trucks, integrated with Toyota's fuel cell drive technology, were operated by UPS, TTSI, SCE, and TLS in revenue service. The demonstration fleet fueled at the S2S hydrogen fueling stations that were built in Ontario, California and Wilmington, California. An additional station at the Port of Long Beach (Portal Station) was available for fueling the fleet. Portal Station was supported by grants from the California Energy Commission (CEC) and South Coast AQMD and used as match funding for the S2S project. POH demonstrated two battery-electric yard tractors, and TLS demonstrated two zero-emission forklifts at their warehouse facility, showcasing elements of the entire supply chain operating on zero-emissions. This project showcased a snapshot of the zero-emission supply chain of the future, providing a model by which freight facilities can support zero-emission operations.

33 ADVANCED PROPULSION SYSTEMS↗

Mitigating Data Center Impact on Grid Stability: A Coordinated Control Strategy Using Verrus StabiliGrid Architecture

Large data centers, which now represent a significant and growing share of the total U.S. grid load, can inadvertently destabilize the electrical grid when they disconnect simultaneously during brief voltage disturbances. The July 10, 2024, Eastern Interconnection incident, in which a sub-100-millisecond transmission fault triggered the cascading loss of approximately 1,500 MW of data center load, illustrates this vulnerability. While commercial battery energy storage systems (BESS) deployed in data centers provide device-level fault ride-through per IEEE 1547, they lack coordination with facility protection logic and uninterruptible power supplies (UPS), limiting their effectiveness as grid-stabilizing assets. This report presents the Verrus StabiliGrid architecture, a coordinated control framework that integrates BESS, UPS, and point-of-interconnection (POI) protection settings to enable data centers to ride through both undervoltage and overvoltage grid contingencies without disconnecting. The four-step strategy encompasses: (1) high-resolution power quality monitoring to detect the grid state during events such as undervoltage, overvoltage, underfrequency, and overfrequency; (2) POI protection settings that allow for extended ride-through and grid-connected operation during grid contingencies; (3) grid state-driven autonomous dispatch of assets to improve grid resilience by reducing power draw during undervoltage or absorbing more power during overvoltage events; and (4) coordinated post-recovery dispatch of data center assets to restore firm load to pre-contingency levels. Validated through controller-hardware-in-the-loop (C-HIL) simulations at the National Laboratory of the Rockies, results show grid import restoration to pre-fault levels within 100 milliseconds of voltage recovery. This work advances the ability of data centers to transition from passive, disturbance-sensitive loads to active participants in grid stability, a capability increasingly required by emerging NERC and ERCOT regulatory frameworks.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Efficient Space–Time Reduced Order Model for Linear Dynamical Systems in Python Using Less than 120 Lines of Code

A classical reduced order model (ROM) for dynamical problems typically involves only the spatial reduction of a given problem. Recently, a novel space–time ROM for linear dynamical problems has been developed [Choi et al., Space–tume reduced order model for large-scale linear dynamical systems with application to Boltzmann transport problems, Journal of Computational Physics, 2020], which further reduces the problem size by introducing a temporal reduction in addition to a spatial reduction without much loss in accuracy. The authors show an order of a thousand speed-up with a relative error of less than 10−5 for a large-scale Boltzmann transport problem. In this work, we present for the first time the derivation of the space–time least-squares Petrov–Galerkin (LSPG) projection for linear dynamical systems and its corresponding block structures. Utilizing these block structures, we demonstrate the ease of construction of the space–time ROM method with two model problems: 2D diffusion and 2D convection diffusion, with and without a linear source term. For each problem, we demonstrate the entire process of generating the full order model (FOM) data, constructing the space–time ROM, and predicting the reduced-order solutions, all in less than 120 lines of Python code. We compare our LSPG method with the traditional Galerkin method and show that the space–time ROMs can achieve O(10−3) to O(10−4) relative errors for these problems. Depending on parameter–separability, online speed-ups may or may not be achieved. For the FOMs with parameter–separability, the space–time ROMs can achieve O(10) online speed-ups. Finally, we present an error analysis for the space–time LSPG projection and derive an error bound, which shows an improvement compared to traditional spatial Galerkin ROM methods.

97 MATHEMATICS AND COMPUTING↗

Exploring ice sheet model sensitivity to ocean thermal forcing and basal sliding using the Community Ice Sheet Model (CISM)

Abstract. Multi-meter sea level rise (SLR) is thought to be possible within the next few centuries, with most of the uncertainty originating from the Antarctic land ice contribution. One source of uncertainty relates to the ice sheet model initialization. Since ice sheets have a long response time (compared to other Earth system components such as the atmosphere), ice sheet model initialization methods can have significant impacts on how the ice sheet responds to future forcings. To assess this, we generated 25 different ice sheet spin-ups, using the Community Ice Sheet Model (CISM) at a 4 km resolution. During each spin-up, we varied two key parameters known to impact the sensitivity of the ice sheet to future forcing: one related to the sensitivity of the ice shelf melt rate to ocean thermal forcing (TF) and the other related to the basal friction. The spin-ups all nudge toward observed thickness and enforce a no-advance calving criterion, such that all final spin-up states resemble observations but differ in their melt and friction parameter settings. Each spin-up was then forced with future ocean thermal forcings from 13 different CMIP6 models under the Shared Socioeconomic Pathway (SSP)5-8.5 emissions scenario and modern climatological surface mass balance data. Our results show that the effects of the ice sheet and ocean parameter settings used during the spin-up are capable of impacting multi-century future SLR predictions by as much as 2 m. By the end of this century, the effects of these choices are more modest, but still significant, with differences of up to 0.2 m of SLR. We have identified a combined ocean and ice parameter space that leads to widespread mass loss within 500 years (low friction and high melt rate sensitivity). To explore temperature thresholds, we also ran a synthetically forced CISM ensemble that is focused on the Amundsen region only. Given certain ocean and ice parameter choices, Amundsen mass loss can be triggered with thermal forcing anomalies between 1.5 and 2 ∘C relative to the spin-up. Our results emphasize the critical importance of considering ice sheet and ocean parameter choices during spin-up for SLR predictions and suggest the importance of including glacial isostatic adjustment in ice sheet simulations.

54 ENVIRONMENTAL SCIENCES↗

Supporting Bioproducts Industry Growth with a System Dynamics Decision-Support Tool

A bio-based economy requires chemical products as well as fuels to be produced from biomass. Although a variety of universities, government agencies, start-ups and established firms have engaged in bioproduct development, many projects have failed to reach the point of commercialization and commercialized bioproducts struggle to capture and maintain market share. To date, there has been no general research into the factors that contribute to bioproduct failure or success. This work presents the Bioproduct Transition Dynamics (BTD) system dynamics model, a decision-support tool that simulates the bioproduct development process from pre-piloting research through construction of the first commercial-scale plant, as well as the processes of obtaining funding from investors and government agencies. The core of the BTD is a feedback loop between bioproduct developers and funders, which relates development progress measured with indicators such as net present value to funders’ decisions to continue investing. External factors such as feedstock prices, market size and growth, and the existence of bioproduct consumers also influence a development project’s chances of receiving follow-on funding. Virtually any bioproduct can be represented with the BTD: direct replacements, performance-advantaged products, niche and commodity markets can all be modeled. The goal of the BTD project is to inform decisions made by developers in both established firms and start-ups, investors, government agencies, and other stakeholders interested in growing the nascent U.S. bioproducts industry. This talk will cover the general structure and functionality of the BTD, and present results from an analysis performed with the BTD to demonstrate its use as a decision-support tool and the insights it can provide. The goal of the analysis is to identify the most critical factors that lead to direct replacement and performance-advantaged bioproduct projects emerging successfully from the “Valley of Death”, and to determine if these factors differ between the two bioproduct types.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Control System Upgrade for Battery State-of-Charge Indications

The Advanced Test Reactor (ATR) Complex at Idaho National Laboratory (INL) relies on Battery Backed Power (BBP) systems and Uninterruptible Power Supplies (UPS) to ensure continuous power supply to critical components. This project aims to enhance the reliability and functionality of the battery monitoring and control systems by updating the State-of-Charge (SOC) system, Programmable Logic Controller (PLC), and Human-Machine Interface (HMI) for the nuclear safety-related battery banks. The current system, while functional, has areas for improvement, particularly in recharging calculations and alarm functions. The project objectives include developing flow charts, programming the new PLC and HMI, conducting bench tests, and updating design documentation. Additionally, the project ensures compliance with safety standards, develops training materials, creates comprehensive documentation, and integrates seamlessly with existing ATR infrastructure. The new SOC system is designed to be scalable for future upgrades, improve efficiency, enhance data accuracy, implement redundancy features, and achieve project goals within budget constraints while considering environmental impact. The methodology involved familiarizing with BBP and UPS systems, collecting current readings, rescaling signals, learning ladder logic, and updating the HMI. The transition from SLC 5/03 PLC using RS Logix 500 to CompactLogix 5380 using Studio 5000 was a key step. Despite challenges in transferring outdated PLC ladder logic and HMI code, starting from scratch led to a more accurate and efficient monitoring system, contributing to improved safety and operational efficiency. The project is currently awaiting approval of the Engineering Calculation and Analysis Report (ECAR) before implementation.

42 - ENGINEERING↗

Accelerating x-ray tracing for exascale systems using Kokkos

The upcoming exascale computing systems Frontier and Aurora will draw much of their computing power from GPU accelerators. The hardware for these systems will be provided by AMD and Intel, respectively, each supporting their own GPU programming model. The challenge for applications that harness one of these exascale systems will be to avoid lock-in and to preserve performance portability. We report here on our results of using Kokkos to accelerate a real-world application on NERSC's Perlmutter Phase 1 (using NVIDIA A100 accelerators) and Crusher, the testbed system for OLCF's Frontier (using AMD MI250X). By porting to Kokkos, we successfully ran the same X-ray tracing code on both systems and achieved speed-ups between 13 % and 66 % compared to the original CUDA code. Finally, these results are a highly encouraging demonstration of using Kokkos to accelerate production science code.

97 MATHEMATICS AND COMPUTING↗

Large language model evaluation for high–performance computing software development

We apply AI-assisted large language model (LLM) capabilities of GPT-3 targeting high-performance computing (HPC) kernels for (i) code generation, and (ii) auto-parallelization of serial code in C ++, Fortran, Python and Julia. Our scope includes the following fundamental numerical kernels: AXPY, GEMV, GEMM, SpMV, Jacobi Stencil, and CG, and language/programming models: (1) C++ (e.g., OpenMP [including offload], OpenACC, Kokkos, SyCL, CUDA, and HIP), (2) Fortran (e.g., OpenMP [including offload] and OpenACC), (3) Python (e.g., numpy, Numba, cuPy, and pyCUDA), and (4) Julia (e.g., Threads, CUDA.jl, AMDGPU.jl, and KernelAbstractions.jl). Kernel implementations are generated using GitHub Copilot capabilities powered by the GPT-based OpenAI Codex available in Visual Studio Code given simple + + prompt variants. To quantify and compare the generated results, we propose a proficiency metric around the initial 10 suggestions given for each prompt. For auto-parallelization, we use ChatGPT interactively giving simple prompts as in a dialogue with another human including simple “prompt engineering” follow ups. Results suggest that correct outputs for C++ correlate with the adoption and maturity of programming models. For example, OpenMP and CUDA score really high, whereas HIP is still lacking. We found that prompts from either a targeted language such as Fortran or the more general-purpose Python can benefit from adding language keywords, while Julia prompts perform acceptably well for its Threads and CUDA.jl programming models. Finally, we expect to provide an initial quantifiable point of reference for code generation in each programming model using a state-of-the-art LLM. Overall, understanding the convergence of LLMs, AI, and HPC is crucial due to its rapidly evolving nature and how it is redefining human-computer interactions.

97 MATHEMATICS AND COMPUTING↗

Results of a Geant4 benchmarking study for bio‐medical applications, performed with the G4‐Med system

Geant4, a Monte Carlo Simulation Toolkit extensively used in bio-medical physics, is in continuous evolution to include newest research findings to improve its accuracy and to respond to the evolving needs of a very diverse user community. In 2014, the G4-Med benchmarking system was born from the effort of the Geant4 Medical Simulation Benchmarking Group, to benchmark and monitor the evolution of Geant4 for medical physics applications. The G4-Med system was first described in our Medical Physics Special Report published in 2021. Results of the tests were reported for Geant4 10.5. Purpose In this work, we describe the evolution of the G4-Med benchmarking system. Methods The G4-Med benchmarking suite currently includes 23 tests, which benchmark Geant4 from the calculation of basic physical quantities to the simulation of more clinically relevant set-ups. New tests concern the benchmarking of Geant4-DNA physics and chemistry components for regression testing purposes, dosimetry for brachytherapy with a 125 I source, dosimetry for external x-ray and electron FLASH radiotherapy, experimental microdosimetry for proton therapy, and in vivo PET for carbon and oxygen beams. Regression testing has been performed between Geant4 10.5 and 11.1. Finally, a simple Geant4 simulation has been developed and used to compare Geant4 EM physics constructors and physics lists in terms of execution times. Results In summary, our EM tests show that the parameters of the multiple scattering in the Geant4 EM constructor G4EmStandardPhysics_option3 in Geant4 11.1, while improving the modeling of the electron backscattering in high atomic number targets, are not adequate for dosimetry for clinical x-ray and electron beams. Therefore, these parameters have been reverted back to those of Geant4 10.5 in Geant4 11.2.1. The x-ray radiotherapy test shows significant differences in the modeling of the bremsstrahlung process, especially between G4EmPenelopePhysics and the other constructors under study (G4EmLivermorePhysics, G4EmStandardPhysics_option3, and G4EmStandardPhysics_option4). These differences will be studied in an in-depth investigation within our Group. Improvement in Geant4 11.1 has been observed for the modeling of the proton and carbon ion Bragg peak with energies of clinical interest, thanks to the adoption of ICRU90 to calculate the low energy proton stopping powers in water and of the Linhard–Sorensen ion model, available in Geant4 since version 11.0. Nuclear fragmentation tests of interest for carbon ion therapy show differences between Geant4 10.5 and 11.1 in terms of fragment yields. In particular, a higher production of boron fragments is observed with Geant4 11.1, leading to a better agreement with reference data for this fragment. Conclusions Based on the overall results of our tests, we recommend to use G4EmStandardPhysics_option4 as EM constructor and QGSP_BIC_HP with G4EmStandardPhysics_option4, for hadrontherapy applications. The Geant4-DNA physics lists report differences in modeling electron interactions in water, however, the tests have a pure regression testing purpose so no recommendation can be formulated.

62 RADIOLOGY AND NUCLEAR MEDICINE↗

Optimizing multigrid reduction-in-time and Parareal coarse-grid operators for linear advection

Parallel-in-time methods, such as multigrid reduction-in-time (MGRIT) and Parareal, provide an attractive option for increasing concurrency when simulating time-dependent partial differential equations (PDEs) in modern high-performance computing environments. While these techniques have been very successful for parabolic equations, it has often been observed that their performance suffers dramatically when applied to advection-dominated problems or purely hyperbolic PDEs using standard rediscretization approaches on coarse grids. In this paper, we apply MGRIT or Parareal to the constant-coefficient linear advection equation, appealing to existing convergence theory to provide insight into the typically nonscalable or even divergent behavior of these solvers for this problem. To overcome these failings, we replace rediscretization on coarse grids with improved coarse-grid operators that are computed by applying optimization techniques to approximately minimize error estimates from the convergence theory. Therefore, one of our main findings is that, in order to obtain fast convergence as for parabolic problems, coarse-grid operators should take into account the behavior of the hyperbolic problem by tracking the characteristic curves. Our approach is tested for schemes of various orders using explicit or implicit Runge–Kutta methods combined with upwind-finite-difference spatial discretizations. In all cases, we obtain scalable convergence in just a handful of iterations, with parallel tests also showing significant speed-ups over sequential time-stepping.

97 MATHEMATICS AND COMPUTING↗

The microstructure of a selective laser melting (SLM)-fabricated NiTi shape memory alloy with superior tensile property and shape memory recoverability

A selective laser melting (SLM)-fabricated NiTi with superior tensile property and shape memory recoverability was obtained by using a unique stripe rotation scanning strategy. Here, the alloy characteristics, formation mechanisms and evolution in terms of twins, dislocations and precipitations of the alloy were systematically studied. Compared to the conventional smelting-followed-by-machining, the SLM fabrication process involves rapid solidification and repeated heating, which confer distinctive characteristics to the microstructures of SLM-fabricated NiTi alloys. Rapid solidification promotes the formation of a supersaturation solid solution matrix containing a high concentration vacancies, which in turn aggregate to generate a high density of dislocations. During subsequent repeated heating stages, these dislocations occur thermal motion along three directions of <001 > , <111> and <110 >, leading to the formation of thermal kinks, helical dislocations and wave morphology. Simultaneously, precipitated particles Ti 3 Ni 4 repeatedly nucleate and heterogeneous grow with the movement of dislocation. Such precipitation behavior, termed repeated precipitation, has not been previously reported in the conventional NiTi alloys, suggesting that it could be a unique characteristic of such alloy. After martensitic transformation, only two twins, {1 1¯ 1} type I twin and compound twin, are detected. The twinning lamellae of these two twins, where precipitation and dislocation pile-ups exist, often have uneven thickness and chaotic arrangement. Besides, the unique self-accommodated microstructures, such as secondary {1 1¯ 1} type I twin and compound twin with “herring-bone’’ lamellae, which often appear in the deformed or nanocrystalline NiTi, can also be observed. These unique microstructures may confer the distinctive properties to the SLM-fabricated NiTi.

36 MATERIALS SCIENCE↗

Model-parallel Fourier neural operators as learned surrogates for large-scale parametric PDEs

Fourier neural operators (FNOs) are a recently introduced neural network architecture for learning solution operators of partial differential equations (PDEs), which have been shown to perform significantly better than comparable deep learning approaches. Once trained, FNOs can achieve speed-ups of multiple orders of magnitude over conventional numerical PDE solvers. However, due to the high dimensionality of their input data and network weights, FNOs have so far only been applied to two-dimensional or small three-dimensional problems. To remove this limited problem-size barrier, we propose a model-parallel version of FNOs based on domain-decomposition of both the input data and network weights. Here, we demonstrate that our model-parallel FNO is able to predict time-varying PDE solutions of over 2.6 billion variables on Perlmutter using up to 512 A100 GPUs and show an example of training a distributed FNO on the Azure cloud for simulating multiphase CO 2 dynamics in the Earth’s subsurface.

58 GEOSCIENCES↗

Exponential time differencing for the tracer equations appearing in primitive equation ocean models

The tracer equations are part of the primitive equations used in ocean modeling and describe the transport of tracers, such as temperature, salinity or chemicals, in the ocean. Depending on the number of tracers considered, several equations may be added to and coupled to the dynamics system. In many relevant situations, the time-step requirements of explicit methods imposed by the transport and mixing in the vertical direction are more restrictive than those for the horizontal, and this may cause the need to use very small time steps if a fully explicit method is employed. To overcome this issue, we propose an exponential time differencing (ETD) solver where the vertical terms (transport and diffusion) are treated with a matrix exponential, whereas the horizontal terms are dealt with in an explicit way. In this work, we investigate numerically the computational speed-ups that can be obtained over other semi-implicit methods, and we analyze the advantages of the method in the case of multiple tracers.

42 ENGINEERING↗

PPINN: Parareal physics-informed neural network for time-dependent PDEs

Physics-informed neural networks (PINNs) encode physical conservation laws and prior physical knowledge into the neural networks, ensuring the correct physics is represented accurately while alleviating the need for supervised learning to a great degree. While effective for relatively short-term time integration, when long time integration of the time-dependent PDEs is sought, the time–space domain may become arbitrarily large and hence training of the neural network may become prohibitively expensive. To this end, we develop a parareal physics-informed neural network (PPINN), hence decomposing a long-time problem into many independent short-time problems supervised by an inexpensive/fast coarse-grained (CG) solver. In particular, the serial CG solver is designed to provide approximate predictions of the solution at discrete times, while initiate many fine PINNs simultaneously to correct the solution iteratively. There is a two-fold benefit from training PINNs with small-data sets rather than working on a large-data set directly, i.e., training of individual PINNs with small-data is much faster, while training the fine PINNs can be readily parallelized. Consequently, compared to the original PINN approach, the proposed PPINN approach may achieve a significant speed-up for long-time integration of PDEs, assuming that the CG solver is fast and can provide reasonable predictions of the solution, hence aiding the PPINN solution to converge in just a few iterations. To investigate the PPINN performance on solving time-dependent PDEs, we first apply the PPINN to solve the Burgers equation, and subsequently we apply the PPINN to solve a two-dimensional nonlinear diffusion–reaction equation. Furthermore, our results demonstrate that PPINNs converge in a few iterations with significant speed-ups proportional to the number of time-subdomains employed.

42 ENGINEERING↗