Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “runtime”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

Analyzing inference workloads for spatiotemporal modeling

Ensuring power grid resiliency, forecasting climate conditions, and optimization of transportation infrastructure are some of the many application areas where data is collected in both space and time. Spatiotemporal modeling is about modeling those patterns for forecasting future trends and carrying out critical decision-making by leveraging machine learning/deep learning. Once trained offline, field deployment of trained models for near real-time inference could be challenging because performance can vary significantly depending on the environment, available compute resources and tolerance to ambiguity in results. Users deploying spatiotemporal models for solving complex problems can benefit from analytical studies considering a plethora of system adaptations to understand the associated performance-quality trade-offs. To facilitate the co-design of next-generation hardware architectures for field deployment of trained models, it is critical to characterize the workloads of these deep learning (DL) applications during inference and assess their computational patterns at different levels of the execution stack. In this paper, we develop several variants of deep learning applications that use spatiotemporal data from dynamical systems. We study the associated computational patterns for inference workloads at different levels, considering relevant models (Long short-term Memory, Convolutional Neural Network and Spatio-Temporal Graph Convolution Network), DL frameworks (Tensorflow and PyTorch), precision (FP16, FP32, AMP, INT16 and INT8), inference runtime (ONNX and AI Template), post-training quantization (TensorRT) and platforms (Nvidia DGX A100 and Sambanova SN10 RDU). Overall, our findings indicate that although there is potential in mixed-precision models and post-training quantization for spatiotemporal modeling, extracting efficiency from contemporary GPU systems might be challenging. Instead, co-designing custom accelerators by leveraging optimized High Level Synthesis frameworks (such as SODA High-Level Synthesizer for customized FPGA/ASIC targets) can make workload-specific adjustments to enhance the efficiency.

97 MATHEMATICS AND COMPUTING↗

DTLMod: A simulation framework for in situ workflow optimization

In situ processing workflows have become essential for coping with the explosion in data volume and velocity in large-scale scientific computing, providing domain scientists with early insights at runtime. Multiple frameworks implement this paradigm through a data transport layer (DTL), offering different data access modes and deployment schemes, but researchers currently lack the appropriate tools to assess design and deployment options before committing to costly real experiments. We introduce DTLMod, an open-source simulated DTL that enables performance evaluation of in situ workflow configurations at scale. Built on SimGrid, it links into any SimGrid-based simulator and is available in C++ and Python. We evaluate DTLMod along four axes: scalability (tens of thousands of simulated processes across interconnected clusters in seconds, with linear memory scaling), versatility (three implementation variants trading fidelity for speed), accuracy (simulated times faithfully reflecting real behavior), and practical utility (two use cases demonstrating evidence-based workflow design decisions).

Suter, Fred [ORNL] (ORCID:0000000319021955)↗

Proxy-based Bayesian inversion of strain tensor data measured during well tests

Recent instrument developments have made it possible to measure the strain tensor caused by injecting or pumping fluid from aquifers or reservoirs, but the full value of these data is limited because the long runtimes of poroelastic forward models makes it impractical to use many inversion schemes. This limits the interpretation of strain data for managing the recovery of resources or storage of wastes in the subsurface. This paper describes a method of inverting deformation data using a poroelastic numerical simulator so the results can be used to manage reservoirs or aquifers. We developed a workflow designed to reduce the number of simulations sufficiently to make it feasible to use DREAMzs, an advanced Bayesian inversion method that translates the uncertainties from different sources into unbiased posterior parameter distributions and uncertainty envelopes around the field data. Using a KNN proxy model for the poroelastic simulator is key to reducing the overall computations, and the workflow includes a strategy for ensuring the proxy model results converge on the results from the simulator. The workflow is tested using an idealized example that verifies the ability to correctly identify parameters and characterize noise used to perturb the data. Field data from an injection test at an oil reservoir near Tulsa, Oklahoma, are also used to evaluate the efficacy of the workflow with a real dataset. The workflow identified 265 history matching solutions out of 1240 total simulation runs (21% acceptance ratio), where the results were used to characterize posterior parameter distribution and evaluate the prediction uncertainty. Furthermore, this workflow is significant because it enables strain tensor, or other geomechanical measurements to be interpreted to guide decision-making during energy and environmental processes in the subsurface.

42 ENGINEERING↗

A tri-level optimization model for interdependent infrastructure network resilience against compound hazard events

Resilient operation of interdependent infrastructures against compound hazard events is essential for maintaining societal well-being. To address consequence assessment challenges in this problem space, we propose a novel policy-guided tri-level optimization model applied to a proof-of-concept case study with fuel distribution and transportation networks – encompassing one realistic network; one fictitious, yet realistic network; as well as networks drawn from three synthetic distributions. Mathematically, our approach takes the form of a defender-attacker-defender (DAD) model—a multi-agent tri-level optimization, comprised of a defender, attacker, and an operator acting in sequence. Here, in this study, our notional operator may choose proxy actions to operate an interdependent system comprised of fuel terminals and gas stations (functioning as supplies) and a transportation network with traffic flow (functioning as demand) to minimize unmet demand at gas stations. A notional attacker aims to hypothetically disrupt normal operations by reducing supply at the supply terminals, and the notional defender aims to identify best proxy defense policy options which include hardening supply terminals or allowing alternative distribution methods such as trucking reserve supplies. We solve our DAD formulation at a metropolitan scale and present practical defense policy insights against hypothetical compound hazards. We demonstrate the generalizability of our framework by presenting results for a realistic network; a fictitious, yet realistic network; as well as for three networks drawn from synthetic distributions. Additionally, we demonstrate the scalability of the framework by investigating runtime performance as a function of the network size. Steps for future research are also discussed.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Benchmarking blockchain-based gene-drug interaction data sharing methods: A case study from the iDASH 2019 secure genome analysis competition blockchain track

Blockchain distributed ledger technology is just starting to be adopted in genomics and healthcare applications. Despite its increased prevalence in biomedical research applications, skepticism regarding the practicality of blockchain technology for real-world problems is still strong and there are few implementations beyond proof-of-concept. We focus on benchmarking blockchain strategies applied to distributed methods for sharing records of gene-drug interactions. We expect this type of sharing will expedite personalized medicine. We generated gene-drug interaction test datasets using the Clinical Pharmacogenetics Implementation Consortium (CPIC) resource. We developed three blockchain-based methods to share patient records on gene-drug interactions: Query Index, Index Everything, and Dual-Scenario Indexing. We achieved a runtime of about 60 s for importing 4,000 gene-drug interaction records from four sites, and about 0.5 s for a data retrieval query. Our results demonstrated that it is feasible to leverage blockchain as a new platform to share data among institutions.

60 APPLIED LIFE SCIENCES↗

Multigrid deflation for Lattice QCD

Computing the trace of the inverse of large matrices is typically addressed through statistical methods. Deflating out the lowest eigenvectors or singular vectors of the matrix reduces the variance of the trace estimator. This work summarizes our efforts to reduce the computational cost of computing the deflation space while achieving the desired variance reduction for Lattice QCD applications. Previous efforts computed the lower part of the singular spectrum of the Dirac operator by using an eigensolver preconditioned with a multigrid linear system solver. Despite the improvement in performance in those applications, as the problem size grows the runtime and storage demands of this approach will eventually dominate the stochastic estimation part of the computation. In this work, we propose to compute the deflation space in one of the following two ways. First, by using an inexact eigensolver on the Hermitian, but maximally indefinite, operator. Second, by exploiting the fact that the multigrid prolongator for this operator is rich in components toward the lower part of the singular spectrum. We show experimentally that the inexact eigensolver can approximate the lower part of the spectrum even for ill-conditioned operators. Also, the deflation based on the multigrid prolongator is more efficient to compute and apply, and, despite its limited ability to approximate the fine level spectrum, it obtains similar variance reduction on the trace estimator as deflating with approximate eigenvectors from the fine level operator.

97 MATHEMATICS AND COMPUTING↗

Improvements to a class of hybrid methods for radiation transport: Nyström reconstruction and defect correction methods

In this study, two modifications are introduced for improving the accuracy, versatility, and robustness of a class of hybrid methods for radiation transport. In general, such methods are constructed by splitting the radiative flux into collided and uncollided components to which low- and high-resolution angular approximations are applied, respectively. In this work we focus on discrete ordinates discretizations of high and low order. The first modification we introduce changes the way in which the collided component is mapped into the uncollided component at the end of each time step in a simulation. The new mapping is a Nyström-type reconstruction that is applicable to arbitrary discrete ordinates quadratures, is guaranteed to preserve positivity of the solution provided that all ordinate weights are positive, is significantly more accurate than previous methods, and can be readily extended to other discretizations such as moment methods, finite element methods, and diffusion approximations. The second modification leverages integral deferred correction (IDC) to iteratively correct for the splitting error introduced by the inconsistency in angular discretization between the collided and uncollided components, in addition to improving the accuracy of the low-order temporal error that is treated by traditional IDC methods. Numerical tests in one- and two-dimensional geometries are used to demonstrate the increased accuracy and efficiency of the proposed modifications. It is found that the two techniques combined yield methods with solution accuracy and memory requirements comparable to that of monolithic discrete ordinates methods while reducing runtime by as much as a factor of between two and ten, depending on the problem.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

DPM: A deep learning PDE augmentation method with application to large-eddy simulation

A framework is introduced that leverages known physics to reduce overfitting in machine learning for scientific applications. The partial differential equation (PDE) that expresses the physics is augmented with a neural network that uses available data to learn a description of the corresponding unknown or unrepresented physics. Training within this combined system corrects for missing, unknown, or erroneously represented physics, including discretization errors associated with the PDE's numerical solution. For optimization of the network within the PDE, an adjoint PDE is solved to provide high-dimensional gradients, and a stochastic adjoint method (SAM) further accelerates training. Additionally, the approach is demonstrated for large-eddy simulation (LES) of turbulence. High-fidelity direct numerical simulations (DNS) of decaying isotropic turbulence provide the training data used to learn sub-filter-scale closures for the filtered Navier–Stokes equations. Out-of-sample comparisons show that the deep learning PDE method outperforms widely-used models, even for filter sizes so large that they become qualitatively incorrect. It also significantly outperforms the same neural network when a priori trained based on simple data mismatch, not accounting for the full PDE. Measures of discretization errors, which are well-known to be consequential in LES, point to the importance of the unified training formulation's design, which without modification corrects for them. For comparable accuracy, simulation runtime is significantly reduced. A relaxation of the typical discrete enforcement of the divergence-free constraint in the solver is also successful, instead allowing the DPM to approximately enforce incompressibility physics. Since the training loss function is not restricted to correspond directly to the closure to be learned, training can incorporate diverse data, including experimental data.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

A gradient-based deep neural network model for simulating multiphase flow in porous media

We report simulation of multiphase flow in porous media is crucial for the effective management of subsurface energy and environment-related activities. The numerical simulators used for modeling such processes rely on spatial and temporal discretization of the governing mass and energy balance partial-differential equations (PDEs) into algebraic systems via finite-difference/volume/element methods. These simulators usually require dedicated software development and maintenance, and suffer low efficiency from a runtime and memory standpoint for problems with multi-scale heterogeneity, coupled-physics processes or fluids with complex phase behavior. Therefore, developing cost-effective, data-driven models can become a practical choice, and in this work, we choose deep learning approaches as they can handle high dimensional data and accurately predict state variables with strong nonlinearity. In this paper, we describe a gradient-based deep neural network (GDNN) constrained by the physics related to multiphase flow in porous media. We tackle the nonlinearity of flow in porous media induced by rock heterogeneity, fluid properties, and fluid-rock interactions by decomposing the nonlinear PDEs into a dictionary of elementary differential operators. We use a combination of operators to handle rock spatial heterogeneity and fluid flow by advection. Since the augmented differential operators are inherently related to the physics of fluid flow, we treat them as first principles prior knowledge to regularize the GDNN training. We use the example of pressure management at geologic CO 2 storage sites, where CO 2 is injected in saline aquifers and brine is produced, and apply GDNN to construct a predictive model that is trained with physics-based simulation data and emulates the physics process. We demonstrate that GDNN can effectively predict the nonlinear patterns of subsurface responses, including the temporal and spatial evolution of the pressure and saturation plumes. We also successfully extend the GDNN to convolutional neural network (CNN), namely gradient-based CNN (GCNN), and validate its capability to improve the prediction accuracy. GDNN has great potential to tackle challenging problems that are governed by highly nonlinear physics and enable the development of data-driven models with higher fidelity.

42 ENGINEERING↗

Neural-network based collision operators for the Boltzmann equation

Kinetic gas dynamics in rarefied and moderate-density regimes have complex behavior associated with collisional processes. These processes are generally defined by convolution integrals over a high-dimensional space (as in the Boltzmann operator), or require evaluating complex auxiliary variables (as in Rosenbluth potentials in Fokker-Planck operators) that are challenging to implement and computationally expensive to evaluate. In this work, we develop a data-driven neural network model that augments a simple and inexpensive BGK collision operator with a machine-learned correction term, which improves the fidelity of the simple operator with a small overhead to overall runtime. The composite collision operator has a tunable fidelity and, in this work, is trained using and tested against a direct-simulation Monte-Carlo (DSMC) collision operator.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Forecasting Multi-Step-Ahead Street-Scale Nuisance Flooding using a seq2seq LSTM Surrogate Model for Real-Time Application in a Coastal-Urban City

In coastal-urban cities facing an elevated risk of nuisance flooding (by rain and tide) due to increased heavy rainfall, sea level rise, urbanization, and aging drainage systems, real-time flood forecasting at the street-scale can provide useful information to transportation decision-makers. Physics-Based Models (PBMs) that offer high accuracy come with high computational runtimes and costs that limit their application for real-time flood forecasting. To address this challenge, Machine Learning (ML) surrogate models trained from PBMs have been proposed to provide street-scale flood forecasts. Previous related studies have focused on using Long Short-Term Memory (LSTM) architectures to model hourly flood depth on streets. While LSTM models can capture input sequences effectively, they fall short in accurately preserving output sequences, limiting their suitability for multi-step-ahead forecasts. The seq2seq LSTM architecture offers a key advantage here by capturing the full sequence of input–output, making it potentially more suitable for multi-step-ahead flood forecasts compared to traditional LSTM models. However, seq2seq LSTM has not been tested for street-scale flood forecasting, particularly for rapidly fluctuating nuisance flooding events which require special attention to its temporal sequences. Hence, in this study, we applied the seq2seq LSTM model to explore multi-step-ahead street-scale nuisance flooding and compared its results to the traditional LSTM model as a benchmark model. LSTM and seq2seq LSTM surrogate models were applied to 22 flood-prone streets in Norfolk, Virginia, as a case study with a 4-hr (short-term) and 8-hr (long-term) lead time. The models were trained with environmental (rainfall and tide) and topographic (elevation, Topographic Wetness Index, and Depth-To-Water) features along with PBM-derived water depths for different storm events. The results demonstrated satisfactory performance of both LSTM and seq2seq LSTM surrogate models throughout the forecast period compared to the PBM. However, the seq2seq LSTM showed lower Mean Absolute Error (MAE)/ Root Mean Square Error (RMSE) and higher Nash–Sutcliffe Efficiency (NSE)/ correlation than the LSTM across most lead times, particularly for long-term forecasting due to its supremacy in handling both input–output sequences together, which is missing in the traditional LSTM. For example, in the long-term, the average RMSE ranges were 0.0268–0.0373 m for LSTM and 0.0226–0.0319 m for seq2seq LSTM, while in the short-term, they were 0.0263–0.0293 m and 0.0261–0.0283 m, respectively. Additionally, while both models exhibited similar performance in distinguishing flooded and non-flooded streets for flood depth ≥ 0.1 m, the seq2seq LSTM model demonstrated superior performance for higher flood depths (such as ≥ 0.2 m and ≥ 0.3 m). Once trained, inference took only 0.09 to 0.11 s (short-term) and 0.30 to 0.35 s (long-term) per storm event for the 22 streets, making the application highly suitable for real-time decision-making during nuisance flood events.

54 ENVIRONMENTAL SCIENCES↗

Intercomparison of flood inundation models across land use types and hydrological flood stages

Flood Inundation Mapping (FIM) model selection is a key operational decision because accurate, rapid mapping underpins early warning and resource allocation. FIM performance is context-dependent and can vary with hydrograph phase, land-use/land-cover (LULC), and the evaluation benchmark. Intercomparison studies typically assess a single near-peak snapshot against one reference dataset. Here, we provide a context-stratified intercomparison across (i) multiple hydrograph phases, (ii) LULC classes, and (iii) benchmark types, for five FIM approaches spanning a wide range of physical complexity and operational cost (TRITON, LISFLOOD-FP, HEC-RAS 2D, ARC-Curve2Flood, and OWP HAND-FIM). We use the Hurricane Matthew flood (2016) in the Neuse River Basin, North Carolina, USA, as a case study. Using high-resolution remote sensing-derived flood inundation maps, hand-labeled points, and building footprints, we assess model skill across two rising and two falling hydrograph limbs and across major LULC types. Results show that model rankings shift systematically across contexts: LISFLOOD-FP ranks highest in three of four flood phases, while TRITON leads during one rising limb phase; LISFLOOD-FP performs best in vegetated areas, whereas HEC-RAS improves relative performance in agricultural and urban areas; and benchmark choice influences conclusions, with LISFLOOD-FP performing best for flooded-building detection in the late falling limb, while TRITON ranks highest against hand-labeled points. We also report representative wall-clock runtimes for each workflow to provide use-case context for operational feasibility. Together, these results offer transferable guidance for model selection and for designing large-scale, benchmark-aware FIM intercomparison studies.

Nikrou, Parvaneh [University of Alabama]↗

Photon detector response function methodology using MCNP and shift hybrid radiation transport code for wide-area contamination assay applications

Here, radiation transport modeling using the Monte Carlo N-Particle (MCNP) radiation transport code and Monte Carlo code, Shift, were employed to model detector responses for a variety of wide-area photon contamination scenarios. In this study, 2" × 2" and 3" × 3" cylindrical NaI(Tl) scintillation detector configurations at source detector-distances of 0.5 cm, 1 cm, 2.54 cm, 10 cm, and 30 cm were modeled. Media of soil, concrete, and steel were evaluated for contamination depths ranging from surface to a depth of an infinite thickness in each medium for photon energies ranging from 20 keV to 3 MeV, which correspond to the energies that current detectors can discern. Monoenergetic photon surface contamination detector responses for each of the media, source–detector distances, and detectors were estimated using MCNP v6.2. Shift was harnessed for improved variance reduction of particle transport in highly attenuating media to obtain average cell fluxes in the two MCNP NaI(Tl) scintillation detector configurations. Average cell flux values in Shift were coupled with detector responses from MCNP to convert average cell flux in a void to energy distribution of pulses in the NaI(Tl) scintillation detector crystal of interest. An optimized detector response function methodology was developed by coupling the MCNP radiation transport method with the Consistent Adjoint Driven Importance Sampling (CADIS) hybrid radiation transport method built into Shift to significantly decrease the runtime of thousands of MCNP pulse height simulations. The methodology may be utilized to quickly and accurately facilitate the assessment of a broad range of wide-area environmental contamination assay and decommissioning cleanup applications.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

A flexible data acquisition system architecture for the Nab experiment

Here, the Nab experiment will measure the electron–neutrino correlation and Fierz interference term in free neutron beta decay to test the Standard Model and probe Beyond the Standard Model physics. Using National Instrument’s PXIe-5171 Reconfigurable Oscilloscope module, we have developed a data acquisition system that is not only capable of meeting Nab’s specifications, but flexible enough to be adapted in situ as the experimental environment dictates. The L1 and L2 trigger logic can be reconfigured to optimize the system for coincidence event detection at runtime through configuration files and LabVIEW controls. This system is capable of identifying L1 triggers at a rate of at least 1 MHz, while reading out a peak signal rate of approximately 2 GB/s. During the commissioning phase of the experiment, the system ran at a sustained readout rate of 400 MB/s of detector signal data originating from roughly 6 kHz L2 triggers, well within the peak performance of the system.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Multiphysics for nuclear energy applications using a cohesive computational framework

With the recent development of advanced numerical algorithms, software design, and low-cost high-performance computer hardware, reliance on coupled multiphysics to predict the behavior of complex physical systems is beginning to become standard practice. This is especially true in nuclear energy applications where strong nonlinear interdependencies exist between reactor physics, radiation transport, multi-scale nuclear fuels performance, thermal fluids, etc. Resolving these nonlinear dependencies requires choices in multiphysics software approaches. Two main multiphysics modeling and simulation approaches have emerged. The first is based upon "code coupling" where disparate physics codes of different software design, code languages, and spatial and temporal integration schemes are coupled together with relatively complex data passing interfaces. The second multiphysics software approach is to employ a "cohesive" framework where all physics applications are developed with a common software design, i.e., data structures, syntax, input format, integrated spatial and temporal discretization schemes, etc. In this paper we present the Multiphysics Object-Oriented Simulation Environment (MOOSE) development and runtime framework and describe the framework's cohesive modeling and simulation multiphysics approach. Then, a "cohesive-like" extension of the MOOSE framework is presented where MOOSE-based physics software applications are efficiently coupled to non-MOOSE (external) physics codes to form multiphysics applications using MOOSE's unique interface capabilities. Finally, several examples of MOOSE's cohesive and cohesive-like multiphysics applications will be demonstrated. These multiphysics demonstrations will incorporate both MOOSE-based applications and external codes, including Nek5000, RELAP-7, TRACE, BISON, and Pronghorn.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

HARD: A performance portable radiation hydrodynamics code based on FleCSI framework

Hydrodynamics And Radiation Diffusion (HARD) is an open-source application for high-performance simulations of compressible hydrodynamics with radiation-diffusion coupling. Built on the FleCSI (Bergen et al., 2021 [1]) (Flexible Computational Science Infrastructure) framework, HARD expresses its computational units as tasks whose execution can be orchestrated by multiple back-end runtimes, including Legion (Bauer et al., 2012 [2]), MPI (Forum, 1994 [3]), and HPX (Kaiser et al., 2020 [4]). Node-level parallelism is handled through Kokkos (Edwards et al., 2014 [5]), providing a single-source, portable code base that runs efficiently on laptops, small homogeneous clusters, and the largest heterogeneous supercomputers currently available. To ensure scientific reliability, HARD includes a regression test suite that automatically reproduces canonical verification problems such as the Sod and LeBlanc shock tubes, and the Sedov blast wave, comparing numerical solutions against known analytical results. The project is distributed under an OSI-approved license, hosted on GitHub, and accompanied by reproducible build scripts and continuous integration workflows. This combination of performance portability, verification infrastructure, and community-focused development makes HARD a sustainable platform for advancing radiation hydrodynamics research across multiple domains.

97 MATHEMATICS AND COMPUTING↗

Enabling Multireference Calculations on Multimetallic Systems with Graphic Processing Units

Modeling multimetallic systems efficiently enables faster prediction of desirable chemical properties and the design of new materials. This work describes an initial implementation for performing multireference wave function method localized active-space self-consistent field (LASSCF) calculations through the use of multiple graphics processing units (GPUs) to accelerate time-to-solution. Density fitting is leveraged to reduce memory requirements, and we demonstrate the ability to fully utilize multi-GPU compute nodes. Performance improvements of 5–10x in total application runtime were observed in LASSCF calculations for multimetallic catalyst systems up to 1200 AOs and an active space of (22e,40o) using up to four NVIDIA A100 GPUs. Furthermore, written with performance portability in mind, a comparable performance is also observed in early runs on the Aurora exascale system using Intel Max Series GPUs.

Algorithms↗

A Scalable Semi-Implicit Barotropic Mode Solver for the MPAS-Ocean

A scalable semi-implicit barotropic mode solver for the ocean component of the model for prediction across scales has been implemented as a competitor to an existing explicit-subcycling scheme to allow faster and more stable simulations while not sacrificing accuracy. The semi-implicit solver adopts the pipelined preconditioned bi-conjugate gradient stabilization algorithm as an iterative solver in conjunction with the restricted additive Schwarz preconditioner that accelerates the convergence rate of the iterative solver. The preconditioner is constructed from a linearized barotropic system that also reorders the system for optimal performance, while the semi-implicit solver deals with the fully nonlinear barotropic system that requires reassembly of the coefficient matrix for every time step. Several numerical experiments, from simple one-dimensional tests to three-dimensional real-world tests, demonstrate that the semi-implicit solver has almost the same accuracy and better parallel scalability compared with the existing scheme while allowing faster and more stable simulations. Furthermore, the semi-implicit solver accelerates the barotropic mode up to 2.9 times faster than the existing scheme on 16,320 processors, leading to an overall runtime speedup of 1.9.

97 MATHEMATICS AND COMPUTING↗