Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Global optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

Multi-objective Bayesian active learning for MeV-ultrafast electron diffraction

Ultrafast electron diffraction using MeV energy beams(MeV-UED) has enabled unprecedented scientific opportunities in the study of ultrafast structural dynamics in a variety of gas, liquid and solid state systems. Broad scientific applications usually pose different requirements for electron probe properties. Due to the complex, nonlinear and correlated nature of accelerator systems, electron beam property optimization is a time-taking process and often relies on extensive hand-tuning by experienced human operators. Algorithm based efficient online tuning strategies are highly desired. Here, we demonstrate multi-objective Bayesian active learning for speeding up online beam tuning at the SLAC MeV-UED facility. The multi-objective Bayesian optimization algorithm was used for efficiently searching the parameter space and mapping out the Pareto Fronts which give the trade-offs between key beam properties. Such scheme enables an unprecedented overview of the global behavior of the experimental system and takes a significantly smaller number of measurements compared with traditional methods such as a grid scan. This methodology can be applied in other experimental scenarios that require simultaneously optimizing multiple objectives by explorations in high dimensional, nonlinear and correlated systems.

43 PARTICLE ACCELERATORS↗

Improving Prediction of Peroxide Value of Edible Oils Using Regularized Regression Models

We present four unique prediction techniques, combined with multiple data pre-processing methods, utilizing a wide range of both oil types and oil peroxide values (PV) as well as incorporating natural aging for peroxide creation. Samples were PV assayed using a standard starch titration method, AOCS Method Cd 8-53, and used as a verified reference method for PV determination. Near-infrared (NIR) spectra were collected from each sample in two unique optical pathlengths (OPLs), 2 and 24 mm, then fused into a third distinct set. All three sets were used in partial least squares (PLS) regression, ridge regression, LASSO regression, and elastic net regression model calculation. While no individual regression model was established as the best, global models for each regression type and pre-processing method show good agreement between all regression types when performed in their optimal scenarios. Furthermore, small spectral window size boxcar averaging shows prediction accuracy improvements for edible oil PVs. Best-performing models for each regression type are: PLS regression, 25 point boxcar window fused OPL spectral information RMSEP = 2.50; ridge regression, 5 point boxcar window, 24 mm OPL, RMSEP = 2.20; LASSO raw spectral information, 24 mm OPL, RMSEP = 1.80; and elastic net, 10 point boxcar window, 24 mm OPL, RMSEP = 1.91. The results show promising advancements in the development of a full global model for PV determination of edible oils.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Design optimization of integrated cooling inserts in modular Fischer-Tropsch reactors

Sustainable production of liquid fuels and feedstocks from atmospheric CO 2 through carbon recycling technologies is necessary to broaden decarbonization efforts and further reduce global emissions. Fischer–Tropsch (FT) reactors can address this challenge by providing storable, high-value liquid hydrocarbon fuels and feedstocks from syngas generated through reductive CO 2 utilization. FT technology, however, is most cost effective at large-scale, fixed-site plants and does not effectively address the smaller, globally distributed CO 2 point sources. Alternatively, smaller, modular reactors can be composed together to the specific scale of the emission source, yielding a flexible solution that enables more widespread deployment. In these modular reactors, thermal management using a finned cooling insert embedded within the catalyst matrix is critical to maintaining the performance and viability. Thus, in this work, topology optimization is used to determine the cooling insert geometry that maximizes reactor productivity while preventing auto-thermal runaway. Optimal designs are generated for varying number of constituent fins and over a range of maximum operating temperatures. Constraints including minimum feature length-scales and prescribed cooling insert material are imposed on the designs to enhance manufacturability. The impact of design features such as insert tapering and increased length scale hierarchy is automatically revealed by the systematic design framework employed. Here, this approach thus generates novel optimal geometries while automating, accelerating, and enhancing the design process compared to traditional heuristic approaches for cooling insert design in modular FT reactors.

30 DIRECT ENERGY CONVERSION↗

Repository of HydroSMADE: Hydropower Site-level Monthly Availability Data Ensemble for 1950-2100 at Existing and Potential Global Sites

This repository presents HydroSMADE—Hydropower Site-level Monthly Availability Data Ensemble, a new open dataset that provides monthly hydropower availability for 1,593 existing and 124,333 potential sites worldwide over the period 1950–2100. The dataset is generated by using a global hydrologic model (Xanthos) with explicit representation of hydropower operation. Specifically, HydroSMADE distinguishes between storage and diversion sites, applies optimized operating rules, and incorporates site-specific characteristics such as generation capacity, maximum turbine flow, and reservoir storage. Driven by bias-corrected meteorological inputs, the data is provided for 30 alternative future scenarios. The scenarios consist of the full factorial combination of three standard CMIP6 atmospheric forcing pathways (SSP1-2.6, SSP3-7.0, and SSP5-8.5) and ten CMIP6 General Circulation Models (GCMs): GFDL-ESM4, IPSL-CM6A-LR, MPI-ESM1-2-HR, MRI-ESM2-0, EC-Earth3, CanESM5, MIROC6, CNRM-ESM2-1, UKESM1-0-LL, and CNRM-CM6-1. The repository contains a total of 122 files: a text file (readme.txt) containing a brief description of the included data, a CSV file containing site attributes, and the remaining 120 files (in CSV) containing site-level monthly hydropower availability. Example Jupyter Notebooks to explore the HydroSMADE dataset are available on GitHub at https://github.com/kamal0013/HydroSMADE More details on the methods and technical validation of HydroSMADE are available in the following paper by the same authors: Chowdhury, A. K., Abeshu, G. W., Zhao, M., Wild, T. B., Hassan, N., Ying, Z., Kim, G. J., Matthew, B., Jonathan, L., & Li, H.-Y. (Submitted). Hydropower Site-level Monthly Availability Data Ensemble for 1950-2100 at Existing and Potential Global Sites.

Existing and Potential Sites↗

ORNL_AISD_NiNb

This dataset describes the nickel-niobium solid solution binary alloy, where the two constituent elements nickel (Ni) and niobium (Nb) are randomly placed on an underlying crystal lattice. This dataset for nickel-niobium (Ni-Nb) alloys available includes the formation energy and bulk modulus for each crystal structure. Each atomic sample has a disordered phase which is obtained starting from an initial regular crystal structure of type body-centered cubic (BCC), face-centered cubic (FCC), or hexagonal compact packed (HCP). The geometry optimization ensures that all the alloy samples reached the equilibrium with negative formation energy. We perform geometry optimizations using the LAMMPS simulation package [1], a flexible simulation tool for particle-based materials modeling at the atomic, meso, and continuum scales. We utilized the embedded atom model (EAM) potential for Ni and Nb developed in a previous study [2]. The potential could describe behaviors of the liquid and solid phases of Ni-Nb alloy. The structural factors and angular distributions of three atoms are well-matched with X-ray and ab initio-based molecular dynamics data. We prepared the three different crystals with different initial lattice parameters (3.52 Ã… for FCC, 3.32 Ã… for BCC, and 3.5 Ã… for HCP). We performed energy minimization in two steps. Firstly, we minimized the structures with an isotropic unit cell to minimize the side effects from our arbitrary lattice parameters for all other compositions. Then, we applied geometry optimization with a triclinic (non-orthogonal) unit cell to fully minimize the stress components to calculate the elastic constants. In this procedure, we chose 10,000 as the maximum number of allowable steps aimed at obtaining fully relaxed atomic geometries. The dataset consists of three sets of crystal structures. The first set contains 46,086 irregular crystal structures, each of them with 54 atoms, obtained through optimization starting from a regular BCC crystal structure. The second set contains 24,543 irregular crystal structures, each of them with 32 atoms, obtained through optimization starting from a regular FCC crystal structure. The third set contains 39,303 irregular crystal structures, each of them with 48 atoms, obtained through optimization starting from a regular HCP crystal structure. The atomic configurations within each set span the possible compositional range. The three sets have been unified in a global dataset, which is extremely heterogeneous in terms of crystal structures, lattice volumes, and atomic configurations. Organization of files inside the dataset: the dataset contains three subdirectories called • BCC_opt • FCC_opt • HCP_opt based on the type of initial regular structure used to start the geometry optimization. Inside each of these folders, every atomic structure is identified by a string “A_B_Câ€, where A denotes the number of Nb in the system, B denotes index of structure with a given Nb number, and C denotes the total number of structures generated with a given Nb number. For each optimized crystal structure identified by the unique string of characters “A_B_Câ€, three files are provided: • A_B_C_opt.xyz: The optimized geometries in xyz format • A_B_C_opt.cfg: The optimized geometries in cfg format. It includes cell information and atomic energy, and forces calculated from LAMMPS. • A_B_C.elastic: Raw data of 21 elastic constants from LAMMPS output. • A_B_C.bulk: Calculated upper and lower bounds of bulk modulus and averaged one based on Voigt-Reuss-Hill approach from *.elastic. References: [1] A. P. Thompson, H. M. Aktulga, R. Berger, D. S. Bolintineanu, W. M. Brown, P. S. Crozier, P. J. in 't Veld, A. Kohlmeyer, S. G. Moore, T. D. Nguyen, R. Shan, M. J. Stevens, J. Tranchida, C. Trott, and S. J. Plimpton. LAMMPS - a flexible simulation tool for particle-based materials modeling at the atomic, meso, and continuum scales. Comp. Phys. Comm., 271:108171, 2022. [2] Y Zhang, R Ashcraft, MI Mendelev, CZ Wang, and KF Kelton. Experimental and molecular dynamics simulation study of structure of liquid and amorphous ni62nb38 alloy. The Journal of chemical physics, 145(20):204505, 2016.

36 MATERIALS SCIENCE↗

A proximal trust-region method for nonsmooth optimization with inexact function and gradient evaluations

Many applications require minimizing the sum of smooth and nonsmooth functions. For example, basis pursuit denoising problems in data science require minimizing a measure of data misfit plus an $\ell^1$-regularizer. Similar problems arise in the optimal control of partial differential equations (PDEs) when sparsity of the control is desired. Here, we develop a novel trust-region method to minimize the sum of a smooth nonconvex function and a nonsmooth convex function. Our method is unique in that it permits and systematically controls the use of inexact objective function and derivative evaluations. When using a quadratic Taylor model for the trust-region subproblem, our algorithm is an inexact, matrix-free proximal Newton-type method that permits indefinite Hessians. We prove global convergence of our method in Hilbert space and demonstrate its efficacy on three examples from data science and PDE-constrained optimization.

97 MATHEMATICS AND COMPUTING↗

Effect of epitope variant co-delivery on the depth of CD8 T cell responses induced by HIV-1 conserved mosaic vaccines

To stop the HIV-1 pandemic, vaccines must induce responses capable of controlling vast HIV-1 variants circulating in the population as well as those evolved in each individual following transmission. Numerous strategies have been proposed, of which the most promising include focusing responses on the vulnerable sites of HIV-1 displaying the least entropy among global isolates and using algorithms that maximize vaccine match to circulating HIV-1 variants by vaccine cocktails of optimized complementing sequences. In this study, we investigated CD8 T cell responses induced by a bi-valent mosaic of highly conserved HIVconsvX regions delivered by a combination of simian adenovirus ChAdOx1 and poxvirus MVA. We compared partially and fully mono- and bi-valent prime-boost regimens and their ability to elicit T cells recognizing natural epitope variants using an interferon-g enzyme-linked immunospot (ELISPOT) assay. We used 11 well-defined CD8 T cell epitopes in two mouse haplotypes and, for each epitope, assessed recognition of the two vaccine forms together with the other most frequent epitope variants in the HIV-1 database. We conclude that for the magnitude and depth of epitope recognition, CD8 T cell responses benefitted in most comparisons from the combined bi-valent mosaic and envisage the main advantage of the bi-valent vaccine during its deployment to diverse populations.

59 BASIC BIOLOGICAL SCIENCES↗

BEYONDPLANCK III. Commander3

We describe the computational infrastructure for end-to-end Bayesian cosmic microwave background (CMB) analysis implemented by the BeyondPlanck Collaboration. The code is called Commander3. It provides a statistically consistent framework for global analysis of CMB and microwave observations and may be useful for a wide range of legacy, current, and future experiments. The paper has three main goals. Firstly, we provide a high-level overview of the existing code base, aiming to guide readers who wish to extend and adapt the code according to their own needs or re-implement it from scratch in a different programming language. Secondly, we discuss some critical computational challenges that arise within any global CMB analysis framework, for instance in-memory compression of time-ordered data, fast Fourier transform optimization, and parallelization and load-balancing. Thirdly, we quantify the CPU and RAM requirements for the current BEYONDPLANCK analysis, finding that a total of 1.5 TB of RAM is required for efficient analysis and that the total cost of a full Gibbs sample for LFI is 170 CPU-hrs, including both low-level processing and high-level component separation, which is well within the capabilities of current low-cost computing facilities. The existing code base is made publicly available under a GNU General Public Library (GPL) license.

79 ASTRONOMY AND ASTROPHYSICS↗

Improving ideal MHD equilibrium accuracy with physics-informed neural networks

We present a novel approach to compute three-dimensional magnetohydrodynamic equilibria with isotropic pressure profiles and nested surfaces by parametrizing Fourier modes with artificial neural networks (NNs). The full nonlinear global force residual of single equilibria across the volume in real space is then minimized with first order optimizers and compared to equilibria computed by conventional solvers. Already, we observe competitive computational cost to arrive at the same minimum residuals computable with existing codes. With increased computational cost, lower minima of the residual are computable with the NNs than with any other tested solver, establishing a new lower bound for the force residual. We use minimally complex NNs, and we expect significant improvements for solving not only single equilibria with NNs, but also for creating NN models valid over continuous distributions of equilibria.

ideal magnetohydrodynamics↗

Constrained Bayesian optimization of criticality experiments

The design of criticality experiments is typically an iterative process that employs a Monte Carlo transport code. The goal is to find a design that optimizes some variable, like the sensitivity of a response to a cross section, while simultaneously ensuring criticality. The high fidelity of the Monte Carlo code is a great asset, but it makes exploring the design space computationally expensive. Herein, we present how a constrained Bayesian optimization algorithm can be used to efficiently design a criticality experiment. It uses Gaussian processes as a surrogate model to probe the design space and to reduce the number of code executions that are needed to find the optimum. Furthermore, we demonstrate constrained Bayesian optimization with a Pu-239/polyethylene solution system and a TEX experiment that is designed for criticality safety validation of a nuclear waste model at the Hanford Site. For both systems, a global optimum was found within 75 Monte Carlo simulations.

42 ENGINEERING↗

Operational optimization for multi-functional charging station with electric and hydrogen-powered vehicles

The rapid adoption of electric vehicles (EVs) and hydrogen fuel cell vehicles (HFCVs), combined with global efforts to reduce carbon emissions, has accelerated the development of EV charging and hydrogen refueling stations. In response to this demand, this paper introduces the concept of Multi-Functional Charging Station (MFCS) that integrates power generation, EV charging, battery swapping, and hydrogen refueling. A comprehensive operational model is developed for the MFCS that couples electricity and hydrogen conversion and storage technologies to enhance infrastructure utilization and improve overall system efficiency. The model also considers multiple revenue streams, including participation in energy and ancillary markets. To validate the effectiveness of the proposed model and evaluate its performance, a series of numerical experiments are conducted with different charger numbers, different electricity purchase limits, and different charger allocations. Numerical results demonstrate that shared charger configurations can lead to 8.11 % improvement in operational profit by improving resource utilization and reducing the number of depleted batteries at the end of operations compared to allocated charger setups. By varying the number of chargers, sensitivity analysis identifies diminishing marginal returns beyond about 45 chargers, suggesting it as an optimal sizing point under current settings. The integration of electricity and hydrogen conversion is also explored under scenarios with limited external electricity purchases. In conclusion, these findings indicate that optimizing charger allocation and energy management can significantly enhance station productivity and profitability, ultimately supporting the broader adoption of electrified and hydrogen-based transportation solutions.

Charging station↗

Evaluation of Thermolytic Hydrogen Generation Rate Models at High-Temperature/High-Hydroxide Regimes

This report describes the results of testing performed to extend the applicable ranges of temperature and hydroxide concentration for use within the Glycolate and Global Total Organic Carbon (TOC) Hydrogen Generation Rate (HGR) expressions. Seven experimental conditions (six simulants of the 242-25H Evaporator system chosen as a D-optimal set of experiments and a single test conducted at an elevated boiling point of 170 °C) were investigated in the presence of sodium glycolate and Xiameter TM AFE-1010. Glycolate was employed to study the extension of the Glycolate Thermolytic HGR expression while Xiameter TM AFE-1010 was employed to study the extension of the Global TOC Thermolytic HGR expression. The following conclusions were derived from this testing: The Glycolate Thermolytic HGR expression may be confidently used to predict thermolytic HGRs from glycolate at temperatures as high as 170 °C and hydroxide concentrations as high as 23 M.; The hydroxide and temperature-dependence predicted by the Global TOC Thermolytic HGR expression has been confirmed at temperatures as high as 170 °C and hydroxide concentrations as high as 23 M, suggesting that the Global TOC Thermolytic HGR expression may be used at these ranges.; Methane was observed from tests with Xiameter TM AFE-1010 at production rates higher than those observed for hydrogen. These rates were observed at temperatures higher than 100 °C.; Preliminary models suggest that increasing hydroxide/temperature causes an increase in Methane Generation Rate (MGR) from Xiameter TM AFE-1010. The following recommendations are based on this testing: The existing equations for thermolytic HGR from glycolate and non-glycolate organics should be used at Concentration, Storage, and Transfer Facilities (CSTF) storage and evaporation conditions, including temperatures and hydroxide concentrations exhibited in the 242-25H Evaporator.; Further investigation should be made into the influence of methylsilanes on CSTF flammability. This investigation should include: determination of the types of methylsilanes historically added to the CSTF, determination of methane formation rates from each type of methylsilane, and determination of the extent of degradation of methylsilanes in CSTF waste.; Characterization techniques should be developed by Savannah River National Laboratory (SRNL) to assist in the speciation of methylsilane-containing waste in the CSTF.; Additional testing with radioactive waste should be performed to determine the MGRs possible in radioactive waste and better inform model predictions made from testing with simulants.

08 HYDROGEN↗

Environmental and socio-economic Pareto-front trade-off analysis of U.S. PET packaging material in a circular economy

Various recycling technologies are emerging to implement circular economy in plasticssupply chain systems. However, the environmental and socio-economic trade-offs of in circular economy are not well understood at a systems level. Particularly, quantifying these trade-offs as a function of end-of-life (EOL) management decisions, including transition of recycling technologies, systems level metrics such as circularity, recycled content, and the need for fossil-derived plastics are not well understood. Here, the present study addressed these research gaps by applying a systems analysis modeling approach that utilizes material flow analysis, life cycle assessment, socioeconomic data, and system optimization techniques for polyethylene terephthalate (PET) packaging supply chains in the United States. Pareto-front trade-offs between conflicting environmental and socio-economic impacts as well as those between socioeconomic impacts and circularity were explored using the epsilon constraint method. The Pareto-front trade-off analysis revealed the transition of EOL management strategies for PET packaging systems, including changes in selection of recycling technologies, to aid decision making process by quantifying studied system metrics. Transitioning from environmentally optimal to socio-economically optimal systems led to increased employment (by 17%), wages (by 26%), and revenues (by 6%) but also led to increased global warming potential (GWP; by 65%), energy consumption (by 59%), and reliance on fossil PET in the system (by 78%). Finally, the results show that there is not a unique set of recycling technologies to achieve a sustainable circular economy of PET packaging system, instead it depends on the decision maker’s objectives and targeted metrics of the system.

54 - ENVIRONMENTAL SCIENCES/GLOBAL CLIMATE CHANGE ↗

On optimal control of hybrid dynamical systems using complementarity constraints

Optimal control for switch-based dynamical systems is a challenging problem in the process control literature. In this study, we model these systems as hybrid dynamical systems with finite number of unknown switching points and reformulate them using non-smooth and non-convex complementarity constraints as a mathematical program with complementarity constraints (MPCC). We utilize a moving finite element based strategy to discretize the differential equation system to accurately locate the unknown switching points at the finite element boundary and achieve high-order accuracy at intermediate non-collocation points. We propose a globalization approach to solve the discretized MPCC problem using a mixed NLP/MILP-based strategy to converge to a non-spurious first-order optimal solution. The method is tested on three dynamic optimization examples, including a gas–liquid tank model and an optimal control problem with a sliding mode solution.

97 MATHEMATICS AND COMPUTING↗

Exploring the Meta-regulon of the CRP/FNR Family of Global Transcriptional Regulators in a Partial-Nitritation Anammox Microbiome

Microbiomes are important contributors to many ecosystems, including ones where nutrient cycling is stimulated by aeration control. Optimizing cyclic aeration helps reduce energy needs and maximize microbiome performance during wastewater treatment; however, little is known about how most microbial community members respond to these alternating conditions.

transcription factors↗

Conservative Numerical Schemes with Optimal Dispersive Wave Relations: Part II. Numerical Evaluations

A new energy and enstrophy conserving scheme (EEC) for the shallow water equations is proposed and evaluated using a suite of test cases over the global spherical or bounded domain. The evaluation is organized around a set of pre-defined properties: accuracy of individual operators, accuracy of the whole scheme, conservation of key quantities, control of the divergence variable, representation of the energy and enstrophy spectra, and simulation of nonlinear dynamics. The results confirm that the scheme is between the first and second order accurate, and conserves the total energy and potential enstrophy up to the time truncation errors. Here, the scheme is capable of producing more physically realistic energy and enstrophy spectra, indicating that it can help prevent the unphysical energy cascade towards the finest resolvable scales. With an optimal representation of the dispersive wave relations, the scheme is able to keep the flow close to being non-divergent, and maintain the geostrophically balanced structures with large-scale geophysical flows over long-term simulations.

54 ENVIRONMENTAL SCIENCES↗

Convergence of micro-geochemistry and micro-geomechanics towards understanding proppant shale rock interaction: A Caney shale case study in southern Oklahoma, USA

As a direct outcome of economic development coupled with an increase in population, global energy demand will continue to rise in the coming decades. Although renewable energy sources are increasingly investigated for optimal production, the immediate needs require focus on energy sources that are currently available and reliable, with a minimal environmental impact; the efficient exploration and production of unconventional hydrocarbon resources is bridging the energy needs and energy aspirations, during the current energy transition period. The main challenges are related to the accurate quantification of the critical rock properties that influence production, their heterogeneity and the multiscale driven physico-chemical nature of rock–fluid interactions. A key feature of shale reservoirs is their low permeability due to dominating nanoporosity of the clay-rich matrix. As a means of producing these reservoirs in a cost-effective manner, a prerequisite is creation of hydraulic fracture networks capable of the highest level of continued conductivity. Fracturing fluid chemical design, formation brine geochemical composition, and rock mineralogy all contribute to swelling-induced conductivity damage. The Caney Shale is an organic-rich, often calcareous mudrock. Many studies have examined the impact that clay has on different kinds of shale productivity but there is currently no data reported on the Caney Shale in relation to horizontal drilling; all reported data on the Caney Shale is on vertical wells which are shallow, compared to an emerging play that is at double the depth. Here, in this work we develop geochemical–geomechanical integration of rock properties at micro-and nanoscales that can provide insights into the potential proppant embedment and its mitigation. The novel methodology amalgamates the following: computed X-ray tomography, scanning electron microscopy, energy dispersive spectroscopy, micro-indentation, and Raman spectroscopy techniques. Here, our results show that due to the multiscale heterogeneity in the Caney Shale, these geochemical and structural properties translate into a variation in mechanical properties that will impact interaction between the proppant and the host shale rock.

03 NATURAL GAS↗

An implicit barotropic mode solver for MPAS-ocean using a modern Fortran solver interface

Here, we demonstrate use of a modern Fortran solver interface to manage solver algorithms for an implicit barotropic mode solver in the Model for Predictions Across Scales-Ocean (MPAS-O). ForTrilinos, a Fortran interface to Trilinos that contains a large collection of solver capabilities written in C++, has been implemented in MPAS-O to provide access to a suite of linear solver options. By virtue of the simplified wrapper and interface generator (SWIG) automation tool that generates modern Fortran interfaces to C++ code, we were able to implement the Fortran solver interface in MPAS-O using a familiar Fortran coding style while minimizing performance degradation. The ForTrilinos solver interface is written within MPAS-O’s time stepping modules as a subroutine in conjunction with MPAS-O code. Applied to an idealized ocean and a high-resolution realistic ocean test case, parallel performance of ForTrilinos solvers is examined. It is found that parallel scalability of the ForTrilinos solvers is highly dependent on the number of global synchronization points per solver iteration in each iterative solver algorithm. ForTrilinos solvers perform best compared to the Fortran hand-crafted (FHC) solver when the amount of work per processor is large enough. However, parallel scalability is better with the FHC solver and so when the work per core is modest FHC outperforms ForTrilinos. The intercomparison between the ForTrilinos and FHC solvers reveals that this performance hit in the ForTrilinos solver mostly comes from the global synchronization process, while suggesting that the matrix-vector multiplication process in the FHC solver needs to be optimized for better performance.

97 MATHEMATICS AND COMPUTING↗