Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “implicit methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Spectrally Stabilized Interface Capturing Formulation and Implementation in Nek5000/NekRS

This report documents the formulation of a novel level-set method for incompressible two-phase flows in the continuous Galerkin (CG) high order spectral element framework. The overall method hinges on a novel implementation of the spectral vanishing viscosity (SVV) operator for the stabilization of linear/non-linear hyperbolic problems. The multidimensional SVV convolution kernels, which in essence, have a similar effect as a high pass filter applied to the derivatives, are formulated by exploiting the tensor product form, analogous to the construction of the usual stiffness matrix system. The resulting kernels are directionally decoupled and ensure a linear, symmetric positive definite, elliptic matrix operator. The SVV formulation is demonstrated to provide a robust stabilizing mechanism through challenging linear and non-linear hyperbolic problems, including problems pertinent to the level-set formulation. The two-phase framework conceptualized herein is based on the conservative level-set (CLS) method which represents the interface between the fluids by the 0.5 iso-contour of the smoothed Heaviside function. The CLS method is augmented with a preconditioning procedure for interface normals using the signed distance function which precludes the manifestation of spurious oscillations in the vicinty of the interface. Further, the existing mixed explicit-implicit approach for the solution of Navier-Stokes equations in Nek5000, as described in Tomboulides et al, is augmented with a pressure coefficient splitting approach for the Poisson equation, which greatly accelerated the convergence of pressure solver for two-phase systems with large density ratio. The robustness and accuracy of the overall two-phase method is demonstrated through canonical challenging problems involving high density and viscosity ratios, with and without surface tension. The two-phase formulation is wholly implemented in Nek5000 and the SVV stabilization method is implemented in NekRS, which is the essential precursor to the two-phase framework, undergoing active development.

97 MATHEMATICS AND COMPUTING↗

MINE: maximally informative next experiment—toward a new GWAS experimental design and methodology

Abstract The computational methodology of Genome Wide Association Studies (GWAS) currently has several limitations: (i) the number of observations (rows) on a quantitative trait tends to be smaller than the number of single nucleotide polymorphisms (SNPs) (columns) in the design matrix; (ii) each SNP is usually modeled separately, failing to acknowledge interaction between each other (ie epistasis); (iii) there is implicit linkage disequilibrium (LD) between neighboring SNPs due to their linkage. To overcome these issues, we developed a tool that uses ensemble methods to fit mixed linear models to GWAS data, and these ensemble methods include the development of a new experimental design approach in GWAS, which uses the resultant models and data to select the next informative experiment over time. This new adaptive and staged approach for GWAS experimental design was developed and tested in a 3 yr adaptive model-guided discovery experiment against a fixed classical design. In Sorghum bicolor a total of 79, 86, and 78 accessions were tested in years 1, 2, and 3, respectively out of 343 accessions available in the Bioenergy Association Panel (BAP) each identified for 232,303 SNPs, 1 every 2–3 kb in the genomes. We demonstrated the feasibility of MINE enacted with 8 people in the field per year over 3 yr vs in 1 large classical design enacted with 20 people in 1 yr. The MINE results for chromosomal regions identified controlling dry weight were confirmed against results from previous sorghum GWAS experiments and 1 large classical design for the BAP panel.

Genetics & Heredity↗

Buffered freepointer management memory system

A system and method of buffered freepointer management to handle burst traffic to fixed size structures in an external memory system. A circular queue stores implicitly linked free memory locations, along with an explicitly linked list in memory. The queue is updated at the head with newly released locations, and new locations from memory are added at the tail. When a freed location in the queue is reused, external memory need not be updated. When the queue is full, the system attempts to release some of the freepointers such as by dropping them if they are already linked, updating the linked list in memory only if those dropped are not already linked. Latency can be further reduced by loading new locations from memory when the queue is nearly empty, rather than waiting for empty condition, and by writing unlinked locations to memory when the queue is nearly full.

Jacob, Philip↗

Robust Adaptive Decentralized Dynamic State Estimation with Unknown Control Inputs using Field PMU Measurements

This paper proposes a robust adaptive decentralized dynamic state estimation method for power system with unknown inputs of the highly detailed synchronous machine model. The temporal and spatial correlations among the unknown inputs are used to derive a vector auto-regressive model. The latter is further integrated together with state transition and measurement models for joint state and unknown inputs estimation. Thanks to the consideration of implicit cross-correlations between the states and the unknown inputs, only generator terminal voltage and current phasors are needed. Test results on the US WECC system using the field PMU measurements show that the proposed method is able to track both the system dynamic states and unknown controller inputs. These information could significantly benefit the validation and calibration of generator controller parameters.

Zhao, Junbo↗

BISON Robustness and Performance Improvements

BISON is a modern finite-element based nuclear fuel performance code that has been under development at the Idaho National Laboratory (USA) since 2009 [1]. The code is applicable to both steady and transient fuel behavior and can be used to analyze 1D (spherically symmetric), 2D (axisymmetric and generalized plane strain) or 3D geometries. BISON is the fuel performance code used within CASL for LWR fuel under both normal operating and accident conditions. BISON is built using the INL Multiphysics ObjectOriented Simulation Environment, or MOOSE [2, 3]. MOOSE is a massively parallel, finite element-based framework to solve systems of coupled non-linear partial differential equations using the Jacobian-Free Newton Krylov (JFNK) method [4]. This enables investigation of computationally large problems, for example a full stack of discrete pellets in a LWR fuel rod, or every rod in a full reactor core. MOOSE supports the use of complex two and three-dimensional meshes and uses implicit time integration, important for the widely varied time scale in nuclear fuel simulation. An object-oriented architecture is employed which greatly minimizes the programming effort required to add new material and behavioral models. The flexibility of the implicit and fully coupled multiphysics approach comes with a need for constructing suitable approximations for the Jacobian matrix of the coupled system used for either preconditioning a Krylov solve or in a direct Newton solve. Preconditioning options for Bison problems need to be revisited with new preconditioning methods becoming available.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Fast solvers for tokamak fluid models with PETSc

Multigrid (MG) is widely recognized as a highly effective solver for the model problem, the Laplacian, but textbook MG fails on most problems of interest. MG methods have been applied to complex, real-world applications with careful consideration of the physical model and discretization. In this work we develop the first step in applying MG methods to science and engineering relevant magnetohydrodynamics (MHD) tokamak models in the M3D-C1 (https://m3dc1.pppl.gov) fusion energy science code. The semi-implicit time integrator in M3D-C1 is composed of many linear solves. The implicit advance of the momentum equation is the most challenging and is the focus of this work. The current production solver in M3D-C1 is a block Jacobi (BJ) preconditioner within a Krylov solver, where blocks group degrees of freedom on planes of constant toroidal coordinate. BJ convergence degrades as the number of planes increases due to the spectral properties of the matrix preconditioned with BJ. The partially magnetic field-aligned, regular toroidal grid structure in M3D-C1 is amenable to semi-coarsening geometric MG in the toroidal direction. This paper develops such a solver and demonstrates competitive performance on a runaway electron model of a SPARC (https://cfs.energy/technology/sparc) disruption, and superior robustness on a stellarator model on which the BJ solver fails to converge.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A fast implicit solver for semiconductor models in one space dimension

Several different approaches are proposed for solving fully implicit discretizations of a simplified Boltzmann-Poisson system with a linear relaxation-type collision kernel. This system models the evolution of free electrons in semiconductor devices under a low-density assumption. At each implicit time step, the discretized system is formulated as a fixed-point problem, which can then be solved with a variety of methods. A key algorithmic component in all the approaches considered here is a recently developed sweeping algorithm for Vlasov-Poisson systems. A synthetic acceleration scheme has been implemented to accelerate the convergence of iterative solvers by using the solution to a drift-diffusion equation as a preconditioner. The performance of four iterative solvers and their accelerated variants has been compared on problems modeling semiconductor devices with various electron mean-free-path.

97 MATHEMATICS AND COMPUTING↗

A Simple, Scalable Large Deformation Solid Mechanics Implementation in the MOOSE Framework

This article describes a large deformation solid mechanics solver implemented as part of the freely available and open source MOOSE finite element simulation framework. The article documents the choices made in developing the solid mechanics framework and describes novel formulations for the gradient operator and constitutive modeling framework made to simplify implementations of different coordinate systems, stabilized gradient operators, and different constitutive model inputs and outputs. In the process, the article describes a new formulation that casts objective integration of the Cauchy stress as a linear transformation of the small stress rate. Finally, the article presents key implementation details and examines the parallel efficiency of the solid mechanics solver implemented in MOOSE. The implementation retains a good weak scaling efficiency beyond 1,000 parallel processes. The article includes a discussion of the factors limiting the parallel efficiency of implicit, large deformation solid mechanics codes on current high-performance computers, with the main current limitation being the scalability of the algebraic multigrid methods used to solve the linearized equilibrium equations.

Applied computing → Computer-aided design↗

Inference-Engine v0.1.0

Given a pre-trained neural network, Inference-Engine performs maps network inputs to outputs by executing the forward pass through the provided network. Although the predominant programming language for machine-learning is Python, most high-performance computing (HPC) applications are written in Fortran, C, or C++. Inference-Engine aims to support HPC programs and is written in Fortran, a language with a large feature set supporting interoperability with C. This software exposes concurrency in a portable way by using standard language features that some modern Fortran compilers can exploit with various optimizations, including offloading computation to a Graphics Processing Unit (GPU). In particular, this software makes extensive use of Fortran's "do concurrent" parallel loop construct, implicitly parallel array statements, and pure procedures that can be invoked inside "do concurrent" blocks. Inference-Engine also supports dynamic choice of inference methods at runtime. Two current options include one method that uses Fortran's "dot_product" intrinsic function inside "do concurrent" blocks and another method that instead uses Fortran' "matmul" array intrinsic function. We plan to investigate automatic compiler offloading of "do concurrent" calculations to GPUs and compile-time substitution of optimized libraries such as the Basic Linear Algebra Library (BLAS) for "matmul" invocations. We also envision the potential for the choice of which method to use could happen at program launch based on in situ performance measurements on any given platform.

Rouson, Damian↗

CU-BENs: A structural modeling finite element library

The present work discusses capabilities within the finite element library CU-BENs. CU-BENs focuses on applying the finite element method to structural mechanics problems encountered within the context of inverse problems and partitioned fluid–structure interaction; thus element formulations are primarily of a structural type — truss, frame, and triangular discrete Kirchhoff theory shells. CU-BENs defaults to the skyline sparse storage scheme for the system matrix, but also supports other storage schemes when using external libraries such as LAPACK and UMFPACK. CU-BENs includes built-in nonlinear solution strategies, such as the Newton Raphson method and modified spherical arc length method, that are available within static analyses as well within the context of transient dynamic analyses involving a generalized-α implementation of the Newmark implicit time integration scheme.

97 MATHEMATICS AND COMPUTING↗

Simulation of crumpled sheets via alternating quasistatic and dynamic representations

In this work, we present a method for simulating the large-scale deformation and crumpling of thin, elastoplastic sheets. Motivated by the physical behavior of thin sheets during crumpling, two different formulations of the governing equations of motion are used: (1) a quasistatic formulation that effectively describes smooth deformations, and (2) a fully dynamic formulation that captures large changes in the sheet's velocity. The former is a differential-algebraic system of equations integrated implicitly in time, while the latter is a set of ordinary differential equations (ODEs) integrated explicitly. We adopt a hybrid integration scheme to adaptively alternate between the quasistatic and dynamic representations as appropriate. Further, we demonstrate the capacity of this method to effectively simulate a variety of crumpling phenomena. Finally, we show that statistical properties, notably the accumulation of creases under repeated loading, as well as the area distribution of facets, are consistent with experimental observations.

97 MATHEMATICS AND COMPUTING↗

PARIS: Predicting application resilience using machine learning

The traditional method to study application resilience to errors in HPC applications uses fault injection (FI), a time-consuming approach. Furthermore, while analytical models have been built to overcome the inefficiencies of FI, they lack accuracy. In this paper, we present PARIS, a machine-learning method to predict application resilience that avoids the time-consuming process of random FI and provides higher prediction accuracy than analytical models. PARIS captures the implicit relationship between application characteristics and application resilience, which is difficult to capture using most analytical models. We overcome many technical challenges for feature construction, extraction, and selection to use machine learning in our prediction approach. Our evaluation on 16 HPC benchmarks shows that PARIS achieves high prediction accuracy. PARIS is up to 450x faster than random FI (49x on average). Compared to the state-of-the-art analytical model, PARIS is at least 63% better in terms of accuracy and has comparable execution time on average.

97 MATHEMATICS AND COMPUTING↗

Scalable Multiphysics Block Preconditioning for Low Mach Number Compressible Resistive MHD with Application to Magnetic Confinement Fusion

This study investigates multiphysics block preconditioners that are critical in devising scalable Newton–Krylov iterative solvers for longer time-scale fully implicit fluid plasma models. The specific model of interest is the visco-resistive, low Mach number, compressible magnetohydrodynamics (MHD) model. This model describes the dynamics of conducting fluids in the presence of electromagnetic fields and can be used to study aspects of astrophysical phenomena, important science and technology applications, and basic plasma physics. The specific application of interest that motivates this study is the macroscopic simulation of longer time-scale stability and disruptions of magnetic confinement fusion devices, specifically the ITER Tokamak. The computational solution of the governing balance equations for mass, momentum, heat transfer, and magnetic induction for resistive MHD systems can be extremely challenging. These difficulties arise from both the strong nonlinear, nonsymmetric coupling of fluid and electromagnetic phenomena as well as the significant range of time and length scales that the interactions of these physical mechanisms produce. To handle the range of time and spatial scales of interest, a fully implicit unstructured variational multiscale finite element formulation is employed. For the scalable solution of the Newton linearized systems, fully coupled block preconditioners are designed to leverage algebraic multigrid subsolves. In conclusion, results are presented for the strong and weak scaling of the method as well as the robustness of these techniques for a large range of Lundquist numbers.

97 MATHEMATICS AND COMPUTING↗

Physics-guided dual implicit neural representations for source separation

Significant challenges exist in efficient data analysis of most advanced experimental and observational techniques because the collected signals often include unwanted contributions, such as background and signal distortions, that can obscure the physically relevant information of interest. To address this, we have developed a self-supervised machine-learning approach for source separation using a dual implicit neural representation framework that jointly trains two neural networks: one for approximating distortions of the physical signal of interest and the other for learning the effective background contribution. Our method learns directly from the raw data by minimizing a reconstruction-based loss function without requiring labeled data or pre-defined dictionaries. We demonstrate the effectiveness of our framework by considering a challenging case study involving large-scale simulated, as well as experimental, momentum-energy-dependent inelastic neutron scattering data in a four-dimensional parameter space, characterized by heterogeneous background contributions and unknown distortions to the target signal. The method is found to successfully separate physically meaningful signals from a complex or structured background even when the signal characteristics vary across all four dimensions of the parameter space. An analytical approach that informs the choice of the regularization parameter is presented. Our method offers a versatile framework for addressing source separation problems across diverse domains, ranging from superimposed signals in astronomical measurements to structural features in biomedical image reconstructions.

47 OTHER INSTRUMENTATION↗

Data-driven linear time advance operators for the acceleration of plasma physics simulation

In this study, we demonstrate the application of data-driven linear operator construction for time advance with a goal of accelerating plasma physics simulation. We apply dynamic mode decomposition (DMD) to data produced by the nonlinear SOLPS-ITER (Scrape-off Layer Plasma Simulator - International Thermonuclear Experimental Reactor) plasma boundary code suite in order to estimate a series of linear operators and monitor their predictive accuracy via online error analysis. We find that this approach defines when these dynamics can be represented by a sequence of approximate linear operators and is essential for providing consistent projections when compared to an unconstrained application. For linear diffusion and advection–diffusion fluid test problems, we construct and apply operators within explicit and implicit time advance schemes, demonstrating that stability can be robustly guaranteed in each case. We further investigate the use of the linear time advance operators within several integration methods including forward Euler, backward Euler, and the matrix exponential. The application of this method to simulation data from SOLPS-ITER, with varying levels of Markov chain Monte Carlo numerical noise, shows that constrained DMD operators yield a capability to identify, extract, and integrate a (slow) subset of the present timescales. Example applications show that for projected speedup factors of [Formula: see text], and [Formula: see text], a mean relative error of 3%, 5%, and 8% and maximum relative error less than 20% are achievable, which appears acceptable for typical SOLPS-ITER steady-state simulations.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Divergence Reduction in Monte Carlo Neutron Transport with On-GPU Asynchronous Scheduling

While Monte Carlo Neutron Transport (MCNT) is near-embarrasingly parallel, the effectively unpredictable lifetime of neutrons can lead to divergence when MCNT is evaluated on GPUs. Divergence is the phenomenon of adjacent threads in a warp executing different control flow paths; on GPUS, it reduces performance because each work group may only execute one path at a time. The process of Thread Data Remapping (TDR) resolves these discrepancies by moving data across hardware such that data in the same warp will be processed through similar paths. A common issue among prior implementations of TDR is the synchronous nature of its remapping and processing cycles, which exhaustively sort data produced by prior processing passes and exhaustively evaluate the sorted data. In another work, we defined a method of remapping data through an asynchronous scheduler which allows for work to be stored in shared memory and deferred arbitrarily until that work is a viable option for low-divergence evaluation. This article surveys a wider set of cases, with the goal of characterizing performance trends across a more comprehensive set of parameters. These parameters include cross sections of scattering/capturing/fission, use of implicit capture, source neutron counts, simulation time spans, and tuned memory allocations. Across these cases, we have recorded minimum and average execution times, as well as a heuristically tuned near-optimal memory allocation size for both synchronous and asynchronous scheduling. Across the collected data, it is shown that the asynchronous method is faster and more memory efficient in the majority of cases, and that it requires less tuning to achieve competitive performance.

Computer Science↗

Thermal Stability of π-Conjugated n -Ethylene-Glycol-Terminated Quaterthiophene Oligomers: A Computational and Experimental Study

This work represents a joint computational and experimental study on a series of $\textit{n}$-ethylene glycol (PEO$\textit{n}$)-terminated quaterthiophene (4T) oligomers for 1 < $\textit{n}$ < 10 to elucidate their self-assembly behavior into a smectic-like lamellar phase. This study builds on an earlier study for $\textit{n}$ = 4 that showed that our model predictions were consistent with experimental data on the melting behavior and structure of the lamellar phase, with the latter consisting of crystal-like 4T domains and liquid-like PEO4 domains. The present study aims to understand how the length of the terminal PEO$\textit{n}$ chains modulates the disordering temperature of the lamellar phase and hence the relative stability of the ordered structure. In this work, a simplified bilayer model, where the 4T domains are not explicitly described, is put forward to efficiently estimate the disordering effect of the PEO domains with increasing n; this method is first validated by correctly predicting that layers of alkyl (PE)-capped 4T oligomers (for 1 < $\textit{n}$ < 10) stay ordered at room temperature. Both 4T-domain implicit and explicit model simulations reveal that the order-disorder temperature decreases with the length of the PEO capping chains, as the associated increase in conformational entropy drives a tendency toward disorder that overtakes the cohesive energy, keeping the ordered packing of the 4T domains.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Semi-implicit continuum kinetic modeling of weakly collisional parallel transport in a magnetic mirror

We present implicit-explicit (IMEX) kinetic simulations of weakly collisional parallel plasma transport in magnetic mirror configurations using the continuum code COGENT. The numerical scheme employs a Jacobian-free Newton–Krylov method with algebraic multigrid preconditioning to overcome the severe time step limitations imposed by strong mirror forces in fully explicit schemes. Applied to parameters relevant to the Wisconsin HTS Axisymmetric Mirror experiment, the IMEX approach enables time steps up to 2.5×10 4 times larger than those permitted by explicit methods, resulting in a 2500× speedup in 1D–2V simulations of parallel transport with kinetic ions and Boltzmann electrons. Additionally, a reduced bounce-averaged model for a square mirror is implemented to support the computationally intensive fully kinetic simulations. The bounce-averaged formulation is used to evaluate the numerical convergence of the velocity-space discretization algorithms and to assess the role of the collision model by comparing simulations employing the nonlinear Fokker–Planck and the simplified Lenard–Bernstein–Dougherty collision operators.

Collision theories↗