Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “IMPLEMENTATION”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Multi-species collisions for delta-f gyrokinetic simulations: Implementation and verification with GENE

we report a multi-species linearized collision operator based on the model developed by Sugama et al. has been implemented in the nonlinear gyrokinetic code, GENE. Such a model conserves particles, momentum, and energy to machine precision, and is shown to have negative definite free energy dissipation characteristics, satisfying Boltzmann’s H-theorem, including for realistic mass ratio. Finite Larmor Radius (FLR) effects have also been implemented into the local version of the code. For the global version of the code, the collision operator has been developed to allow for block-structured velocity space grids, allowing for computationally tractable collisional global simulations. The validity of the collision operator has been demonstrated by relaxation and conservation tests, as well as appropriate benchmarks. The newly implemented operator shall be used in future simulations to study magnetically confined fusion plasma turbulence and transport in more extreme regions with higher collisionality.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

A higher-order finite-element implementation of the nonlinear Fokker–Planck collision operator for charged particle collisions in a low density plasma

Collisions between particles in a low density plasma are described by the Fokker–Planck collision operator. In applications, this nonlinear integro-differential operator is often approximated by linearised or ad-hoc model operators due to computational cost and complexity. In this work, we present an implementation of the nonlinear Fokker–Planck collision operator written in terms of Rosenbluth potentials in the Rosenbluth–MacDonald–Judd (RMJ) form. The Rosenbluth potentials may be obtained either by direct integration or by solving partial differential equations (PDEs) similar to Poisson's equation: we optimise for performance and scalability by using sparse matrices to solve the relevant PDEs. We represent the distribution function using a tensor-product continuous-Galerkin finite-element representation and we derive and describe the implementation of the weak form of the collision operator. We present tests demonstrating a successful implementation using an explicit time integrator and we comment on the speed and accuracy of the operator. Finally, we speculate on the potential for applications in the current and next generation of kinetic plasma models.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

LuGo: An enhanced quantum phase estimation implementation

Quantum Phase Estimation (QPE) is a cardinal algorithm in quantum computing that plays a crucial role in various applications, including cryptography, molecular simulation, and solving systems of linear equations. However, the standard implementation of QPE faces challenges related to time complexity and circuit depth, which limit its practicality for large-scale computations. We introduce LuGo, a novel framework designed to enhance the performance of QPE by reducing circuit duplication, as well as using parallelization techniques to achieve faster generation of the QPE circuit and gate reduction. We validate the effectiveness of our framework by generating quantum linear solver circuits, which require both QPE and inverse QPE, to solve linear systems of equations. LuGo achieves significant improvements in both computational efficiency and hardware requirements without compromising on accuracy. Compared to a standard QPE implementation, LuGo reduces time consumption to generate a circuit that solves a 2 6 × 2 6 system matrix by a factor of 50.68 and over 31× reduction of quantum gates and circuit depth, with no fidelity loss on an ideal quantum simulator. Furthermore, we demonstrated the versatility and scalability of LuGo enabled HHL algorithm by simulating a canonical Hele-Shaw fluid problem using a quantum simulator. With these advantages, LuGo paves the way for more efficient implementations of QPE, enabling broader applications across several quantum computing domains.

Quantum algorithm↗

A robust spectral element implementation of the $k - τ$ RANS model in Nek5000/NekRS

The $k - ω$ Reynolds Averaged Navier Stokes (RANS) model is one of the industry standard approaches for modeling of turbulent flows. It performs better than the $k - ϵ$ model for low Reynolds number flows and is also more suitable for boundary layers with adverse pressure gradients. Major drawback of the model, however, is that the asymptotic value of $ω$ at the walls is singular, necessitating the use of a contrived “sufficiently” large value for $ω$ as the boundary condition for its transport equation. Here, this invariably leads to the solution being sensitive to near wall grid spacing. While an acceptable solution for low order (finite volume) methods, the excessive near wall gradients lead to persistent numerical stability issues in high order codes. To alleviate the problem, specifically in the context of the high order spectral element code Nek5000, a regularized $k - ω$ approach was formulated in our prior work (Tomboulides et al., 2018). The formulation, however, relies on the use of wall distance and its gradients for modeling the closure terms and can pose problems for simulations in complex geometries. This work presents a novel implementation of the $k - τ$ RANS model in Nek5000, where $τ = 1/ω$, eliminating the need for regularization, owing to the asymptotically bounded behavior of the source terms in the $τ$ transport equation, and also eliminating dependence on wall distance. Robustness and stability of the $k - τ$ model is ensured through implicit treatment of the source terms and their careful numerical implementation and demonstrated through several cases aimed at verification and validation. Studies include both canonical and engineering relevant problems, viz., turbulent channel flow, pipe flow, backward facing step, flow over NACA 0012 airfoil and flow in a T-junction. Results from the $k - τ$ model are shown to be consistent with regularized $k - ω$ model and also with the $k - ω$ SST model in OpenFOAM (for select studies). Comparison with experimental data is also shown, where available, to bolster validation efforts for the $k - τ$ model implementation through prediction of key turbulent quantities of interest.

Nek5000↗

Implementation of dietary methionine restriction using casein after selective, oxidative deletion of methionine

Dietary methionine restriction (MR) is normally implemented using diets formulated from elemental amino acids (AA) that reduce methionine content to 0.17%. However, translational implementation of MR with elemental AA-based diets is intractable due to poor palatability. To solve this problem and restrict methionine using intact proteins, casein was subjected to mild oxidation to selectively reduce methionine. Diets were then formulated using oxidized casein, adding back methionine to produce a final concentration of 0.17%. The biological efficacy of dietary MR using the oxidized casein (Ox Cas) diet was compared with the standard elemental MR diet in terms of the behavioral, metabolic, endocrine, and transcriptional responses to the four diets. The Ox Cas MR diet faithfully reproduced the expected physiological, biochemical, and transcriptional responses in liver and inguinal white adipose tissue. Collectively, these findings demonstrate that dietary MR can be effectively implemented using casein after selective oxidative reduction of methionine.

59 BASIC BIOLOGICAL SCIENCES↗

Semi-Empirical Shadow Molecular Dynamics: A PyTorch Implementation

Here, extended Lagrangian Born–Oppenheimer molecular dynamics (XL-BOMD) in its most recent shadow potential energy version has been implemented in the semiempirical PyTorch-based software PySeQM. The implementation includes finite electronic temperatures, canonical density matrix perturbation theory, and an adaptive Krylov subspace approximation for the integration of the electronic equations of motion within the XL-BOMB approach (KSA-XL-BOMD). The PyTorch implementation leverages the use of GPU and machine learning hardware accelerators for the simulations. The new XL-BOMD formulation allows studying more challenging chemical systems with charge instabilities and low electronic energy gaps. The current public release of PySeQM continues our development of modular architecture for large-scale simulations employing semi-empirical quantum-mechanical treatment. Applied to molecular dynamics, simulation of 840 carbon atoms, one integration time step executes in 4 s on a single Nvidia RTX A6000 GPU.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Cholesky Decomposition-Based Implementation of Relativistic Two-Component Coupled-Cluster Methods for Medium-Sized Molecules

A Cholesky decomposition (CD)-based implementation of relativistic two-component coupled-cluster (CC) and equation-of-motion CC (EOM-CC) methods using an exact two-component Hamiltonian augmented with atomic-mean-field integrals (the X2CAMF scheme) is reported. Furthermore, the present CD-based implementation of X2CAMF-CC and EOM-CC methods employs atomic-orbital-based algorithms to avoid the construction of two-electron integrals and intermediates involving three and four virtual indices. The CD-based implementation extends the applicability of X2CAMF-CC and EOMCC methods to medium-sized molecules with the correlation of around 1000 spinors. Benchmark calculations for uranium-containing small molecules have been performed to assess the dependence of CC results with respect to the Cholesky threshold. A Cholesky threshold of 10 –4 is shown to maintain chemical accuracy. Example calculations to illustrate the capability of the CD-based relativistic CC methods are reported for the bond dissociation energy of the uranium hexafluoride molecule, UF 6 , with up to quadruple-zeta basis sets and the lowest excitation energy in solvated uranyl ion [UO 2 2+ (H 2 O) 12 ].

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Static Subspace Approximation for Random Phase Approximation Correlation Energies: Implementation and Performance

Developing theoretical understanding of complex reactions and processes at interfaces requires using methods that go beyond semilocal density functional theory to accurately describe the interactions between solvent, reactants and substrates. Methods based on many-body perturbation theory, such as the random phase approximation (RPA), have previously been limited due to their computational complexity. However, this is now a surmountable barrier due to the advances in computational power available, in particular through modern GPU-based supercomputers. In this work, we describe the implementation of RPA calculations within BerkeleyGW and show its favorable computational performance on large complex systems relevant for catalysis and electrochemistry applications. Our implementation builds off of the static subspace approximation which, by employing a compressed representation of the frequency dependent polarizability, enables the evaluation of the RPA correlation energy with significant acceleration and systematically controllable accuracy. We find that the computational cost of calculating the RPA correlation energy scales only linearly with system size for systems containing up to 50 thousand bands, and is expected to scale quadratically thereafter. We also show excellent strong scaling results across several supercomputers, demonstrating the performance and portability of this implementation.

algorithmic development↗

Quantum reservoir computing implementation on coherently coupled quantum oscillators

Quantum reservoir computing is a promising approach for quantum neural networks, capable of solving hard learning tasks on both classical and quantum input data. However, current approaches with qubits suffer from limited connectivity. We propose an implementation for quantum reservoir that obtains a large number of densely connected neurons by using parametrically coupled quantum oscillators instead of physically coupled qubits. We analyze a specific hardware implementation based on superconducting circuits: with just two coupled quantum oscillators, we create a quantum reservoir comprising up to 81 neurons. We obtain state-of-the-art accuracy of 99% on benchmark tasks that otherwise require at least 24 classical oscillators to be solved. Our results give the coupling and dissipation requirements in the system and show how they affect the performance of the quantum reservoir. Beyond quantum reservoir computing, the use of parametrically coupled bosonic modes holds promise for realizing large quantum neural network architectures, with billions of neurons implemented with only 10 coupled quantum oscillators.

97 MATHEMATICS AND COMPUTING↗

Low cost, flexible, and distribution level universal grid analyser platform: designs and implementations

This study presents the designs and implementations of a distribution level open-universal grid analyser (Open-UGA) platform. The proposed Open-UGA platform consists of distribution-level phasor measurement units (PMUs), a standard signal generator, a router, and a server. Firstly, an overall introduction for the software, hardware, and server architectures of the Open-UGA platform is given. To give a detailed design, the software, hardware, server block diagrams, flowcharts, and printed circuit board photo of the Open-UGA platform are presented in detail. Then, four different types of distribution level PMU algorithms are introduced and implemented in the Open-UGA platform to verify the flexibility and reconfigurability. The flowcharts and functionalities of these four UGAs with different PMU algorithms are given as example implementations. Lastly, a performance comparison is conducted with both quantitative and illustrative results.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Real Time implementation of Artificial Intelligence compression algorithm for High-Speed Streaming Readout signals

The new generation of high-energy physics experiments plans to acquire data in streaming mode. With this approach, it is possible to access the information of the whole detector (organized in time slices) for optimal and lossless triggering of data acquisitions. With this approach, data rates, especially in large detectors, are often very high, and the network is likely to be the bottleneck for the entire Streaming Read Out system. The aim of this work is to study the implementation of a lossy compression algorithm based on Artificial Intelligence: an Autoencoder. With Machine Learning it is possible to achieve a high compression ratio and fast inference time with only a small degradation of the signals, almost negligible for the specific application. This work explores different configurations of the Autoencoder and the implementation on different hardware. Different Autoencoder configurations are explored to find the best trade-off between compression ratio and reconstruction loss, both for signals and energy spectrum. Different hardware implementations are also explored to find the best platform to achieve real-time performance for the specific application.

Rossi, Fabio (ORCID:0009000385713885)↗

First implementation of gyrokinetic exact linearized Landau collision operator and comparison with models

Gyrokinetic simulations are fundamental to understanding and predicting turbulent transport in magnetically confined fusion plasmas. Previous simulations have used model collision operators with approximate field-particle terms of unknown accuracy and/or have neglected collisional finite Larmor radius (FLR) effects. We have implemented the linearized Fokker–Planck collision operator with exact field-particle terms and full FLR effects in a gyrokinetic code (GENE). The new operator, referred to as “exact” in this paper, allows the accuracy of model collision operators to be assessed. The conservative Landau form is implemented because its symmetry underlies the conservation laws and the H-theorem, and enables numerical methods to preserve this conservation, independent of resolution. The implementation utilizes the finite-volume method recently employed to discretize the Sugama collision model in GENE, allowing direct comparison between the two operators. Results show that the Sugama model appears accurate for the growth rates of trapped electron modes (TEMs) driven only by density gradients, but appreciably underestimates the growth rates as the collisionality and electron temperature gradient increase. The TEM turbulent fluxes near the nonlinear threshold using the exact operator are similar to the Sugama model for the n e = d ln T e /d ln n e = 0 case, but substantially larger than the Sugama model for the n e = 1 case. The FLR effects reduce the growth rates increasingly with wavenumber deepening a “valley” at the intermediate binormal wavenumber as the unstable mode extends from the TEM regime to the electron temperature gradient instability regime. Application to the Hinton–Rosenbluth problem shows that zonal flows decay faster as the radial wavenumber increases and the exact operator yields weaker decay rates.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Analytical derivatives of the individual state energies in ensemble density functional theory. II. Implementation on graphical processing units (GPUs)

Conical intersections control excited state reactivity, and thus, elucidating and predicting their geometric and energetic characteristics are crucial for understanding photochemistry. Locating these intersections requires accurate and efficient electronic structure methods. Unfortunately, the most accurate methods (e.g., multireference perturbation theories such as XMS-CASPT2) are computationally challenging for large molecules. The state-interaction state-averaged restricted ensemble referenced Kohn–Sham (SI-SA-REKS) method is a computationally efficient alternative. The application of SI-SA-REKS to photochemistry was previously hampered by a lack of analytical nuclear gradients and nonadiabatic coupling matrix elements. We have recently derived analytical energy derivatives for the SI-SA-REKS method and implemented the method effectively on graphical processing units. We demonstrate that our implementation gives the correct conical intersection topography and energetics for several examples. Furthermore, our implementation of SI-SA-REKS is computationally efficient, with observed sub-quadratic scaling as a function of molecular size. This demonstrates the promise of SI-SA-REKS for excited state dynamics of large molecular systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Cluster perturbation theory X: A parallel implementation of Lagrangian perturbation series for the coupled cluster singles and doubles ground-state energy through fifth order

We describe an efficient implementation of cluster perturbation and Møller-Plesset Lagrangian energy series through fifth order that target the coupled cluster singles and doubles energy utilizing the resolution of the identity approximation. We illustrate the computational performance of the implementation by performing ground state energy calculations on systems with up to 1200 basis functions using a single node and by comparison to conventional CCSD calculations. We further show that our hybrid MPI/OMP parallel implementation that also utilizes graphical processing units can be used to obtain fifth order energies on systems with almost 1200 basis functions with a 90 minute "time to solution" running on Frontier at Oak Ridge National Laboratory.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Implementation of the D1S Methodology for Shutdown Dose Rate Calculations in the OpenMC Monte Carlo Particle Transport Code

We present an implementation of the direct one-step (D1S) methodology for shutdown dose rate (SDR) calculations in the OpenMC Monte Carlo particle transport code. In addition to being the first fully open-source D1S implementation, it is also the first to require no ad hoc source code or nuclear data library modifications. The code can seamlessly switch between production of prompt and decay photons based on a user input parameter, and the decay data needed for decay photon generation are made available through a depletion chain file, which is already used for OpenMC’s built-in depletion/activation solver. A set of Python functions significantly eases the burden of computing and applying time correction factors needed to properly account for the time dependence of radionuclide activity. To assess the accuracy of the D1S implementation, SDR calculations have been carried out for three problems: a prism of iron irradiated by 14-MeV neutrons, the ITER port plug computational benchmark, and the Frascati Neutron Generator (FNG) ITER dose rate benchmark problem from the Shielding INtegral Benchmark Archive and Database (SINBAD). For each of these problems, comparisons were made to calculations using the rigorous two-step (R2S) method. The results on the iron prism problem illustrate how the D1S method achieves superior spatial resolution compared to the R2S method without the need for spatial discretization of the activation regions. The D1S and R2S results for the ITER port plug benchmark agree well with previously reported results in the literature. While the D1S results are 10% to 15% lower than the R2S results, this may be due to stochastic uncertainty and/or spatial discretization in the R2S calculations. On the FNG dose rate benchmark problem, the D1S method produces dose rate estimates that are within 4% of the dose rates predicted using a cell-based R2S workflow. The D1S estimates of the SDR are also in reasonable agreement with the experimental measurements and show the same basic trends that have been observed in previous works. A qualitative analysis of the execution time and uncertainty for the R2S and D1S workflows suggests that the D1S method would attain a higher figure of merit.

D1S method↗

An implementation of the phase-field model based on coupled thermomechanical finite element solvers for large-strain twinning, explicit dynamic fracture and the classical Stefan problem

The implementation of a phase-field model in finite elements usually requires significant expertise and involves the development of a user element with additional degrees of freedom. An alternative implementation of the phase-field model within a thermo-mechanical finite element simulation package was presented in (Cho et al 2012 Int. J. Solids Struct. 49 1973–1992), where the phase-field variable is treated as the temperature degree of freedom. However, this approach has only been used for small strain phase-field modelling of martensitic transformations and quasistatic phase-field modelling of fracture. Here, we present a phase-field finite element implementation via the temperature degree of freedom for several additional cases from the literature: (i) the large-strain phase-field description of deformation twinning presented in (Clayton and Knap 2011 Physica D 240 841–858), (ii) phase-field description of brittle fracture with inertial effects based on the theory from (Molnár and Gravouil 2017, Finite Elem. Anal. Des. 130 27–38) and (Miehe et al 2010 Int. J. Numer. Methods Eng. 83 1273–1311) and (iii) the classical Stefan problem of solidification presented in (Mackenzie and Robertson 2002 J. Comput. Phys. 181 526–544). The last problem involves the temperature and phase-field variables as unknowns.

36 MATERIALS SCIENCE↗

Improving reproducibility in synchrotron tomography using implementation-adapted filters

For reconstructing large tomographic datasets fast, filtered backprojection-type or Fourier-based algorithms are still the method of choice, as they have been for decades. These robust and computationally efficient algorithms have been integrated in a broad range of software packages. The continuous mathematical formulas used for image reconstruction in such algorithms are unambiguous. However, variations in discretization and interpolation result in quantitative differences between reconstructed images, and corresponding segmentations, obtained from different software. This hinders reproducibility of experimental results, making it difficult to ensure that results and conclusions from experiments can be reproduced at different facilities or using different software. In this paper, a way to reduce such differences by optimizing the filter used in analytical algorithms is proposed. These filters can be computed using a wrapper routine around a black-box implementation of a reconstruction algorithm, and lead to quantitatively similar reconstructions. Use cases for this approach are demonstrated by computing implementation-adapted filters for several open-source implementations and applying them to simulated phantoms and real-world data acquired at the synchrotron. Our contribution to a reproducible reconstruction step forms a building block towards a fully reproducible synchrotron tomography data processing pipeline.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

A Performance-Portable MultiGPU Implementation of 3D Euler Equations using ProtoX and IRIS

Computational scientists often face challenges when developing and optimizing code for high-performance computing (HPC), especially when trying to leverage GPUs. Given the heterogeneity of the nodes that comprise many modern HPC facilities, considerable demand exists for performance portable solutions for the core computational kernels used in many scientific computing libraries. In this work, we demonstrate a fourth-order finite volume method–based implementation of the Euler equations, which are an integral part of computational fluid dynamics. Our performance-portable multiGPU implementation for Euler equations uses ProtoX to generate kernels and IRIS for portability. ProtoX is a domain-specific language that uses a structured-grid partial differential equation library called Proto as its front end and the SPIRAL code generation system as its back end to generate optimized kernels for different architectures. Optimized kernels generated by ProtoX are orchestrated through the IRIS intelligent runtime system to provide portability. Two levels of optimizations within the IRIS runtime— directed acyclic graph fusion and task fusion—are explored to efficiently utilize computing resources in a multiGPU environment. Performance improvement through these optimizations is showcased by comparing the base ProtoX-IRIS implementation on AMD GPUs (Frontier node) and on NVIDIA GPUs (NVIDIA DGX-1).

Mankad, Het↗