Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “IMPLEMENTATION”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

A new procedure for implementing the modified inherent strain method with improved accuracy in predicting both residual stress and deformation for laser powder bed fusion

As a metal additive manufacturing (AM) process, laser powder bed fusion (L-PBF) has been widely used to produce parts with complex geometries. The large thermal gradient caused by the fast, intense, and repeated laser scanning induces significant residual deformation and stress to the as-built parts, which increase manufacturing difficulty and geometrical inaccuracy as a result. The modified inherent strain (MIS) method exploiting multiscale process simulations was developed to simulate residual deformation accurately and efficiently. However, the existing procedure of implementing the MIS method is found to give inaccurate residual stress prediction. Here in this work, a new implementation procedure for the MIS method is proposed to improve the simulation accuracy of residual stress without degrading the residual deformation prediction. The new procedure concerns the application of inherent strains to the part-scale layer-by-layer finite element model to obtain residual stress and deformation field. While the existing implementation of the part-scale MIS model involves only mechanical properties at ambient temperature, the new procedure adds one more solution step employing mechanical properties at an elevated temperature determined from the inherent strain extraction step. Both numerical and experimental studies are conducted to validate the proposed new implementation procedure. It shows that by using the new procedure, the MIS-based simulation can predict both residual stress and deformation of as-built L-PBF metal parts with good accuracy.

36 MATERIALS SCIENCE↗

Implementation of Triply Periodic Minimal Surfaces (TPMS) as surface objects in OpenMC

Triply Periodic Minimal Surfaces (TPMS) represent a promising geometry for future fuel designs due to their significant surface-to-volume ratio, which facilitates efficient cooling of nuclear fuel, a crucial factor for safety and efficiency. Demonstrating the remarkable capabilities of TPMS fuel requires initial modeling and simulation. This paper presents an implementation of TPMS in the Monte Carlo code OpenMC, enabling reactor physics modeling of TPMS. Here, the primary advantages over traditional methods using CAD files include reduced memory requirements for computations and high-fidelity implementation. This implementation has been tested against CAD files loaded in Serpent2, yielding promising results with low biases in the $k_{\textrm{eff}}$, comparable to biases in the material balance sheet. The implementation presented in this work will be used in future reactor physics computations related to new reactor designs involving TPMS-based fuels.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Thermodynamic modeling with uncertainty quantification using the modified quasichemical model in quadruplet approximation: Implementation into PyCalphad and ESPEI

The modified quasichemical model in the quadruplet approximation (MQMQA) considers the first- and the second-nearest-neighbor coordination and interactions, particularly useful in describing short-range ordering (SRO) in complex liquids such as molten salts, slag in metal processing, and electrolytic solutions. Here, the present work implements the MQMQA into the Python based open-source software PyCalphad for thermodynamic calculations. This endeavor facilitates the development of MQMQA-based thermodynamic database with uncertainty quantification (UQ) and propagation (UP) using the open-source software ESPEI. A new database structure based on Extensible Markup Language (XML) is proposed for ESPEI evaluation of MQMQA model parameters. Using the KF-NiF 2 , KCl-NaCl-MgCl 2 , and CaCl 2 -CaF 2 -LiCl-LiF salt systems as examples, we demonstrate the successful implementation of MQMQA in PyCalphad through thermodynamic calculations of Gibbs energy, equilibrium quadruplet fractions, and phase diagram, as well as database development with UQ and UP using ESPEI. Furthermore, as an application of the present implementation, both the LiF–TbF 3 and LiF-HoF 3 systems have been modeled by MQMQA for the first time, which are in good agreement with experiments. The present implementation hence offers an open-source capability for performing CALPHAD modeling for complex liquids with SRO using MQMQA plus a new XML database structure.

36 MATERIALS SCIENCE↗

Implementing a neural network interatomic model with performance portability for emerging exascale architectures

The two main thrusts of computational science are increasingly accurate predictions and faster calculations; to this end, the zeitgeist in molecular dynamics (MD) simulations is pursuing machine learned and data driven interatomic models, e.g. neural network potentials, and novel hardware architectures, e.g. GPUs. Current implementations of neural network potentials are orders of magnitude slower than traditional interatomic models and while looming exascale computing offers the ability to run large, accurate simulations with these models, achieving portable performance for MD with new and varied exascale hardware requires rethinking traditional algorithms, using novel data structures, and library solutions. We re-implement a neural network interatomic model in CabanaMD, an MD proxy application, built on libraries developed for performance portability. Our implementation shows significantly improved thread scaling in this complex kernel as compared to a current LAMMPS implementation, across both strong and weak scaling. Our single-source solution enables simulations up to 20 million atoms on a single CPU node and 4 million atoms with improved performance on a single GPU. Furthermore, we also explore parallelism and data layout choices (using flexible data structures called AoSoAs) and their effect on performance, seeing up to ~50% and ~5% improvements in performance on a GPU by choosing the right level of parallelism and data layout respectively.

97 MATHEMATICS AND COMPUTING↗

Analysis of US Industrial Assessment Centers (IACs) implementation

Industrial energy assessments are a fundamental action toward developing a decarbonization strategy at any level. They provide an understanding of where energy efficiency opportunities exist and help to make informed business decisions about the costs and benefits of implementing sustainable policies and practices. This paper examines the effectiveness of the Department of Energy's Industrial Assessment Centers Program at providing useful energy-efficiency audits as well as the barriers faced by the program and plans for future growth and improvement. This paper presents an analysis of the program between 1981 and 2022, covering 20,290 industrial assessments and 151,198 recommendations, with $2.6 billion of recommended savings. The analysis includes a breakdown of the IAC recommendations and implemented projects based on both the industrial subsectors according to the Standard Industrial Classification code, energy type, and systems evaluated. The results include a 47 % implementation rate, a gap analysis to understand the missed opportunities, and a discussion about reasons for recommendation rejection. A comparison of the IAC Program with other country-level energy auditing or assessment programs was conducted, and suggestions for improving implementation rates were mentioned.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Implementation and verification of PyNE R2S with DAG-OpenMC

The mesh-based Rigorous-Two-Step (R2S) method has been widely used in the accurate estimation of the shutdown dose rate (SDR) of fusion systems. Several mesh-based R2S code has been implemented based on Monte Carlo particle transport code MCNP5 or DAG-MCNP5 and inventory calculation code such as FISPACT-II, ALARA or ACAB. Recently, the OpenMC, a community-developed Monte Carlo particle transport code, implemented the photon particle transport capability, which enables OpenMC to be used in R2S. In this paper, PyNE R2S is extended with the support of DAG-OpenMC, i.e., OpenMC with DAGMC. Modifications that allow PyNE to read the flux result in the state point file of OpenMC has been implemented to PyNE R2S workflow. A custom source routine of OpenMC that reads the photon source file from PyNE R2S and sampling photon particle in OpenMC photon transport has been implemented in OpenMC. With these modifications, PyNE R2S is now capable of performing R2S calculations with both DAG-MCNP5 and DAG-OpenMC. The ITER FNG dose rate benchmark problem has been used to validate the code. The FNG neutron source has been modified for neutron transport with DAG-OpenMC. The shutdown dose rate of 19 cooling times has been calculated and compared with the experimental data and the computational results of other R2S codes. The results of PyNE R2S with DAG-OpenMC show satisfactory agreement with the experimental data and other computational results. Therefore, we consider PyNE R2S with DAG-OpenMC is a reliable code to calculate the SDR of fusion systems.

fusion↗

Biaxial steel plated concrete constitutive models for composite structures: Implementation and validation

The Steel-plated Concrete (SC) technique is an alternative construction technique with faster onsite construction speed, reduced construction time, and increased structural performance. Aiming to predict the force transfer mechanism of SC elements, an innovative constitutive model package is proposed by implementing experimental-based biaxial steel plate concrete models into the nonlinear finite element (FE) model “Membrane Model of SC (MM-CS).” First, the formulation and implementation of the MM-SC is illustrated in detail, including the equilibrium and compatibility equations and the implementation of constitutive models. Next, various experimental data of SC members are selected and simulated using the proposed MM-SC model. In this research, two types of structures, SC panels and framed SC shear walls, and three types of loading conditions, including uniaxial compression, pure shear, and combined axial-flexural-shear tests, are analyzed respectively. Good agreements were obtained between the reported results and the FE simulation results in terms of yield capacity and ultimate capacity, proving the reliability of the implemented analytical constitutive material models and the biaxial membrane model formulations.

42 ENGINEERING↗

gRASPA

GPU Monte Carlo Simulation Code with a taste of RASPA We present enhancements in Monte Carlo simulation speed and functionality within an open-source code, gRASPA, which uses graphical processing units (GPUs) to achieve significant performance improvements compared to serial, CPU implementations of Monte Carlo. The code supports a wide range of Monte Carlo simulations, including canonical ensemble (NVT), grand canonical, NVT Gibbs, Widom test particle insertions, and continuous-fractional component Monte Carlo. Implementation of grand canonical transition matrix Monte Carlo (GC-TMMC) and a novel feature to allow different moves for the different components of metal-organic framework (MOF) structures exemplify the capabilities of gRASPA for precise free energy calculations and enhanced adsorption studies, respectively. The introduction of a High-Throughput Computing (HTC) mode permits many Monte Carlo simulations on a single GPU device for accelerated materials discovery. The code can incorporate machine learning (ML) potentials. The open-source nature of gRASPA promotes reproducibility and openness in science, and users may add features to the code and optimize it for their own purposes. The code is written in CUDA/C++ and SYCL/C++ to support different GPU vendors. The gRASPA code is publicly available at https://github.com/snurr-group/gRASPA.

Li, Zhao [Purdue/Northwestern/Notre Dame Universit↗

Implementation of McMurchie–Davidson Algorithm for Gaussian AO Integrals Suited for SIMD Processors

We report an implementation of the McMurchie− Davidson evaluation scheme for 1- and 2-particle Gaussian AO integrals designed for processors with Single Instruction Multiple Data (SIMD) instruction sets. Like in our recent MD implementation for graphical processing units (GPUs) [Asadchev, A.; Valeev, E. F.. J. Chem. Phys. 2024, 160, 244109.], variable-sized batches of shellsets of integrals are evaluated at a time. By optimizing for the floating point instruction throughput rather than minimizing the number of operations, this approach achieves up to 50% of the theoretical hardware peak FP64 performance for many common SIMD-equipped platforms (AVX2, AVX512, NEON), which translates to speedups of up to 30 over the state-of-the-art one-shellset-at-a-time implementation of Obara−Saika-type schemes in Libint for a variety of primitive and contracted integrals. As with our previous work, we rely on the standard C++ programming language such as the std::simd standard library feature to be included in the 2026 ISO C++ standard without any explicit code generation to keep the code base small and portable. The implementation is part of the open source LibintX library freely available at https://github.com/ValeevGroup/libintx.

Basis sets↗

SHarmonic: A fast and accurate implementation of spherical harmonics for electronic-structure calculations

The authors present SHarmonic, a new implementation of the spherical harmonics targeted for electronic-structure calculations. Their approach is to use explicit formulas for the harmonics written in terms of normalized Cartesian coordinates. This approach results in a code that is as precise as other implementations while being at least one order of magnitude more computationally efficient. The library can run on graphics processing units as well, achieving an additional order of magnitude in execution speed. This new implementation is simple to use and is provided under an open-source license; it can be readily used by other codes to avoid the error-prone and cumbersome implementation of the spherical harmonics.

Mathematics and Computing↗

Feasibility Study on Implementing a Staggered-Grid Finite Volume Method for System Analysis Code Development Under the MOOSE Framework

Here, this work summarizes a feasibility study on testing numerical algorithms that are suitable and efficient for advanced system analysis code development under the mutli-physics framework, MOOSE. The key to the test bed is the implementation of high-order one-dimensional staggered-grid finite volume method (SG-FVM), and its direct interaction with the linear/nonlinear solver, PETSc. The test bed utilized a more flexible code structure to enable the finite volume method implementation and direct interacting with the solver package, instead of using the natively supported finite element method by the framework. Using a suite of selected test problems with different problem sizes and levels of complexity, the implemented SG-FVM demonstrated superior performance improvement against a direct finite element method implementation through MOOSE. On two computer systems, the speedup was observed to be significant, with at least one order of magnitude of solving time reduction. For a complex reactor model, transient simulation was performed using the newly developed finite volume method code, the results of which agree very well with the reference results from the finite element method code. Overall, this study demonstrates a successful feasibility study on the proposed numerical algorithms and software structure to support advanced system analysis tool development.

MOOSE↗

Comparison of steady-state analytical wake models implemented in wind farm analysis software

A common set of mathematical wind turbine wake models are implemented in a few, well-adopted computational tools for wind farm wake modelling. Although the referenced mathematical formulations are common, implementation details may lead to differences in results. This study presents a systematic comparison of the implementation of mathematical wake models in open source, Python-based wind turbine wake modelling software, and a set of the models are directly compared. Despite aligning only the mathematical model parameters and retaining the default computational model parameters, good agreement is found across most of the model implementations, and additional agreement is expected upon further parameters alignment.

17 WIND ENERGY↗

Integration of Ag-CBRAM crossbars and Mott ReLU neurons for efficient implementation of deep neural networks in hardware

In-memory computing with emerging non-volatile memory devices (eNVMs) has shown promising results in accelerating matrix-vector multiplications. However, activation function calculations are still being implemented with general processors or large and complex neuron peripheral circuits. Here, we present the integration of Ag-based conductive bridge random access memory (Ag-CBRAM) crossbar arrays with Mott rectified linear unit (ReLU) activation neurons for scalable, energy and area-efficient hardware (HW) implementation of deep neural networks. We develop Ag-CBRAM devices that can achieve a high ON/OFF ratio and multi-level programmability. Compact and energy-efficient Mott ReLU neuron devices implementing ReLU activation function are directly connected to the columns of Ag-CBRAM crossbars to compute the output from the weighted sum current. We implement convolution filters and activations for VGG-16 using our integrated HW and demonstrate the successful generation of feature maps for CIFAR-10 images in HW. Our approach paves a new way toward building a highly compact and energy-efficient eNVMs-based in-memory computing system.

Mott insulators↗

Analyzing the Multiscale Impacts of Implementing Energy-Efficient HVAC Improvements Through Energy Audits and Economic Input–Output Analysis

Abstract Heating, ventilation, and air-conditioning (HVAC) systems are usually an industry’s highest consumer of energy, most of which goes toward space cooling in buildings. Industrial energy-efficiency audits not only benefit manufacturers but also generate significant economic and environmental benefits to localities, states, and the nation. This article analyzes the micro- and macro scale impacts of implementing energy-efficient HVAC systems by integrating the industrial building energy data with the macroeconomic regional economic flow model. Micro-scale data include 10 years of historical energy, cost, and carbon dioxide savings achieved from energy-efficient HVAC implementation offered to manufacturers through industrial energy audits. The data were integrated into the macroeconomic modeling framework to illuminate the cascading regional economic impacts of implementing energy-efficient HVAC recommendations in manufacturing facilities. Results show that if recommendations had been implemented throughout all manufacturers in the region, $656 M energy costs would have been directly saved, 7.8 million metric tons of carbon dioxide emissions would have been avoided, and 4387 jobs could have been created, resulting in a total annual economic impact of $899 M stemming from direct, indirect, and induced impacts. The results offer insight into how industrial energy systems can be designed and provide models for how communities can accomplish a net-zero society.

Energy & Fuels↗

Implementing Arbitrary/Common Concurrent Writes of CRCW PRAM

The Parallel Random Access Machines (PRAM) abstraction is the simplest and most elegant algorithmic model for the design and analysis of parallel algorithms. It consists of different models categorized based on the underlying memory access mode used, the most powerful of which is the Concurrent Read Concurrent Write (CRCW) model. A PRAM algorithm describes a series of rounds, each of which consists of a collection of operations that can be executed concurrently within the same time step. However, the lack of support for concurrent memory accesses and the prevalence of asynchronous programming models led to the belief that implementing CRCW PRAM algorithms is unattainable and prompted many to avoid this model except for theoretical studies of optimal performance.In this work, we study the arbitrary and common concurrent writes in the CRCW PRAM model and explore implementation challenges on general-purpose systems. Moreover, we examine current practices for implementing common/arbitrary concurrent writes and propose a new efficient lightweight and thread-safe method to implement concurrent writes through leveraging atomic instructions. To demonstrate the efficacy of our method, we developed OpenMP kernels for classical CRCW PRAM algorithms and provide experimental results and comparisons based on run time performance measured over the x86 multicore architecture. Our results show a performance speedup compared to current practices up to 4.5x across all our benchmarks.

Ghanim, Fady↗

Quantum Algorithm Implementations for Beginners

As quantum computers become available to the general public, the need has arisen to train a cohort of quantum programmers, many of whom have been developing classical computer programs for most of their careers. While currently available quantum computers have less than 100 qubits, quantum computing hardware is widely expected to grow in terms of qubit count, quality, and connectivity. This review aims at explaining the principles of quantum programming, which are quite different from classical programming, with straightforward algebra that makes understanding of the underlying fascinating quantum mechanical principles optional. We give an introduction to quantum computing algorithms and their implementation on real quantum hardware. We survey 20 different quantum algorithms, attempting to describe each in a succinct and self-contained fashion. We show how these algorithms can be implemented on IBM’s quantum computer, and in each case, we discuss the results of the implementation with respect to differences between the simulator and the actual hardware runs. This article introduces computer scientists, physicists, and engineers to quantum algorithms and provides a blueprint for their implementations.

97 MATHEMATICS AND COMPUTING↗

Direction-optimizing Label Propagation Framework for Structure Detection in Graphs: Design, Implementation, and Experimental Analysis

Label Propagation is not only a well-known machine learning algorithm for classification but also an effective method for discovering communities and connected components in networks. We propose a new Direction-optimizing Label Propagation Algorithm (DOLPA) framework that enhances the performance of the standard Label Propagation Algorithm (LPA), increases its scalability, and extends its versatility and application scope. As a central feature, the DOLPA framework relies on the use of frontiers and alternates between label push and label pull operations to attain high performance. It is formulated in such a way that the same basic algorithm can be used for finding communities or connected components in graphs by only changing the objective function used. Additionally, DOLPA has parameters for tuning the processing order of vertices in a graph to reduce the number of edges visited and improve the quality of solution obtained. We present the design and implementation of the enhanced algorithm as well as our shared-memory parallelization of it using OpenMP. We also present an extensive experimental evaluation of our implementations using the LFR benchmark and real-world networks drawn from various domains. Compared with an implementation of LPA for community detection available in a widely used network analysis software, we achieve at most five times the F-Score while maintaining similar runtime for graphs with overlapping communities. We also compare DOLPA against an implementation of the Louvain method for community detection using the same LFR-graphs and show that DOLPA achieves about three times the F-Score at just 10% of the runtime. For connected component decomposition, our algorithm achieves orders of magnitude speedups over the basic LP-based algorithm on large-diameter graphs, up to 13.2× speedup over the Shiloach-Vishkin algorithm, and up to 1.6× speedup over Afforest on an Intel Xeon processor using 40 threads.

97 MATHEMATICS AND COMPUTING↗

Breathe Well, Live Well: Implementing an Adult Asthma Self-Management Education Program

Asthma remains a significant health problem in the United States. Adults with poorly controlled asthma can affect their community in a number of ways, from lost productivity in the workplace to health care costs to premature death. Asthma self-management education helps individuals achieve better control of their asthma and is critical for the overall health and well-being of individuals with asthma. While there are numerous programs and initiatives targeting children with asthma, there is a lack of comparable focus on the needs of adults with asthma. The American Lung Association developed Breathe Well, Live Well, an adult asthma self-management education program, and launched it nationwide in 2007. The program for adults has a flexible delivery format for community-based implementation. This article describes the development, dissemination, and transformation of the program. Each stage of implementation showed positive changes in asthma self-management practices that contribute to better asthma control, and one local implementation additionally showed decreased reports of missed work and unscheduled health care visits among participants. The findings from the three evaluations support the use of Breathe Well, Live Well for broad community-based implementation to improve asthma self-management efficacy and behaviors.

Gardner, Emily A.↗