Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “computation time”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Homomorphic data compression for real time photon correlation analysis

The construction of highly coherent X-ray sources, combined with next-generation detectors that are larger and faster, has enabled new research opportunities across the scientific landscape. Among the techniques that benefit most from these advancements is X-ray photon correlation spectroscopy (XPCS), where faster acquisition unlocks the ability to study faster dynamics within samples. However, faster acquisition on larger detectors also introduces unprecedented challenges for online data processing and offline data storage. Such challenges are particularly prominent for XPCS, where real time analyses require simultaneous calculation of all the previously acquired data in the time series. We present a homomorphic compression scheme to effectively reduce the computational time and memory space required for XPCS analysis. Leveraging similarities in the mathematical expression between a matrix-based compression algorithm and the correlation calculation, our approach allows direct operation on the compressed data without their decompression. The offline compression scheme extends storage capacity by a factor of 40 while preserving key features in the lossy compressed data. Meanwhile, the online compression scheme reduces the computational time to below 1 ms, enabling real time calculation of the correlation functions at kHz framerate. Our demonstration of a homomorphic compression of scientific data provides an effective solution to the big data challenge at coherent light sources. Beyond the example shown in this work, the framework can be extended to facilitate real-time operations directly on a compressed data stream for other techniques.

36 MATERIALS SCIENCE↗

Apodization Specific Fitting for Improved Resolution, Charge Measurement, and Data Analysis Speed in Charge Detection Mass Spectrometry

Short-time Fourier transforms with short segment lengths are typically used to analyze single ion charge detection mass spectrometry (CDMS) data either to overcome effects of frequency shifts that may occur during the trapping period or to more precisely determine the time at which an ion changes mass or charge, or enters an unstable orbit. The short segment lengths can lead to scalloping loss unless a large number of zero-fills are used, making computational time a significant factor in real-time analysis of data. Apodization specific fitting leads to a 9-fold reduction in computation time compared to zero-filling to a similar extent of accuracy. This makes possible real-time data analysis using a standard desktop computer. Rectangular apodization leads to higher resolution than the more commonly used Gaussian or Hann apodization and makes it possible to separate ions with similar frequencies, a significant advantage for experiments in which the masses of many individual ions are measured simultaneously. Equally important is a >20% increase in S/N obtained with rectangular apodization compared to Gaussian or Hann, which directly translates to a corresponding improvement in accuracy of both charge measurements and ion energy measurements that rely on the amplitudes of the fundamental and harmonic frequencies. Finally, combined with computing the fast Fourier transform in a lower-level language, this fitting procedure eliminates computational barriers and should enable real-time processing of CDMS data on a laptop computer.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Status of Multiple Channel Fuel Performance Capabilities Within the SAS4A/SASSYS-1 Safety Analysis Software

SAS4A/SASSYS-1 (SAS) is a fast-running simulation tool used to perform deterministic analysis of anticipated events as well as design basis and beyond design basis accidents for advanced liquid-metal-cooled nuclear reactors. It is a critical element of safety analysis capabilities for the U.S. Department of Energy and is utilized within industry to perform the transient safety analyses required to support the licensing of Liquid Metal-cooled Fast Reactors (LMFRs). Although SAS is exceptionally fast for most transient scenarios, fuel performance calculations, along with the associated pre-transient characterization of the fuel pin, may be required for transient scenarios where fuel pin failure is hypothesized. Both the pre-transient characterization and the transient fuel performance calculation are necessary to properly quantify margins to potential fuel failure and assess the time spent potentially exceeding such margins during events. While safety analysis calculations with fuel performance models provide a more detailed characterization of the reactor during a transient, the pre-transient characterization can be time-consuming and computationally expensive. Often, large numbers of fuel pins have been exposed to similar pre-transient irradiation conditions. Similarly, the same pre-transient fuel characterization may be applicable to numerous transient conditions. This provides an opportunity to optimize the SAS computational framework such that pre-transient fuel characterization can be shared across multiple channels (fuel pins) and across multiple simulations, thus dramatically reducing overall computational costs. This report summarizes progress toward enhancing the SAS computational framework to support shared, multiple channel fuel performance characterizations intended to significantly reduce computational costs. Preliminary testing has shown that the computational time saved by using the pre-transient sharing capability is approximately equal to the time it takes to perform the pre-transient characterization.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Speedup of UEDGE Parameter Scans Using Machine-Learning Optimized OpenMP Parallelization and a Continuation Solver

This article presents the OpenMP parallelization of the preconditioning Jacobian assembly and right‐hand side residual evaluation in UEDGE. A continuation algorithm, utilizing the internal NKSOL implicit Jacobian‐Free Newton‐Krylov solver to efficiently scan physical parameters, is also presented. The implemented parallelization reduces the computational time for a benchmark scan run on 32 threads by compared to the serial version when using trained random forest regression models to identify the optimal decomposition of the system of equations. Random forest regression models applied to the UEDGE time‐dependent and continuation solver algorithms did not yield meaningful improvement in computational performance. A benchmark DIII‐D gas injection rate scan in the 0.35–0.75 kA interval, performed on a test cluster using the parallelized code and continuation solver, produced 1066 steady‐state solutions with a 22 s average wall‐clock computational time per steady‐state solution.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Linear stability analysis via simulated annealing and accelerated relaxation

Simulated annealing (SA) is a kind of relaxation method for finding equilibria of Hamiltonian systems. A set of evolution equations is solved with SA, which is derived from the original Hamiltonian system so that the energy of the system changes monotonically while preserving Casimir invariants inherent to noncanonical Hamiltonian systems. The energy extremum reached by SA is an equilibrium. Since SA searches for an energy extremum, it can also be used for stability analysis when initiated from a state where a perturbation is added to an equilibrium. The procedure of the stability analysis is explained, and some examples are shown. Because the time evolution is computationally time consuming, efficient relaxation is necessary for SA to be practically useful. An acceleration method is developed by introducing time dependence in the symmetric kernel used in the double bracket, which is part of the SA formulation described here. An explicit formulation for low-beta reduced magnetohydrodynamics (MHD) in cylindrical geometry is presented. In conclusion, since SA for low-beta reduced MHD has two advection fields that relax, it is important to balance the orders of magnitude of these advection fields.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A Pulsar-Inspired Timing Framework for Power System: Optimization and Performance Evaluation

Due to their excellent stability, neutron pulsar stars are considered promising candidate timing sources for power system applications. However, the complexity of pulsar signals necessitates advanced processing algorithms to provide accurate timing references. This paper presents the foundational framework for pulsar signal processing, serving as the basis for further optimization. To enhance the timing accuracy and computation efficiency in pulsar period searches, three algorithms are proposed as the initial optimization step: wavelet de-noising, fast folding, and cross-correlation for profile evaluation. Wavelet de-noising improves signal-to-noise ratio (SNR) by 36%–70%. Fast folding reduces computation time from hundreds of seconds to mere milliseconds. Cross-correlation works better than traditional SNR-based methods by effectively identifying the optimal period. The performance of the proposed algorithms is evaluated using observation data from telescopes. Together, these algorithms significantly improve pulsar timing performance, reducing the error of the Pulse Per Second (PPS) signal from hundreds to tens of microseconds.

Wu, Ori [ORNL] (ORCID:0000000326723410)↗

Parallel transport sweeps on two-dimensional cartesian and hexagonal grids

This paper aims to provide a proof of concept for parallel transport sweeps on two-dimensional hexagonal grids for the discrete ordinates transport equation. While the method is an extension of the popular and well-established Koch-Baker-Alcoulffe (KBA) algorithm, there are significant differences between the cartesian and hexagonal grid and thereafter sweep. The most important is the three-way connectivity of hexagons within the grid which creates greater dependencies between the elements. The KBA method in structured orthogonal grids was first implemented in the DRAGON5 code and the method is first described here. The differences in implementation for the hexagonal grid are also described. Benchmark results are also presented, showing roughly 10 times speedup in computational times with roughly 100 processors, in both cases. (authors)

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Towards a more general understanding of the algorithmic utility of recurrent connections

Lateral and recurrent connections are ubiquitous in biological neural circuits. Yet while the strong computational abilities of feedforward networks have been extensively studied, our understanding of the role and advantages of recurrent computations that might explain their prevalence remains an important open challenge. Foundational studies by Minsky and Roelfsema argued that computations that require propagation of global information for local computation to take place would particularly benefit from the sequential, parallel nature of processing in recurrent networks. Such “tag propagation” algorithms perform repeated, local propagation of information and were originally introduced in the context of detecting connectedness, a task that is challenging for feedforward networks. Here, we advance the understanding of the utility of lateral and recurrent computation by first performing a large-scale empirical study of neural architectures for the computation of connectedness to explore feedforward solutions more fully and establish robustly the importance of recurrent architectures. In addition, we highlight a tradeoff between computation time and performance and construct hybrid feedforward/recurrent models that perform well even in the presence of varying computational time limitations. We then generalize tag propagation architectures to propagating multiple interacting tags and demonstrate that these are efficient computational substrates for more general computations of connectedness by introducing and solving an abstracted biologically inspired decision-making task. Our work thus clarifies and expands the set of computational tasks that can be solved efficiently by recurrent computation, yielding hypotheses for structure in population activity that may be present in such tasks.

59 BASIC BIOLOGICAL SCIENCES↗

The DESC stellarator code suite Part 3: Quasi-symmetry optimization

The DESC stellarator optimization code takes advantage of advanced numerical methods to search the full parameter space much faster than conventional tools. Only a single equilibrium solution is needed at each optimization step thanks to automatic differentiation, which efficiently provides exact derivative information. A Gauss–Newton trust-region optimization method uses second-order derivative information to take large steps in parameter space and converges rapidly. With just-in-time compilation and GPU portability, high-dimensional stellarator optimization runs take orders of magnitude less computation time with DESC compared to other approaches. This paper presents the theory of the DESC fixed-boundary local optimization algorithm along with demonstrations of how to easily implement it in the code. Example quasi-symmetry optimizations are shown and compared to results from conventional tools. Three different forms of quasi-symmetry objectives are available in DESC, and their relative advantages are discussed in detail. In the examples presented, the triple product formulation yields the best optimization results in terms of minimized computation time and particle transport. This paper concludes with an explanation of how the modular code suite can be extended to accommodate other types of optimization problems.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Multiparticle cumulant mapping for Coulomb explosion imaging: Calculations and algorithm

We present a versatile cumulant mapping algorithm for analyzing correlated particle emission, offering insights into complex electronic and nuclear dynamics. Recently, we have demonstrated the use of cumulant mapping to extract information-rich correlations between the momenta of multiple fragments produced in Coulomb explosion imaging experiments [C. Cheng et al., Phys. Rev. Lett. 130, 093001 (2023)]. We define cumulant mapping in terms of histograms, enabling fast computation of linear (additive) observables. However, applying the same algorithm to nonlinear (nonadditive) observables poses challenges, as the computation time of conventional estimators scales nonlinearly with data size. To overcome this, we develop estimators and an accompanying algorithm to enable computationally efficient estimation of the cumulant of interest. Comparisons of computation times and signal-to-noise ratios reveal the superior performance of our approach. This method is demonstrated on the (D+, D+, C+, O+) dissociation channel of CD 2 ⁢O 4+ produced in a strong-field ionization experiment. Additionally, Poisson statistics are used to simulate the two methods and provide insights into the efficiency of our algorithm. The proposed methodology unlocks efficient computation of cumulant mapping for a broader range of complex systems and observables, such as the laser pulse dependence of ionization dynamics.

74 ATOMIC AND MOLECULAR PHYSICS↗

RESOLVING THE ELECTROCHEMICAL EQUATIONS OF A SOLID OXIDE FUEL CELL FOR USE IN TRANSIENT SIMULATION AND INTEGRATION INTO CYBER-PHYSICAL SYSTEMS

A major challenge with complex cyber-physical systems stems from long model computational time that creates a mismatch between the model system and the physical system. The numerical modeling of solid oxide fuel cells (SOFCs) presents particular challenges due to the highly coupled nature of the underlying equations and the multiphysics needed to fully resolve their behavior during a transient event. To this end current approaches revolve around splitting the computational efforts into resolving temperature effects and resolving electrochemical effects. Current methods employed for the transient simulation of an SOFC for implementation in the Hybrid Performance (HyPer) facility cyberphysical plant at the National Energy Technology Laboratory reveal a distinct need for accelerated results with a high degree of stability. To this aim, an investigation into the computational time for the code reveals that the underlying electrochemical algorithm takes an order of magnitude more time than its thermal counterpart and has a tendency to vary in terms of iteration time and as such a rework of the underlying system is proposed. The primary method for accelerated electrochemical algorithm solutions is to employ higher order root finding recipes for the resolution of the highly coupled electrochemical equations. This is done with the intention to reduce the overall number of subiterations necessary for resolving voltage, current density, and species concentration, properties of the fuel cell that are all directly coupled and require nested iterative approaches. The overall objective of this approach is an order of magnitude reduction in calculation time without sacrificing stability and increasing accuracy. Specific approaches involve using both bounded and unbounded techniques, such as the False Position method and the Secant method (or if applicable Newton-Raphson) respectively, the drawbacks being slower convergence for False Position and instability for the Secant or Newton-Raphson methods. Current preliminary results on simplified versions of the parent functions involved for electrochemical calculations indicate a reduction in computational steps by a factor of two for the secant method and a factor of three for Newton-Raphson. When implemented into new modified electrochemical algorithms, the results indicate a possible order of magnitude reduction in calculation time.

Arias, Jesus↗

A deep learning-guided automated workflow in LipidOz for detailed characterization of fungal fatty acid unsaturation by ozonolysis

Understanding fungal lipid biology and metabolism is critical for antifungal target discovery as lipids play central roles in cellular processes. Nuances in lipid structural differences can significantly impact their functions, making it necessary to characterize lipids in detail to enable and understanding of their roles in these complex systems. In particular, lipid double bond (DB) locations are an important component of lipid structure that can only be determined using a few specialized analytical techniques. Ozone-induced dissociation mass spectrometry (OzID-MS) is one such technique that uses ozone to break lipid DBs, producing pairs of characteristic fragments that allow the determination of DB positions. In this work we apply OzID-MS and LipidOz software to analyze the complex lipids of Saccharomyces cerevisiae yeast strains transfected with different fatty acid desaturases from Histoplasma capsulatum to determine the specific unsaturated lipids produce. The automated data analysis in LipidOz made the determination of DB positions from this large dataset more practical, but manual verification for all targets was still time-consuming. The DL model reduces manual involvement in data analysis, but since it was trained using mammalian lipid extracts, the prediction accuracy on yeast-derived data was reduced. We addressed both shortcomings by retraining the DL model to act as a pre-filter to prioritize targets for automated analysis, providing confident manually verified results but requiring less computational time and manual effort. Our workflow resulted in the determination of novel DB positions and enzymatic specificity.

mass spectrometry, deep learning, Lipidomics, doub↗

Accelerated deep self-supervised ptycho-laminography for three-dimensional nanoscale imaging of integrated circuits

Three-dimensional inspection of nanostructures such as integrated circuits is important for security and reliability assurance. Two scanning operations are required: ptychographic to recover the complex transmissivity of the specimen, and rotation of the specimen to acquire multiple projections covering the 3D spatial frequency domain. Two types of rotational scanning are possible: tomographic and laminographic. For flat, extended samples, for which the full 180° coverage is not possible, the latter is preferable because it provides better coverage of the 3D spatial frequency domain compared to limited-angle tomography. It is also because the amount of attenuation through the sample is approximately the same for all projections. However, both techniques are time consuming because of extensive acquisition and computation time. Here, we demonstrate the acceleration of ptycho-laminographic reconstruction of integrated circuits with 16 times fewer angular samples and 4.67 times faster computation by using a physics-regularized deep self-supervised learning architecture. We check the fidelity of our reconstruction against a densely sampled reconstruction that uses full scanning and no learning. As already reported elsewhere [ Opt. Express 28 , 12872 ( 2020 ) OPEXFF 1094-4087 10.1364/OE.379200 ], we observe improvement of reconstruction quality even over the densely sampled reconstruction, due to the ability of the self-supervised learning kernel to fill the missing cone.

47 OTHER INSTRUMENTATION↗

A Novel Framework for Performance Evaluation and Design Optimization of PCM Embedded Heat Exchangers for the Built Environment

This research sheds light on the performance evaluation and design optimization of PCM-HXs for the built environment, addressing several barriers to practical issues to PCM-HX commercialization such as modeling aspects (i.e., modeling expertise and computational / time investment, etc.), manufacturing aspects (i.e., at-scale manufacturing, cost assessments, etc.) and experimental performance assessment (i.e., reliable experimental data, assessment of multiple PCM-working fluid combinations, etc.). We present a novel, comprehensive, and experimentally-validated design optimization framework for PCM-HXs capable of simulating any PCM-HX geometry with reasonable accuracy and significant computational time savings when compared to traditional CFD-based design practices. The framework was validated for a wide range of PCM-HX configurations, including a design optimization for a domestic hot water heater application where TES partially replaces electrical heating input. The resulting PCM-HXs were found to deliver 34-68% of the total daily hot water supply with only 5-10% package volume increase from the water heater, thus within U.S. DOE targets for TES systems. To identify the most promising HXs for PCM applications, first-order geometry and cost analyses were conducted based on off-the-shelf HX products. As part of this work, 9 PCM-HX prototypes were manufactured using additive and conventional manufacturing methods. Detailed economy-of-scale assessments were conducted for the most promising PCM-HXs and were found to have a good outlook for the next 5-10 years. The PCM-HX design optimization framework was validated through comprehensive in-house experimental testing using newly-developed PCM-to-fluid test facilities. In total,10 total in-house component-level experiments were conducted using these prototypes, including 9 with water and 1 with refrigerant (R410A) as the working fluid. It was found that the framework can successfully predict experimental thermal-hydraulic performance within ±10-20% the first time without manual design changes, eliminating the need for time-consuming and expensive prototyping efforts as part of the design process. As part of this work, a publicly-available PCM web tool was released which includes a PCM property database (531 PCMs) and PCM-HX modeling tool to assist the design community on common PCM-HX use-cases, e.g., single/multiple flow path(s) fluid-to-PCM and air-to-fluid-to-PCM configurations (https://ceeeweb.umd.edu/pcmapp/). This work will accelerate the design and time to market for next generation PCM-HXs.

25 ENERGY STORAGE↗

Generating Skeletal Chemical Reaction Mechanisms for Post-Detonation Flows

This report documents the generation of a skeletal chemical reaction mechanism for use with hemispherical pentaerythritol tetranitrate charges. Skeletal mechanisms can substantially reduce computation time while maintaining accuracy. The methodology within uses faster running sample simulations to build a representative thermodynamic state space. These thermodynamic states are used with a constant-volume reactor analysis and a reaction flow analysis to remove unimportant species and reactions from a full chemical reaction mechanism. For the given test case, this results in a 6x speedup in computation time for directly comparable simulations in 2D axisymmetric simulations. We see a 30x speedup in simulations in 3D Cartesian coordainates when compared to a prior full kinetics simulation. There is strong agreement between temperature and species mass fraction profiles between the full and skeletal chemical reaction mechanisms. These methodologies can be applied to any explosive, given the availability of sample simulations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Mobiliti v1.0

Mobiliti is a software platform designed to emulate the dynamics of a regional transportation road network. It is built on open-source software that provides parallel discrete-event simulation. The software is transformative in the area of transportation network simulation because of the geospatial scale and fidelity of the network model and the computational time it takes to model a full day of travel demand. For example, it runs a simulation of the entire San Francisco Bay Area, with a network model of ~1M links and a population that completes ~19M trips in ~5 minutes. This scale of simulation has not been attempted with existing simulation models due to the complexity of the model and the computational time it would take to complete. The intent of the software is to create a digital twin capability for cities to evaluate consequences of infrastructure or policy changes on road network dynamics.

Macfarlane, Jane↗

Comparison of excess free energy at an interface according to the applied interpolation scheme for elasticity: A phase-field method

Phase-field modeling is an effective simulation technique for modeling microstructure evolution of elastically anisotropic systems. To introduce the elastic energy contribution in a phase field model, an interpolation scheme is used to define the mechanical properties within the phases and across the continuous interface. Several existing interpolation schemes introduce a potential excess elastic energy at the interface, which undesirable effect on microstructure evolution needs to be evaluated. In this study, we focused on three interpolation schemes including Khachaturyan’ scheme (KHS), Voigt–Taylor’s scheme (VTS), and Steinbach–Apel’s scheme (SAS). Comparisons of these schemes’ performances were performed in three configuration types using the MOOSE (Multiphysics Object-Oriented Simulation Environment) framework: bi-crystal, isotropic particle-matrix and anisotropic particle-matrix. The contribution of excess elastic energy on the interface energy as a function of interface width and the computational time to steady-state were evaluated in these three configurations. SAS introduces the lowest excess elastic energy contribution and the VTS has the biggest contribution amongst the considered schemes. Moreover, when modeling precipitation in an anisotropic elastic material, the SAS approach seems to predict more physical convex shapes during growth, making it preferable to KHS and VTS. Finally, as currently implemented, SAS requires the largest computational time and KHS requires the smallest time to reach steady-state amongst the considered schemes.

36 MATERIALS SCIENCE↗

Towards Efficient Alternating Current Optimal Power Flow Analysis on Graphical Processing Units

We present a solution of sparse ACOPF analysis on GPU. In particular, we discuss the performance bottlenecks and detail our efforts to accelerate the linear solver, a core component of ACOPF that dominates the computational time. ACOPF solutions of two large-scale systems, synthetic Northeast (25,000 buses) and Eastern (70,000 buses) \cite{birchfield2017tamu-cases} on GPU show promising speed-up compared to CPU based solution using a state-of-the-art solver. To our knowledge, this is the first result demonstrating acceleration of sparse ACOPF on GPUs.

Power grid analysis, GPU↗