Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Incremental Computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

A Novel Approach for Bounding the Stress Experienced by the Core of Utility-Scale Printed Circuit Heat Exchangers Under Thermohydraulic Loads

Printed circuit heat exchangers (PCHEs) have rapidly gained popularity since being introduced nearly three decades ago, and they are currently widely deployed in the petrochemical and aviation industry. Their compactness, thermohydraulic efficiency, inherent suitability for high temperature/pressure fluids containment, and demonstrated durability are some of the reasons the nuclear industry is seeking to adopt this technology as well. However, the relatively strict nuclear-related regulatory design code, especially when classified as critical to the safety of the reactors, is posing challenges to adopting the technology. From stress analysis point of view, one undesirable feature of PCHEs is their geometrical complexity, which is implied by their multilength-scale features. As a result, a full-scale model of a utility-scale exchanger cannot simply be solved on a computer because meshing such components results in a vast number of degrees-of-freedom. Here, this work seeks to address the challenge of stress analyses to PCHEs by presenting a method to simplify the geometry of PCHE designs. The models proposed by this work can be practically analyzed on a standard computer and provide a path for implementing ASME design rules. The analyses presented herein are divided into five separate investigations. Each is carried out to incrementally simplify the analyzed model by addressing features such as the shapes of the flow passages, the complex distribution of stress in large components, the three-dimensionality of the stress and strain, the thermal stresses caused by thermohydraulic operation observed experimentally and more.

42 ENGINEERING↗

Adaptive language model training for molecular design

Abstract The vast size of chemical space necessitates computational approaches to automate and accelerate the design of molecular sequences to guide experimental efforts for drug discovery. Genetic algorithms provide a useful framework to incrementally generate molecules by applying mutations to known chemical structures. Recently, masked language models have been applied to automate the mutation process by leveraging large compound libraries to learn commonly occurring chemical sequences (i.e., using tokenization) and predict rearrangements (i.e., using mask prediction). Here, we consider how language models can be adapted to improve molecule generation for different optimization tasks. We use two different generation strategies for comparison, fixed and adaptive. The fixed strategy uses a pre-trained model to generate mutations; the adaptive strategy trains the language model on each new generation of molecules selected for target properties during optimization. Our results show that the adaptive strategy allows the language model to more closely fit the distribution of molecules in the population. Therefore, for enhanced fitness optimization, we suggest the use of the fixed strategy during an initial phase followed by the use of the adaptive strategy. We demonstrate the impact of adaptive training by searching for molecules that optimize both heuristic metrics, drug-likeness and synthesizability, as well as predicted protein binding affinity from a surrogate model. Our results show that the adaptive strategy provides a significant improvement in fitness optimization compared to the fixed pre-trained model, empowering the application of language models to molecular design tasks.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

SCALE 6.2.4 Validation: Radiation Shielding

For safe and reliable use of computer codes by the community, accuracy must be clearly evaluated. In particular, the nuclear reactor engineering and licensing field needs accurate tools for radiation shielding modeling. Monaco with Automated Variance Reduction using Importance Calculations (MAVRIC) is one such tool, with built-in variance reduction methods distributed within the SCALE code, and its validity is demonstrated in this report for the released version 6.2.4. Representative benchmarks corresponding to shielding analysis are selected for the validation study. Typical experimental results analyzed from those benchmarks include neutron fluxes, detector count rates, detector energy response functions, neutron and gamma doses, foil neutron activation rates and activities, neutron leakage fluxes, and skyshine dose rates. Thousands of points of comparison between experiment and calculation are presented in this work. Other than rare outliers typically explained by either a lack of information or large uncertainties in the experiment conditions, material, or dimensions, MAVRIC agrees well with the experiment results. MAVRIC is also compared to Monte Carlo N-Particle (MCNP) calculations when available, and both codes generally produce good agreements within estimated uncertainties. The selected benchmarks are obtained from reliable sources such as the International Criticality Safety Benchmark Evaluation Project Handbook (ICSBEP Handbook), the Shielding Integral Benchmark Archive & Database (SINBAD), and other shielding validation work found in the literature. Additional datapoints and benchmarks will be added to future versions of this report to incrementally expand the shielding validation suite incrementally.

61 RADIATION PROTECTION AND DOSIMETRY↗

Software stewardship and advancement of a high-performance computing scientific application: QMCPACK

Here, we provide an overview of the software engineering efforts and their impact in QMCPACK, a production-level ab-initio Quantum Monte Carlo open-source code targeting high-performance computing (HPC) systems. Aspects included are: (i) strategic expansion of continuous integration (CI) targeting CPUs, using GitHub Actions own runners, and NVIDIA and AMD GPUs used in pre-exascale systems, (ii) incremental reduction of memory leaks using sanitizers, (iii) incorporation of Docker containers for CI and reproducibility, and (iv) refactoring efforts to improve maintainability, testing coverage, and memory lifetime management. We quantify the value of these improvements by providing metrics to illustrate the shift towards a predictive, rather than reactive, maintenance approach. Our goal, in documenting the impact of these efforts on QMCPACK, is to contribute to the body of knowledge on the importance of research software engineering (RSE) for the stewardship and advancement of community HPC codes to enable scientific discovery at scale.

97 MATHEMATICS AND COMPUTING↗

Implementation of Plot File Testing in the DYNA3D/ParaDyn Software Quality Assurance Suite

Automated testing of DYNA3D/ParaDyn plot files was added to the DYNA3D/ParaDyn software quality assurance (SQA) test suite. The new capability extracts select data from the plot files generated during each verification run and compares it to the same baseline answers used to verify the problem. Deviations between baseline answers and plot file values are reported in the same manner as solution discrepancies, and differences in precision levels between the baseline answers and plot file results are accounted for. The new testing leverages the existing SQA test suite framework and test problems and the Python Mili reader and minimally increases the overall run time (< 5%) of the SQA test suite. This new capability provides incremental end-toend testing of the most common DYNA3D/ParaDyn simulation workflows.

42 ENGINEERING↗

Learning constitutive relations using symmetric positive definite neural networks

In this work, we present a new neural-network architecture, called the Cholesky-factored symmetric positive definite neural network (SPD-NN), for modeling constitutive relations in computational mechanics. Instead of directly predicting the stress of the material, the SPD-NN trains a neural network to predict the Cholesky factor of the tangent stiffness matrix, based on which the stress is calculated in incremental form. As a result of this special structure, SPD-NN weakly imposes convexity on the strain energy function, satisfies the second order work criterion (Hill's criterion) and time consistency for path-dependent materials, and therefore improves numerical stability, especially when the SPD-NN is used in finite element simulations. Depending on the types of available data, we propose two training methods, namely direct training for strain and stress pairs and indirect training for loads and displacement pairs. We demonstrate the effectiveness of SPD-NN on hyperelastic, elasto-plastic, and multiscale fiber-reinforced plate problems from solid mechanics. The generality and robustness of SPD-NN make it a promising tool for a wide range of constitutive modeling applications.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

I-GCN: A Graph Convolutional Network Accelerator with Runtime Locality Enhancement through Islandization

In this paper, we propose a novel hardware accelerator for GCN inference called I-GCN that significantly improves data locality and reduces unnecessary computation through a new online graph restructuring algorithm we refer to as islandization. The proposed algorithm finds clusters of nodes with strong internal but weak external connections. The islandization process yields two major benefits. First, by processing islands rather than individual nodes, there is better on-chip data reuse and fewer off-chip memory accesses. Second, there is less redundant computation as aggregation for common/shared neighbors in an island can be reused. The parallel search, identification, and leverage of graph islands are all handled purely in hardware at runtime working in an incremental pipelined manner. This is done without any preprocessing of the graph data or adjustment of the GCN model structure.

Geng, Tong↗

Software engineering to sustain a high-performance computing scientific application: QMCPACK

We provide an overview of the software engineering efforts and their impact in QMCPACK, a production-level ab-initio Quantum MonteCarlo open-source code targeting high-performance computing (HPC) systems. Aspects included are: (i) strategic expansion ofcontinuous integration (CI) targeting CPU, using GitHub Actions runners, and graphics processing units (GPU) in pre-exascalesystems, using self-hosted hardware; (ii) incremental reduction of memory leaks using sanitizers, (iii) incorporation of Dockercontainers for CI and reproducibility, and (iv) refactoring efforts to improve maintainability, testing coverage, and memory lifetime management. We quantify the value of these improvements by providing metrics to illustrate the shift towards a predictive, rather than reactive, sustainable maintenance approach. Our goal, in documenting the impact of these efforts on QMCPACK, is to contribute to the body of knowledge on the importance of research software engineering (RSE) for the sustainability of community HPC codes and scientific discovery at scale.

Godoy, William↗

AMM: Adaptive Multilinear Meshes

Adaptive representations are increasingly indispensable for reducing the in-memory and on-disk footprints of large-scale data. Usual solutions are designed broadly along two themes: reducing data precision, e.g., through compression, or adapting data resolution, e.g., using spatial hierarchies. Additionally, recent research suggests that combining the two approaches, i.e., adapting both resolution and precision simultaneously, can offer significant gains over using them individually. However, there currently exist no practical solutions to creating and evaluating such representations at scale. In this work, we present a new resolution-precision-adaptive representation to support hybrid data reduction schemes and offer an interface to existing tools and algorithms. Through novelties in spatial hierarchy, our representation, Adaptive Multilinear Meshes (AMM), provides considerable reduction in the mesh size. AMM creates a piecewise multilinear representation of uniformly sampled scalar data and can selectively relax or enforce constraints on conformity, continuity, and coverage, delivering a flexible adaptive representation. AMM also supports representing the function using mixed-precision values to further the achievable gains in data reduction. We describe a practical approach to creating AMM incrementally using arbitrary orderings of data and demonstrate AMM on six types of resolution and precision datastreams. By interfacing with state-of-the-art rendering tools through VTK, we demonstrate the practical and computational advantages of our representation for visualization techniques. With an open-source release of our tool to create AMM, we make such evaluation of data reduction accessible to the community, which we hope will foster new opportunities and future data reduction schemes.

97 MATHEMATICS AND COMPUTING↗

Endogenizing Probabilistic Resource Adequacy Risks in Deterministic Capacity Expansion Models

In this work, we demonstrate how power system capacity expansion models can understate the stochastic effects of thermal outages when considering resource availabilities on an hourly expected value basis, yielding system designs with multiple orders of magnitude more shortfall risk than stated adequacy targets. We develop a novel approximation approach to efficiently endogenize awareness of this risk in a deterministic, linear capacity expansion framework. We compare this approach to exogenous tuning of an energy reserve margin, the leading alternative method to compensate for unmodeled probabilistic shortfall risk. Empirical results from a test system show that the new endogenous method cost-effectively meets all regional reliability targets with a single optimization solve, and produces a near-identical system design as the incumbent method without the need for repeated re-optimizations to find an appropriate reserve level. The endogenous method may also use iterative re-optimizations to further improve solution quality, although these incremental benefits were modest in the system studied.

capacity expansion modeling↗

In situ compression artifact removal in scientific data using deep transfer learning and experience replay

The massive amount of data produced during simulation on high-performance computers has grown exponentially over the past decade, exacerbating the need for streaming compression and decompression methods for efficient storage and transfer of this data---key to realizing the full potential of large-scale computational science. Lossy compression approaches such as JPEG when applied to scientific simulation data realized as a stream of images can achieve good compression rates but at the cost of introducing compression artifacts and loss of information. This paper develops a unified framework for in situ compression artifact removal in which the fully convolutional neural network architectures are combined with scalable training, transfer learning, and experience replay to achieve superior accuracy and efficiency while significantly decreasing the storage footprint as compared with the traditional optimization-based approaches. We demonstrate the proposed approach and compare it with compressed sensing postprocessing and other baseline deep learning models using climate simulations and nuclear reactor simulations, both of which are driven by hyperbolic partial differential equations. Our approach when applied to remove the compression artifacts on the JPEG-compressed nuclear reactor simulation data (using a transfer-trained model that was pretrained on the climate simulation data and updated incrementally as the nuclear reactor simulation progressed), achieved a significant improvement---mean peak signal-to-noise ratio of 42.438 as compared with 27.725 obtained with the compressed sensing approach.

97 MATHEMATICS AND COMPUTING↗

Interpreting and Stabilizing Machine-Learning Parametrizations of Convection

Neural networks are a promising technique for parameterizing subgrid-scale physics (e.g., moist atmospheric convection) in coarse-resolution climate models, but their lack of interpretability and reliability prevents widespread adoption. For instance, it is not fully understood why neural network parameterizations often cause dramatic instability when coupled to atmospheric fluid dynamics. This paper introduces tools for interpreting their behavior that are customized to the parameterization task. First, we assess the nonlinear sensitivity of a neural network to lower-tropospheric stability and the midtropospheric moisture, two widely studied controls of moist convection. Second, we couple the linearized response functions of these neural networks to simplified gravity wave dynamics, and analytically diagnose the corresponding phase speeds, growth rates, wavelengths, and spatial structures. To demonstrate their versatility, these techniques are tested on two sets of neural networks, one trained with a super-parameterized version of the Community Atmosphere Model (SPCAM) and the second with a near-global cloud-resolving model (GCRM). Additionally, even though the SPCAM simulation has a warmer climate than the cloud-resolving model, both neural networks predict stronger heating/drying in moist and unstable environments, which is consistent with observations. Moreover, the spectral analysis can predict that instability occurs when GCMs are coupled to networks that support gravity waves that are unstable and have phase speeds larger than 5 m s -1 . In contrast, standing unstable modes do not cause catastrophic instability. Using these tools, differences between the SPCAM-trained versus GCRM-trained neural networks are analyzed, and strategies to incrementally improve both of their coupled online performance unveiled.

54 ENVIRONMENTAL SCIENCES↗

ECRAM Materials, Devices, Circuits and Architectures: A Perspective

Abstract Non‐von‐Neumann computing using neuromorphic systems based on two‐terminal resistive nonvolatile memory elements has emerged as a promising approach, but its full potential has not been realized due to the lack of materials and devices with the appropriate attributes. Unlike memristors, which require large write currents to drive phase transformations or filament growth, electrochemical random access memory (ECRAM) decouples the “write” and “read” operations using a “gate” electrode to tune the conductance state through charge‐transfer reactions, and every electron transferred through the external circuit in ECRAM corresponds to the migration of ≈1 ion used to store analogue information. Like static dopants in traditional semiconductors, electrochemically inserted ions modulate the conductivity by locally perturbing a host's electronic structure; however, ECRAM does so in a dynamic and reversible manner. The resulting change in conductance can span orders of magnitude, from gradual increments needed for analog elements, to large, abrupt changes for dynamically reconfigurable adaptive architectures. In this in‐depth perspective, the history of ECRAM, the recent progress in devices spanning organic, inorganic, and 2D materials, circuits, architectures, the rich portfolio of challenging, fundamental questions, and how ECRAM can be harnessed to realize a new paradigm for low‐power neuromorphic computing are discussed.

Talin, A. Alec↗

NREL's Journey with HPC in the Cloud and Hybrid Computing

This is a planned lightning talk at the NLIT Summit 2025 conference. This would serve as somewhat of a progress update to the presentation I gave at re:Invent 2024 back in November which can be seen here: https://www.youtube.com/watch?t=2133&v=NMq3kL9qObU&feature=youtu.be (my section begins at the included timestamp value). This presentation discusses our usage of Cloud-hosted HPC systems, and in what circumstances they benefit our researchers strategically. We have been making incremental progress in this area since that recording, so for this presentation I would include our latest experiences and observations as we are beginning to implement a hybrid HPC solution. We're in the midst of a cross-team effort of implementing a prototype hybridization solution which would allow users to strategically burst jobs to the cloud. In this talk for NLIT, I would detail lessons-learned, non-starters, architecture diagrams, and other implementation details that may benefit those interested as we continue our experimentation. Our prototype may not be complete by the time of this presentation, but even in the discovery phase of our anticipated design we've discovered a lot of information from others who have worked on hybrid solutions that are worth sharing.

97 MATHEMATICS AND COMPUTING↗

Streaming Generalized Canonical Polyadic Tensor Decompositions

In this paper, we develop a method which we call OnlineGCP for computing the Generalized Canonical Polyadic (GCP) tensor decomposition of streaming data. GCP differs from traditional canonical polyadic (CP) tensor decompositions as it allows for arbitrary objective functions which the CP model attempts to minimize. This approach can provide better fits and more interpretable models when the observed tensor data is strongly non-Gaussian. In the streaming case, tensor data is gradually observed over time and the algorithm must incrementally update a GCP factorization with limited access to prior data. In this work, we extend the GCP formalism to the streaming context by deriving a GCP optimization problem to be solved as new tensor data is observed, formulate a tunable history term to balance reconstruction of recently observed data with data observed in the past, develop a scalable solution strategy based on segregated solves using stochastic gradient descent methods, describe a software implementation that provides performance and portability to contemporary CPU and GPU architectures and integrates with Matlab for enhanced usability, and demonstrate the utility and performance of the approach and software on several synthetic and real tensor data sets.

97 MATHEMATICS AND COMPUTING↗

First simultaneous observation of co- and counter-current fast-ion losses in the ASDEX Upgrade tokamak

In ITER and future fusion power plants, the source of the fusion born alpha particles is almost isotropic in pitch angle, thus having co- and counter-current populations. For trapped ions, the co-current side of the orbit corresponds to its outer leg, while the counter-current side corresponds to the inner leg. Understanding the mechanisms responsible for the fast-ion losses (FILs) is critical for future magnetically confined fusion power plants. To further study the interplay of fast ions with plasma instabilities, a double pinhole collimator has been developed for a Fast-Ion Loss Detector (FILD) in the ASDEX Upgrade tokamak (AUG). This new FILD opens the operational window to simultaneous measurements of the co- and counter-current ion velocity-space. In this paper, the first results for the AUG double collimator FILD detector are shown. The commissioning of this new probe is carried out in H-mode plasmas with an on-axis magnetic field $B_0 = -2.5$ T, and a plasma current $I_{\textrm{p}} = 0.7$ MA. Simultaneous co- and counter-current FILs have been measured. Both have shown a similar dependence on Ion Cyclotron Resonance Heating (ICRH) power, where the main difference is the intensity of the losses, with the co-losses being an order of magnitude larger. Toroidal Alfvén eigenmode-coherent ICRH-only losses have been identified for the co-current ions. Additionally, the presence of Edge Localized Modes during the discharge were shown to increment Neutral Beam Injection prompt losses, while partially mitigating ICRH-driven losses on both co- and counter- sides of the velocity-space. Finally, a very trapped and high gyroradius losses, with an unclear origin, have been measured in the co- and counter-current velocity-space. The computed ion trajectories show that these ions remain permanently near the vessel wall, suggesting that they are accelerated within the scrape-off layer.

ELM↗

Micropolar Elastoplasticity Using a Fast Fourier Transform‐Based Solver

ABSTRACT This work presents a micromechanical spectral formulation for obtaining the full‐field and homogenized response of elastoplastic micropolar composites. A closed‐form radial‐return mapping is derived from thermodynamics‐based micropolar elastoplastic constitutive equations to determine the increment of plastic strain necessary to return the generalized stress state to the yield surface, and the algorithm implementation is verified using the method of numerically manufactured solutions. Then, size‐dependent material response and micro‐plasticity are shown as features that may be efficiently simulated in this micropolar elastoplastic framework. The computational efficiency of the formulation enables the generation of large datasets in reasonable computing times.

42 ENGINEERING↗

Hybrid data-driven and model-informed online tool wear detection in milling machines

Precision machining tool wear is responsible for low product throughput and quality. Monitoring the tool wear online is vital to prevent degradation in machining quality. However, direct real-time tool wear measurement is not practical. This paper presents residual-based anomaly detection models, combining a hybrid model comprised of a physics-based model and a data-driven model (a decision tree or a neural network) to predict signals of interest (e.g., power or forces) under nominal conditions, followed by Page’s cumulative sum test for detecting tool wear on-line using the computer numerical control machine measurements. The most informative features are ranked using dynamic programming and its approximation variants from real-time measurements and machine settings, such as the width of cut, depth of cut, feed rate and spindle speed, that serve as inputs to the predictive models. The baseline nominal model is incrementally updated with experimental data via a gradient boosted adaptation model to generate the residuals that account for discrepancies between the actual machine data under normal conditions and the baseline nominal model predictions. The hybrid model is validated against 20 Mazak milling machine experimental tests and one Haas run-to-failure experiment. The proposed anomaly detector is applied to synthetic data from simulations of the physics-based model at different operating conditions, measurement noise levels, and tool wear levels, and the methods were able to achieve an overall 92% accuracy in data with 1% noise. The anomaly detection methods based on hybrid model reduced the false alarms of either the data-driven or physical-based models alone, and are found to be capable of good online detection of tool wear.

42 ENGINEERING↗