Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Incremental Computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Addressing quantum’s “fine print” with efficient state preparation and information extraction for quantum algorithms and geologic fracture networks

Abstract Quantum algorithms provide an exponential speedup for solving certain classes of linear systems, including those that model geologic fracture flow. However, this revolutionary gain in efficiency does not come without difficulty. Quantum algorithms require that problems satisfy not only algorithm-specific constraints, but also application-specific ones. Otherwise, the quantum advantage carefully attained through algorithmic ingenuity can be entirely negated. Previous work addressing quantum algorithms for geologic fracture flow has illustrated core algorithmic approaches while incrementally removing assumptions. This work addresses two further requirements for solving geologic fracture flow systems with quantum algorithms: efficient system state preparation and efficient information extraction. Our approach to addressing each is consistent with an overall exponential speed-up.

58 GEOSCIENCES↗

Advanced Computational Modeling of High-Level Waste Vitrification at the Hanford Site

The U.S. Department of Energy (DOE) has selected vitrification for stabilizing legacy tank waste at the Hanford site, where radioactive waste from plutonium production was historically stored in underground tanks. This waste will be separated into low-activity waste (LAW) and high-level waste (HLW) fractions and processed at the Waste Treatment and Immobilization Plant (WTP). At WTP, glass melters are used for the vitrification of radioactive tank waste, transforming it into a stable borosilicate glass form for safe long-term storage. The melter vessel is constructed from highly durable and heat-resistant materials, where the vitrification process occurs. The main regions that are modeled are the melt pool, plenum, cold cap, riser/discharge chamber, and surrounding structure with insulation layers. Forced convection induced by air bubblers at the base of the melter ensure uniform temperature distribution and provide heat to the cold cap layer. The cold cap is a region of reacting batch feed that floats on top of the molten glass and is where the batch-to-glass reactions occur. Joule heating provided by electrodes mounted along the vertical walls of the melter and immersed directly in the glass, generates the necessary heat for the net endothermic conversion processes that occur in the cold cap. The high temperatures, radioactivity, and opaque nature of the glass prevent direct observation inside the melters. Therefore, computational models are essential for providing insight into factors that affect melter throughput. Thermocouples in the plenum provide operators with plenum temperature measurements. Operational adjustments include bubbling rate, voltage supplied to the electrodes, feed adjustments, and glass removal rate. Different computational fluid dynamics (CFD) models have been developed, each serving a specific purpose. There are CFD models of different scale melters, as well as models that capture the two-phase flow interfaces of rising bubbles in the molten glass or models with a simplified molten glass region so that the surrounding structure and plenum can be feasibly incorporated. Pilot-scale melter models have been developed to serve as validation of the methods employed in the simulation of the full-scale WTP melters. Models incorporating resolved bubbling are used to develop momentum source terms to implement into a single phase, multi-region, steady-state flow model that is being validated by measured process parameters such as glass production rate, voltage, input power, plenum temperatures, etc. The resolved bubbling model uses the multiphase volume of fluid approach to model the system with a high-resolution interface capturing scheme to maintain sharp interfaces between the molten glass and the air phase. The suite of CFD models is continually being improved to incorporate more realistic physics and achieve faster turnaround time. For example, an incremental controller is implemented to automatically adjust electrode voltage within the simulation to a molten glass set point temperature of 1150°C. Newer models feature improved meshes to ensure conformal meshes between regions and eliminate unnecessary mesh refinement in areas that are not of interest (such as boundary layers in offgas ports). Instead of explicitly modeling the structural, refractory, and insulation layers of the melter, a thermal resistance approach is used with published correlations used for boundary conditions. The development of robust and efficient CFD models will be instrumental in enabling the WTP to successfully fulfill its mission of safely stabilizing legacy nuclear waste.

12 - MGMT OF RADIOACTIVE AND NON-RADIOACTIVE WASTE↗

An Incremental Tensor Train Decomposition Algorithm

We present a new algorithm for incrementally updating the tensor train decomposition of a stream of tensor data. This new algorithm, called the tensor train incremental core expansion (TT-ICE) improves upon the current state-of-the-art algorithms for compressing in tensor train format by developing a new adaptive approach that incurs significantly slower rank growth and guarantees compression accuracy. This capability is achieved by limiting the number of new vectors appended to the TT-cores of an existing accumulation tensor after each data increment. These vectors represent directions orthogonal to the span of existing cores and are limited to those needed to represent a newly arrived tensor to a target accuracy. We provide two versions of the algorithm: TT-ICE and TT-ICE accelerated with heuristics (TT-ICE*). Here, we provide a proof of correctness for TT-ICE and empirically demonstrate the performance of the algorithms in compressing large-scale video and scientific simulation datasets. Compared to existing approaches that also use rank adaptation, TT-ICE* achieves 57× higher compression and up to 95% reduction in computational time.

97 MATHEMATICS AND COMPUTING↗

Primal interface debonding formulation for finite strain isotropic plasticity

In this work, a framework is developed for modeling ductile damage of nonlinear materials whose plastic deformation is characterized using rate independent classical plasticity. This method relies on the assumption that the free energy can be decomposed into elastic, plastic and damage parts. A thermodynamically consistent method is derived which satisfies the second law of thermodynamics in the Clausius–Duhem inequality form. The dissipation associated with plasticity takes place in the domain only, while damage dissipation is localized to the interface. The method is developed using Variational Multiscale ideas to obtain definitions of the interface fluxes within a primal formulation analogous to the Discontinuous Galerkin method, which ensures weakly vanishing interface gap prior to reaching a damage initiation criterion. The local nonlinear problem to calculate both plastic deformation gradient and damage variable follows an incremental approach similar to classical plasticity return mapping algorithm. This elastoplastic damage formulation is developed for material undergoing finite strain, and it naturally accommodates a trapezoidal traction separation law (TSL) whose shape can be varied to model either ductile interface behavior or brittle interface behavior. The formulation's performance is assessed through modeling a patch test and a compact tension specimen.

42 ENGINEERING↗

Scalable Pattern Matching in Metadata Graphs via Constraint Checking

Pattern matching is a fundamental tool for answering complex graph queries. Unfortunately, existing solutions have limited capabilities: They do not scale to process large graphs and/or support only a restricted set of search templates or usage scenarios. Moreover, the algorithms at the core of the existing techniques are not suitable for today’s graph processing infrastructures relying on horizontal scalability and shared-nothing clusters, as most of these algorithms are inherently sequential and difficult to parallelize. In this article we present an algorithmic pipeline that bases pattern matching on constraint checking. The key intuition is that each vertex and edge participating in a match has to meet a set of constraints implicitly specified by the search template. These constraints can be verified independently and typically are less expensive to compute than searching the full template. The pipeline we propose generates these constraints and iterates over them to eliminate all the vertices and edges that do not participate in any match, thus reducing the background graph to a subgraph that is the union of all template matches—the complete set of all vertices and edges that participate in at least one match. Additional analysis can be performed on this annotated, reduced graph, such as full match enumeration, match counting, or computing vertex/edge centrality. Furthermore, a vertex-centric formulation for constraint checking algorithms exists, and this makes it possible to harness existing high-performance, vertex-centric graph processing frameworks. This technique (i) enables highly scalable pattern matching in metadata (labeled) graphs; (ii) supports arbitrary patterns with 100% precision; (iii) enables tradeoffs between precision and time-to-solution, while always selects all vertices and edges that participate in matches, thus offering 100% recall; and (iv) supports a set of popular data analytics scenarios. We implement our approach on top of HavoqGT, an open-source asynchronous graph processing framework, and demonstrate its advantages through strong and weak scaling experiments on massive scale real-world (up to 257 billion edges) and synthetic (up to 4.4 trillion edges) labeled graphs, respectively, and at scales (1,024 nodes / 36,864 cores), orders of magnitude larger than used in the past for similar problems. This article serves two purposes: First, it synthesises the knowledge accumulated during a long-term project. Second, it presents new system features, usage scenarios, optimizations, and comparisons with related work that strengthen the confidence that pattern matching based on iterative pruning via constraint checking is an effective and scalable approach in practice. The new contributions include the following: (i) We demonstrate the ability of the constraint checking approach to efficiently support two additional search scenarios that often emerge in practice, interactive incremental search and exploratory search. (ii) We empirically compare our solution with two additional state-of-the-art systems, Arabsque and TriAD. (iii) We show the ability of our solution to accommodate a more diverse range of datasets with varying properties, e.g., scale, skewness, label distribution, and match frequency. (iv) We introduce or extend a number of system features (e.g., work aggregation, load balancing, and the ability to cap the generated traffic) and design optimizations and demonstrate their advantages with respect to improving performance and scalability. (v) We present bottleneck analysis and insights into artifacts that influence performance. (vi) We present a theoretical complexity argument that motivates the performance gains we observe.

97 MATHEMATICS AND COMPUTING↗

Snow ALbedo eVOlution (SALVO) Campaign Broadband Albedo from April - June, 2024 in Utqiagivk, AK level a1

A field-portable broadband (285 – 2800 nm) albedometer was used to make spatially distributed albedo measurements on tundra and sea ice surfaces. The albedometer consists of paired upward-looking and downward-looking pyranometers, which were both connected to a data logger. The instrument was mounted approximately 1 m above the surface using a tripod and was placed on a 1.4 m-long boom to minimize the impacts of shading from the operator and to observe surfaces undisturbed by footprints (see Appendix for photos of measurement setup and uncertainty assessment). Albedo measurements were taken parallel to the 200-m albedo lines at 5-m increments (41 measurements) ~1.2 m south of the line. On the operator’s end of the boom, there was a bubble level that was aligned with the bubble level on the upward-looking pyranometer. To take a measurement, the operator first relocated the tripod to the measurement location, then leveled the instrument and held it level for at least twice the pyranometers’ response time (5 or 15 seconds, see below), and finally depressed a trigger on the data logger. The data logger recorded the instantaneous voltage on both pyranometers, the measurement number, and the time. The data logger also converted the voltages to irradiances, and from these computed the ratio (outgoing/incoming) for albedo, which could be checked in the field. The operator recorded in a field notebook the measurement number that corresponded with the locations on the line and any pertinent notes (e.g., invalid measurements). With this setup, a trained operator could measure a 200-m albedo line (41 measurements) in approximately 30 minutes. Measurements were made within 3 hours of solar noon. The data logger had sufficient storage capacity to record all measurements from the campaign, but data were downloaded to a computer after each measurement day.

54 ENVIRONMENTAL SCIENCES↗

Logic in Memory Emulator

Logic in Memory Emulator (LiME) is a hardware/software tool specially designed for memory system evaluation and experiment. Emerging memories display a wide range of bandwidths, latencies, and capacities, making it challenging for the computer architect to navigate the design space of potential memory configurations, and for the application developer to assess performance implications of using such memories. With the LiME framework, architectural ideas can be prototyped in great detail yet with sufficient performance to support realistic evaluation on long running applications. LiME consists of two fundamental components: 1) the hardware and OS infrastructure for the emulator, and 2) a suite of benchmark applications to assist in characterizing the performance of current and future computer architectures. Some of the applications have been collected from other open source projects. Uses: Logging, replay and analysis of an application's memory behavior Evaluate impact of emerging memory technology on application performance. Emulate complex memory interactions in whole applications orders of magnitude faster than software simulation. Emulate acceleration hardware co-located with the memory subsystem. Features: Capture and log external memory accesses to a separate off-chip memory device without affecting application execution. Memory traces include the address, length, timestamp, and optionally the data for each transaction. Captured trace data can be saved to an SD card for off-line analysis. Configure a wide range of memory latencies in sub-nanosecond increments that encompass highbandwidth and storage class memories. Specify regions of interest (ROI) in applications to reduce the amount of trace data captured for analysis. Currently supports execution on Xilinx Zynq SoC which integrates an ARM processor with FPGA logic on a single device. Applications can be run under Linux or in bare metal mode on the ARM cores.

Jain, AbhishekK↗

Phase-Field Modeling of Damage Evolution in Ceramic Matrix Composite (CMC) and Environmental Barrier Coating (EBC)

Ceramic matrix composites (CMCs) protected by environmental barrier coatings (EBCs) present a promising materials solution for next generation gas turbines. Developments of more robust and efficient EBCs and mechanically tougher CMCs are thus of significant technological importance. Here we develop a phase-field modeling framework that incorporates the thermally grown oxide (TGO), recognized as a critical factor for degradation and failure of EBCs. We simulate crack growth in the TGO and the potential extension into the bond coat / CMC substrate. The model efficiently takes account of the large inelastic deformation induced by the severe volume expansion of TGO, thanks to our recently developed, so-called incremental realization of inelastic deformation (IRID) algorithm. A phase-field model is built for damage evolution in CMCs including crack growth and interfacial sliding. The effects of fiber layout and interfacial sliding on the macroscopic toughness of CMCs are revealed by large-scale simulations and compared to experiments.

advanced energy systems and materials↗

Quantifying local rearrangements in three-dimensional granular materials: Rearrangement measures, correlations, and relationship to stresses

Quantifying the ways in which local particle rearrangements contribute to macroscopic plasticity is one of the fundamental pursuits of granular mechanics and soft matter physics. Here we examine local rearrangements that occur naturally during the deformation of three samples of 3D granular materials subjected to distinct boundary conditions by employing in situ x-ray measurements of particle-resolved structure and stress. We focus on five distinct rearrangement measures, their statistics, interrelationships, contributions to macroscopic deformation, repeatability, and dependence on local structure and stress. Our most significant findings are that local rearrangements (1) are correlated on a scale of three to four particle diameters, (2) exhibit volumetric strain-shear strain and nonaffine displacement-rotation coupling, (3) exhibit correlations that suggest either rearrangement repeatability or that rearrangements span multiple steps of incremental sample strain, and (4) show little dependence on local stress but correlate with quantities describing local structure, such as porosity. Our results are presented in the context of relevant plasticity theories and are consistent with recent findings suggesting that local structure may play at least as important of a role as local stress in determining the nature of local rearrangements.

58 GEOSCIENCES↗

Analysis of the SiMPL Method for Density-Based Topology Optimization

We present a rigorous convergence analysis of a new method for density-based topology optimization that provides pointwise bound-preserving design updates and faster convergence than other popular first-order topology optimization methods. Due to its strong bound preservation, the method is exceptionally robust, as demonstrated in numerous examples here and in the companion article [D. Kim et al., Struct. Multidiscip. Optim., 68 (2025), 74]. Furthermore, it is easy to implement with clear structure and analytical expressions for the updates. Our analysis covers two versions of the method, characterized by the employed line search strategies. We consider a modified Armijo backtracking line search and a Bregman backtracking line search. For both line search algorithms, our algorithm delivers a strict monotone decrease in the objective function and further intuitive convergence properties, e.g., strong and pointwise convergence of the density variables on the active sets, norm convergence to zero of the increments, convergence of the Lagrange multipliers, and more. In addition, the numerical experiments demonstrate apparent mesh-independent convergence of the algorithm. Here, we refer to the new algorithm as the SiMPL method (pronounced “simple”), which stands for Sigmoidal Mirror descent with a Projected Latent variable.

97 MATHEMATICS AND COMPUTING↗

How efficiently can AI recognize Wireless Devices?

This poster presents a hardware benchmarking methodology for a 3-layer CNN waveform classifier deployed using ONNX Runtime on an NVIDIA Jetson AGX Orin. The dataset consist of 9 signal types, -30 to +30 dB SNR with 5dB increments. Benchmarking on the Jetson AGX Orin gave an accuracy of 91.9% and GPU throughput of 107,120 predictions/sec (23× faster than CPU). The Jetson GPU reached approximately 27M samples/sec with stable performance but fell below the 40 MHz rate needed for real-time radio feeds. Sustained testing of 5 minutes confirmed stable performance with no memory leaks, establishing a reproducible benchmarking baseline for future edge-deployment optimization.

99 - GENERAL AND MISCELLANEOUS↗

Airfoil Computational Fluid Dynamics - 2k shapes, 25 AoA's, 3 Re numbers

This dataset contains aerodynamic quantities - including flow field values (momentum, energy, and vorticity) and summary values (coefficients of lift, drag, and momentum) - for 1,830 airfoil shapes computed using the HAM2D CFD (computational fluid dynamics) model. The airfoil shapes were designed using the separable shape tensor parameterization that encodes two-dimensional shapes as elements of the Grassmann manifold. This data-driven approach learns two independent spaces of parameter from a collection of sample airfoils. The first captures large-scale, linear perturbations, and the second defines small-scale, higher-order perturbations. For this dataset, we used the G2Aero database of over 19,000 airfoil shapes to learn a parameter space that captured a wide array of shape characteristics. We sampled airfoil designs over both parameter spaces to explore the full range of possible shape variations. The aerodynamic quantities for the generated airfoil were obtained using the HAM2D code, which is a finite-volume Reynolds-averaged Navier-Stokes (RANS) flow solver. We employ a fifth-order WENO scheme for spatial reconstruction with Roe's flux difference scheme for inviscid flux and second-order central differencing for viscous flux. A preconditioned GMRES method is applied for implicit integration. The Spalart-Allmaras 1-eq turbulence model is used for the turbulence closure, and the Medida-Baeder 2-eq transition model is applied to account for the effects of laminar turbulent transition. The airfoil grid is generated with a total of 400 points on the airfoil surface, the initial wall-normal spacing of y+ = 1, and an outer boundary located at 300 chord lengths away from the wall. The CFD simulations are performed at a freestream Mach number of 0.1, for or three different Reynolds' numbers (3M, 6M, and 9M), and for 25 angles of attack from -4 deg. to 20 deg. with 1 degree increments. Across all these various parameters, this dataset includes the results from over 250,000 CFD simulations. The simulations were performed using the Bridges-2 system at the Pittsburgh Supercomputing Center in February 2023 as part of the INTEGRATE project funded by the Advanced Research Projects Agency - Energy, in the U.S. Department of Energy. The data was collected, reformatted, and preprocessed for this OEDI submission in July 2023 under the Foundational AI for Wind Energy project funded by the U.S. Department of Energy Wind Energy Technologies Office. This dataset is intended to serve as a benchmark against which new artificial intelligence (AI) or machine learning (ML) tools may be tested. Baseline AI/ML methods for analyzing this dataset have been implemented, and a link to their repository containing those models has been provided. The .h5 data file structure can be found in the GitHub Repository resource under explore_airfoil_2k_data.ipynb.

2k↗

Devices and methods for increasing the speed and efficiency at which a computer is capable of modeling a plurality of random walkers using a density method

A method for increasing a speed or energy efficiency at which a computer is capable of modeling a plurality of random walkers. The method includes defining a virtual space in which a plurality of virtual random walkers will move among different locations in the virtual space, wherein the virtual space comprises a plurality of vertices and wherein the different locations are ones of the plurality of vertices. A corresponding set of neurons in a spiking neural network is assigned to a corresponding vertex such that there is a correspondence between sets of neurons and the plurality of vertices, wherein a spiking neural network comprising a plurality of sets of spiking neurons is established. A virtual random walk of the plurality of virtual random walkers is executed using the spiking neural network, wherein executing includes tracking how many virtual random walkers are at each vertex at a given time increment.

Aimone, James Bradley↗

High-dimensional discrete Fourier transform gates with a quantum frequency processor

The discrete Fourier transform (DFT) is of fundamental interest in photonic quantum information, yet the ability to scale it to high dimensions depends heavily on the physical encoding, with practical recipes lacking in emerging platforms such as frequency bins. In this article, we show that d -point frequency-bin DFTs can be realized with a fixed three-component quantum frequency processor (QFP), simply by adding to the electro-optic modulation signals one radio-frequency harmonic per each incremental increase in d . We verify gate fidelity F W > 0.9997 and success probability P W > 0.965 up to d = 10 in numerical simulations, and experimentally implement the solution for d = 3, utilizing measurements with parallel DFTs to quantify entanglement and perform tomography of multiple two-photon frequency-bin states. Our results furnish new opportunities for high-dimensional frequency-bin protocols in quantum communications and networking.

97 MATHEMATICS AND COMPUTING↗

Snow ALbedo eVOlution (SALVO) Campaign Spectral Albedo and Related Measurements from April - June, 2024 in Utqiagivk, AK

A field-portable spectroradiometer, referred to herein as an ‘ASD’, was used to make spatially-distributed spectral albedo (350 – 2500 nm) measurements on tundra and sea ice surfaces. The ASD detector is carried in a backpack and controlled via a computer mounted on the front of the operator (see Figure 1). The ASD measures the spectral irradiance from a fiber optic cable that is routed from the backpack to a custom, gooseneck cosine collector mounted on the end of a 1-m long boom (Grenfell and Perovich, 2008). The boom was held at hip height (approximately 1 m) and had an integrated bubble level for levelling. To make an albedo measurement, first the operator collect an incident (down-welling) irradiance, followed by a reflected (up-welling) measurement. The time between incident and reflected measurements was typically between 11 and 26 seconds (interquartile range). For each measurement, 10 spectra are averaged together. Albedo is calculated as the ratio of the reflected to incident measurement, which obviates the need for absolute radiometric calibration. Albedo measurements were taken parallel to the 200-m albedo lines at 5-m increments (41 measurements) ~1 m south of the line. While the ASD operator was making measurements, an assistant kept notes on the scan number associated with each measurement, the surface type (see below), and collected photos of each measurement (see companion oblique photos data archive). Measurements were made within 3 hours of solar noon.

ASD Spectroradiometer↗

MPAS-Seaice (v1.0.0): sea-ice dynamics on unstructured Voronoi meshes

Abstract. We present MPAS-Seaice, a sea-ice model which uses the Model for Prediction Across Scales (MPAS) framework and spherical centroidal Voronoi tessellation (SCVT) unstructured meshes. As well as SCVT meshes, MPAS-Seaice can run on the traditional quadrilateral grids used by sea-ice models such as CICE. The MPAS-Seaice velocity solver uses the elastic–viscous–plastic (EVP) rheology and the variational discretization of the internal stress divergence operator used by CICE, but adapted for the polygonal cells of MPAS meshes, or alternatively an integral (“finite-volume”) formulation of the stress divergence operator. An incremental remapping advection scheme is used for mass and tracer transport. We validate these formulations with idealized test cases, both planar and on the sphere. The variational scheme displays lower errors than the finite-volume formulation for the strain rate operator but higher errors for the stress divergence operator. The variational stress divergence operator displays increased errors around the pentagonal cells of a quasi-uniform mesh, which is ameliorated with an alternate formulation for the operator. MPAS-Seaice shares the sophisticated column physics and biogeochemistry of CICE and when used with quadrilateral meshes can reproduce the results of CICE. We have used global simulations with realistic forcing to validate MPAS-Seaice against similar simulations with CICE and against observations. We find very similar results compared to CICE, with differences explained by minor differences in implementation such as with interpolation between the primary and dual meshes at coastlines. We have assessed the computational performance of the model, which, because it is unstructured, runs with 70 % of the throughput of CICE for a comparison quadrilateral simulation. The SCVT meshes used by MPAS-Seaice allow removal of equatorial model cells and flexibility in domain decomposition, improving model performance. MPAS-Seaice is the current sea-ice component of the Energy Exascale Earth System Model (E3SM).

58 GEOSCIENCES↗

A Kaczmarz-inspired approach to accelerate the optimization of neural network wavefunctions

Neural network wavefunctions optimized using the variational Monte Carlo method have been shown to produce highly accurate results for the electronic structure of atoms and small molecules, but the high cost of optimizing such wavefunctions prevents their application to larger systems. We propose the Subsampled Projected-Increment Natural Gradient Descent (SPRING) optimizer to reduce this bottleneck. SPRING combines ideas from the recently introduced minimum-step stochastic reconfiguration optimizer (MinSR) and the classical randomized Kaczmarz method for solving linear least-squares problems. We demonstrate that SPRING outperforms both MinSR and the popular Kronecker-Factored Approximate Curvature method (KFAC) across a number of small atoms and molecules, given that the learning rates of all methods are optimally tuned. For example, on the oxygen atom, SPRING attains chemical accuracy after forty thousand training iterations, whereas both MinSR and KFAC fail to do so even after one hundred thousand iterations.

97 MATHEMATICS AND COMPUTING↗