Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Operator inference”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Design-Technology Co-Optimization for NVM-based Neuromorphic Processing Elements

An emerging use-case of machine learning (ML) is to train a model on a high-performance system and deploy the trained model on energy-constrained embedded systems. Neuromorphic hardware platforms, which operate on principles of the biological brain, can significantly lower the energy overhead of a machine learning inference task, making these platforms an attractive solution for embedded ML systems. In this paper, we present a design-technology tradeoff analysis to implement such inference tasks on the processing elements (PEs) of a Non-Volatile Memory (NVM)-based neuromorphic hardware. Through detailed circuit-level simulations at scaled process technology nodes, we show the negative impact of technology scaling on the information-processing latency, which impacts the quality-of-service (QoS) of an embedded ML system. At a finer granularity, the latency inside a PE depends on 1) the delay introduced by parasitic components on its current paths, and 2) the varying delay to sense different resistance states of its NVM cells. Based on these two observations, we make the following three contributions. First, on the technology front, we propose an optimization scheme where the NVM resistance state that takes the longest time to sense is set on current paths having the least delay, and vice versa, reducing the average PE latency, which improves the QoS. Second, on the architecture front, we introduce isolation transistors within each PE to partition it into regions that can be individually power-gated, reducing both latency and energy. Finally, on the system-software front, we propose a mechanism to leverage the proposed technological and architectural enhancements when implementing a machine-learning inference task on neuromorphic PEs of the hardware. Evaluations with a recent neuromorphic hardware architecture show that our proposed design-technology co-optimization approach improves both performance and energy efficiency of machine-learning inference tasks without incurring high cost-per-bit.

42 ENGINEERING↗

WellPINN: Accurate Well Representation for Transient Fluid Pressure Diffusion in Subsurface Reservoirs With Physics‐Informed Neural Networks

Accurate representation of pumping wells is essential for reliable reservoir characterization and simulation of operational scenarios in subsurface flow models. Physics-informed neural networks (PINNs) are emerging as a promising alternative to numerical models for reservoir modeling, offering seamless integration of monitoring data and governing physical equations. However, existing PINN-based studies face major challenges in capturing fluid pressure near wells when using a source/sink term, particularly during the early stages after pumping begins. We address this problem by introducing WellPINN, a workflow in which an initially trained PINN infers fluid pressure across the entire reservoir domain using a large equivalent well radius. This initial PINN solution is then locally refined around the well by a set of subdomain PINNs that are trained for smaller equivalent well radii. Continuity across these subdomain interfaces as well as at the initial condition is ensured by hard-constraining each PINN on its subdomain boundary. Our results demonstrate WellPINN as the first workflow of its kind to focus on accurate inference of fluid pressure from pumping rates throughout the entire injection period, significantly advancing the potential of PINNs for inverse modeling and operational scenario simulations. All data and code for this paper are openly available at https://doi.org/10.20350/DIGITALCSIC/17260.

58 GEOSCIENCES↗

Empirical radius formulas for canonical neutron stars from bidirectionally selecting features of equations of state in extended Bayesian analyses of observational data

Significant advancement in Bayesian inference of nuclear equation of state (EOS) from gravitational wave and x-ray observations of neutron stars (NSs) has been made by the nuclear astrophysics community especially since GW170817. By extending the traditional Bayesian analysis which normally ends at presenting the marginalized posterior probability distribution functions (PDFs) of individual EOS parameters and their correlations (or sometimes only the Pearson correlation coefficients which are only reliably useful when the variables are linearly correlated while they are actually often not), we search for a data-driven and robust empirical formula for the radius 𝑅 1.4 of canonical NSs in terms of the characteristic EOS parameters (features). We also identify the single most important but currently poorly known EOS parameter for determining the 𝑅 1.4 . Using three regression-model-building methodologies: bidirectional stepwise feature selection, least absolute shrinkage selection operator (LASSO) regression, and neural network regression on a large set of posterior EOSs and the corresponding 𝑅 1.4 values inferred from earlier comprehensive Bayesian analyses of NS observational data, we systematically and rigorously develop the most probable 𝑅 1.4 formulas with varying statistical accuracy and technical complexity. Here, the most important EOS parameters for determining 𝑅 1.4 are found consistently in each of the feature selection processes to be (in order of decreasing importance): curvature 𝐾 sym , slope 𝐿, skewness 𝐽 sym of nuclear symmetry energy, skewness 𝐽 0 , incompressibility 𝐾 0 of symmetric nuclear matter, and the magnitude 𝐸 sym ⁡(𝜌 0 ) of symmetry energy at the saturation density 𝜌 0 of nuclear matter.

Bayesian methods↗

Physics-based modeling and information-theoretic sensor and settings selection for tool wear detection in precision machining

Precision machining of metals is an energy intensive process with applications and impacts across the manufacturing industry. The energy efficiency, product yield, and maintenance of the precision machine require a digital twin that can assist with prognostics and health management. In this report a physics-based model is developed and validated against face milling data, and then used for the timely and precise inference of machining faults that cannot be measured directly. Computer numerical control (CNC) measurements of power and force are used through this physics-based machining model to predict deviations of the outputs of power consumption and cutting forces during normal operation. A model-based fault detection and isolation methodology is applied to determine the optimal (traditional and available) sensor suite and the test settings (admissible input values) that improve the inference of tool wear in face milling. The optimal sensor suite and input test settings are obtained by solving a mixed integer non-linear program that optimizes information-theoretic metrics relevant to the detection and isolation of tool wear from steady-state or transient machining measurements. Dynamic time warping and k—NN classification are then used to validate the robustness of the optimal design for fault detection test design, including the optimal sensor suite.

42 ENGINEERING↗

nPINNs: Nonlocal physics-informed neural networks for a parametrized nonlocal universal Laplacian operator. Algorithms and applications

Physics-informed neural networks (PINNs) are effective in solving inverse problems based on differential and integro-differential equations with sparse, noisy, unstructured, and multifidelity data. PINNs incorporate all available information, including governing equations (reflecting physical laws), initial-boundary conditions, and observations of quantities of interest, into a loss function to be minimized, thus recasting the original problem into an optimization problem. In this paper, we extend PINNs to parameter and function inference for integral equations such as nonlocal Poisson and nonlocal turbulence models, and we refer to them as nonlocal PINNs (nPINNs). The contribution of the paper is three-fold. First, we propose a unified nonlocal Laplace operator, which converges to the classical Laplacian as one of the operator parameters, the nonlocal interaction radius $\delta$ goes to zero, and to the fractional Laplacian as $\delta$ goes to infinity. This universal operator forms a super-set of classical Laplacian and fractional Laplacian operators and, thus, has the potential to fit a broad spectrum of data sets. We also provide theoretical convergence rates with respect to $\delta$ and verify them via numerical experiments. Second, we use nPINNs to estimate the two parameters, $\delta$ and $\alpha$, characterizing the kernel of the unified operator. The strong non-convexity of the loss function yielding multiple (good) local minima reveals the occurrence of the operator mimicking phenomenon, that is, different pairs of estimated parameters could produce multiple solutions of comparable accuracy. Third, we propose another nonlocal operator with spatially variable order $\alpha(y)$, which is more suitable for modeling turbulent Couette flow. Our results show that nPINNs can jointly infer this function as well as $\delta$. More importantly, these parameters exhibit a universal behavior with respect to the Reynolds number, a finding that contributes to our understanding of nonlocal interactions in wall-bounded turbulence.

97 MATHEMATICS AND COMPUTING↗

In-situ sensor monitoring of multi-class gas porosity formation in laser powder bed fusion using convolutional neural network

In-situ monitoring of defect formation remains a significant challenge in the laser powder bed fusion (LPBF) process. Recent advances have enabled real-time defect detection with machine learning and in-situ sensing technologies; however, most studies focus on binary classification of keyhole pores, limiting nuanced multi-class pore differentiation and formation mechanisms. This work introduces a multi-class pore detection framework (no pore, small pores < 15 µm, and large pores > 15 µm) by leveraging photodiode sensor data alongside high-fidelity synchrotron X-ray imaging. The 15 µm threshold is selected to distinguish between two fundamentally different defect mechanisms, following the physical size-mechanism boundary established by prior high-resolution synchrotron X-ray characterization of Al6061 LPBF. Distinguishing these classes is critical because large keyhole pores are structurally detrimental, whereas small gas pores are often benign, requiring different process control strategies. Thermal emission monitoring data collected simultaneously with high-speed X-ray imaging at the Stanford Synchrotron Radiation Lightsource (SSRL), are correlated with subsurface melt pool dynamics to establish ground truth. Continuous Wavelet Transform (CWT) with optimized parameters converts the photodiode time-series signals into time–frequency images, facilitating feature extraction. Convolutional Neural Networks (CNN) are then applied for real-time multi-class pore classification in an average inference time of 1 ms per signal window. It achieves 79% accuracy and an Area Under the Receiver Operating Characteristic curve (AUC ROC) score of 0.89 with five-fold cross-validation. The results demonstrate that coupling CWT-based feature engineering with CNN architecture enables reliable multi-class pore detection in Al6061 builds using affordable in-situ sensors. This approach advances scalable and affordable quality assurance in additive manufacturing by moving beyond binary defect detection toward more nuanced classification of porosity mechanisms with in-situ sensors and machine learning.

Laser powder bed fusion, Multi-class pores, In-sit↗

Data-driven models of nonautonomous systems

Nonautonomous dynamical systems are characterized by time-dependent inputs, which complicates the discovery of predictive models describing the spatiotemporal evolution of the state variables of quantities of interest from their temporal snapshots. When dynamic mode decomposition (DMD) is used to infer a linear model, this difficulty manifests itself in the need to approximate the time-dependent Koopman operators. Our approach is to approximate the original nonautonomous system with a modified system derived via a local parameterization of the time-dependent inputs. The modified system comprises a sequence of local parametric systems, which are subsequently approximated by a parametric surrogate model using the DRIPS (dimension reduction and interpolation in parameter space) framework. The offline step of DRIPS relies on DMD to build a linear surrogate model, endowed with reduced-order bases for the observables mapped from training data. The online step interpolates on suitable manifolds to construct a sequence of iterative parametric surrogate models; the target/test parameter points on these manifolds are specified by a local parameterization of the test time-dependent inputs. Here, we use numerical experimentation to demonstrate the robustness of our method and compare its performance with that of deep neural networks.

97 MATHEMATICS AND COMPUTING↗

CHMMPY: A python package for constrained Hidden Markov Models

SAND2025-11909O chmmpy software analyzes multivariate timeseries data to detect patterns. It uses a Hidden Markov Model (HMM) and application-specific constraints that reflect known relationships among hidden states to accomplish this. The chmmpy software provides a generic framework for expressing application-specific constraints and supporting constrained HMM inference using optimization solvers. chmmpy is available on GitHub. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Hart, William↗

A Machine Learning Framework to Deconstruct the Primary Drivers for Electricity Market Price Events

As the electricity grid is moving towards a 100% Renewable Energy Source Bulk Power Grid, the overall operations of the power system operations and electricity markets are changing. The electricity markets are not only dispatching resources economically but also taking into account various controllable actions like renewable curtailment, transmission congestion mitigation, and energy storage optimization to make sure the grid is operating reliably. As a result, price formations in electricity markets have become quite complex. Traditional root cause analysis and statistical approaches are rendered inapplicable to analyze and infer the main drivers behind price formation in the modern grid and markets with variable renewable energy (VRE). In this paper, we propose a machine learning analysis framework to deconstruct some primary drivers for price formation in modern electricity markets with high renewable energy and the outcomes can be utilized for various critical aspects of market design, renewable dispatch and curtailment, operations, and cyber-security applications. The framework can be applied to any ISO or market data and in this paper it is applied to open-source publicly available datasets from California Independent System Operator (CAISO) and ISO New England.

machine learning (ML), electricity markets, Renewa↗

A Bimodal Diagnostic Cloud Fraction Parameterization. Part I: Motivating Analysis and Scheme Description

Cloud fraction parameterizations are beneficial to regional, convection-permitting numerical weather prediction. For its operational regional midlatitude forecasts, the Met Office uses a diagnostic cloud fraction scheme that relies on a unimodal, symmetric subgrid saturation-departure distribution. This scheme has been shown before to underestimate cloud cover and hence an empirically based bias correction is used operationally to improve performance. This first of a series of two papers proposes a new diagnostic cloud scheme as a more physically based alternative to the operational bias correction. The new cloud scheme identifies entrainment zones associated with strong temperature inversions. Additionally, for model grid boxes located in this entrainment zone, collocated moist and dry Gaussian modes are used to represent the subgrid conditions. The mean and width of the Gaussian modes, inferred from the turbulent characteristics, are then used to diagnose cloud water content and cloud fraction. It is shown that the new scheme diagnoses enhanced cloud cover for a given gridbox mean humidity, similar to the current operational approach. It does so, however, in a physically meaningful way. Using observed aircraft data and ground-based retrievals over the southern Great Plains in the United States, it is shown that the new scheme improves the relation between cloud fraction, relative humidity, and liquid water content. An emergent property of the scheme is its ability to infer skewed and bimodal distributions from the large-scale state that qualitatively compare well against observations. A detailed evaluation and resolution sensitivity study will follow in Part II.

54 ENVIRONMENTAL SCIENCES↗

ML–Enabled FPGA Framework for Fast Quantum State Discrimination in Mid-Circuit Measurement Regimes

Accurate and low-latency quantum state discrimination is essential for protocols involving mid-circuit measurement (MCM) and conditional feed-forward. In superconducting quantum systems, conventional readout pipelines transfer measurement data to host processors for post-processing, introducing millisecond-scale delays that far exceed qubit coherence times. To overcome this bottleneck, we present an in-situ machine learning (ML) inference engine implemented on an FPGA for real-time quantum state discrimination. Our design performs inference directly on digitized readout signals with 40 ns latency, supports both qubit and qutrit readout, and enables conditional operations without host-side intervention. This capability is critical for MCM and for feedback-driven protocols such as quantum error correction. We validate the system on superconducting transmon hardware, demonstrating robust discrimination fidelity across multiple qubit and qutrit channels. We further demonstrate conditional qutrit logic driven by FPGA-resident classification, highlighting the potential of low-latency ML-on-FPGA control for NISQ applications and scalable fault-tolerant quantum computing.

Vora, Neel [Lawrence Berkeley National Laboratory ↗

Inferring the scrape-off layer heat flux width in a divertor with a low degree of axisymmetry

Plasma facing components (PFCs) in the next generation of tokamak devices will operate in challenging environments, with heat loads predicted to exceed 10 MWm -2 . The magnitude of these heat loads is set by the width of the channel, the ‘scrape-off layer’ (SOL), into which heat is exhausted, and can be characterised by an e-folding length scale for the decay of heat flux across the channel. It is expected this channel will narrow as tokamaks move towards reactor relevant conditions. Understanding the processes involved in setting the SOL heat flux width is imperative to be able to predict the heat loads PFCs must handle in future devices. Measurements of the SOL width are performed on the high-field spherical tokamak, ST40, using a newly commissioned infrared thermography system. With its high on-axis toroidal magnetic field (≥1.5 T) ST40 is uniquely positioned to investigate the influence of toroidal field on the heat flux width in spherical tokamaks, whilst also extending measurements of the SOL width in spherical tokamaks to increased poloidal field (≥0.3 T). Due to the divertor on ST40 having a low degree of axisymmetry, it is necessary for a set of radial measurements of the heat flux to be taken across the divertor, made possible using an automated toolchain that fully incorporates its 3D geometry. These radial profiles are combined with the magnetic geometry of the plasma to infer the width of the SOL, with both Eich and double exponential profiles of heat flux observed. A reduction in the heat flux is observed toroidally across part of the divertor, along with increased heat loads observed locally around the edges of the tiles. Future work in characterising the impact of tile misalignment and uncertainties in the reconstructed divertor magnetic geometry is required in order to further understand the observed heat flux patterns, as are additional investigations into the role potentially being played by an inhomogeneous sheath electric field.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Modeling of particle transport, neutrals and radiation in magnetically-confined plasmas with Aurora

In this work, we present Aurora, an open-source package for particle transport, neutrals and radiation modeling in magnetic confinement fusion plasmas. Aurora's modern multi-language interface enables simulations of 1.5D impurity transport within high-performance computing frameworks, particularly for the inference of particle transport coefficients. A user-friendly Python library allows simple interaction with atomic rates from the Atomic Data and Atomic Structure database as well as other sources. This enables a range of radiation predictions, both for power balance and spectroscopic analysis. We discuss here the superstaging approximation for complex ions, as a way to group charge states and reduce computational cost, demonstrating its wide applicability within the Aurora forward model and beyond. Aurora also facilitates neutral particle analysis, both from experimental spectroscopic data and other simulation codes. Leveraging Aurora's capabilities to interface SOLPS-ITER results, we demonstrate that charge exchange is unlikely to affect the total radiated power from the ITER core during high performance operation. Finally, we describe the ImpRad module in the one modeling framework for integrated task framework, developed to enable experimental analysis and transport inferences on multiple devices using Aurora.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Distributed-Memory Sparse Deep Neural Network Inference Using Global Arrays

Partitioned Global Address Space (PGAS) models exhibit tremendous promise in developing efficient and productive distributed-memory parallel applications. They have been used extensively in scientific computations due to conveniently offering a ``shared-memory''-like model and convenient interfaces that separate communication with synchronization. Traditionally, PGAS communication models have been applied to dense/contiguously distributed data, but most modern applications depict varied levels of sparsity. Existing PGAS models require certain adaptations to support distributed sparse computations, since associated computations often require matrix arithmetic, in addition to data movement. The Global Arrays toolkit from Pacific Northwest National Laboratory (PNNL) is one of the earliest PGAS models to combine one-sided data communication and distributed matrix operations and is still used in the popular NWChem quantum chemistry suite. Recently, we have expanded the Global Arrays toolkit to support common sparse operations, like sparse matrix-dense matrix multiplies (SpMM), sparse matrix-sparse matrix multiplication (SpGEMM) and Sampled Dense-Dense Matrix Multiplication (SDDMM). As it turns out, these operations are the bedrock of sparse Deep Learning (DL); sparse deep neural networks and Graph Neural Networks (GNNs) have gained increasing attention recently in achieving speedups on training and inference with reduced memory footprints. Unlike scientific applications in High Performance Computing (HPC), modern (distributed-memory capable) DL toolkits often rely on non-standardized and closed-source vendor software optimizations, creating challenges in software-hardware co-design at scale. Our goal is to support a variety of distributed-memory sparse matrix operations and helper functions in the newly created Sparse Global Arrays (SGA), such that it is possible to build portable and productive Machine Learning scenarios for algorithm/software and hardware codesign purposes. Contemporary data-parallel schemes for training/inference are undergoing a major overhaul since model replication limits scalability and causes resource inefficiencies. As such, we have adopted tensor parallelism in decomposing the model and inputs, to mitigate memory issues. Current implementation is built on top of MPI and uses CPUs to maximize the portability across the platforms.

Distributed computing, machine learning↗

Projecting Future Energy Production from Operating Wind Farms in North America. Part I: Dynamical Downscaling

Abstract New simulations at 12-km grid spacing with the Weather and Research Forecasting (WRF) Model nested in the MPI Earth System Model (ESM) are used to quantify possible changes in wind power generation potential as a result of global warming. Annual capacity factors (CF; measures of electrical power production) computed by applying a power curve to hourly wind speeds at wind turbine hub height from this simulation are also used to illustrate the pitfalls in seeking to infer changes in wind power generation directly from low-spatial-resolution and time-averaged ESM output. WRF-derived CF are evaluated using observed daily CF from operating wind farms. The spatial correlation coefficient between modeled and observed mean CF is 0.65, and the root-mean-square error is 5.4 percentage points. Output from the MPI-WRF Model chain also captures some of the seasonal variability and the probability distribution of daily CF at operating wind farms. Projections of mean annual CF (CF A ) indicate no change to 2050 in the southern Great Plains and Northeast. Interannual variability of CF A increases in the Midwest, and CF A declines by up to 2 percentage points in the northern Great Plains. The probability of wind droughts (extended periods with anomalously low production) and wind bonus periods (high production) remains unchanged over most of the eastern United States. The probability of wind bonus periods exhibits some evidence of higher values over the Midwest in the 2040s, whereas the converse is true over the northern Great Plains. Significance Statement Wind energy is playing an increasingly important role in low-carbon-emission electricity generation. It is a “weather dependent” renewable energy source, and thus changes in the global atmosphere may cause changes in regional wind power production (PP) potential. We use PP data from operating wind farms to demonstrate that regional simulations exhibit skill in capturing actual power production. Projections to the middle of this century indicate that over most of North America east of the Rocky Mountains annual expected PP is largely unchanged, as is the probability of extended periods of anomalously high or low production. Any small declines in annual PP are of much smaller magnitude than changes due to technological innovation over the last two decades.

Meteorology & Atmospheric Sciences↗

Deep Learning for Subsurface Flow: A Comparative Study of U‐Net, Fourier Neural Operators, and Transformers in Underground Hydrogen Storage

Subsurface flow research is essential for the sustainable management of natural resources and the environment. Deep learning (DL) has significantly advanced this field by developing efficient and accurate surrogate models to replace computationally expensive physics‐based simulations. These surrogate models are commonly used to predict the spatiotemporal evolution of state variables, such as gas saturation and reservoir pressure, in heterogeneous geological formations. Despite the various DL models applied to this task, there is a lack of studies systematically comparing their performance. This absence of comparative analysis leads to somewhat arbitrary DL model selection in subsurface flow research, resulting in suboptimal performance and potentially inaccurate predictions. To bridge this gap, we conduct a systematic comparison study of three popular DL architectures—U‐Net, Fourier Neural Operators (FNO), and Segmentation Transformer (SETR)—in surrogate modeling of underground hydrogen storage (UHS). We focus on UHS due to its promise of enhancing clean energy resilience and its cyclic operational conditions that represent common scenarios in various subsurface applications. We evaluate the models based on accuracy, training cost, and inference speed. The comparison shows that U‐Net achieves the highest accuracy, followed by SETR and FNO. Despite its lower accuracy, FNO has the highest inference speed. SETR offers competitive accuracy with the least training memory usage, demonstrating the potential of transformers in learning subsurface flow. Our results provide guidance for selecting DL models for surrogate modeling in a wide range of subsurface flow problems.

42 ENGINEERING↗

Structure-aware Initialization via Numerical Continuation and Informed Priors

Scientific machine learning (SciML) often operates in ill-conditioned, weakly identifiable regimes due to limited data or indirect observations. In such settings, optimization and inference are highly sensitive to the starting point, making initialization--often under-reported--a consequential degree of freedom. Random initialization is not a neutral default as it induces an implicit prior over candidate solutions and can systematically bias the result, producing large run-to-run variability. Here, we formalize this view by treating initialization as a hidden confounder in SciML and develop a unifying theory for structure-aware initialization via numerical continuation, constructing warm starts from related problem instances. Across representative tasks, including physics-informed neural networks, maximum likelihood estimation, and variational inference, warm starts have been shown to consistently reduce optimization effort and improve reliability.

Data integrity↗

epi_inference v. 1.0

SAND2021-0503 O The epi_inference software supports the prediction of disease transmission parameters from data. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Hart, William↗