Engineering PapersSearch

SEARCH · Engineering Papers

Results for “black-box models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Empirical Comparison of Machine Learning Approaches for Black-Box Modeling of Power Conversion System Dynamics

Inverter-based resources are key components in modern power systems, but accurately modeling their complex behavior can be challenging. Standard, generic converter models often oversimplify inverter dynamics, leading to significant errors in predicting performance. In this work, we compare several data-driven machine learning (ML) approaches for inverter modeling, performing experiments on power conversion systems, systematically varying input conditions, and recording the resulting voltages and currents. The ML models were then trained on this measured data to capture the inverter's dynamic response and to predict the inverter's output current. A performance comparison between the four ML models under study is conducted, laying the foundation for future work on hardware implementation for real-time inference.

30 DIRECT ENERGY CONVERSION

A Data-Driven Method for Modeling Creep-Fatigue Stress- Strain Behavior Using Neural ODEs

In this paper, we introduce a data-driven machine learning approach for modeling one-dimensional stress–strain behavior under cyclic loading, utilizing experimental data from the nickel-based Alloy 617. The study employs uniaxial creep–fatigue test data acquired under various loading histories and compares two distinct neural network-based ODE models. The first model, known as the black-box model, comprehensively describes the strain–stress relationship using a Neural ODE equation. To interpret this black-box model, we apply the Sparse Identification of Nonlinear Dynamical Systems (SINDy) technique, transforming the black-box model into an equation-based model using symbolic regression. The second model, the Neural flow rule model, incorporates Hooke’s Law for the linear elastic component, with the nonlinear part characterized by a Neural ODE. Both models are trained with experimental data to accurately reflect the observed stress–strain behavior. We conduct a detailed comparison with the standard Chaboche model, which includes three back stresses. Our results demonstrate that the neural network-based ODE models precisely capture the experimental creep–fatigue mechanical behavior, exceeding the standard Chaboche model’s accuracy. Furthermore, an interpretable model derived from the black-box neural ODE model through symbolic regression achieves accuracy comparable to the Chaboche model, enhancing its interpretability. The results highlight the potential of neural network-based ODE models to depict complex creep–fatigue behavior, eliminating the necessity for experts to define a specific, material-focused model form.

creep-fatigue

Physics vs structure: A systematic benchmark of learning strategies for multi-zone building thermal dynamics

Recent advances in physics-informed and data-driven machine learning promise improved thermal models for advanced building control, yet there is limited quantitative evidence on when added physics structure and architectural complexity are beneficial. Here, this work presents a systematic benchmark of five representative system identification methods for modeling multi-zone building thermal dynamics: linear state-space models, multi-layer perceptrons, neural state-space models, neural ordinary differential equations, and physically-consistent neural networks. The methods are evaluated across multiple data regimes and zone coupling strategies. Using a high-fidelity multi-zone commercial building emulator, we examine short-term and long-term prediction accuracy, computational efficiency, and ease of development. Our results reveal critical trade-offs between prediction performance, model complexity, and physical consistency. We demonstrate that decoupled, nonlinear black-box models consistently outperform coupled physics-constrained architectures in both predictive accuracy and out-of-distribution robustness in majority of the test cases for the building type considered in the study. Our findings quantify the cost of complexity in building thermal modeling and provide concrete, actionable, scenario-based guidelines for selecting model classes for control-oriented applications.

Building thermal modeling

Surrogate model evaluation and building energy benchmarking for commercial buildings

Building energy consumption benchmarking involves challenges associated with various energy patterns for different building types; heating, ventilating, and air-conditioning (HVAC) system types; and climates. Given significant variation in energy use patterns, accurate prediction of long-term energy use using surrogate models remains challenging. Multiple linear regression (MLR) is commonly used for building energy benchmarking because of its simple structure; however, it lacks accuracy compared to other black-box models. Although many studies have compared surrogate models and offer guidance on model selection based on metrics, they do not provide detailed analysis on improving the surrogate model accuracy. In this paper, we implement a surrogate model using polynomial ridge regression (i.e., MLR with interaction terms combined with ridge regularization) for small office and retail strip mall buildings across six HVAC system types and all climate zones, for electricity and natural gas in baseline and proposed scenarios. A simulation workflow is developed using OpenStudio TM /EnergyPlus TM to generate simulation data using measures over a wide range of efficiency inputs. Enhancements based on statistical insights are used for improving the model accuracy using filters, input transformations, and change points. Surrogate models achieved average coefficient of variation of the root mean squared error (CVRMSE) values of 2.17, 1.06, 2.05, and 3.26 for proposed electricity, proposed natural gas, baseline electricity, and baseline natural gas, respectively, with enhancements reducing CVRMSE by an average of 14.9% across all combinations. We provide model interpretation via Shapley additive explanations to determine which input variables most influence energy consumption and provide supportive arguments for enhancements.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Study of the Protection Improvements for a Weak Grid Area With High Inverter-Based Resources (IBRs)

This project designs enhanced protection scheme for the real-world weak grid area with a high penetration of IBRs. As the existing protection schemes are originally designed for traditional synchronous machines, we first evaluate if the protection scheme will continue to operate reliably in systems with high levels of IBRs. Hardware relays are tested using a controller-hardware-in-the-loop setup. PSCAD electromagnetic transient simulation with IBR original equipment manufacturer black-box models is used to perform fault studies and generate COMTRADE data, which are replayed by a real-time digital simulator (RTDS) to feed input to the hardware relays. Three scenarios are analyzed: normal operation, an N-1 contingency, and an IBR-only scenario. The evaluation results reveal the following: 1) the protection scheme remains reliable under normal conditions and N-1 contingencies and 2) in IBR-only scenarios, differential protection (87L) continues to operate reliably, whereas local protection elements, such as distance and directional elements, fail because of the lack of regulated negative sequence current contributed by IBRs. Enhanced protection is designed to address the challenge of lack of negative sequence current from IBRs, including increased restraining factors a2 and k2 to block 32Q or using V instead QV ORDER for ground faults, enhanced mho distance element with voltage and phase angle supervision for L-L faults. The efficacy of enhanced protection logic is validated and proven to work reliably. Additionally, IEEE Std. 2800-2022 negative sequence current compliant GFL and GFM IBRs from another vendor are tested and proven to work reliably without need for enhanced logic. Therefore, this work provides valuable decision-making for utilities facing protection system challenges due to IBRs, either designing enhanced protection scheme or requesting their IBRs being IEEE Std. 2800-2022 compliant to produce regulated negative sequence current for protection relay to make correct decision.

24 POWER TRANSMISSION AND DISTRIBUTION

Evaluating Protection System Performance for a Real-World Weak Grid Area With High Inverter-Based Resources

This paper evaluates an existing protection scheme implemented in a real-world weak grid area with a high penetration of inverter-based resources (IBRs). The study aims to assess the reliability and adequacy of protection schemes originally designed for traditional synchronous machine systems and determine whether they can continue to operate reliably in systems with high levels of IBRs. Hardware relays are tested using a controller-hardware-in-the-loop setup. PSCAD electromagnetic transient simulation with an IBR original equipment manufacturer black-box model is used to perform fault studies and generate COMTRADE data, which are replayed by a realtime digital simulator (RTDS) to feed input to the hardware relays. Three scenarios are analyzed: normal operation, an N-1 contingency, and an IBR-only scenario. The evaluation results reveal the following: 1) the protection scheme remains reliable under normal conditions and N-1 contingencies and 2) in IBRonly scenarios, differential protection (87L) continues to operate reliably, whereas local protection elements, such as distance and directional elements, fail because of the lack of regulated negative sequence current contributed by IBRs. These findings provide utilities with valuable insights for improving their protection systems in high-IBRs.

24 POWER TRANSMISSION AND DISTRIBUTION

Evaluating Protection System Performance for a Real-World Weak Grid Area With High Inverter-Based Resources: Preprint

This paper evaluates an existing protection scheme implemented in a real-world weak grid area with a high penetration of inverter-based resources (IBRs). The study aims to assess the reliability and adequacy of protection schemes originally designed for traditional synchronous machine systems and determine whether they can continue to operate reliably in systems with high levels of IBRs. Hardware relays are tested using a controller-hardware-in-the-loop setup. PSCAD electromagnetic transient simulation with an IBR original equipment manufacturer black-box model is used to perform fault studies and generate COMTRADE data, which are replayed by a real-time digital simulator (RTDS) to feed input to the hardware relays. Three scenarios are analyzed: normal operation, an N-1 contingency, and an IBR-only scenario. The evaluation results reveal the following: 1) the protection scheme remains reliable under normal conditions and N-1 contingencies and 2) in IBR-only scenarios, differential protection (87L) continues to operate reliably, whereas local protection elements, such as distance and directional elements, fail because of the lack of regulated negative sequence current contributed by IBRs. These findings provide utilities with valuable insights for improving their protection systems in high-IBRs.

24 POWER TRANSMISSION AND DISTRIBUTION

A Systematic Framework for Tuning Open-Source Multifunctional IBR Models To Emulate OEM Black-Box Fault Dynamics

This paper presents a systematic framework to tune a generic IBR EMT model to match with an OEM provided balckbox inverter model based on the fault current responses. The key learnings and findings are summarized as follows: The tunable key parameters include inner control loops and current limiters to align the fault current magnitude, sequence content, and phase trajectories with the OEM models across diverse fault type and locations. The tuned model's fidelity is validated through comparative analysis with an OEM blackbox model, assessing both the fault current response and the responses of multiple relay elements. The results demonstrate the tuned generic model can trigger relay decision logic that is identical or near identical to that of the OEM model, thus generating very good match model for fault studies.

24 POWER TRANSMISSION AND DISTRIBUTION

End-To-End Decentralized Transmission Line Protection in IBR-Dominated Weak Grids Using Interpretable Data-Driven Methods

Traditional transmission line protection relies on predictable synchronous-based fault signatures, which frequently fail under the non-standard, current-limited fault characteristics of Inverter-Based Resources (IBRs). This study investigates how to achieve secure, communication-free fault isolation in IBR-dominated weak grids without relying on opaque, computationally heavy "black-box" machine learning algorithms. To address this, we propose a novel, standalone, and inherently interpretable data-driven protection framework. Unlike centralized methods requiring multi-terminal communication, this decentralized approach relies solely on local measurements using a hierarchical linear-kernel Support Vector Machine (SVM). The methodology decomposes the protection task into four sequential stages that mimic traditional protection elements: fault detection and fault direction identification, fault type classification, zone classification, and location estimation. This multi-stage architecture allows for specialized feature engineering at each stage, combining high computational efficiency with logic traceability. The framework's end-to-end performance was validated via C-code and PSCAD/EMTDC co-simulation, utilizing a real-world utility network and an OEM black-box IBR model. The proposed relay achieves 97.2% overall accuracy and provides a reliable trip decision within a 2.5-cycle window. The results confirm 100% accuracy in fundamental fault detection, reliable zone selectivity across low to moderate fault resistances, and robust security against non-fault transients, proving its immediate viability for integration into commercial numerical relays.

24 POWER TRANSMISSION AND DISTRIBUTION

Traceable Black-Box Watermarks For Federated Learning

Due to the distributed nature of Federated Learning (FL) systems, each local client has access to the global model, which poses a critical risk of model leakage. Existing works have explored injecting watermarks into local models to enable intellectual property protection. However, these methods either focus on non-traceable watermarks or traceable but white-box watermarks. We identify a gap in the literature regarding the formal definition of traceable black-box watermarking and the formulation of the problem of injecting such watermarks into FL systems. In this work, we first formalize the problem of injecting traceable black-box watermarks into FL. Based on the problem, we propose a novel server-side watermarking method, TraMark, which creates a traceable watermarked model for each client, enabling verification of model leakage in black-box settings. To achieve this, TraMark partitions the model parameter space into two distinct regions: the main task region and the watermarking region. Subsequently, a personalized global model is constructed for each client by aggregating only the main task region while preserving the watermarking region. Each model then learns a unique watermark exclusively within the watermarking region using a distinct watermark dataset before being sent back to the local client. Extensive results across various FL systems demonstrate that TraMark ensures the traceability of all watermarked models while preserving their main task performance.

Xu, Jiahao [University of Nevada, Reno]

Semi-analytic solutions to the Noh problem with a black box EoS

The objective of this paper is to derive a method of constructing semi-analytic solutions to the Noh problem when the equation of state is a black box. Such solutions can be used for verification tests of hydrodynamics codes. We present the underlying theory, the method for finding solutions, and several examples of derived semi-analytic solutions. We end by performing a classic verification convergence test comparing numerical results from a hydrodynamics code against a non-trivial semi-analytic solution.

97 MATHEMATICS AND COMPUTING

Ripening of Rh Nanoparticle Catalysts in Reverse Water–Gas Shift via a Data-Driven Model Combining Physics, Theory, and Experiment

Degradation via sintering is an ongoing challenge that impedes the broad commercial success of supported metallic nanoparticle catalysts. To mitigate degradation via informed catalyst design and process operations, here we aim to disambiguate the underlying mechanisms of sintering by combining theory and experiment in a quantitative framework. While mechanistic sintering models exist, they only model a single sintering pathway, even though multiple sintering mechanisms can occur simultaneously or dominate at different stages of the process. Data-driven machine learning models have emerged as a means to represent complex processes through data regression. However, machine learning models have very large data needs and lack mechanistic insights due to their black-box encoding. To develop an interpretive model of catalyst degradation via sintering, we constructed a hybrid model combining mechanistic “physics-based” models and data-driven methods to obtain both reliable predictions and mechanistic insights regarding experimentally observed sintering phenomena. Focusing on nanoparticle sintering in the Rh–TiO 2 catalyst for the reverse water–gas shift (RWGS) reaction, the hybrid model couples a mechanistic term for Ostwald ripening with energy values calculated via density functional theory (DFT) with a parametric, data-driven discrepancy function term for unmodeled mechanisms. The hybrid model is trained using Bayesian inference with data collected from small-angle X-ray scattering (SAXS) in situ experiments wherein average nanoparticle diameter versus time was measured at three relevant operating temperatures. The calibrated hybrid model results show that an Ostwald ripening-only model parameterized with fixed DFT energies does not fully capture the time and temperature dependence of the SAXS-observed sintering kinetics, and that an additional functional contribution, or DFT energy calibration, is required to reconcile simulation and experiment. Analysis of the hybrid-model error confirms that the hybrid model outperforms both the purely mechanistic and purely data-driven alternatives in terms of expected predictive accuracy for time-evolving average particle sizes. Furthermore, the results support the hypothesis that the Ostwald ripening mechanism is less important for explaining the sintering phenomena as operating temperature increases under an assumed fixed DFT parameterization. This could be explained in one of two ways: either latent, unmodeled sintering mechanisms dominate at higher temperatures, or the DFT uncertainty increases with temperature. The proposed modeling approach directly links theory to experiments and simulations via a statistical hybrid modeling framework and can be extended to other catalytic systems to improve predictive models and mechanistic understanding.

Bayesian hybrid modeling

Systematic Construction of Time-Dependent Hamiltonians for Microwave-Driven Josephson Circuits

Time-dependent electromagnetic drives are fundamental for controlling complex quantum systems, including superconducting Josephson circuits. In these devices, accurate time-dependent Hamiltonian models are imperative for predicting their dynamics and designing high-fidelity quantum operations. Existing numerical methods, such as black-box quantization (BBQ) and energy-participation ratio (EPR), excel at modeling the static Hamiltonians of Josephson circuits. However, these techniques do not fully capture the behavior of driven circuits stimulated by external microwave drives, nor do they include a generalized approach to account for the inevitable noise and dissipation that enter through microwave ports. Here, we introduce numerical techniques that leverage classical microwave simulations, efficiently executable in finite-element solvers, to obtain the time-dependent Hamiltonian of microwave-driven superconducting circuits with arbitrary geometries under charge, flux, or mixed electromagnetic modulation. Importantly, our techniques do not rely on a lumped-element description of the superconducting circuit, in contrast to previous approaches to tackling this problem. We demonstrate the versatility of our approach by characterizing the driven properties of realistic circuit devices in complex electromagnetic environments, including coherent dynamics due to charge and flux modulation, as well as drive-induced relaxation and dephasing. Our techniques offer a powerful toolbox for optimizing circuit designs and advancing practical applications in superconducting quantum computing.

Lu, Yao [Yale U.; Yale U. (main); Fermilab] (ORCID

Chemical Recommender System: Replacement Suggestions for Small Molecules

The Chemical Recommender System (CRS) is an open-source, high-performance toolkit that enables real-time similarity searches across the complete PubChem database (over 50 million molecules) using commodity hardware. The CRS addresses critical limitations in existing chemical informatics platforms through a novel vector database infrastructure, extensible model integration capabilities, and complete algorithmic transparency. The system implements a vector database deployment with partitioned indexing that achieves a ~60x speedup over traditional approaches. A containerized model integration framework allows researchers to seamlessly incorporate custom predictive models into the full-scale search and scoring pipeline, while complete configurability of search parameters, filtering logic, and scoring functions provides capabilities not available in existing black-box solutions. Beyond structural similarity, the CRS integrates OPERA QSAR models for thermophysical and toxicity predictions, RDKit synthetic accessibility scoring, and user-defined models to compute weighted final replacement scores. The complete system is accessible through an interactive web application supporting real-time progress monitoring, post-processing score re-weighting, automated PDF reporting, and batch processing capabilities.

Nair, Parthiv Anand [Sandia National Laboratories

Adaptive Computing and Multi-Fidelity Learning

We describe our ongoing research in adaptive computing. Our goal is to use a combination of low- and high-fidelity simulation models to enable computationally efficient optimization and uncertainty quantification. We develop optimization formulations that take into account the compute resources currently available, which act as a constraint with regards to the fidelity level simulation we can run while maximizing information gain. We will discuss a few application examples that can benefit from this approach, especially when considering challenges arising in scaling up experiments and simulations.

97 MATHEMATICS AND COMPUTING

Degenerate coupled-cluster theory

A size-extensive, converging, black-box, ab initio coupled-cluster (ΔCC) ansatz is introduced that computes the energies and wave functions of states from any degenerate or nondegenerate Slater-determinant references with any numbers of α- and β-spin electrons, any patterns of orbital occupancy, any spin multiplicities, and any spatial symmetries. For a nondegenerate reference, it reduces to the single-reference coupled-cluster ansatz. For a degenerate multireference, it is a natural coupled-cluster extension of degenerate Møller–Plesset perturbation (ΔMP) theory. For ionized and electron-attached references, it is a coupled-cluster Green’s function, although the present theory is convergent toward the full-configuration-interaction limits, while the Feynman–Dyson many-body Green’s function (MBGF) theory generally is not. Its single-excitation instance is a projection Hartree–Fock theory as per the Thouless theorem, which may be useful for core ionizations, high-spin states, and possibly electron affinities. Additionally, a new multireference coupled-cluster theory for a general model space is developed. This quasidegenerate coupled-cluster (QCC) theory is exactly converging, but not black-box, and intended for strong correlation. Determinant-based, general-order algorithms of ΔCC and QCC theories are implemented and compared with configuration-interaction (CI) and equation-of-motion coupled-cluster (EOM-CC) theories through octuple excitations and with ΔMP and MBGF theories up to the nineteenth order. An algebraic, optimal-scaling algorithm of the ΔCC theory is computer-synthesized at the levels of single excitations (ΔCCS) and of single and double excitations (ΔCCSD). As a result, the order of performance is QCC ≈ ΔCC > EOM-CC > CI at the same order or QCC ≈ ΔCC > ΔMP > MBGF at the same cost scaling.

Hirata, So [University of Illinois at Urbana-Champ

Gradient-informed Hamiltonian Monte Carlo for multicomponent CALPHAD model optimization and uncertainty quantification

CALPHAD model parameter optimization is inherently challenging due to non-smooth objective functions, high-dimensional parameter spaces, and the need for uncertainty quantification (UQ). Traditional weighted nonlinear least squares approaches are computationally efficient but local, whereas black-box global optimizers and ensemble Markov Chain Monte Carlo (MCMC) methods provide broader exploration at substantial computational cost. The objective of this work is to combine the global exploration capability of gradient-informed Hamiltonian Monte Carlo – specifically the No-U-Turn Sampler (NUTS) – with local deterministic refinement using BFGS to efficiently optimize multicomponent CALPHAD models with minimal manual intervention. Analytic gradients are computed via the Jansson derivative framework. The methodology is demonstrated on the Cr—Fe binary system and extended to the Cr—Fe—Ni ternary system with 32 degrees of freedom. For Cr—Fe, NUTS achieves comparable or superior optimality relative to ensemble MCMC while requiring over an order-of-magnitude fewer likelihood evaluations. Parameter uncertainties are quantified through NUTS sampling and propagated to thermodynamic observables using local expansion, demonstrating a novel modular approach that combines binary and ternary parameter subsets without requiring global relaxation. These results establish gradient-informed exploration as a scalable strategy for multicomponent CALPHAD optimization and provide a practical route towards efficient higher-order database development with quantified uncertainty.

36 MATERIALS SCIENCE

Predicting U 3 O 8 powder processing conditions: An AI/ML approach analyzing deep learning embeddings of SEM micrographs

High-resolution SEM images of uranium-oxide powders encode micro- and nanoscale clues to their synthesis route and calcination temperature. We trained a ResNet-50 model on 11 commercial-scale U₃O₈ classes, ammonium diuranate (ADU) or uranyl peroxide (H₂O₂) precursors calcined at temperatures ranging from 400 to 750 °C and added a 256-D projection head before the classifier to analyze the learned representation. The best of eight seeds reached 92.4 % accuracy on reserved testing data, but our focus is the structure of the embedding space rather than the accuracy and labels. We quantify class relatedness in the original 256-D space using centroid similarity and distributional distances, and we use Uniform Manifold Approximation Projection (UMAP) for visualization. ‘Unknown’ images from different preparation methods, SEM operators, and from the literature localized near the expected classes under a nearest-centroid analysis without retraining, as well as clustered in similar UMAP space. In conclusion, this embedding-centered workflow complements black-box classification by providing quantitative, similarity-based comparisons of U₃O₈ morphologies and reduces storage space by up to 98 % for image data used in millisecond vector search comparisons.

36 MATERIALS SCIENCE