Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “network embeddings”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Symbolic diagnostics to interpret and analyze neural network models

Embedded machine-learned models (EMLMs) have the promise to improve the predictive accuracy of engineering simulators in environments of national interest. EMLMs often comprise complex input-output maps (e.g., neural networks), which make them unamenable to rigorous analysis and generally difficult to interpret. In the face of decades of theory, this lack of interpretability is a significant barrier to building confidence in these models. This work outlines an approach to interpret EMLMs using sparse polynomial regression for comparison with theoretical understanding. To do so, we build on the concept of Locally Interpretable Model-agnostic Explanations (LIME) using physics-informed clustering, prototype selection, and library construction. While general, we demonstrate our method on tensor-basis neural networks used in Reynolds-Averaged Navier-Stokes simulations of hypersonic fluid flows. Results are presented for a simulated toy model and for direct numerical simulations (DNS) of turbulent flows over a flat plate.

97 MATHEMATICS AND COMPUTING↗

Transfer learning nonlinear plasma dynamic transitions in low dimensional embeddings via deep neural networks

Deep learning algorithms provide a new paradigm to study high-dimensional dynamical behaviors, such as those in fusion plasma systems. Development of novel, data-driven model reduction methods, coupled with detection of abnormal modes with plasma physics, opens a unique opportunity to identify plasma instabilities through automated construction of parsimonious models that can be tuned to balance accuracy and cost. Our fusion transfer learning (FTL) model demonstrates success in rapidly reconstructing nonlinear kink mode structures by learning from a limited amount of nonlinear simulation data. The knowledge transfer process leverages a pre-trained neural encoder–decoder network, initially trained on linear simulations, to effectively capture nonlinear dynamics. The low-dimensional embeddings extract the coherent structures of interest, while preserving the inherent dynamics of the complex system. Experimental results highlight FTL’s capacity to capture transitional behaviors and dynamical features in plasma dynamics—a task often challenging for conventional methods. The model developed in this study is generalizable and can be extended broadly through transfer learning to address various magnetohydrodynamics modes.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Codon2Vec v1.0

Background: Codon2Vec is an embedding neural network that predicts 'high' or 'low' gene expression directly from the protein-coding sequences. Embedding neural networks are commonly used for natural language processing (NLP) applications. Analogous to how an English sentence is a string of words, a gene can be thought of as a string of codons. Similar to how NLP neural networks model English sentences as a non-random sequence of words, we considered a coding sequence as a non-random non-overlapping array of codons (k-mers of length = 3). Value Proposition: - Codon2Vec achieved a high median AUC-ROC score of 83.8% when trained and applied to transcriptomic data from 300 fungal species - Unlike Codo2Vec, conventional methods predicting for expression based on codon usage rely on a priori knowledge of optimal codons or a set of reference genes. - Unlike Codon2vec, these methods do not account for the effect of codon order on gene expression. - Codon2Vec neural network bypasses the need for artisanal feature selection step that is necessary for traditional machine learning models.

Wint, Rhondene↗

A Mixed integer linear programming‐based distributed energy management for networked microgrids considering network operational objectives and constraints

Abstract Mixed integer linear programming (MILP)–based distributed energy management for networked microgrids embedded modern distribution systems is proposed. Considering the diverse ownership of microgrids, distributed energy resources (DERs) that interface directly with utilities and responsive loads, an alternating direction method of multipliers–based distributed framework was formulated for the scheduling of networked microgrids embedded modern distribution systems by adjusting nodal price signals iteratively. In addition, to make the formulated optimization problems resolvable through more accessible and popular MILP solvers, different linearisation techniques were employed to transform the nonlinear terms into linear or mixed integer linear formats. The proposed MILP‐based distributed method preserves all participants' autonomy (e.g., microgrids, DERs that interface directly with utilities and responsive loads), while incentivising them to actively participate in the distribution system operation with price signals. The proposed method is validated with results of numerical simulation using a modern distribution system consisting of multiple networked microgrids, DERs that interface directly with utilities, as well as responsive loads.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Modeling and simulation of vascular tumors embedded in evolving capillary networks

Here in this work, we present a coupled 3D–1D model of solid tumor growth within a dynamically changing vascular network to facilitate realistic simulations of angiogenesis. Additionally, the model includes erosion of the extracellular matrix, interstitial flow, and coupled flow in blood vessels and tissue. We employ continuum mixture theory with stochastic Cahn–Hilliard type phase-field models of tumor growth. The interstitial flow is governed by a mesoscale version of Darcy’s law. The flow in the blood vessels is controlled by Poiseuille flow, and Starling’s law is applied to model the mass transfer in and out of blood vessels. The evolution of the network of blood vessels is orchestrated by the concentration of the tumor angiogenesis factors (TAFs); blood vessels grow towards the increasing TAFs concentrations. This process is not deterministic, allowing random growth of blood vessels and, therefore, due to the coupling of nutrients in tissue and vessels, makes the growth of tumors stochastic. We demonstrate the performance of the model by applying it to a variety of scenarios. Numerical experiments illustrate the flexibility of the model and its ability to generate satellite tumors. Simulations of the effects of angiogenesis on tumor growth are presented as well as sample-independent features of cancer.

3D–1D coupled blood flow models↗

Embedding hard physical constraints in neural network coarse-graining of three-dimensional turbulence

In recent years, deep learning approaches have shown much promise in modeling complex systems in the physical sciences. A major challenge in deep learning of partial differential equations is enforcing physical constraints and boundary conditions. In this work, we propose a general framework to directly embed the notion of an incompressible fluid into convolutional neural networks, and apply this to coarse-graining of turbulent flow. These physics-embedded neural networks leverage interpretable strategies from numerical methods and computational fluid dynamics to enforce physical laws and boundary conditions by taking advantage the mathematical properties of the underlying equations. Here, we demonstrate results on three-dimensional fully developed turbulence, showing that this technique drastically improves local conservation of mass, without sacrificing performance according to several other metrics characterizing the fluid flow.

97 MATHEMATICS AND COMPUTING↗

Embedded training of neural-network subgrid-scale turbulence models

We report that the weights of a deep neural-network model are optimized in conjunction with the governing flow equations to provide a model for subgrid-scale stresses in a temporally developing plane turbulent jet at Reynolds number Re 0 = 6000 . The objective function for training is first based on the instantaneous filtered velocity fields from a corresponding direct numerical simulation, and the training is by a stochastic gradient descent method, which uses the adjoint Navier-Stokes equations to provide the end-to-end sensitivities of the model weights to the velocity fields. In-sample and out-of-sample testing on multiple dual-jet configurations show that its required mesh density in each coordinate direction for prediction of mean flow, Reynolds stresses, and spectra is half that needed by the dynamic Smagorinsky model for comparable accuracy. The same neural-network model trained directly to match filtered subgrid-scale stresses, without the constraint of being embedded within the flow equations during the training, fails to provide a qualitatively correct prediction. The coupled formulation is generalized to train based only on mean-flow and Reynolds stresses, which are more readily available in experiments. The mean-flow training provides a robust model, which is important, though a somewhat less accurate prediction for the same coarse meshes, as might be anticipated due to the reduced information available for training in this case. The anticipated advantage of the formulation is that the inclusion of resolved physics in the training increases its capacity to extrapolate. This is assessed for the case of passive scalar transport, for which it outperforms established models due to improved mixing predictions.

42 ENGINEERING↗

A Graph Neural Network Surrogate Model for hls4ml

Recent advancements in use of machine learning (ML) techniques on field-programmable gate arrays (FPGAs) have allowed for the implementation of embedded neural networks with extremely low latency. This is invaluable for particle detectors at the Large Hadron Collider, where latency and used area are strictly bounded. The hls4ml framework is a procedure that converts trained ML model software to a synthesis result to can be used on an FPGA. However, running the pipeline is a time-consuming procedure, and there is a strong risk of failure. In particular, it may not be possible to successfully convert a model into a synthesis result, or the resource consumption of the model may exceed the resources of the target FPGA. To aid with this development, we introduce wa-hls4ml, a surrogate model using a graph neural network to emulate the structure of the source models. The goal is to estimate the chance of success and resource consumption of a given model when passed through the hls4ml pipeline, without needing to run the pipeline.

Plotnikov, Dennis↗

Toward an embedded training tool for Deep Space Network operations

There are three issues to consider when building an embedded training system for a task domain involving the operation of complex equipment: (1) how skill is acquired in the task domain; (2) how the training system should be designed to assist in the acquisition of the skill, and more specifically, how an intelligent tutor could aid in learning; and (3) whether it is feasible to incorporate the resulting training system into the operational environment. This paper describes how these issues have been addressed in a prototype training system that was developed for operations in NASA's Deep Space Network (DSN). The first two issues were addressed by building an executable cognitive model of problem solving and skill acquisition of the task domain and then using the model to design an intelligent tutor. The cognitive model was developed in Soar for the DSN's Link Monitor and Control (LMC) system; it led to several insights about learning in the task domain that were used to design an intelligent tutor called REACT that implements a method called 'impasse-driven tutoring'. REACT is one component of the LMC training system, which also includes a communications link simulator and a graphical user interface. A pilot study of the LMC training system indicates that REACT shows promise as an effective way for helping operators to quickly acquire expert skills.

Hill, Randall W., Jr.↗

Toward an Embedded Training Tool for Deep Space Network Operations

There are three issues to consider when building an embedded training system for a task domain involving the operation of complex equipment: (1) how skill is acquired in the task domain; (2) how the training system should be designed to assist in the acquisition of the skill, and more specifically, how an intelligent tutor could aid in learning; and (3) whether it is feasible to incorporate the resulting training system into the operational environment. This paper describes how these issues have been addressed in a prototype training system that was developed for operations in NASA's Deep Space Network (DSN). The first two issues were addressed by building an executable cognitive model of problem solving and skill acquisition of the task domain and then using the model to design an intelligent tutor.

Johnson, W. Lewis↗

Performance Characteristics of the BlueField-2 SmartNIC

High-performance computing (HPC) researchers have long envisioned scenarios where application workflows could be improved through the use of programmable processing elements embedded in the network fabric. Recently, vendors have introduced programmable Smart Network Interface Cards (SmartNICs) that enable computations to be offloaded to the edge of the network. There is great interest in both the HPC and high-performance data analytics (HPDA) communities in understanding the roles these devices may play in the data paths of upcoming systems. This paper focuses on characterizing both the networking and computing aspects of NVIDIA’s new BlueField-2 SmartNIC when used in a 100Gb/s Ethernet environment. For the networking evaluation we conducted multiple transfer experiments between processors located at the host, the SmartNIC, and a remote host. These tests illuminate how much effort is required to saturate the network and help estimate the processing headroom available on the SmartNIC during transfers. For the computing evaluation we used the stress-ng benchmark to compare the BlueField-2 to other servers and place realistic bounds on the types of offload operations that are appropriate for the hardware. Our findings from this work indicate that while the BlueField-2 provides a flexible means of processing data at the network’s edge, great care must be taken to not overwhelm the hardware. While the host can easily saturate the network link, the SmartNIC’s embedded processors may not have enough computing resources to sustain more than half the expected bandwidth when using kernel-space packet processing. From a computational perspective, encryption operations, memory operations under contention, and on-card IPC operations on the SmartNIC perform significantly better than the general-purpose servers used for comparisons in our experiments. Therefore, applications that mainly focus on these operations may be good candidates for offloading to the SmartNIC.

97 MATHEMATICS AND COMPUTING↗

Enhancing Fault Isolation for Health Monitoring of Electric Aircraft Propulsion by Embedding Failure Mode and Effect Analysis into Bayesian Networks

This paper describes a fault isolation approach for electric powertrains of unmanned aerial vehicles. The approach leverages the combination of failure mode and effect analysis (FMEA) and Bayesian networks, thus introducing depend-ability structures into a diagnostic framework. Faults and failure events from the FMEA are mapped within a Bayesian network, where network edges replicate the links embedded within FMEAs. This framework helps the fault isolation process by identifying the probability of occurrence of specific faults or root causes given evidence observed through sensor signals. The framework is applied to an electric power-train system of a small, rotary-wing unmanned aerial vehicle, demonstrating how a Bayesian network enhanced by FMEA helps disambiguate between root causes of incipient failures, which would otherwise be considered as equally probable.

Fault Isolation↗

Even Higher-Level Synthesis: An Exploration of AI Hardware Accelerators using HLS4ML

With the rise of artificial intelligence, the popularization of deep learning, and a constantly evolving industry, the demand for flexible and efficient tools has never been greater. As algorithms grow more complex, their runtime and energy consumption increase exponentially. Customized hardware accelerators, long used for specific mathematical operations, remain essential for managing modern applications' computational and power demands. Hardware accelerators can speed up complex computations by orders of magnitude, but their manual design and verification processes are often challenging and time-consuming. High-Level Synthesis (HLS) provides a solution by transforming high-level algorithm descriptions, typically written in C++ or SystemC, into synthesizable RTL suitable for hardware implementation. This approach reduces development time for RTL engineers while offering flexibility beyond what traditional handwritten RTL can provide. We extended this capability to the machine-learning domain with the open-source framework hls4ml, which allows neural networks trained in Python frameworks like Tensorflow or PyTorch to be synthesized into efficient hardware representations for the traditional FPGA and ASIC flows. This breakthrough addresses the growing need for reduced design turnaround and easy verification of ML hardware accelerators with low latency and power efficiency constraints. During this tutorial, we will demonstrate how Python complements HLS by simplifying the ML design process, bridging the gap between software and hardware development. Attendees will explore how we translate neural networks modeled in Python into fixed-point C++ models suitable for HLS workflows. We will dive into strategies like Value-Range Analysis and Quantization-Aware Training, which optimize these designs for deployment and evaluate their accuracy, power consumption, and energy efficiency. To exemplify these concepts, experts from Fermilab will share their experiences applying this technology to high-energy physics experiments, where real-time, low-latency processing is critical. Over the years, Fermilab engineers have demonstrated how deep neural networks, optimized for hardware using hls4ml, can meet the stringent requirements of trigger systems at the CERN Large Hadron Collider. These systems rely on rapid decision-making to process immense data volumes while retaining only the most relevant events for further analysis. The application of hls4ml has also been extended to innovative technologies like smart pixel arrays. These smart pixels integrate ML inference capabilities directly into sensor devices, enabling localized data processing at the pixel level. This approach drastically reduces the need to transmit raw data to external processing units, significantly decreasing power consumption and latency. By embedding neural networks within the pixel architecture, the smart pixels can identify and prioritize relevant data in real time, providing a highly efficient solution for edge computing in scenarios such as particle detectors and imaging systems. Fermilab's work highlights the potential of hardware-accelerated ML in scenarios where both speed and power efficiency are mission-critical. Through this tutorial, attendees will gain valuable insights into the challenges and solutions of deploying ML in hardware. Understanding how HLS and hls4ml streamline the development of neural network-based hardware accelerators is fundamental for the industry's future. Participants will learn how these technologies are shaping the future of AI and scientific computing.

Di Guglielmo, Giuseppe [Fermilab]↗

Probabilistic Physics-Informed Graph Convolutional Network for Active Distribution System Voltage Prediction

Here this letter proposes a novel data-driven probabilistic physics-informed graph convolutional network (GCN) for active distribution system voltage prediction with PVs and EVs. It leverages both measurements and network topology to accurately and efficiently predict node voltages without the need for an accurate distribution system power flow model. The dropout-enabled Bayesian inference is developed to achieve uncertainty quantification of the voltage prediction. Thanks to the network model embedding, it also has robustness against topology changes, a key difference with existing machine learning-based approaches. Comparison results with other state-of-the-art machine learning methods on a realistic 759-node distribution system demonstrate that the proposed method can achieve better accuracy and robustness under different scenarios.

24 POWER TRANSMISSION AND DISTRIBUTION↗