Performance Evaluation of Network Flow and Device Classification using Network Features and Device Embeddings
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
In recent years, deep learning approaches have shown much promise in modeling complex systems in the physical sciences. A major challenge in deep learning of partial differential equations is enforcing physical constraints and boundary conditions. In this work, we propose a general framework to directly embed the notion of an incompressible fluid into convolutional neural networks, and apply this to coarse-graining of turbulent flow. These physics-embedded neural networks leverage interpretable strategies from numerical methods and computational fluid dynamics to enforce physical laws and boundary conditions by taking advantage the mathematical properties of the underlying equations. Here, we demonstrate results on three-dimensional fully developed turbulence, showing that this technique drastically improves local conservation of mass, without sacrificing performance according to several other metrics characterizing the fluid flow.
We report that the weights of a deep neural-network model are optimized in conjunction with the governing flow equations to provide a model for subgrid-scale stresses in a temporally developing plane turbulent jet at Reynolds number Re 0 = 6000 . The objective function for training is first based on the instantaneous filtered velocity fields from a corresponding direct numerical simulation, and the training is by a stochastic gradient descent method, which uses the adjoint Navier-Stokes equations to provide the end-to-end sensitivities of the model weights to the velocity fields. In-sample and out-of-sample testing on multiple dual-jet configurations show that its required mesh density in each coordinate direction for prediction of mean flow, Reynolds stresses, and spectra is half that needed by the dynamic Smagorinsky model for comparable accuracy. The same neural-network model trained directly to match filtered subgrid-scale stresses, without the constraint of being embedded within the flow equations during the training, fails to provide a qualitatively correct prediction. The coupled formulation is generalized to train based only on mean-flow and Reynolds stresses, which are more readily available in experiments. The mean-flow training provides a robust model, which is important, though a somewhat less accurate prediction for the same coarse meshes, as might be anticipated due to the reduced information available for training in this case. The anticipated advantage of the formulation is that the inclusion of resolved physics in the training increases its capacity to extrapolate. This is assessed for the case of passive scalar transport, for which it outperforms established models due to improved mixing predictions.
Recent advancements in use of machine learning (ML) techniques on field-programmable gate arrays (FPGAs) have allowed for the implementation of embedded neural networks with extremely low latency. This is invaluable for particle detectors at the Large Hadron Collider, where latency and used area are strictly bounded. The hls4ml framework is a procedure that converts trained ML model software to a synthesis result to can be used on an FPGA. However, running the pipeline is a time-consuming procedure, and there is a strong risk of failure. In particular, it may not be possible to successfully convert a model into a synthesis result, or the resource consumption of the model may exceed the resources of the target FPGA. To aid with this development, we introduce wa-hls4ml, a surrogate model using a graph neural network to emulate the structure of the source models. The goal is to estimate the chance of success and resource consumption of a given model when passed through the hls4ml pipeline, without needing to run the pipeline.
High-performance computing (HPC) researchers have long envisioned scenarios where application workflows could be improved through the use of programmable processing elements embedded in the network fabric. Recently, vendors have introduced programmable Smart Network Interface Cards (SmartNICs) that enable computations to be offloaded to the edge of the network. There is great interest in both the HPC and high-performance data analytics (HPDA) communities in understanding the roles these devices may play in the data paths of upcoming systems. This paper focuses on characterizing both the networking and computing aspects of NVIDIA’s new BlueField-2 SmartNIC when used in a 100Gb/s Ethernet environment. For the networking evaluation we conducted multiple transfer experiments between processors located at the host, the SmartNIC, and a remote host. These tests illuminate how much effort is required to saturate the network and help estimate the processing headroom available on the SmartNIC during transfers. For the computing evaluation we used the stress-ng benchmark to compare the BlueField-2 to other servers and place realistic bounds on the types of offload operations that are appropriate for the hardware. Our findings from this work indicate that while the BlueField-2 provides a flexible means of processing data at the network’s edge, great care must be taken to not overwhelm the hardware. While the host can easily saturate the network link, the SmartNIC’s embedded processors may not have enough computing resources to sustain more than half the expected bandwidth when using kernel-space packet processing. From a computational perspective, encryption operations, memory operations under contention, and on-card IPC operations on the SmartNIC perform significantly better than the general-purpose servers used for comparisons in our experiments. Therefore, applications that mainly focus on these operations may be good candidates for offloading to the SmartNIC.
With the rise of artificial intelligence, the popularization of deep learning, and a constantly evolving industry, the demand for flexible and efficient tools has never been greater. As algorithms grow more complex, their runtime and energy consumption increase exponentially. Customized hardware accelerators, long used for specific mathematical operations, remain essential for managing modern applications' computational and power demands. Hardware accelerators can speed up complex computations by orders of magnitude, but their manual design and verification processes are often challenging and time-consuming. High-Level Synthesis (HLS) provides a solution by transforming high-level algorithm descriptions, typically written in C++ or SystemC, into synthesizable RTL suitable for hardware implementation. This approach reduces development time for RTL engineers while offering flexibility beyond what traditional handwritten RTL can provide. We extended this capability to the machine-learning domain with the open-source framework hls4ml, which allows neural networks trained in Python frameworks like Tensorflow or PyTorch to be synthesized into efficient hardware representations for the traditional FPGA and ASIC flows. This breakthrough addresses the growing need for reduced design turnaround and easy verification of ML hardware accelerators with low latency and power efficiency constraints. During this tutorial, we will demonstrate how Python complements HLS by simplifying the ML design process, bridging the gap between software and hardware development. Attendees will explore how we translate neural networks modeled in Python into fixed-point C++ models suitable for HLS workflows. We will dive into strategies like Value-Range Analysis and Quantization-Aware Training, which optimize these designs for deployment and evaluate their accuracy, power consumption, and energy efficiency. To exemplify these concepts, experts from Fermilab will share their experiences applying this technology to high-energy physics experiments, where real-time, low-latency processing is critical. Over the years, Fermilab engineers have demonstrated how deep neural networks, optimized for hardware using hls4ml, can meet the stringent requirements of trigger systems at the CERN Large Hadron Collider. These systems rely on rapid decision-making to process immense data volumes while retaining only the most relevant events for further analysis. The application of hls4ml has also been extended to innovative technologies like smart pixel arrays. These smart pixels integrate ML inference capabilities directly into sensor devices, enabling localized data processing at the pixel level. This approach drastically reduces the need to transmit raw data to external processing units, significantly decreasing power consumption and latency. By embedding neural networks within the pixel architecture, the smart pixels can identify and prioritize relevant data in real time, providing a highly efficient solution for edge computing in scenarios such as particle detectors and imaging systems. Fermilab's work highlights the potential of hardware-accelerated ML in scenarios where both speed and power efficiency are mission-critical. Through this tutorial, attendees will gain valuable insights into the challenges and solutions of deploying ML in hardware. Understanding how HLS and hls4ml streamline the development of neural network-based hardware accelerators is fundamental for the industry's future. Participants will learn how these technologies are shaping the future of AI and scientific computing.
Here this letter proposes a novel data-driven probabilistic physics-informed graph convolutional network (GCN) for active distribution system voltage prediction with PVs and EVs. It leverages both measurements and network topology to accurately and efficiently predict node voltages without the need for an accurate distribution system power flow model. The dropout-enabled Bayesian inference is developed to achieve uncertainty quantification of the voltage prediction. Thanks to the network model embedding, it also has robustness against topology changes, a key difference with existing machine learning-based approaches. Comparison results with other state-of-the-art machine learning methods on a realistic 759-node distribution system demonstrate that the proposed method can achieve better accuracy and robustness under different scenarios.
Fault detection and diagnosis is critical to power plant operation to ensure attaining high reliability while reducing operation cost. As more renewable power is introduced to the power grid, traditional fossil power plants take on the extra burden of excessive load cycling to compensate the generation variability from renewable power. Such load cycling will pose more reliability challenges to power plant operation. There are a number of challenges faced by today’s asset health management system in coal- fired (or gas) power plants: 1) high-dimensional nonlinear interaction among multiple time series measurements; 2) high measurement variance induced by operational conditions/modes; 3) variation among asset types and plant configurations; and 4) a small number of faulty events to learn from. To cope with these challenges, today’s fielded asset health management systems rely heavily on manual efforts from domain experts and hand-crafted features or rules based on domain knowledge. Despite its role in plant reliability, such a practice is costly and hinders its scalability and sustainability, particularly when a plant undergoes modifications. The objective of this project is to develop a novel end-to-end AI learning system that is trainable (i.e., the AI representation of a complex system behavior can be directly learned from properly labeled data) for accurate fault detection and root cause analysis. The ability to create a fault detection model directly from time series could alleviate the efforts associated with today’s asset management solution development. In the course of this project, we have achieved the following: Created an AI model development environment incorporating state-of-the-art neural network architectures for rapid model development and evaluation; Developed novel learning strategies for training of fault detection model; Developed special-purpose neural network architecture embedded with variable association graph aiming for better interpretability; Developed a learning strategy to leverage a small number of faulty events for enhanced fault detection capability; Conducted detailed experimental study based on public benchmark datasets and demonstrated the effectiveness of the proposed solution; and Validated the developed system with data from both a coal-fired plant boiler dynamic simulation model and real-world coal-fired power plant covering multiple asset and fault types. Overall, the project attained a technology readiness level of TRL 5 from TRL 2 at the beginning of the project.
Abstract Although drug combinations in cancer treatment appear to be a promising therapeutic strategy with respect to monotherapy, it is arduous to discover new synergistic drug combinations due to the combinatorial explosion. Deep learning technology holds immense promise for better prediction of in vitro synergistic drug combinations for certain cell lines. In methods applying such technology, omics data are widely adopted to construct cell line features. However, biological network data are rarely considered yet, which is worthy of in-depth study. In this study, we propose a novel deep learning method, termed PRODeepSyn, for predicting anticancer synergistic drug combinations. By leveraging the Graph Convolutional Network, PRODeepSyn integrates the protein–protein interaction (PPI) network with omics data to construct low-dimensional dense embeddings for cell lines. PRODeepSyn then builds a deep neural network with the Batch Normalization mechanism to predict synergy scores using the cell line embeddings and drug features. PRODeepSyn achieves the lowest root mean square error of 15.08 and the highest Pearson correlation coefficient of 0.75, outperforming two deep learning methods and four machine learning methods. On the classification task, PRODeepSyn achieves an area under the receiver operator characteristics curve of 0.90, an area under the precision–recall curve of 0.63 and a Cohen’s Kappa of 0.53. In the ablation study, we find that using the multi-omics data and the integrated PPI network’s information both can improve the prediction results. Additionally, the case study demonstrates the consistency between PRODeepSyn and previous studies.
Predictions of hydrologic variables across the entire water cycle have significant value for water resources management as well as downstream applications such as ecosystem and water quality modeling. Recently, purely data-driven deep learning models like long short-term memory (LSTM) showed seemingly insurmountable performance in modeling rainfall runoff and other geoscientific variables, yet they cannot predict untrained physical variables and remain challenging to interpret. Here, we show that differentiable, learnable, process-based models (called δ models here) can approach the performance level of LSTM for the intensively observed variable (streamflow) with regionalized parameterization. We use a simple hydrologic model HBV as the backbone and use embedded neural networks, which can only be trained in a differentiable programming framework, to parameterize, enhance, or replace the process-based model's modules. Without using an ensemble or post-processor, δ models can obtain a median Nash-Sutcliffe efficiency of 0.732 for 671 basins across the USA for the Daymet forcing data set, compared to 0.748 from a state-of-the-art LSTM model with the same setup. For another forcing data set, the difference is even smaller: 0.715 versus 0.722. Meanwhile, the resulting learnable process-based models can output a full set of untrained variables, for example, soil and groundwater storage, snowpack, evapotranspiration, and baseflow, and can later be constrained by their observations. Both simulated evapotranspiration and fraction of discharge from baseflow agreed decently with alternative estimates. The general framework can work with models with various process complexity and opens up the path for learning physics from big data.
Abstract Climate models are essential to understand and project climate change, yet long‐standing biases and uncertainties in their projections remain. This is largely associated with the representation of subgrid‐scale processes, particularly clouds and convection. Deep learning can learn these subgrid‐scale processes from computationally expensive storm‐resolving models while retaining many features at a fraction of computational cost. Yet, climate simulations with embedded neural network parameterizations are still challenging and highly depend on the deep learning solution. This is likely associated with spurious non‐physical correlations learned by the neural networks due to the complexity of the physical dynamical system. Here, we show that the combination of causality with deep learning helps removing spurious correlations and optimizing the neural network algorithm. To resolve this, we apply a causal discovery method to unveil causal drivers in the set of input predictors of atmospheric subgrid‐scale processes of a superparameterized climate model in which deep convection is explicitly resolved. The resulting causally‐informed neural networks are coupled to the climate model, hence, replacing the superparameterization and radiation scheme. We show that the climate simulations with causally‐informed neural network parameterizations retain many convection‐related properties and accurately generate the climate of the original high‐resolution climate model, while retaining similar generalization capabilities to unseen climates compared to the non‐causal approach. The combination of causal discovery and deep learning is a new and promising approach that leads to stable and more trustworthy climate simulations and paves the way toward more physically‐based causal deep learning approaches also in other scientific disciplines.
Phycobilisome (PBS) structures are elaborate antennae in cyanobacteria and red algae 1,2 . These large protein complexes capture incident sunlight and transfer the energy through a network of embedded pigment molecules called bilins to the photosynthetic reaction centres. However, light harvesting must also be balanced against the risks of photodamage. A known mode of photoprotection is mediated by orange carotenoid protein (OCP), which binds to PBS when light intensities are high to mediate photoprotective, non-photochemical quenching 3-6 . Here we use cryogenic electron microscopy to solve four structures of the 6.2 MDa PBS, with and without OCP bound, from the model cyanobacterium Synechocystis sp. PCC 6803. The structures contain a previously undescribed linker protein that binds to the membrane-facing side of PBS. For the unquenched PBS, the structures also reveal three different conformational states of the antenna, two previously unknown. The conformational states result from positional switching of two of the rods and may constitute a new mode of regulation of light harvesting. Further, only one of the three PBS conformations can bind to OCP, which suggests that not every PBS is equally susceptible to non-photochemical quenching. In the OCP-PBS complex, quenching is achieved through the binding of four 34kDa OCPs organized as two dimers. The complex reveals the structure of the active form of OCP, in which an approximately 60Å displacement of its regulatory carboxy terminal domain occurs. Finally, by combining our structure with spectroscopic properties 7 , we elucidate energy transfer pathways within PBS in both the quenched and light-harvesting states. Collectively, our results provide detailed insights into the biophysical underpinnings of the control of cyanobacterial light harvesting. The data also have implications for bioengineering PBS regulation in natural and artificial light-harvesting systems.
This article reports the synergy between ceramic nanofibers and a polymer, and the enhanced interfacial Li-ion transport along the nanofiber/polymer interface in a solid-state ceramic/polymer composite electrolyte, in which a three-dimensional (3D) electrospun aluminum-doped Li 0.33 La 0.557 TiO 3 (LLTO) nanofiber network is embedded in a polyvinylidene fluoride-hexafluoropropylene (PVDF-HFP) matrix. Strong chemical interaction occurs between the nanofibers and the polymer matrix. Addition of the ceramic nanofibers into the polymer matrix results in the dehydrofluorination of the PVDF chains, deprotonation of the –CH 2 moiety and amorphization of the polymer matrix. Solid-state nuclear magnetic resonance (NMR) spectra reveal that lithium ions transport via three pathways: (i) intra-polymer transport, (ii) intra-nanofiber transport, and (iii) interfacial polymer/nanofiber transport. In addition, lithium phosphate is coated on the LLTO nanofiber surface before the nanofibers are embedded into the polymer matrix. The presence of lithium phosphate at the LLTO/polymer interface further enhances the chemical interaction between the nanofibers and the polymer, which promotes the lithium ion transport along the polymer/nanofiber interface. This in turn improves the ionic conductivity and electrochemical cycling stability of the nanofiber/polymer composite. As a result, the flexible LLTO/Li 3 PO 4 /polymer composite electrolyte membrane exhibits an ionic conductivity of 5.1 × 10 -4 S cm -1 at room temperature and an electrochemical stability window of 5.0 V vs. Li/Li + . A symmetric Li|electrolyte|Li half-cell shows a low overpotential of 50 mV at a constant current density of 0.5 mA cm -2 for more than 800 h. In addition, a full cell is constructed by sandwiching the composite electrolyte between a lithium metal anode and a LiFePO 4 -based cathode. Such an all-solid-state lithium metal battery exhibits excellent cycling performance and rate capability.
Integration and operation of distributed generation (DG) and energy storage (ES) in power distribution systems are enabled by communication networks and embedded sensor and control devices that increase the vulnerability of the systems to cyber-threats, broadening the attack surface and making adversary actions more unpredictable. This paper proposes a methodology that uses ellipsoidal approximations to quantify the potential damage caused by successful attacks that affect, directly or indirectly, the desired operation setpoints and may drive the power distribution operation to unsafe states by violating the limits of voltage or line flows. More specifically, a new methodology is introduced to find the optimal non-symmetric operating constraints that can be imposed to each DG and ES in order to guarantee that the power distribution system is resilient to any malicious setpoints. The proposed method takes as inputs the system topology, DG and ES capabilities, and load limits to solve a convex optimization problem formulated using linear matrix inequalities (LMIs) and the power flow equations. The proposed solution is agnostic to the attacker's action or load profile and it does not require any assumption about the location or means of the attack. The numerical results on a test distribution feeder with several DG and ES illustrate how the proposed resilient operating constraints guarantee the security of the power distribution system under setpoint attacks.
Modeling real-world phenomena to any degree of accuracy is a challenge that the scientific research community has navigated since its foundation. Lack of information and limited computational and observational resources necessitate modeling assumptions which, when invalid, lead to model-form error (MFE). The work reported herein explored a novel method to represent model-form uncertainty (MFU) that combines Bayesian statistics with the emerging field of universal differential equations (UDEs). The fundamental principle behind UDEs is simple: use known equational forms that govern a dynamical system when you have them; then incorporate data-driven approaches – in this case neural networks (NNs) – embedded within the governing equations to learn the interacting terms that were underrepresented. Utilizing epidemiology as our motivating exemplar, this report will highlight the challenges of modeling novel infectious diseases while introducing ways to incorporate NN approximations to MFE. Prior to embarking on a Bayesian calibration, we first explored methods to augment the standard (non-Bayesian) UDE training procedure to account for uncertainty and increase robustness of training. In addition, it is often the case that uncertainty in observations is significant; this may be due to randomness or lack of precision in the measurement process. This uncertainty typically manifests as “noisy” observations which deviate from a true underlying signal. To account for such variability, the NN approximation to MFE is endowed with a probabilistic representation and is updated using available observational data in a Bayesian framework. By representing the MFU explicitly and deploying an embedded, data-driven model, this approach enables an agile, expressive, and interpretable method for representing MFU. In this report we will provide evidence that Bayesian UDEs show promise as a novel framework for any science-based, data-driven MFU representation; while emphasizing that significant advances must be made in the calibration of Bayesian NNs to ensure a robust calibration procedure.
Recent advancements in use of machine learning techniques on field-programmable gate arrays (FPGAs) have allowed for implementation of embedded neural networks with extremely low latency. This is invaluable for particle detectors at the Large Hadron Collider, where latency and used area must be strictly bounded. The hls4ml framework is a procedure for converting from trained machine learning model software, to a synthesis result that can be used on an FPGA. However, running the pipeline is a time-consuming procedure, and there is a strong risk of failure. In particular, it is possible that the model is unable to be converted into a synthesis result, or that the resource consumption of the model will exceed the resources of the target FPGA. To aid with this development, we introduce wa-hls4ml, a surrogate model which uses a graph neural network to emulate the structure of the source models. The goal is to estimate the chance of success and resource consumption of an arbitrary model when passed through the hls4ml procedure, without the time consumption of actually running the pipeline.
Design under uncertainty has significantly grown in research developments during the past decade. Additionally, machine learning (ML) and explainable ML (XML) have offered various opportunities to provide reliable predictable models. The current article investigates the use of finite element modeling (FEM), ML and XML predictions, and uncertain-based design of carbon-carbon (C-C) composites for use in ultra-high temperatures. A C-C composite concentrating solar power (CSP) as a microvascular receiver is considered as a case study. These C-C composites are fiber composites with directly integrated carbonized microchannels to form a lightweight, high-absorptivity material that includes an embedded microvascular network of channels. The topology of these microchannels is engineered to optimize heat transfer to a supercritical carbon dioxide (sCO2) heat transfer fluid. The mechanical characterization of C-C composites is highly challenging. Thus, designing every component made of C-C composites for ultra-high temperature applications needs an uncertainty-based analysis. As a part of a comprehensive project on the development of a novel carbonized microvascular C-C composite, this paper explores C-C composite sensitivity analysis, FEM, ML prediction, and XML analysis. The resulting composite can then be carbonized and coated with an oxidation-resistant coating to form a thermally efficient and mechanically robust C-C composite. An ANSYS 3-D-FE model was used to analyze the CSP’s stress/strain. To consider the variability in the mechanical and thermal properties of C-C composites, various mechanical properties are considered as the ANSYS FEM’s input. A synthetic dataset from 730 ANSYS runs was produced to feed into the ML and XML algorithms for uncertainty analysis and prediction. The ML and XML algorithms could accurately predict the CSP stresses/strains.
A nanofibrous catalyst and method of manufacture. A precursor solution of a transition metal based material is formed into a plurality of interconnected nanofibers by electro-spinning the precursor solution with the nanofibers converted to a catalytically active material by a heat treatment. Selected subsequent treatments can enhance catalytic activity.