Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “edge inference”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Even Higher-Level Synthesis: An Exploration of AI Hardware Accelerators using HLS4ML

With the rise of artificial intelligence, the popularization of deep learning, and a constantly evolving industry, the demand for flexible and efficient tools has never been greater. As algorithms grow more complex, their runtime and energy consumption increase exponentially. Customized hardware accelerators, long used for specific mathematical operations, remain essential for managing modern applications' computational and power demands. Hardware accelerators can speed up complex computations by orders of magnitude, but their manual design and verification processes are often challenging and time-consuming. High-Level Synthesis (HLS) provides a solution by transforming high-level algorithm descriptions, typically written in C++ or SystemC, into synthesizable RTL suitable for hardware implementation. This approach reduces development time for RTL engineers while offering flexibility beyond what traditional handwritten RTL can provide. We extended this capability to the machine-learning domain with the open-source framework hls4ml, which allows neural networks trained in Python frameworks like Tensorflow or PyTorch to be synthesized into efficient hardware representations for the traditional FPGA and ASIC flows. This breakthrough addresses the growing need for reduced design turnaround and easy verification of ML hardware accelerators with low latency and power efficiency constraints. During this tutorial, we will demonstrate how Python complements HLS by simplifying the ML design process, bridging the gap between software and hardware development. Attendees will explore how we translate neural networks modeled in Python into fixed-point C++ models suitable for HLS workflows. We will dive into strategies like Value-Range Analysis and Quantization-Aware Training, which optimize these designs for deployment and evaluate their accuracy, power consumption, and energy efficiency. To exemplify these concepts, experts from Fermilab will share their experiences applying this technology to high-energy physics experiments, where real-time, low-latency processing is critical. Over the years, Fermilab engineers have demonstrated how deep neural networks, optimized for hardware using hls4ml, can meet the stringent requirements of trigger systems at the CERN Large Hadron Collider. These systems rely on rapid decision-making to process immense data volumes while retaining only the most relevant events for further analysis. The application of hls4ml has also been extended to innovative technologies like smart pixel arrays. These smart pixels integrate ML inference capabilities directly into sensor devices, enabling localized data processing at the pixel level. This approach drastically reduces the need to transmit raw data to external processing units, significantly decreasing power consumption and latency. By embedding neural networks within the pixel architecture, the smart pixels can identify and prioritize relevant data in real time, providing a highly efficient solution for edge computing in scenarios such as particle detectors and imaging systems. Fermilab's work highlights the potential of hardware-accelerated ML in scenarios where both speed and power efficiency are mission-critical. Through this tutorial, attendees will gain valuable insights into the challenges and solutions of deploying ML in hardware. Understanding how HLS and hls4ml streamline the development of neural network-based hardware accelerators is fundamental for the industry's future. Participants will learn how these technologies are shaping the future of AI and scientific computing.

Di Guglielmo, Giuseppe [Fermilab]↗

Latest Word on Retreat of the West Antarctic Ice Sheet

The West Antarctic ice sheet during the Last Glacial Maximum (LGM) is estimated to have been three times its present volume and to have extended close to the edge of the continental shelf Holocene retreat of this ice sheet in the Ross Sea began between 11,000 and 12,000 years ago. This history implies an average contribution of this ice sheet to sea level of 0.9 mm/a. Evidence of dateable past grounding line positions in the Ross sector are broadly consistent with a linear retreat model. However, inferred rates of retreat for some of these grounding line positions are not consistent with a linear retreat model. More rapid retreat approximately 7600 years ago and possible near-stability in the Ross Sea sector at present suggest a slow rate of initial retreat followed by a more rapid-than-average retreat during the late Holocene, returning to a near-zero rate of retreat currently. This model is also consistent with the mid-Holocene high stand observations of eustatic sea level. Recent compilation of Antarctic bed elevations (BEDMAP) illustrates that the LGM and present grounding lines occur in the shallowest waters, further supporting the model of a middle phase of rapid retreat bracketed by an older and a more recent phase of modest retreat. Extension of these hypotheses into the future make subsequent behavior of the West Antarctic ice sheet more difficult to predict but suggest that if it loses its hold on the present shallow bed, the final retreat of the ice sheet could be very rapid.

Bindschadler, R.↗

Peak Prediction Using Multi Layer Perceptron (MLP) for Edge Computing ASICs Targeting Scientific Applications

High data rate detectors play an integral part in scientific research and their development is actively pursued at High Energy Physics (HEP) facilities around the world. Edge Machine Learning (ML) offers the ability to reduce data rates by integrating ML algorithms into Application Specific Integrated Circuits (ASICs) on the front end electronics. In this work, we explore a set of neural network architectures for predicting the peak amplitudes in the detector's sensor response. We have designed and synthesized several MLP based neural networks comparing their inference accuracy, power consumption, and area targeting for minimal latency. The neural networks are synthesized in a commercial 65nm process. The effect of quantizing the network's weights and biases on hardware performance and area is reported. We also conduct design space exploration to compare between design alternatives

47 OTHER INSTRUMENTATION↗

Identification of a characteristic doping for charge order phenomena in Bi-2212 cuprates via RIXS

Identifying quantum critical points (QCPs) and their associated fluctuations may hold the key to unraveling the unusual electronic phenomena observed in cuprate superconductors. Recently, signatures of quantum fluctuations associated with charge order (CO) have been inferred from the anomalous enhancement of CO excitations that accompany the reduction of the CO order parameter in the superconducting state. Furthermore, to gain more insight about the interplay between CO and superconductivity, we investigate the doping dependence of this phenomenon throughout the Bi-2212 cuprate phase diagram using resonant inelastic x-ray scattering (RIXS) at the Cu L 3 edge. As doping increases, the CO wavevector decreases, saturating near a commensurate value of 0.25 r.l.u. beyond a characteristic doping p c , where the correlation length becomes shorter than the apparent periodicity (4a 0 ). Such behavior is indicative of the fluctuating nature of the CO; and the proliferation of CO excitations in the superconducting state also appears strongest at p c , consistent with expected behavior at a CO QCP. Intriguingly, p c appears to be near optimal doping, where the superconducting transition temperature T c is maximal.

36 MATERIALS SCIENCE↗

Sources of field-aligned currents in the auroral plasma

Data from the Dynamics Explorer 1 High Altitude Plasma Instrument (HAPI) and magnetometer are used to investigate the sources of field-aligned currents in the nightside auroral zone. It is found that the formula developed by S. Knight predicts the field-aligned current density fairly accurately in regions where a significant potential drop can be inferred from the HAPI data; there are, however, regions in which the proportionality between potential drop and field-aligned current does not hold. In particular, occurrences of strong upward field-aligned current associated not with inverted-V events but instead with suprathermal bursts are noted. In addition, upward field-aligned currents are often observed to peak near the edges of inverted-V events, rather than in the center as would be predicted by Knight.

Marshall, J. A.↗

Fragmentation of the edge of a terminated Cu nanolayer within a Nb matrix upon annealing

Here, we investigate the evolution of the tip of a terminated Cu nano-layer embedded within a Nb matrix upon annealing at 800°C. The tip of the terminated Cu layer fragments progressively into a series of single-crystal Cu particles upon continued annealing. This finding shows that layer fragmentation is the second step, after layer pinchoff, in the thermal coarsening of nanolayered metal composites. Comparison to prior phase-field modeling permits us to infer that the mobility at 800°C of Cu-Nb, as modeled by the Cahn-Hillard equation, is M(c)=(M 0 = 7.3 ± 3.6 nm 5 / s • eV) |1 – c 2 |, where the order parameter, c, is 1 for Cu and -1 for Nb.

36 MATERIALS SCIENCE↗

Constructing Regulatory Networks to Compare Axenic and Interspecies Microbial Gene Transcription

In this preliminary study, we constructed gene regulatory networks (GRNs) from transcriptional expression data of axenic and interspecies microbial cultures with the goal of predicting how cocultivation affected greenhouse gas respiration by these species. The specific strains of Methylotuvimicrobium alkaliphilum 20Z, a methylotroph, and Cyanobacterium stanieri HL-69, a phototroph, were chosen for their viability in industrial bioprocessing. We ranked directed interactions between gene pairs based on the ability of the input gene to predict the expression of a target gene relative to their transcriptomes. While we were able to identify topological differences between conditions, our initial findings require validation through experimental analysis and further modeling. We aimed to develop a systematic thresholding approach to optimize the accuracy of our networks. We filtered out trial networks separately from top gene interactions of the scored rankings. Parameters of unfiltered and filtered networks were used to test and develop thresholding approaches. Knee point detection of edge weight distributions was explored as an approach for separating significant interactions from insignificant interactions in unfiltered networks. While knee detection failed to produce analogous networks for broad cross-condition comparisons, the results informed us about the proportions of significant edges present in unfiltered networks. We also calculated the average mean degree for nodes in a selection of trial networks to find a thresholding value characteristic to all groups. While we did not reach a definitive conclusion, we gained insight into the coregulatory structures of our groups and made critical evaluations of systematic methods for filtering networks. We recommend an iterative process for the inference of GRNs, where the most significant results from preliminary explorations are used to improve the efficiency with which regulatory motifs are chosen for experimental characterization. Experimental results can then inform the framework of adjusted models to improve broad interpretations of GRNs.

59 BASIC BIOLOGICAL SCIENCES↗

A digital scene matching technique for geometric image correction and autonomous navigation

A technique is described for precise registration of two images of the same area, taken under different conditions. The technique, called AUTO-MATCH, involves digital preprocessing of the images to extract edge contours, followed by a correlation of corresponding edge end points. The availability of an array of endpoints makes possible subpixed registration accuracy. The technique is applied to a system for assessment of geometric image quality, to be installed at NASA-Goddard. This system measures the registration vectors over an array of window pairs from LANDSAT images, so as to determine the distortion between them. Similar measurements could be used to infer relative positioning of the spacecraft. The technique is considered for updating the inertial navigation systems of missiles or aircraft. Scene matching between an image obtained aboard the vehicle and a stored reference can eliminate the drift of an inertial platform in order to demonstrate this capability in real time. A laboratory demonstration of AUTO-MATCH is described.

Tisdale, G. E.↗

Understanding the Scalability of Bayesian Network Inference using Clique Tree Growth Curves

Bayesian networks (BNs) are used to represent and efficiently compute with multi-variate probability distributions in a wide range of disciplines. One of the main approaches to perform computation in BNs is clique tree clustering and propagation. In this approach, BN computation consists of propagation in a clique tree compiled from a Bayesian network. There is a lack of understanding of how clique tree computation time, and BN computation time in more general, depends on variations in BN size and structure. On the one hand, complexity results tell us that many interesting BN queries are NP-hard or worse to answer, and it is not hard to find application BNs where the clique tree approach in practice cannot be used. On the other hand, it is well-known that tree-structured BNs can be used to answer probabilistic queries in polynomial time. In this article, we develop an approach to characterizing clique tree growth as a function of parameters that can be computed in polynomial time from BNs, specifically: (i) the ratio of the number of a BN's non-root nodes to the number of root nodes, or (ii) the expected number of moral edges in their moral graphs. Our approach is based on combining analytical and experimental results. Analytically, we partition the set of cliques in a clique tree into different sets, and introduce a growth curve for each set. For the special case of bipartite BNs, we consequently have two growth curves, a mixed clique growth curve and a root clique growth curve. In experiments, we systematically increase the degree of the root nodes in bipartite Bayesian networks, and find that root clique growth is well-approximated by Gompertz growth curves. It is believed that this research improves the understanding of the scaling behavior of clique tree clustering, provides a foundation for benchmarking and developing improved BN inference and machine learning algorithms, and presents an aid for analytical trade-off studies of clique tree clustering using growth curves.

Mengshoel, Ole Jakob↗

Analog Systems for Edge Optimization

Over the past decade, analog computing has the subject of substantial research interest providing a path toward improved computational efficiency in the post-Dennard era. Analog matrix vector multiplication (MVM) accelerators provide a popular approach given the ubiquity of MVM operations in numerous applications. However, historically analog computing systems can struggle with applications requiring high precision due to the inherent susceptibility of these systems to analog non-idealities. Therefore, prior work on analog systems has focused either on applications known to be tolerant of limited precision (e.g., neural network inference), or using expensive techniques to emulate high-precision using many analog MVM operations. In this work, we propose an alternative approach. Motivated by recent advances in inexact nonlinear solvers and optimizers, we explore the potential of co-designing optimization algorithms which can take full advantage of the fundamentally inexact analog MVM operations. To enable these co-designed algorithms we also develop a general mathematical theory of the precision and energy efficiency of analog operations, and a new system architecture for tightly-coupled analog and digital computation. Finally, we examine the applicability of analog computing to a wider class of symmetric positive definite systems and find potential in using analog operations as a sparse approximate inverse preconditioner. With these core innovations, this project provides a path toward effectively implementing optimization algorithms on power-constrained autonomous and semi-autonomous systems.

97 MATHEMATICS AND COMPUTING↗

Autonomous nondestructive evaluation of resistance spot welded joints

The application of non-destructive evaluation approaches has attracted strong interests in modern automotive industries. Here, we present an autonomous deep-computing framework to analyze raw videos from infrared systems and to predict weld nugget shape and size with unprecedented accuracy and speed. In a comprehensive training and testing experiment with 90 videos (seven sets of welding material stack-ups), a new method was developed to assemble sufficient datasets for neural network training. Our framework successfully predicts all the nugget shapes with F1 scores that range from 0.84 to 0.92. The total training time on Nvidia DGX station takes less than 10 min for each set of welding material stack-up. The real inference time of an individual dataset (with 30 video frames) takes about 0.005 s. The procedure and methods developed in the study can be applied to other image-based weld property prediction, as well as other manufacturing processes. Furthermore, our well-trained neural networks take limited memory resources (2.3 MB) and are suitable for embedded microprocessors for in-situ welding quality control as edge computing within an intelligent welding framework.

42 ENGINEERING↗

Development of Whole System Digital Twins for Advanced Reactors: Leveraging Graph Neural Networks and SAM Simulations

Here, in this work, we introduce a novel method to develop whole system digital twins (DTs) for advanced nuclear reactors. This method treats a complex reactor system as a heterogeneous graph: with the system components as different types of graph nodes and their physical interconnections as edges. Based on the heterogeneous graph, a graph neural network combining graph convolution and temporal node attention is developed as the DT, facilitating a comprehensive understanding of the system's dynamic behavior. By utilizing the System Analysis Module (SAM) code for simulating various operational transients, we develop a graph-based database that trains the DT. This DT is characterized by two primary functions: It can infer the entire system's status using sparse node information, and it can predict the progress of transients based on current and historical system information. Our approach is validated through case studies on the Experimental Breeder Reactor II (EBR-II) system and a generic Fluoride-salt-cooled High-temperature Reactor (gFHR), demonstrating the DT's accuracy in forecasting operational transients. The DT's rapid computation capabilities enhance its potential for supporting advanced reactor operations, offering benefits in intelligent simulation, autonomous control, and anomaly detection, paving the way for improved safety analysis and intelligent component health management for advanced reactor systems and reducing their operations and maintenance cost.

EBR-II↗

Effect of an oscillating flow direction on leading edge heat transfer

An experimental investigation was conducted to examine the effect of a periodic variation in the angle of attack on heat transfer at the leading edge of a gas turbine blade. A circular cylinder was used as a large-scale model of the leading edge region. The cylinder was placed in a wind tunnel and was oscillated rotationally about its axis. The incident flow Reynolds number and the Strouhal number of oscillation were chosen to model an actual turbine condition. Incident turbulence levels up to 4.9 percent were produced by grids placed upstream of the cylinder. The transfer rate was measured using a mass transfer technique and heat transfer rates inferred from the results. A direct comparison of the unsteady and steady results indicate that the effect is dependent on the Strouhal number, turbulence level, and the turbulence length scale, but that the largest observed effect was only a 10 percent augmentation at the nominal stagnation position.

Marziale, M. L.↗

Hopkins ultraviolet telescope observations of the far-ultraviolet spectrum of NGC 4151

Observations of the FUV spectrum of the Seyfert galaxy NGC 4151 from 912 to 1860 A during the flight of Astro-1 aboard the space shuttle Columbia in December 1990 are reported. Broad emission lines with full-width at half-maximum of about 8500 km/s dominate the spectrum. Numerous absorption features modify the continuum shape, particularly at wavelengths shortward of Ly-alpha. The continuum turns over sharply below 1000 A and disappears by 924 A, well above the redshifted Lyman edge of NGC 4151 at 915 A. The absorption lines have intrinsic widths of about 1000 km/s and are blueshifted relative to the system velocity of NGC 4151 by 200-1300 km/s. Absorption of the continuum by the converging higher-order Lyman lines explains the sharp turnover of the continuum below 1000 A. The blueshifts of the absorption lines, their large intrinsic widths, and the inferred high densities are all consistent with outflowing material originating in the broad-line region.

Kriss, G. A.↗

Toward an Autonomous Workflow for Single Crystal Neutron Diffraction

The operation of the neutron facility relies heavily on beamline scientists. Some experiments can take one or two days with experts making decisions along the way. Leveraging the computing power of HPC platforms and AI advances in image analyses, here we demonstrate an autonomous workflow for the single-crystal neutron diffraction experiments. The workflow consists of three components: an inference service that provides real-time AI segmentation on the image stream from the experiments conducted at the neutron facility, a continuous integration service that launches distributed training jobs on Summit to update the AI model on newly collected images, and a frontend web service to display the AI tagged images to the expert. Ultimately, the feedback can be directly fed to the equipment at the edge in deciding the next-step experiment without requiring an expert in the loop. With the analyses of the requirements and benchmarks of the performance for each component, this effort serves as the first step toward an autonomous workflow for real-time experiment steering at ORNL neutron facilities.

Yin, Junqi↗

Pretest Computational Assessment of Boundary Layer Transition in the NASA Juncture Flow Model with an NACA 0015-Based Wing

The first two phases of the NASA Juncture Flow experiment were carried out on a DLR-F6 swept-wing model and were designed to provide “CFD validation-quality” data toward the assessment and improvement of existing CFD turbulence models in predicting onset and extent of three-dimensional separated flow near the wing-juncture trailing-edge region. The next phase of experiments will involve an NACA 0015-based swept wing, as prior risk reduction experiments had indicated that this wing shape resulted in reduced separation near the juncture region than the DLR-F6 wing, thus providing a better option to evaluate the ability of CFD models to predict incipient turbulent separation. The NACA 0015 measurements will also include IR thermography to infer the variation of transition front with respect to an increasing angle of attack. The primary objective of this work is to computationally make a preliminary assessment of the transition front on both surfaces of the NACA 0015 wing at a crank-chord-based Reynolds number of 2.4 x 106 for four different angles of attack, (0°, 2.5°, 5°, and 7.5°) and to determine the dominant mechanisms responsible for transition. This assessment includes both RANS-based transition models from NASA’s OVERFLOW 2.3b flow solver and linear parabolized stability equations (PSE) stability analysis based on the Langley Stability and Transition Analysis code, LASTRAC. Linear PSE results indicate that the upper surface of the wing is dominated by Tollmien- Schlichting (TS) instabilities, and that the laminar flow region shrinks from about 50% chord to a very small region just downstream of the attachment line as the angle of attack is increased from 0° to 7.5°. Consequently, the transition fronts predicted by the Spalart- Allmaras-based amplification factor transport (AFT-2017b) equation model (which accounts for the TS instabilities alone) and the Menter’s shear-stress transport equation (SST2003)- based Langtry-Menter transition model with ability to account for both TS and crossflow effects (LM2015) compare well with those predicted using linear PSE. On the lower surface of the wing, stationary crossflow (CF) instabilities begin to appear on the inboard portion of the wing in addition to the TS-instabilities for the larger angles of attack (5° and 7.5°), further reducing the laminar flow extent within the inboard region. The LM2015 model that accounts for CF effects is able to replicate this trend but appears to predict a slightly earlier transition. The outcome of this effort will inform the experiment and, when the actual experimental data become available, provide further opportunity to assess and improve the various transition models.

CFD modeling↗

The effectiveness of D 2 pellet injection in reducing intra-ELM and inter-ELM tungsten divertor erosion rates in DIII-D during the Metal Rings Campaign

Abstract Edge localized modes (ELMs) in H-mode plasmas erode plasma-facing components (PFCs) and lead to impurities in the core, reducing confinement. This study analyzes D 2 pellet injection on the DIII-D fusion experiment used as an ELM mitigation technique applied during the 2016 tungsten Metal Rings Campaign to reduce W erosion during ELMs. The 400.9 nm photon wavelength line emission intensity of tungsten atoms (WI) filterscope channels and Langmuir probes were used to infer the gross erosion rate of tungsten-coated tiles installed in the divertor of DIII-D. D 2 mass injection rates ranging from 34 to 41 arbitrary units (A.U.) and no D 2 injection resulted in a similar total W erosion rate during ELMs (intra-ELM). On average, results show a 29% increase in the total gross W erosion rate with intermediate mass injection rates (∼13–23 A.U.) compared to the no pellets and the highest injection rate cases. On average, the fast D 2 mass injection rate cases had 15% less erosion in the inter-ELM phase than the case with no pellets. Generally, higher D 2 mass injection rates increased the ELM frequency, and the highest injection rates reduced the average erosion per ELM and fractional carbon impurities at the top of the pedestal by nearly 40% when compared to the no-pellet case. As expected, a higher D 2 pellet injection rate led to a higher plasma density and lower plasma temperature in the divertor. Additionally, an increasing divertor inter-ELM plasma electron density directly correlated to more frequent pellet injection and a decrease in both the average gross intra-ELM W erosion and the total gross intra-ELM W erosion rate. Simulations of intra-ELM erosion using the ‘free-streaming plus recycling model’ (FSRM) underestimate W erosion during pellet injection by about 30% on average. The discrepancies between the experimental measurements and the FSRM intra-ELM W erosion predictions are postulated to be due to C/W material mixing. A simple analytic mixed-material model is presented and results in better agreement with the experimental data. These results highlight the importance of incorporating the effects of a mixed-material layer in the analysis of PFC erosion.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Design and performance of the halogen occultation experiment (HALOE) remote sensor

HALOE is an optical remote sensor that measures extinction of solar radiation caused by the earth's atmosphere in eight channels, ranging in wavelength from 2.5 to 10.1 microns. These measurements, which occur twice each satellite orbit during solar occultation, are inverted to yield vertical distributions of middle atmosphere ozone (O3), water vapor, nitrogen dioxide, nitric oxide, hydrogen fluoride, hydrogen chloride, and methane. A channel located in the 2.7 region is used to infer the tangent point pressure by measuring carbon dioxide absorption. The HALOE instrument consists of a two-axis gimbal system, telescope, spectral discrimination optics and a 12-bit data system. The gimbal system tracks the solar radiometric centroid in the azimuthal plane and tracks the solar limb in the elevation plane, placing the instrument's instantaneous field-of-view 4 arcmin down from the solar top edge. The instrument gathers data for tangent altitudes ranging from 150 km to the earth's horizon. Prior to an orbital sunset and after an orbital sunrise, HALOE automatically performs calibration sequences to enhance data interpretation. The instrument is presently being tested at the NASA Langley Research Center in preparation for launch on the Upper Atmosphere Research Satellite near the end of this decade. This paper describes the instrumenmt design, operation, and functional performance.

Baker, R. L.↗