Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “edge inference”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

An ontology-based knowledge graph for representing interactions involving RNA molecules

The "RNA world" represents a novel frontier for the study of fundamental biological processes and human diseases and is paving the way for the development of new drugs tailored to each patient's biomolecular characteristics. Although scientific data about coding and non-coding RNA molecules are constantly produced and available from public repositories, they are scattered across different databases and a centralized, uniform, and semantically consistent representation of the "RNA world" is still lacking. We propose RNA-KG, a knowledge graph (KG) encompassing biological knowledge about RNAs gathered from more than 60 public databases, integrating functional relationships with genes, proteins, and chemicals and ontologically grounded biomedical concepts. To develop RNA-KG, we first identified, pre-processed, and characterized each data source; next, we built a meta-graph that provides an ontological description of the KG by representing all the bio-molecular entities and medical concepts of interest in this domain, as well as the types of interactions connecting them. Finally, we leveraged an instance-based semantically abstracted knowledge model to specify the ontological alignment according to which RNA-KG was generated. RNA-KG can be downloaded in different formats and also queried by a SPARQL endpoint. A thorough topological analysis of the resulting heterogeneous graph provides further insights into the characteristics of the "RNA world". RNA-KG can be both directly explored and visualized, and/or analyzed by applying computational methods to infer bio-medical knowledge from its heterogeneous nodes and edges. The resource can be easily updated with new experimental data, and specific views of the overall KG can be extracted according to the bio-medical problem to be studied.

59 BASIC BIOLOGICAL SCIENCES↗

Physics consistent machine learning framework for inverse modeling with applications to ICF capsule implosions

In high energy density physics (HEDP) and inertial confinement fusion (ICF), predictive modeling is complicated by uncertainty in parameters that characterize various aspects of the modeled system, such as those characterizing material properties, equation of state (EOS), opacities, and initial conditions. Typically, however, these parameters are not directly observable. What is observed instead is a time sequence of radiographic projections using X-rays. In this work, we define a set of sparse hydrodynamic features derived from the outgoing shock profile and outer material edge, which can be obtained from radiographic measurements, to directly infer such parameters. Our machine learning (ML)-based methodology involves a pipeline of two architectures, a radiograph-to-features network (R2FNet) and a features-to-parameters network (F2PNet), that are trained independently and later combined to approximate a posterior distribution for the parameters from radiographs. We show that the machine learning architectures are able to accurately infer initial conditions and EOS parameters, and that the estimated parameters can be used in a hydrodynamics code to obtain density fields, shocks, and material interfaces that satisfy thermodynamic and hydrodynamic consistency. Finally, we demonstrate that features resulting from an unknown EOS model can be successfully mapped onto parameters of a chosen analytical EOS model, implying that network predictions are learning physics, with a degree of invariance to the underlying choice of EOS model. To the best of our knowledge, our framework is the first demonstration of recovering both thermodynamic and hydrodynamic consistent density fields from noisy radiographs.

97 MATHEMATICS AND COMPUTING↗

Pedestal main ion particle transport inference through gas puff modulation with experimental source measurements

Abstract Transport in the DIII-D high confinement mode (H-mode) pedestal is investigated through a periodic edge gas puff modulation (GPM) which perturbs the deuterium density and source profiles. By using absolutely calibrated experimental edge ionization profile measurements, radial profiles of diffusion ( D ) and convection ( v ) are calculated into the pedestal region without depending on modeling the edge ionization source. An analytic approach with closed-form expressions for the D and v profiles and a more advanced Bayesian approach show evidence of an inward particle convection on the order of 1 m s −1 extending to normalized poloidal flux ( Ψ N ) of 0.98. Meanwhile, diffusion reaches a minimum value of ( 0.03 ± 0.02 ) m 2 s −1 in the pedestal region. Notably, the Bayesian approach, which utilizes the Aurora 1.5 D forward model inside the IMPRAD OMFIT module, provides radially resolved transport profiles with associated uncertainty without requiring an explicit form for the perturbation to the density profile or source. The combination of experimental ionization measurements and Bayesian inference provides an enhanced robust framework for investigating edge particle transport coefficients to experimentally test transport physics in order to improve predictive capabilities in the tokamak edge.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Tree rings reveal the transient risk of extinction hidden inside climate envelope forecasts

Given the importance of climate in shaping species’ geographic distributions, climate change poses an existential threat to biodiversity. Climate envelope modeling, the predominant approach used to quantify this threat, presumes that individuals in populations respond to climate variability and change according to species-level responses inferred from spatial occurrence data—such that individuals at the cool edge of a species’ distribution should benefit from warming (the “leading edge”), whereas individuals at the warm edge should suffer (the “trailing edge”). Using 1,558 tree-ring time series of an aridland pine (Pinus edulis) collected at 977 locations across the species’ distribution, we found that trees everywhere grow less in warmer-than-average and drier-than-average years. Ubiquitous negative temperature sensitivity indicates that individuals across the entire distribution should suffer with warming—the entire distribution is a trailing edge. Species-level responses to spatial climate variation are opposite in sign to individual-scale responses to time-varying climate for approximately half the species’ distribution with respect to temperature and the majority of the species’ distribution with respect to precipitation. These findings, added to evidence from the literature for scale-dependent climate responses in hundreds of species, suggest that correlative, equilibrium-based range forecasts may fail to accurately represent how individuals in populations will be impacted by changing climate. A scale-dependent view of the impact of climate change on biodiversity highlights the transient risk of extinction hidden inside climate envelope forecasts and the importance of evolution in rescuing species from extinction whenever local climate variability and change exceeds individual-scale climate tolerances.

54 ENVIRONMENTAL SCIENCES↗

Knowledge Graph of RB-Tnseq Data from Fitness Browser (KP-DP1)

Motivation: Predicting microbial gene fitness across environmental conditions remains a central challenge for predictive phenomics and autonomous experimentation. Fitness assays generate large volumes of genotype–phenotype measurements difficult to integrate with experimental metadata and biological function in a form that supports mechanistic reasoning. Knowledge graphs offer a semantic framework for unifying modalities and enabling context-aware inference. Results: We build GIMME (Graph Inference for Microbial Metabolism Exploration), a semantically grounded knowledge graph that unifies gene fitness measurements spanning 10 Pseudomonas species with experimental metadata and biological context. Media are decomposed into chemical components and experiments carry structured links to natural-language descriptions. The resulting graph supports two inference modes: (1) symbolic graph traversal to surface candidate gene–environment and gene–chemical associations, and (2) learned inference using heterogeneous graph neural networks that propagate information across neighborhoods. We formulate link regression over (gene, media, experiment) triplets, combining learned gene embeddings with pretrained LLM sourced text embeddings of node descriptions to predict gene fitness. We then augment a baseline MLP with an auxiliary message-passing encoder (GraphSAGE/GAT) that propagates information over gene–protein–function and media–chemical subgraphs, and fuse the two pathways with a gated residual connection. This approach produces strong agreement with held-out fitness measurements (GraphSAGE Pearson r 0.74) while also highlighting inference challenges in extreme-fitness regimes. We aggregate GAT edge-attention weights by relation type and layer to estimate which biological and environmental relations most influence fitness predictions. Conclusion: This work explores using knowledge graphs as “context graphs” for microbial phenotype prediction. They provide a rich substrate which enables explainable retrieval of supporting evidence, and provides a natural bridge to autonomous workflows that prioritize the next experiment.

59 BASIC BIOLOGICAL SCIENCES↗

Community detection in hypergraphs via mutual information maximization

Abstract The hypergraph community detection problem seeks to identify groups of related vertices in hypergraph data. We propose an information-theoretic hypergraph community detection algorithm which compresses the observed data in terms of community labels and community-edge intersections. This algorithm can also be viewed as maximum-likelihood inference in a degree-corrected microcanonical stochastic blockmodel. We perform the compression/inference step via simulated annealing. Unlike several recent algorithms based on canonical models, our microcanonical algorithm does not require inference of statistical parameters such as vertex degrees or pairwise group connection rates. Through synthetic experiments, we find that our algorithm succeeds down to recently-conjectured thresholds for sparse random hypergraphs. We also find competitive performance in cluster recovery tasks on several hypergraph data sets.

97 MATHEMATICS AND COMPUTING↗

Sentinel

Network intrusion detection systems (NIDS) are commonplace in network security but they frequently employ algorithms that are computational demanding requiring hardware and software with significant power requirements. Two examples of such resource-intensive algorithms used for network security are regular expression matching and broader signature pattern matching which are commonly used in deep packet inspection (DPI). Network security algorithms that have large power requirements may be a challenge for low-power internet-of-things (IoT) environments, which generally lack the power resources to implement complex security measures like computationally expensive DPI at the edge. Furthermore, IoT environments incorporating 5G standalone networks have network latency constraints beyond just power that make DPI at the edge even more difficult. Programmable logic is ideally suited for machine learning inference for DPI because of its deep instruction level parallelism and single-cycle memory access. Machine learning approaches for DPI have been explored before using the programmable logic of field programmable gate arrays (FPGA) as a potential solution for NIDS approaches that would be power-suitable for IoT. However, those previous programmable logic NIDS approaches utilize either a supervised or unsupervised learning model. Sentinel utilizes the ensemble of these two machine learning approaches known as a semi-supervised approach which has shown promise in NIDS implementations. Sentinel provides a programmable logic implementation of a semi-supervised approach for DPI which operates at much lower power and latency than a GPU implementation with negligible loss of accuracy due to quantization through a logistic regressor.

Anderson, MatthewW [Idaho National Laboratory (INL↗

Disentangling the Black Hole Mass Spectrum with Photometric Microlensing Surveys

Abstract From the formation mechanisms of stars and compact objects to nuclear physics, modern astronomy frequently leverages surveys to understand populations of objects to answer fundamental questions. The population of dark and isolated compact objects in the Galaxy contains critical information related to many of these topics, but is only practically accessible via gravitational microlensing. However, photometric microlensing observables are degenerate for different types of lenses, and one can seldom classify an event as involving either a compact object or stellar lens on its own. To address this difficulty, we apply a Bayesian framework that treats lens type probabilistically and jointly with a lens population model. This method allows lens population characteristics to be inferred despite intrinsic uncertainty in the lens class of any single event. We investigate this method’s effectiveness on a simulated ground-based photometric survey in the context of characterizing a hypothetical population of primordial black holes (PBHs) with an average mass of 30 M ⊙ . On simulated data, our method outperforms current black hole (BH) lens identification pipelines and characterizes different subpopulations of lenses while jointly constraining the PBH contribution to dark matter to ≈25%. Key to robust inference, our method can marginalize over population model uncertainty. We find the lower mass cutoff for stellar origin BHs, a key observable in understanding the BH mass gap, particularly difficult to infer in our simulations. This work lays the foundation for cutting-edge PBH abundance constraints to be extracted from current photometric microlensing surveys.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Analyzing Data Privacy for Edge Systems

Internet-of-Things (IoT)-based streaming applications are all around us. Currently, we are transitioning from IoT processing being performed on the cloud to the edge. While moving to the edge provides significant networking efficiency benefits, IoT edge computing creates significant data privacy concerns. We propose a methodology that can successfully privacy protect the continual data streams generated by sensors on the edge device. We implement local differential privacy on streaming data and incorporate Bayesian inference and Gaussian process to evaluate the privacy policy. We demonstrate our methodology on a real-world smart meter testbed and identify the optimal privacy protection settings.

Kotevska, Olivera↗

Avalanche statistics of fluctuation-induced fluxes from the SLPM and the W7-AS stellarator

Measurements of fluctuating floating potentials and ion saturation currents at different radial locations in the Santander Linear Plasma Machine (Castellanos et al 2005 Plasma Phys. Control. Fusion47 2067) and at the edge of the W7-AS stellarator by means of radially movable Langmuir probes allow to infer the corresponding fluctuation-induced radial flux temporal series. Avalanche-like transport events are identified in the time series and statistically characterized in terms of avalanche size/duration/quiet-time distributions and size-duration scaling relations. Transport is diffusive in the inner and intermediate radial region of the SLPM r < r tr ≈ 2.6 cm, undergoing a transition at r tr , becoming non-diffusive in the outermost region of the device, r > r tr . Here, the results obtained at the edge of the W7-AS stellarator are similar to those found in SLPM for r > r tr , i.e. consistent with what would be expected for scale-free, self-similar plasma transport dynamics near a critical state.

Avalanches↗

Cross Inference of Throughput Profiles Using Micro Kernel Network Method

Dedicated network connections are being increasingly deployed in cloud, centralized and edge computing and data infrastructures, whose throughput profiles are critical indicators of the underlying data transfer performance. Due to the cost and disruptions to physical infrastructures, network emulators, such as Mininet, are often used to generate measurements needed to estimate throughput profiles, typically expressed as a function of the connection round trip time. The profiles estimated using measurements from such emulated networks are usually inaccurate for high bandwidth and high latency connections, since they do not accurately reflect the critical network transport dynamics mainly due to computing and memory constraints of the host. We present a machine learning (ML) method to estimate the throughput profiles using emulation measurements to closely match the testbed and production network profiles. In particular, we propose a micro Kernel Network (mKN) that provides baseline throughput measurements on the host running Mininet emulations, which are used to learn a regression map that converts them to the corresponding testbed measurement estimates. Once initially learned, this map is applied to measurements from subsequent network emulations on the same host. We present experimental measurements to illustrate this approach, and derive generalization equations for the proposed mKN-ML method. Using a four-site scenario emulation, we show the effectiveness of this method in providing accurate concave throughput profiles from inaccurate convex or non-smooth ones indicated by Mininet emulation.

Rao, Nageswara↗

Differentially Private Synthesis and Sharing of Network Data Via Bayesian Exponential Random Graph Models

Abstract Network data often contain sensitive relational information. One approach to protecting sensitive information while offering flexibility for network analysis is to share synthesized networks based on the information in originally observed networks. We employ differential privacy (DP) and exponential random graph models (ERGMs) and propose the DP-ERGM method to synthesize network data. We apply DP-ERGM to two real-world networks. We then compare the utility of synthesized networks generated by DP-ERGM, the DyadWise Randomized Response (DWRR) approach, and the Synthesis through Conditional distribution of Edge given nodal Attribute (SCEA) approach. In general, the results suggest that DP-ERGM preserves the original information significantly better than two other approaches in network structural statistics and inference for ERGMs and latent space models. Furthermore, DP-ERGM satisfies node DP through modeling the global network structure with ERGM, a stronger notion of privacy than the edge DP under which DWRR and SCEA operate.

graph synthesis↗

Experimental study of the edge radial electric field in different drift configurations and its role in the access to H-mode at ASDEX Upgrade

The formation of the equilibrium radial electric field (Er) has been studied experimentally at ASDEX Upgrade (AUG) in L-modes of “favorable” (ion ∇ B-drift toward primary X-point) and “unfavorable” (ion ∇ B-drift away from primary X-point) drift configurations, in view of its impact on H-mode access, which changes with drift configurations. Edge electron and ion kinetic profiles and impurity velocity and mean-field Er profiles across the separatrix are investigated, employing new and improved measurement techniques. The experimental results are compared to local neoclassical theory as well as to a simple 1D scrape-off layer (SOL) model. It is found that in L-modes of matched heating power and plasma density, the upstream SOL Er and the main ion pressure gradient in the plasma edge are the same for either drift configurations, whereas the Er well in the confined plasma is shallower in unfavorable compared to the favorable drift configuration. The contributions of toroidal and poloidal main ion flows to Er, which are inferred from local neoclassical theory and the experiment, cannot account for these observed differences. Furthermore, it is found that in the L-mode, the intrinsic toroidal edge rotation decreases with increasing collisionality and it is co-current in the banana-plateau regime for all different drift configurations at AUG. This gives rise to a possible interaction of parallel Pfirsch–Schlüter flows in the SOL with the confined plasma. Thus, the different H-mode power threshold for the two drift configurations cannot be explained in the same way at AUG as suggested by LaBombard et al. [Phys. Plasmas 12, 056111 (2005)] for Alcator C-Mod. Finally, comparisons of Er profiles in favorable and unfavorable drift configurations at the respective confinement transitions show that also the Er gradients are all different, which indirectly indicates a different type or strength of the characteristic edge turbulence in the two drift configurations.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Towards On-Chip Learning for Low Latency Reasoning with End-to-End Synthesis

The Software Defined Architectures (SODA) Synthesizer is an open-source compiler-based tool able to automatically generate domain-specialized systems targeting Application-Specific Integrated Circuits (ASICs) or Field Programmable Gate Arrays (FPGAs) starting from high-level programming. SODA is composed of a frontend, SODA-OPT, which leverages the multilevel intermediate representation (MLIR) framework to interface with productive programming tools (e.g., machine learning frame-works), identify kernels suitable for acceleration, and perform high-level optimizations, and of a state-of-the-art high-level synthesis backend, Bambu from the PandA framework, to generate custom accelerators. One specific application of the SODA Synthesizer is the generation of accelerators to enable ultra-low latency inference and control on autonomous systems for scientific discovery (e.g., electron microscopes, sensors in particle accelerators, etc.). This paper provides an overview of the flow in the context of the generation of accelerators for edge processing to be integrated in transmission electron microscopy (TEM) devices, focusing on use cases from precision material synthesis. We show the tool in action with an example of design space exploration for inference on reconfigurable devices with a conventional deep neural network model (LeNet). Finally, we discuss the research directions and opportunities enabled by SODA in the area of autonomous control for scientific experimental workflows.

Castellana, Vito G.↗

X-ray absorption spectroscopy study of Mn reference compounds for Mn speciation in terrestrial surface environments

Abstract X-ray absorption spectroscopy (XAS) offers great potential to identify and quantify Mn species in surface environments by means of linear combination fit (LCF), fingerprint, and shell-fit analyses of bulk Mn XAS spectra. However, these approaches are complicated by the lack of a comprehensive and accessible spectrum library. Additionally, molecular-level information on Mn coordination in some potentially important Mn species occurring in soils and sediments is missing. Therefore, we investigated a suite of 32 natural and synthetic Mn reference compounds, including Mn oxide, oxyhydroxide, carbonate, phosphate, and silicate minerals, as well as organic and adsorbed Mn species, by Mn K-edge X-ray absorption near edge structure (XANES) and extended X-ray absorption fine structure (EXAFS) spectroscopy. The ability of XAS to infer the average oxidation state (AOS) of Mn was assessed by comparing XANES-derived AOS with the AOS obtained from redox titrations. All reference compounds were studied for their local (<5 Å) Mn coordination environment using EXAFS shell-fit analysis. Statistical analyses were employed to clarify how well and to what extent individual Mn species (groups) can be distinguished by XAS based on spectral uniqueness. Our results show that LCF analysis of normalized XANES spectra can reliably quantify the Mn AOS within ~0.1 v.u. in the range +2 to +4. These spectra are diagnostic for most Mn species investigated, but unsuitable to identify and quantify members of the manganate and Mn(III)-oxyhydroxide groups. First-derivative XANES fingerprinting allows the unique identification of pyrolusite, ramsdellite, and potentially lithiophorite within the manganate group. However, XANES spectra of individual Mn compounds can vary significantly depending on chemical composition and/or crystallinity, which limits the accuracy of XANES-based speciation analyses. In contrast, EXAFS spectra provide a much better discriminatory power to identify and quantify Mn species. Principal component and cluster analyses of k2-weighted EXAFS spectra of Mn reference compounds implied that EXAFS LCF analysis of environmental samples can identify and quantify at least the following primary Mn species groups: (1) Phyllo- and tectomanganates with large tunnel sizes (2 × 2 and larger; hollandite sensu stricto, romanèchite, todorokite); (2) tectomanganates with small tunnel sizes (2 × 2 and smaller; cryptomelane, pyrolusite, ramsdellite); (3) Mn(III)-dominated species (nesosilicates, oxyhydroxides, organic compounds, spinels); (4) Mn(II) species (carbonate, phosphate, and phyllosilicate minerals, adsorbed and organic species); and (5) manganosite. All Mn compounds, except for members of the manganate group (excluding pyrolusite) and adsorbed Mn(II) species, exhibit unique EXAFS spectra that would allow their identification and quantification in mixtures. Therefore, our results highlight the potential of Mn K-edge EXAFS spectroscopy to assess bulk Mn speciation in soils and sediments. A complete XAS-based speciation analysis of bulk Mn in environmental samples should preferably include the determination of Mn valences following the “Combo” method of Manceau et al. (2012), EXAFS LCF analyses based on principal component and target transformation results, as well as EXAFS shell-fit analyses for the validation of LCF results. For this purpose, all 32 XAS reference spectra are provided in the Online Materials1 for further use by the scientific community.

Geochemistry & Geophysics↗

A Voltage Inference Framework for Real-Time Observability in Active Distribution Grids

Active distribution grids are gaining traction to meet the growing environmental, socio-economic, and sustainability targets. Various advanced smart grid technologies facilitate the integration of Distributed Energy Resources (DERs) by supporting the bi-directional power flow. The limited observability of distribution grids, primarily related to their location at the very edge of power system infrastructure, brings challenges to optimal grid management. Moreover, only a limited number of measurements at regular intervals are usually available. This paper presents a novel inference framework, referred to as “Voltage Inference”, to overcome the observability issues. The proposed framework employs a prediction step based on the Multivariate Taylor series approximation, followed by a corrector step that minimizes the estimation error to infer the otherwise unknown voltages from the available measurements. Furthermore, numerical results on the IEEE 13-bus test feeder validate the accuracy and computational performance of the proposed framework.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Toward an AI-Powered Software Pipeline for Real-Time Tracking and Analysis of Wildfire and Smoke

Real-time tracking of wildfires and smoke is crucial for effective response, minimizing damage, protecting lives, and efficiently managing resources during fire emergencies. We develop a web-based AI-powered pipeline that detects wildfires in aerial video and estimates deployment-relevant behavior metrics, including cumulative burned area, burned-area growth rate, fire spread direction, and smoke dispersion. The system combines a YOLO-based detector with YCbCr-based fire segmentation, HSV-based smoke segmentation, Farneback optical flow, and centroid-based spatiotemporal tracking. Using ground sampling distance (GSD), pixel-level fire masks are converted to physical burned-area measurements by correlating fire pixel counts with camera altitude and tilt angle. We benchmark YOLO variants and non-YOLO baselines (GoogLeNet, CNN, DBN, Autoencoder, U-Net, and AlexNet) on the IEEE FLAME dataset and a newly created aerial frame dataset, Wildfire-DB. Cross-dataset evaluation uses a strict threshold-transfer protocol: decision thresholds are selected on FLAME validation and transferred unchanged to Wildfire-DB to quantify generalization under domain shift. YOLOv6 achieves the strongest cross-dataset frame-level fire detection on Wildfire-DB (ROC-AUC 0.8200, PR-AUC 0.8044, and transferred-threshold F1 0.7596). For tracking-oriented deployment requiring oriented localization, YOLO11-OBB provides the most reliable cross-dataset behavior among OBB-capable models while remaining computationally feasible. To analyze the feasibility of UAV deployment, we further measure inference efficiency using synchronized GPU and CPU power logs on a fixed workload of 1569 frames. YOLO-family models process the video in 5.73–12.47 seconds with net energy of 1247.28–1775.39 J, substantially lower latency and energy than heavier classification and reconstruction baselines. Overall, model optimality depends on operational objectives: YOLOv6 is best for cross-dataset detection robustness, whereas YOL...

Color segmentation↗

Real-Time Inference For MI/RR Deblending

The Fermilab Main Injector (MI) and Recycler Ring (RR) share a common beam loss monitor (BLM) system, making loss events difficult to attribute to their source machine when beam is present in both simultaneously. The Real-time Edge AI for Distributed Systems (READS) project addresses this by deblending BLM readings in real time using machine learning (ML). The current FPGA based implementation meets the sub-3 ms latency requirement but carries a resource intensive hls4ml development cycle, motivating exploration of GPU based deployment. This paper characterizes inference latency on an NVIDIA Jetson Orin Nano and introduces a packet organization scheme for assembling synchronized event frames from seven distributed BLM DAQ streams. Using a Python based DAQ simulation with injected timing jitter in place of unavailable live beam data, the pipeline achieved an average end to end latency of 0.456 ms (σ = 0.122 ms) across 167,000 test frames, comfortably meeting the timing constraint. Early outliers were attributed to TensorRT warm-up rather than steady state limitations, suggesting GPU based inference is a viable alternative to the existing FPGA implementation.

Yu, Kellen [Cornell U.]↗