Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “network acceleration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Machine learning accelerated discrete element modeling of granular flows

Granular flows are widely encountered in many industrial processes and natural phenomena. Discrete Element Modeling (DEM) is a useful tool for understanding and troubleshooting devices, which handle granular materials. However, its applicability is significantly limited by the huge computational cost associated with detecting and computing collisions. In this research, the computation speed of DEM was accelerated by orders of magnitude using a convolutional neural network to replace the direct calculation of particle-particle and particle-boundary collisions. The MFiX software was used to generate the training and testing dataset. Additionally, a GPU accelerated TensorFlow model was used to train the neural network and test the results. The model fluctuations caused by different training steps were reduced with a multi-scale loss function. The accuracy was improved with more frames within one training step. The modeling of a rotating drum and a hopper demonstrated the accuracy and efficiency of this machine learning accelerated DEM in the simulation of granular flows.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

CSB-RNN: A Faster-Than-Realtime RNN Acceleration Framework with Compressed Structured Blocks

Recurrent Neural Networks (RNN) is widely applied to temporal sequence analysis, where real-time performance is usually in demand. However, RNN suffers a heavy computational workload as the model comes with a large weight matrix. To alleviate the pain, model compression (pruning) schemes have been proposed for RNN that pruning the redundant (near-zero) weight-values. On the one hand, the non-structured pruning methods achieve a considerable pruning rate while bringing the computational irregularity, which is un-friendly to parallel-hardware. On the other hand, the existing structured pruning methods consider the hardware parallelism; However, they suffer a poor pruning rate due to the restrict constraints on pruning structure. This paper presents CSB-RNN, an optimized full-stack RNN framework with the novel compressed structured block (CSB) technique. The CSB-pruned RNN model comes with both fine-granularity that benefits the pruning rate and regular structure that facilitates the hardware-parallelism. Further, we propose a novel hardware architecture for inferencing the CSB-pruned model. Different from conventional parallel hardware, this architecture solves the block-workload imbalance issue and achieves an over 95% hardware utilization. With the experiments on 10 RNN models in 5 application domains, the CSB-RNN realizes 7×-20× lossless compression and up to 50× acceptable lossy-compression, which is 2×-7× to the prior art. With the addition of the novel hardware, the compressed-RNN inference reaches a super real-time latency of 10-400µs with FPGA implementation.

Shi, Runbin↗

Neural message-passing for objective-based uncertainty quantification and optimal experimental design

Various real-world scientific applications involve the mathematical modeling of complex uncertain systems with numerous unknown parameters. Accurate parameter estimation is often practically infeasible in such systems, as the available training data may be insufficient and the cost of acquiring additional data may be high. In such cases, based on a Bayesian paradigm, we can design robust operators retaining the best overall performance across all possible models and design optimal experiments that can effectively reduce uncertainty to enhance the performance of such operators maximally. While objective-based uncertainty quantification (objective-UQ) based on MOCU (mean objective cost of uncertainty) provides an effective means for quantifying uncertainty in complex systems, the high computational cost of estimating MOCU has been a challenge in applying it to real-world scientific/engineering problems. In this work, we propose a novel scheme to reduce the computational cost for objective-UQ via MOCU based on a data-driven approach. We adopt a neural message-passing model for surrogate modeling, incorporating a novel axiomatic constraint loss that penalizes an increase in the estimated system uncertainty. As an illustrative example, we consider the optimal experimental design (OED) problem for uncertain Kuramoto models, where the goal is to predict the experiments that can most effectively enhance robust synchronization performance through uncertainty reduction. We show that our proposed approach can accelerate MOCU-based OED by four to five orders of magnitude, without any visible performance loss compared to the state-of-the-art. The proposed approach applies to general OED tasks, beyond the Kuramoto model.

97 MATHEMATICS AND COMPUTING↗

0BGRaman: Graph Network based Simulator for Forecasting Molecular Polarizability

This report presents the work performed under the GRaman project, sponsored by the PCSD LDRD Seed program. The project aimed at accelerating ab initio molecular dynamics simulation using Graph Networks. The Graph Network framework is a ML framework that has been successfully employed to simulate the dynamics of several physical systems: including water splashing in a container and flags moving with the wind. In this effort, we performed a data collection campaign for 3 different molecules of interest. We have built tools for preprocessing the trajectories obtained by simulating Raman Spectroscopy with NWChem and translating them into a suitable format for training. We have developed a training algorithm to train the Graph Network based simulators based on our data and developed a simulator that produces trajectories in the same NWChem format. While the tool has improved with each iteration of development and subsequent experiments, the current state of the tool does not allow to directly incorporate the technology within the NWChem framework because the trajectories produced by the tool are not yet accurate enough. However, the technology has proved to have good potential and it is certainly worth further research and development.

97 MATHEMATICS AND COMPUTING↗

Launch Alaska Transportation and Energy Accelerator (LATEA)

The Launch Alaska Transportation and Energy Accelerator (LATEA), funded through the U.S. Department of Energy Office of Technology Commercialization’s Energy Program for Innovation Clusters (EPIC),advanced deployment of innovative and efficient transportation and energy technology in Alaska from October 2021 through June 2025. The project was designed to leverage Launch Alaska’s accelerator model to identify, recruit, and support transportation technology companies with novel solutions to market needs while building the stakeholder networks, demonstration opportunities, and institutional capacity necessary to accelerate commercialization in one of the most challenging operating environments in the United States.

08 HYDROGEN↗

Electronic Bottleneck Suppression in Next‐Generation Networks with Integrated Photonic Digital‐to‐Analog Converters

Digital‐to‐analog converters (DAC) are indispensable functional units in signal processing instrumentation and wide‐band telecommunication links for both civil and military applications. As photonic systems are capable of high data throughput and low latency, an increasingly found system limitation stems from the required domain crossing such as digital to analog and electronic to optical. A photonic DAC implementation, in contrast, enables a seamless signal conversion with respect to both energy efficiency and short signal delay, often requiring bulky discrete optical components and electric–optic transformation, hence introducing inefficiencies. Herein, a novel coherent parallel photonic DAC concept along with a 4‐bit experimental prototype capable of performing this DAC without optic–electric–optic domain crossing is introduced. This new paradigm guarantees a linear intensity weighting among bits when operating at high sampling rates (50 GHz), featuring an exceptional sampling efficiency (> 100 GS ) and small footprint (≈1 mm 2 ) in an 8‐bit implementation. Importantly, this photonic DAC enables seamless interfaces of next‐generation data processing hardware with high relevance in data centers, task‐specific compute accelerators such as neuromorphic engines, and network edge processing applications.

Meng, Jiawei↗

Efficient Interdependent Systems Recovery Modeling with DeepONets

Modeling the recovery of interdependent critical infrastructure is a key component of quantifying and optimizing societal resilience to disruptive events. However, simulating the recovery of large-scale interdependent systems under random disruptive events is computationally expensive. Therefore, we propose the application of Deep Operator Networks (DeepONets) in this paper to accelerate the recovery modeling of interdependent systems. DeepONets are ML architectures which identify mathematical operators from data. The form of governing equations DeepONets identify and the governing equation of interdependent systems recovery model are similar. Therefore, we hypothesize that DeepONets can efficiently model the interdependent systems recovery with little training data. We applied DeepONets to a simple case of four interdependent systems with sixteen states. DeepONets, overall, performed satisfactorily in predicting the recovery of these interdependent systems for out of training sample data when compared to reference results.

97 MATHEMATICS AND COMPUTING↗

Creating a Research Enterprise Framework for Transdisciplinary Networking to Address the Food–Energy–Water Nexus

Urbanization, population growth, and the accelerating consumption of food, energy, and water (FEW) resources bring unprecedented challenges for economic, environmental, and social (EES) sustainability. It is imperative to understand the potential impacts of FEW systems on the realization of the United Nation’s Sustainable Development Goals (SDGs) as the world transitions from natural ecosystems to managed ecosystems at an accelerating rate. A major obstacle is the complexity and emergent behavior of FEW systems and associated networks, for which no single discipline can generate a holistic understanding or meaningful projections. We propose a research enterprise framework for promoting transdisciplinarity and top-down quantification of the interrelationships between FEW and EES systems. Relevant enterprise efforts would emphasize increasing FEW resource accessibility by improving coordinated interplays across sectors and scales, expanding and diversifying supply-chain networks, and innovating technologies for efficient resource utilization. This framework can guide the development of strategic solutions for diminishing the competition among FEW-consuming sectors in a region or country, and for minimizing existing inequalities in FEW availability when a sustainable development agenda is implemented.

42 ENGINEERING↗

FPGA Acceleration of GCN in Light of the Symmetry of Graph Adjacency Matrix

Graph Convolutional Neural Networks (GCNs) are widely used to process large-scale graph data. Different from deep neural networks (DNNs), GCNs are sparse, irregular, and unstructured, posing unique challenges to hardware acceleration with regular processing elements (PEs). In particular, the adjacency matrix of a GCN is extremely sparse, leading to frequent but irregular memory access, low spatial/temporal data locality and poor data reuse. Furthermore, a realistic graph usually consists of unstructured data (e.g., unbalanced distributions), creating significantly different processing times and imbalanced workload for each node in GCN acceleration. To overcome these challenges, we propose an end-to-end hardware-software co-design to accelerate GCNs on resource-constrained FPGAs with the features including: (1) A custom dataflow that leverages symmetry along the diagonal of the adjacency matrix to accelerate feature aggregation for undirected graphs. We utilize either the upper or the lower triangular matrix of the adjacency matrix to perform aggregation in GCN to improve data reuse. (2) Unified compute cores for both aggregation and transform phases, with full support to the symmetry-based dataflow. These cores can be dynamically reconfigured to the systolic mode for transformation or as individual accumulators for aggregation in GCN processing. (3) Preprocessing of the graph in software to rearrange the edges and features to match the custom dataflow. This step improves the regularity in memory access and data reuse in the aggregation phase. Moreover, we quantize the GCN precision from FP32 to INT8 to reduce the memory footprint without losing the inference accuracy. We implement our accelerator design in Intel Stratix10 MX FPGA board with HBM2, and demonstrate 1.3x-110.5x improvement in end-to-end GCN latency as compared to the state-of the-art FPGA implementations, on the graph datasets of Cora, Pubmed, Citeseer and Reddit.

Nair, Gopikrishnan R.↗

Autoencoder Neural Network for chemically reacting systems

Incorporating detailed chemical kinetic models is critical for accurate simulations of reacting flows. However, detailed models involve a large number of thermochemical (TC) state variables. Solving the governing equations to evolve these TC variables becomes impractical for real-world applications. In this work, we propose an autoencoder (AE) neural network (NN)-based reduced model to accelerate such simulations. The AE NN is first trained to find a low-dimensional latent representation of the TC states. Then, the evolving state of a chemical system can be tracked by solving the equations of the latent variables instead of the original TC equations. We demonstrate the reduced model in a syngas CO/H 2 combustion system, using training data collected from canonical perfectly stirred reactors (PSRs). It is found that the AE model can reduce the dimension of the combustion system from 12 to 2 while maintaining low reconstruction error and excellent elemental mass conservation for the test dataset. In the a posteriori test, the combustion states obtained from solving the two latent equations are compared to those from solving the 12 equations of the full model. The AE reduced method is found to be able to capture the diverse combustion states on the top two branches of the S-curve well including the extinction turning point, but with higher prediction errors for states near the ignition turning point.

97 MATHEMATICS AND COMPUTING↗

Ionizing radiation effects in SONOS-based neuromorphic inference accelerators

Here, we evaluate the sensitivity of neuromorphic inference accelerators based on Silicon-Oxide-Nitride-Oxide-Silicon (SONOS) charge trap memory arrays to total ionizing dose (TID) effects. Data retention statistics were collected for 16 Mbit of 40 nm SONOS digital memory exposed to ionizing radiation from a Co-60 source, showing good retention of the bits up to the maximum dose of 500 krad(Si). Using this data, we formulate a rate-equation-based model for the TID response of trapped charge carriers in the ONO stack, and predict the effect of TID on intermediate device states between ‘program’ and ‘erase’. This model is then used to simulate arrays of low-power, analog SONOS devices that store 8-bit neural network weights and support in situ matrix-vector multiplication. We evaluate the accuracy of the irradiated SONOS-based inference accelerator on two image recognition tasks – CIFAR-10 and the challenging ImageNet dataset – using state-of-the-art convolutional neural networks, such as ResNet-50. We find that across the datasets and neural networks evaluated, the accelerator tolerates a maximum TID between 10 krad(Si) and 100 krad(Si), with deeper networks being more susceptible to accuracy losses due to TID.

43 PARTICLE ACCELERATORS↗

Quantum-Based Molecular Dynamics Simulations Using Tensor Cores

Tensor cores, along with tensor processing units, represent a new form of hardware acceleration specifically designed for deep neural network calculations in artificial intelligence applications. Tensor cores provide extraordinary computational speed and energy efficiency but with the caveat that they were designed for tensor contractions (matrix–matrix multiplications) using only low-precision floating-point operations. Despite this perceived limitation, we demonstrate how tensor cores can be applied with high efficiency to the challenging and numerically sensitive problem of quantum-based Born–Oppenheimer molecular dynamics, which requires highly accurate electronic structure optimizations and conservative force evaluations. The interatomic forces are calculated on-the-fly from an electronic structure that is obtained from a generalized deep neural network, where the computational structure naturally takes advantage of the exceptional processing power of the tensor cores and allows for high performance in excess of 100 Tflops on a single Nvidia A100 GPU. Stable molecular dynamics trajectories are generated using the framework of extended Lagrangian Born–Oppenheimer molecular dynamics, which combines computational efficiency with long-term stability, even when using approximate charge relaxations and force evaluations that are limited in accuracy by the numerically noisy conditions caused by the low-precision tensor core floating-point operations. A canonical ensemble simulation scheme is also presented, where the additional numerical noise in the calculated forces is absorbed into a Langevin-like dynamics.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Abstraction hierarchy to define biofoundry workflows and operations for interoperable synthetic biology research and applications

Lack of standardization in biofoundries limits the scalability and efficiency of synthetic biology research. Here, we propose an abstraction hierarchy that organizes biofoundry activities into four interoperable levels: Project, Service/Capability, Workflow, and Unit Operation, effectively streamlining the Design‑Build‑Test‑Learn (DBTL) cycle. This framework enables more modular, flexible, and automated experimental workflows. It improves communication between researchers and systems, supports reproducibility, and facilitates better integration of software tools and artificial intelligence. Our approach lays the foundation for a globally interoperable biofoundry network, advancing collaborative synthetic biology and accelerating innovation in response to scientific and societal challenges.

Kim, Haseong↗

Roadmap on artificial intelligence and big data techniques for superconductivity

This paper presents a roadmap to the application of AI techniques and big data (BD) for different modelling, design, monitoring, manufacturing and operation purposes of different superconducting applications. To help superconductivity researchers, engineers, and manufacturers understand the viability of using AI and BD techniques as future solutions for challenges in superconductivity, a series of short articles are presented to outline some of the potential applications and solutions. These potential futuristic routes and their materials/technologies are considered for a 10–20 yr time-frame.

machine learning, neural network↗

Fine-tuning machine-learned particle-flow reconstruction for new detector geometries in future colliders

We demonstrate transfer learning capabilities in a machine-learned algorithm trained for particle-flow reconstruction in high energy particle colliders. This paper presents a cross-detector fine-tuning study, where we initially pretrain the model on a large full simulation dataset from one detector design, and subsequently fine-tune the model on a sample with a different collider and detector design. Specifically, we use the Compact Linear Collider detector (CLICdet) model for the initial training set and demonstrate successful knowledge transfer to the CLIC-like detector (CLD) proposed for the Future Circular Collider in electron-positron mode. We show that with an order of magnitude less samples from the second dataset, we can achieve the same performance as a costly training from scratch, across particle-level and event-level performance metrics, including jet and missing transverse momentum resolution. Furthermore, we find that the fine-tuned model achieves comparable performance to the traditional rule-based particle-flow approach on event-level metrics after training on 100,000 CLD events, whereas a model trained from scratch requires at least 1 million CLD events to achieve similar reconstruction performance. To our knowledge, this represents the first full-simulation cross-detector transfer learning study for particle-flow reconstruction. These findings offer valuable insights towards building large foundation models that can be fine-tuned across different detector designs and geometries, helping to accelerate the development cycle for new detectors and opening the door to rapid detector design and optimization using machine learning.

43 PARTICLE ACCELERATORS↗

Distributed Small-Signal Stability Conditions for Inverter-Based Unbalanced Microgrids

The proliferation of inverter-based generation and advanced sensing, controls, and communication infrastructure have facilitated accelerated deployment of microgrids. A coordinated network of microgrids can maintain reliable power delivery to critical facilities during extreme events. Low-inertia offered by the power-electronics–interfaced energy resources, however, can present significant challenges to ensuring stable operation of the microgrids. In this work, distributed small-signal stability conditions for inverter-based microgrids are developed that involve the droop-controller parameters and the network parameters (e.g. line impedances, loads). The distributed closed-form parametric stability conditions derived in this paper can be verified in a computationally efficient manner, facilitating reliable design and operations of networks of microgrids. Dynamic phasor models have been used to capture the effects of electromagnetic transients. Furthermore, numerical results are presented, along with PSCAD simulations, to validate the analytical stability conditions. Effects of design choices, such as the conductor types, and inverter sizes, on the small-signal stability of inverter-based microgrids are investigated to derive useful engineering insights.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Deeply learning deep inelastic scattering kinematics

We study the use of deep learning techniques to reconstruct the kinematics of the neutral current deep inelastic scattering (DIS) process in electron–proton collisions. In particular, we use simulated data from the ZEUS experiment at the HERA accelerator facility, and train deep neural networks to reconstruct the kinematic variables Q 2 and x. Our approach is based on the information used in the classical construction methods, the measurements of the scattered lepton, and the hadronic final state in the detector, but is enhanced through correlations and patterns revealed with the simulated data sets. We show that, with the appropriate selection of a training set, the neural networks sufficiently surpass all classical reconstruction methods on most of the kinematic range considered. Rapid access to large samples of simulated data and the ability of neural networks to effectively extract information from large data sets, both suggest that deep learning techniques to reconstruct DIS kinematics can serve as a rigorous method to combine and outperform the classical reconstruction methods.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

National Solar Jobs Accelerator (Final Technical Report (FTR))

The aptitudes and experiences gained through military service—such as dynamic leadership, teamwork and critical thinking skills, technical specialization, and a mission-completion work ethic—make veterans exceptional candidates for a wide range of solar energy careers. The solar industry offers a highly collaborative and purpose-driven work environment that resonates with service members and veterans looking to rise to their next challenge, and solar employers are eager to tap into this valuable talent pool. From October 2019 - February 2023, The National Solar Jobs Accelerator (publicly the Solar Ready Vets Network TM (SRV Network; SRVN)) enhanced and streamlined options for military service members and veterans to pursue solar training, certification, and employment, while advancing solar employers’ efforts and capacity to invest in military talent as part of a long-term workforce development strategy. The SRVN was led by the Interstate Renewable Energy Council (IREC) in partnership with the Solar Energy Industries Association (SEIA), the US Chamber of Commerce Foundation’s Hiring Our Heroes program (HOH) and the North American Board of Certified Energy Practitioners (NABCEP). Through several direct-impact and indirect, high-impact capacity building initiatives aligned with six key objectives, the SRV Network strengthened solar career pathways, and promoted increased representation of military talent across all levels and sectors of the solar workforce. A work-based learning Corporate Fellowship model connected transitioning service members with on-the-job experience in leadership roles with solar employers nationwide. The project advanced broader veteran recruitment and talent development by expanding GI Bill eligibility and streamlining veterans’ pathways for solar training and credentialing, supported direct connections to jobs with top solar employers, and led coordination among key education and industry partners to advance registered apprenticeships aligned with solar career pathways. To ensure that the project best served the needs of all stakeholders, an Advisory Committee of military-connected solar professionals, solar employers, and training providers met biannually to guide project activities and sustainability plans. The project team engaged the broader “SRV Network” (comprised of over 2,000 veterans, employers and training organizations) through regular newsletters and targeted outreach to share resources, hiring fairs, webinars, and other opportunities for engagement. The work done under this award builds on the previous iterations of the Department of Energy’s Solar Ready Vets ® program. As the solar industry continues to grow rapidly over the next decade, the military community will continue to be a highly valuable source of talent. The relationships established and work accomplished through this project will have an enduring positive impact well beyond the funding period.

14 SOLAR ENERGY↗