Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “network acceleration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

(NOICE) Neural Optical Image Categorizer for the E-log

The Fermilab Accelerator Division Electronic logbook (E-log) is a record of all the activities and events in the Division for the past 10 years and more. The E-log search function is a valuable resource and the institutional memory of the accelerator complex. About 300,000 files are stored in the E-log, of which the vast majority are images attached to entries and comments. The visual information contained in the images is not presently searchable. The goal of team NOICE (Neural Optical Image Categorizer for the E-log) was to design a neural network able to produce label categories for these images for use by searches. The group developed a dataset and trained a convolutional neural network (CNN) with optimized hyperparameter

43 PARTICLE ACCELERATORS↗

Leveraging prior mean models for faster Bayesian optimization of particle accelerators

Tuning particle accelerators is a challenging and time-consuming task that can be automated and carried out efficiently using suitable optimization algorithms, such as model-based Bayesian optimization techniques. One of the major advantages of Bayesian algorithms is the ability to incorporate prior information about beam physics and historical behavior into the model used to make control decisions. In this work, we examine incorporating prior accelerator physics information into Bayesian optimization algorithms by utilizing fast executing, neural network models trained on simulated or historical datasets as prior mean functions in Gaussian process models. We show that in ideal cases, this technique substantially increases convergence speed to optimal solutions in high-dimensional tuning parameter spaces. Additionally, we demonstrate that even in non-ideal cases, where prior models of beam dynamics do not exactly match experimental conditions, the use of this technique can still enhance convergence speed. Finally, we demonstrate how these methods can be used to improve optimization in practical applications, such as transferring information gained from beam dynamics simulations to online control of the LCLS injector, and transferring knowledge gained from experimental measurements across different operating modes, such as accelerating different ion species at the ATLAS heavy ion accelerator.

43 PARTICLE ACCELERATORS↗

Universal progression of structure and dynamics in colloidal nanocrystal gels during salt-accelerated aging

Controlling the structure and function of colloidal gels requires a detailed understanding of how the various components govern network formation and aging. In particular, molecular additives like salts are widely used to tune interparticle interactions, yet their influence on gelation pathways in complex systems such as colloidal nanocrystal gels remains inadequately understood. Here, we investigate how noncoordinating salts modulate the evolution of gels formed using chemically linked tin-doped indium oxide nanocrystals. Through combined structural, dynamic, and kinetic analyses, we demonstrate that increasing salt concentration accelerates gelation. When rescaled by salt-dependent characteristic times, the evolution collapses onto universal trajectories, revealing a time-salt superposition principle. The universality extends across length scales, suggesting a consistent salt-dependent mechanism that controls both local structuring and macroscopic network formation. This observed salt modulation of structure and dynamics provides a predictive basis for controlling the kinetics of nonequilibrium nanocrystal gel assembly, enhancing the rational design of functional nanomaterials with tunable properties.

36 MATERIALS SCIENCE↗

Network structure in alteration layer of boroaluminosilicate glass formed by aqueous corrosion

Exogenously-added LiCl has been shown to slightly accelerate the corrosion rate of a boroaluminosilicate glass called International Simple Glass (ISG) in aqueous solutions over forward- and residual-rate regimes, while KCl and CsCl impede. To understand the effect of exogenously added electrolytes on resulting hydrous species and the network structure of alteration layers, infrared spectroscopy was implemented. It was found that the fraction of molecular water relative to the surface-bound hydroxyl species is lower in the KCl and CsCl conditions compared to the LiCl and pure water conditions. An approximation for the spectral features of the thin surface films from an experimentally-obtained specular-reflectance infrared (SR-IR) spectrum was proposed; results indicate no significant difference in the Si-O bonding network of the alteration layers formed in the presence of exogenously added LiCl, KCl and CsCl. Furthermore, the observed change in corrosion rates might be linked to the relative abundance of molecular water species in the porous network, rather than the silicate bonding structure.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

LaserNetUS at the Extreme Light Laboratory. Final report

This is the final report for LaserNetUS, Grant # DE-SC0019419. This project covered the first two annual cycles (2018-2020) of LaserNetUS experiments conducted at the Extreme Light Laboratory, University of Nebraska-Lincoln. The project provided students and scientists from four institutions (BYU, Stanford, UNR, and ARFL) with access to a world-class high-intensity laser facility. Experimental results were obtained on the topics of Nonlinear Thomson scattering, Relativistic vacuum acceleration, and electron beams in relativistic high-energy-density plasma to study x-ray line emission and radio frequencies of ultrashort relativistic electron beam interactions. Another benefit was the training of 10 students (undergraduate or graduate) and 6 young scientists (postdoc or associate/research professors) in critical areas to the future development of high energy density science and high-power laser technology.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Modeling Analog Tile-Based Accelerators Using SST

Analog computing has been widely proposed to improve the energy efficiency of multiple important workloads including neural network operations, and other linear algebra kernels. To properly evaluate analog computing and explore more complex workloads such as systems consisting of multiple analog data paths, system level simulations are required. Moreover, prior work on system architectures for analog computing often rely on custom simulators creating signficant additional design effort and complicating comparisons between different systems. To remedy these issues, this report describes the design and implementation of a flexible tile-based analog accelerator element for the Structural Simulation Toolkit (SST). The element focuses on heavily on the tile controller—an often neglected aspect of prior work—that is sufficiently versatile to simulate a wide range of different tile operations including neural network layers, signal processing kernels, and generic linear algebra operations without major constraints. The tile model also interoperates with existing SST memory and network models to reduce the overall development load and enable future simulation of heterogeneous systems with both conventional digital logic and analog compute tiles. Finally, both the tile and array models are designed to easily support future extensions as new analog operations and applications that can benefit from analog computing are developed.

97 MATHEMATICS AND COMPUTING↗

Towards Generalizable and Efficient Circuit Topology Design: A Graph-Transformer-based Surrogate Model with Curriculum Learning

Unlike circuit parameter and sizing optimizations, the automated design of analog circuit topologies poses significant challenges for learning-based approaches. One challenge arises from the combinatorial growth of the topology space with circuit size, which limits the topology optimization efficiency. Moreover, traditional circuit evaluation methods are time-consuming, while the presence of data discontinuity in the topology space makes the accurate prediction of circuit performance exceptionally difficult for unseen topologies. To tackle these challenges, we design a novel Graph-Transformer-based Network (GTN) as the surrogate model for circuit evaluation, offering a substantial acceleration in the speed of circuit topology optimization without sacrificing performance. Our GTN model architecture is designed to embed voltage changes in circuit loops and current flows in connected devices, enabling accurate performance predictions for circuits with unseen topologies. To address the cold start problem when scaling GTN to large-scale circuits, we further introduce a curriculum learning strategy that progressively trains GTN from small-scale to large-scale circuits. This approach enables the model to first learn fundamental physical principles from simpler topologies and gradually adapt to complex configurations, effectively bridging the circuit complexity gap and improving prediction accuracy. Taking the power converter circuit design as an experimental task, our GTN model significantly outperforms an analytical approach and baseline methods directly utilizing graph neural networks. Furthermore, GTN achieves less than 5% relative error and 196× speed-up compared with high-fidelity simulation. Notably, our GTN surrogate model empowers an automatic circuit design framework to discover circuits of comparable quality to those identified through high-fidelity simulation while reducing the time required by up to 98.2%. With curriculum learning, the enhanced GTN achieves a 51% improvement for performance prediction of large-scale circuits compared to the GTN model without this strategy. These advancements establish GTN as a scalable framework for automated analog circuit design across varying circuit complexity levels.

Lu, Haoshu [New Jersey Institute of Technology (NJ↗

Structure and mechanism of Staphylococcus aureus oleate hydratase (OhyA)

Flavin adenine dinucleotide (FAD)-dependent bacterial oleate hydratases (OhyAs) catalyze the addition of water to isolated fatty acid carbon–carbon double bonds. Staphylococcus aureus uses OhyA to counteract the host innate immune response by inactivating antimicrobial unsaturated fatty acids. Mechanistic information explaining how OhyAs catalyze regiospecific and stereospecific hydration is required to understand their biological functions and the potential for engineering new products. In this study, we deduced the catalytic mechanism of OhyA from multiple structures of S. aureus OhyA in binary and ternary complexes with combinations of ligands along with biochemical analyses of relevant mutants. The substrate-free state shows Arg81 is the gatekeeper that controls fatty acid entrance to the active site. FAD binding engages the catalytic loop to simultaneously rotate Glu82 into its active conformation and Arg81 out of the hydrophobic substrate tunnel, allowing the fatty acid to rotate into the active site. FAD binding also dehydrates the active site, leaving a single water molecule connected to Glu82. This active site water is a hydronium ion based on the analysis of its hydrogen bond network in the OhyA•PEG400•FAD complex. We conclude that OhyA accelerates acid-catalyzed alkene hydration by positioning the fatty acid double bond to attack the active site hydronium ion, followed by the addition of water to the transient carbocation intermediate. Structural transitions within S. aureus OhyA channel oleate to the active site, curl oleate around the substrate water, and stabilize the hydroxylated product to inactivate antimicrobial fatty acids.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Experimental study of UV induced tensile properties deterioration and chemical aging of polyurea–POSS composites

Ultraviolet (UV) radiation present in natural sunlight degrades the chemical and mechanical properties of polymeric matrices in composites. Polyurea possess a unique set of chemical and mechanical properties due to its complex microstructure comprising of hard and soft segments (phases). The primary objective of the work presented here is to characterize the chemical and mechanical degradation of pure polyurea and polyurea - polyhedral oligomeric silsesquioxane (POSS) nanocomposites subjected to UV radiation exposure for 45 days. Control specimen in as-received condition (reference specimen) and UV-aged tensile test specimen are tested in the study. Analysis of Fourier transformed infrared (FTIR) spectra for UV-aged polyurea and polyurea-POSS composites reveals that the addition of POSS nanoparticles accelerate the deterioration of polyurea matrix by breaking the hard phase network and altering the mechanism of degradation for the soft phase present in polyurea microstructure. The deterioration of hard and soft segments cause a significant reduction in tensile properties of polyurea-POSS composites.

Materials Science↗

Scaling neural simulations in STACS

Abstract As modern neuroscience tools acquire more details about the brain, the need to move towards biological-scale neural simulations continues to grow. However, effective simulations at scale remain a challenge. Beyond just the tooling required to enable parallel execution, there is also the unique structure of the synaptic interconnectivity, which is globally sparse but has relatively high connection density and non-local interactions per neuron. There are also various practicalities to consider in high performance computing applications, such as the need for serializing neural networks to support potentially long-running simulations that require checkpoint-restart. Although acceleration on neuromorphic hardware is also a possibility, development in this space can be difficult as hardware support tends to vary between platforms and software support for larger scale models also tends to be limited. In this paper, we focus our attention on Simulation Tool for Asynchronous Cortical Streams (STACS), a spiking neural network simulator that leverages the Charm++ parallel programming framework, with the goal of supporting biological-scale simulations as well as interoperability between platforms. Central to these goals is the implementation of scalable data structures suitable for efficiently distributing a network across parallel partitions. Here, we discuss a straightforward extension of a parallel data format with a history of use in graph partitioners, which also serves as a portable intermediate representation for different neuromorphic backends. We perform scaling studies on the Summit supercomputer, examining the capabilities of STACS in terms of network build and storage, partitioning, and execution. We highlight how a suitably partitioned, spatially dependent synaptic structure introduces a communication workload well-suited to the multicast communication supported by Charm++. We evaluate the strong and weak scaling behavior for networks on the order of millions of neurons and billions of synapses, and show that STACS achieves competitive levels of parallel efficiency.

59 BASIC BIOLOGICAL SCIENCES↗

Community Choice Aggregation(CCA) Data Collection Webinar for Status and Trends in the Voluntary Market Report (2024 Data) [Slides]

We have subcontracted LEAN Energy US, to help us improve our CCA data collection effort for the Annual Voluntary Energy Markets Data Report. LEAN Energy US (Local Energy Aggregation Network) is a national 501(c)3 non-profit organization dedicated to accelerating the country's transition to clean and renewable power, supporting competition and customer choice in the energy sector, and maintaining affordable electricity rates. We work in partnership with a range of organizations to actively support the formation and operational success of Community Choice Aggregation (CCA) programs around the country. This webinar, hosted in partnership with LEAN Energy US, is intended to introduce their members to our data collection effort and encourage CCAs in their network to participate.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Open Science principles for accelerating trait-based science across the Tree of Life

Synthesizing trait observations and knowledge across the Tree of Life remains a grand challenge for biodiversity science. Species traits are widely used in ecological and evolutionary science, and new data and methods have proliferated rapidly. Yet accessing and integrating disparate data sources remains a considerable challenge, slowing progress toward a global synthesis to integrate trait data across organisms. Trait science needs a vision for achieving global integration across all organisms. In this perspective, we outline how the adoption of key Open Science principles—open data, open source and open methods—is transforming trait science, increasing transparency, democratizing access and accelerating global synthesis. To enhance widespread adoption of these principles, we introduce the Open Traits Network (OTN), a global, decentralized community welcoming all researchers and institutions pursuing the collaborative goal of standardizing and integrating trait data across organisms. We demonstrate how adherence to Open Science principles is key to the OTN community and outline five activities that can accelerate the synthesis of trait data across the Tree of Life, thereby facilitating rapid advances to address scientific inquiries and environmental issues. Lessons learned along the path to a global synthesis of trait data will provide a framework for addressing similarly complex data science and informatics challenges.

59 BASIC BIOLOGICAL SCIENCES↗

Even Higher-Level Synthesis: An Exploration of AI Hardware Accelerators using HLS4ML

With the rise of artificial intelligence, the popularization of deep learning, and a constantly evolving industry, the demand for flexible and efficient tools has never been greater. As algorithms grow more complex, their runtime and energy consumption increase exponentially. Customized hardware accelerators, long used for specific mathematical operations, remain essential for managing modern applications' computational and power demands. Hardware accelerators can speed up complex computations by orders of magnitude, but their manual design and verification processes are often challenging and time-consuming. High-Level Synthesis (HLS) provides a solution by transforming high-level algorithm descriptions, typically written in C++ or SystemC, into synthesizable RTL suitable for hardware implementation. This approach reduces development time for RTL engineers while offering flexibility beyond what traditional handwritten RTL can provide. We extended this capability to the machine-learning domain with the open-source framework hls4ml, which allows neural networks trained in Python frameworks like Tensorflow or PyTorch to be synthesized into efficient hardware representations for the traditional FPGA and ASIC flows. This breakthrough addresses the growing need for reduced design turnaround and easy verification of ML hardware accelerators with low latency and power efficiency constraints. During this tutorial, we will demonstrate how Python complements HLS by simplifying the ML design process, bridging the gap between software and hardware development. Attendees will explore how we translate neural networks modeled in Python into fixed-point C++ models suitable for HLS workflows. We will dive into strategies like Value-Range Analysis and Quantization-Aware Training, which optimize these designs for deployment and evaluate their accuracy, power consumption, and energy efficiency. To exemplify these concepts, experts from Fermilab will share their experiences applying this technology to high-energy physics experiments, where real-time, low-latency processing is critical. Over the years, Fermilab engineers have demonstrated how deep neural networks, optimized for hardware using hls4ml, can meet the stringent requirements of trigger systems at the CERN Large Hadron Collider. These systems rely on rapid decision-making to process immense data volumes while retaining only the most relevant events for further analysis. The application of hls4ml has also been extended to innovative technologies like smart pixel arrays. These smart pixels integrate ML inference capabilities directly into sensor devices, enabling localized data processing at the pixel level. This approach drastically reduces the need to transmit raw data to external processing units, significantly decreasing power consumption and latency. By embedding neural networks within the pixel architecture, the smart pixels can identify and prioritize relevant data in real time, providing a highly efficient solution for edge computing in scenarios such as particle detectors and imaging systems. Fermilab's work highlights the potential of hardware-accelerated ML in scenarios where both speed and power efficiency are mission-critical. Through this tutorial, attendees will gain valuable insights into the challenges and solutions of deploying ML in hardware. Understanding how HLS and hls4ml streamline the development of neural network-based hardware accelerators is fundamental for the industry's future. Participants will learn how these technologies are shaping the future of AI and scientific computing.

Di Guglielmo, Giuseppe [Fermilab]↗

Keeping LAMMPS cutting edge

Since its inception 30 years ago, LAMMPS has grown to be a world-class molecular dynamics code and a cornerstone of computational materials science research. This project aimed to keep LAMMPS at the forefront of molecular dynamics simulations by adapting LAMMPS to the latest developments in machine learning technology and hardware. Initially, the project set out to provide a unified implementation of active learning for efficient training data generation in LAMMPS, but the research trajectory pivoted to address more immediate and impactful opportunities. On the hardware side, recent record-breaking molecular dynamics simulations were developed on the Cerebras wafer-scale AI chip, and this project has developed an interface between LAMMPS and the hardware-specific molecular dynamics code to accelerate and simplify development and user adoption. On the software side, PyTorch’s Ahead-of-Time (AOT) compilation features promised increased performance for state-of-the-art equivariant neural network potentials, and this project laid the groundwork for their adoption in LAMMPS, resulting in a nearly 20x acceleration in extreme cases. Combined with a comprehensive benchmark study of LAMMPS across all current exascale systems, this project has reinforced LAMMPS’s role as a versatile, high-performance tool for current and future materials science applications.

36 MATERIALS SCIENCE↗

Southeast Regional CO 2 Utilization and Storage Acceleration Partnership (SECARB-USA): Initial Inventory of Non-Technical Challenges to CCUS Deployment

The “Southeast Regional CO 2 Utilization and Storage Acceleration Partnership” (SECARBUSA) project supports the U.S. Department of Energy (DOE) Office of Fossil Energy's (FE) mission to help the United States meet its need for secure, affordable, and environmentally sound fossil energy supplies by utilizing the advancements made by the current Regional Carbon Sequestration Partnership (RCSP) Initiative to continue to identify and address knowledge gaps. The primary project objective is to identify and address regional onshore storage and transport challenges facing commercial deployment of carbon dioxide (CO 2 ) capture, utilization, and storage (CCUS) technologies. The Research Partners and a selected industry network of experienced CCUS project developers and operators will coordinate their capabilities to accelerate CCUS deployment by achieving four primary research objectives: 1) address key technical challenges; 2) facilitate data collection, sharing and analysis; 3) assess transportation and distribution infrastructure; and 4) promote regional technology transfer and dissemination of knowledge. The SECARB-USA Region includes the states of Alabama, Arkansas, Florida, Georgia, Louisiana, Mississippi, North Carolina, South Carolina, Tennessee, and Virginia and portions of Kentucky, Missouri, Oklahoma, Texas, and West Virginia. Under subtask 5.2: Non-Technical Challenges to CCUS Deployment, the Southern States Energy Board (SSEB) will define and identify An Inventory of Non-Technical Challenges to CCUS Deployment. As an initial step, SSEB organized an Industry and Non-Governmental Organization (NGO) Working Group comprised of knowledgeable market participants to assist in the development of an initial list of non-technical challenges to CCUS development.

42 ENGINEERING↗

Integration of Ag-CBRAM crossbars and Mott ReLU neurons for efficient implementation of deep neural networks in hardware

In-memory computing with emerging non-volatile memory devices (eNVMs) has shown promising results in accelerating matrix-vector multiplications. However, activation function calculations are still being implemented with general processors or large and complex neuron peripheral circuits. Here, we present the integration of Ag-based conductive bridge random access memory (Ag-CBRAM) crossbar arrays with Mott rectified linear unit (ReLU) activation neurons for scalable, energy and area-efficient hardware (HW) implementation of deep neural networks. We develop Ag-CBRAM devices that can achieve a high ON/OFF ratio and multi-level programmability. Compact and energy-efficient Mott ReLU neuron devices implementing ReLU activation function are directly connected to the columns of Ag-CBRAM crossbars to compute the output from the weighted sum current. We implement convolution filters and activations for VGG-16 using our integrated HW and demonstrate the successful generation of feature maps for CIFAR-10 images in HW. Our approach paves a new way toward building a highly compact and energy-efficient eNVMs-based in-memory computing system.

Mott insulators↗

Accelerated Data Analytics for Power-System Time-Series (ADAPT) [Final Project Report]

Over the past decade, electric utilities have made significant progress in deploying networks of phasor measurement units (PMUs). The high-speed measurements provided by PMUs are valuable for offline analysis, but the sheer volume of data presents a challenge for utilities. The Accelerated Data Analytics for Power-System Time-Series (ADAPT) project was created to address this challenge by developing SciSync, an open-source software tool that enables utilities to rapidly read, process, analyze, and review grid measurements. SciSync was designed to provide the capabilities of Archive Walker, a research tool developed by PNNL, along with the reliability and performance of commercial-grade software.

24 POWER TRANSMISSION AND DISTRIBUTION↗

INTEGRATE - Inverse Network Transformations for Efficient Generation of Robust Airfoil and Turbine Enhancements

The INTEGRATE (Inverse Network Transformations for Efficient Generation of Robust Airfoil and Turbine Enhancements) project is developing a new inverse-design capability for the aerodynamic design of wind turbine rotors using invertible neural networks. This AI-based design technology can capture complex non-linear aerodynamic effects while being 100 times faster than design approaches based on computational fluid dynamics. This project enables innovation in wind turbine design by accelerating time to market through higher-accuracy early design iterations to reduce the levelized cost of energy. INVERTIBLE NEURAL NETWORKS Researchers are leveraging a specialized invertible neural network (INN) architecture along with the novel dimension-reduction methods and airfoil/blade shape representations developed by collaborators at the National Institute of Standards and Technology (NIST) learns complex relationships between airfoil or blade shapes and their associated aerodynamic and structural properties. This INN architecture will accelerate designs by providing a cost-effective alternative to current industrial aerodynamic design processes, including: - Blade element momentum (BEM) theory models: limited effectiveness for design of offshore rotors with large, flexible blades where nonlinear aerodynamic effects dominate - Direct design using computational fluid dynamics (CFD): cost-prohibitive - Inverse-design models based on deep neural networks (DNNs): attractive alternative to CFD for 2D design problems, but quickly overwhelmed by the increased number of design variables in 3D problems AUTOMATED COMPUTATIONAL FLUID DYNAMICS FOR TRAINING DATA GENERATION - MERCURY FRAMEWORK The INN is trained on data obtained using the University of Marylands (UMD) Mercury Framework, which has with robust automated mesh generation capabilities and advanced turbulence and transition models validated for wind energy applications. Mercury is a multi-mesh paradigm, heterogeneous CPU-GPU framework. The framework incorporates three flow solvers at UMD, 1) OverTURNS, a structured solver on CPUs, 2) HAMSTR, a line based unstructured solver on CPUs, and 3) GARFIELD, a structured solver on GPUs. The framework is based on Python, that is often used to wrap C or Fortran codes for interoperability with other solvers. Communication between multiple solvers is accomplished with a Topology Independent Overset Grid Assembler (TIOGA). NOVEL AIRFOIL SHAPE REPRESENTATIONS USING GRASSMAN SPACES We developed a novel representation of shapes which decouples affine-style deformations from a rich set of data-driven deformations over a submanifold of the Grassmannian. The Grassmannian representation as an analytic generative model, informed by a database of physically relevant airfoils, offers (i) a rich set of novel 2D airfoil deformations not previously captured in the data , (ii) improved low-dimensional parameter domain for inferential statistics informing design/manufacturing, and (iii) consistent 3D blade representation and perturbation over a sequence of nominal shapes. TECHNOLOGY TRANSFER DEMONSTRATION - COUPLING WITH NREL WISDEM Researchers have integrated the inverse-design tool for 2D airfoils (INN-Airfoil) into WISDEM (Wind Plant Integrated Systems Design and Engineering Model), a multidisciplinary design and optimization framework for assessing the cost of energy, as part of tech-transfer demonstration. The integration of INN-Airfoil into WISDEM allows for the design of airfoils along with the blades that meet the dynamic design constraints on cost of energy, annual energy production, and the capital costs. Through preliminary studies, researchers have shown that the coupled INN-Airfoil + WISDEM approach reduces the cost of energy by around 1% compared to the conventional design approach. This page will serve as a place to easily access all the publications from this work and the repositories for the software developed and released through this pr...

aerodynamics↗