Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “network acceleration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Leveraging prior mean models for faster Bayesian optimization of particle accelerators

Tuning particle accelerators is a challenging and time-consuming task that can be automated and carried out efficiently using suitable optimization algorithms, such as model-based Bayesian optimization techniques. One of the major advantages of Bayesian algorithms is the ability to incorporate prior information about beam physics and historical behavior into the model used to make control decisions. In this work, we examine incorporating prior accelerator physics information into Bayesian optimization algorithms by utilizing fast executing, neural network models trained on simulated or historical datasets as prior mean functions in Gaussian process models. We show that in ideal cases, this technique substantially increases convergence speed to optimal solutions in high-dimensional tuning parameter spaces. Additionally, we demonstrate that even in non-ideal cases, where prior models of beam dynamics do not exactly match experimental conditions, the use of this technique can still enhance convergence speed. Finally, we demonstrate how these methods can be used to improve optimization in practical applications, such as transferring information gained from beam dynamics simulations to online control of the LCLS injector, and transferring knowledge gained from experimental measurements across different operating modes, such as accelerating different ion species at the ATLAS heavy ion accelerator.

43 PARTICLE ACCELERATORS↗

Universal progression of structure and dynamics in colloidal nanocrystal gels during salt-accelerated aging

Controlling the structure and function of colloidal gels requires a detailed understanding of how the various components govern network formation and aging. In particular, molecular additives like salts are widely used to tune interparticle interactions, yet their influence on gelation pathways in complex systems such as colloidal nanocrystal gels remains inadequately understood. Here, we investigate how noncoordinating salts modulate the evolution of gels formed using chemically linked tin-doped indium oxide nanocrystals. Through combined structural, dynamic, and kinetic analyses, we demonstrate that increasing salt concentration accelerates gelation. When rescaled by salt-dependent characteristic times, the evolution collapses onto universal trajectories, revealing a time-salt superposition principle. The universality extends across length scales, suggesting a consistent salt-dependent mechanism that controls both local structuring and macroscopic network formation. This observed salt modulation of structure and dynamics provides a predictive basis for controlling the kinetics of nonequilibrium nanocrystal gel assembly, enhancing the rational design of functional nanomaterials with tunable properties.

36 MATERIALS SCIENCE↗

Network structure in alteration layer of boroaluminosilicate glass formed by aqueous corrosion

Exogenously-added LiCl has been shown to slightly accelerate the corrosion rate of a boroaluminosilicate glass called International Simple Glass (ISG) in aqueous solutions over forward- and residual-rate regimes, while KCl and CsCl impede. To understand the effect of exogenously added electrolytes on resulting hydrous species and the network structure of alteration layers, infrared spectroscopy was implemented. It was found that the fraction of molecular water relative to the surface-bound hydroxyl species is lower in the KCl and CsCl conditions compared to the LiCl and pure water conditions. An approximation for the spectral features of the thin surface films from an experimentally-obtained specular-reflectance infrared (SR-IR) spectrum was proposed; results indicate no significant difference in the Si-O bonding network of the alteration layers formed in the presence of exogenously added LiCl, KCl and CsCl. Furthermore, the observed change in corrosion rates might be linked to the relative abundance of molecular water species in the porous network, rather than the silicate bonding structure.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

LaserNetUS at the Extreme Light Laboratory. Final report

This is the final report for LaserNetUS, Grant # DE-SC0019419. This project covered the first two annual cycles (2018-2020) of LaserNetUS experiments conducted at the Extreme Light Laboratory, University of Nebraska-Lincoln. The project provided students and scientists from four institutions (BYU, Stanford, UNR, and ARFL) with access to a world-class high-intensity laser facility. Experimental results were obtained on the topics of Nonlinear Thomson scattering, Relativistic vacuum acceleration, and electron beams in relativistic high-energy-density plasma to study x-ray line emission and radio frequencies of ultrashort relativistic electron beam interactions. Another benefit was the training of 10 students (undergraduate or graduate) and 6 young scientists (postdoc or associate/research professors) in critical areas to the future development of high energy density science and high-power laser technology.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Modeling Analog Tile-Based Accelerators Using SST

Analog computing has been widely proposed to improve the energy efficiency of multiple important workloads including neural network operations, and other linear algebra kernels. To properly evaluate analog computing and explore more complex workloads such as systems consisting of multiple analog data paths, system level simulations are required. Moreover, prior work on system architectures for analog computing often rely on custom simulators creating signficant additional design effort and complicating comparisons between different systems. To remedy these issues, this report describes the design and implementation of a flexible tile-based analog accelerator element for the Structural Simulation Toolkit (SST). The element focuses on heavily on the tile controller—an often neglected aspect of prior work—that is sufficiently versatile to simulate a wide range of different tile operations including neural network layers, signal processing kernels, and generic linear algebra operations without major constraints. The tile model also interoperates with existing SST memory and network models to reduce the overall development load and enable future simulation of heterogeneous systems with both conventional digital logic and analog compute tiles. Finally, both the tile and array models are designed to easily support future extensions as new analog operations and applications that can benefit from analog computing are developed.

97 MATHEMATICS AND COMPUTING↗

Towards Generalizable and Efficient Circuit Topology Design: A Graph-Transformer-based Surrogate Model with Curriculum Learning

Unlike circuit parameter and sizing optimizations, the automated design of analog circuit topologies poses significant challenges for learning-based approaches. One challenge arises from the combinatorial growth of the topology space with circuit size, which limits the topology optimization efficiency. Moreover, traditional circuit evaluation methods are time-consuming, while the presence of data discontinuity in the topology space makes the accurate prediction of circuit performance exceptionally difficult for unseen topologies. To tackle these challenges, we design a novel Graph-Transformer-based Network (GTN) as the surrogate model for circuit evaluation, offering a substantial acceleration in the speed of circuit topology optimization without sacrificing performance. Our GTN model architecture is designed to embed voltage changes in circuit loops and current flows in connected devices, enabling accurate performance predictions for circuits with unseen topologies. To address the cold start problem when scaling GTN to large-scale circuits, we further introduce a curriculum learning strategy that progressively trains GTN from small-scale to large-scale circuits. This approach enables the model to first learn fundamental physical principles from simpler topologies and gradually adapt to complex configurations, effectively bridging the circuit complexity gap and improving prediction accuracy. Taking the power converter circuit design as an experimental task, our GTN model significantly outperforms an analytical approach and baseline methods directly utilizing graph neural networks. Furthermore, GTN achieves less than 5% relative error and 196× speed-up compared with high-fidelity simulation. Notably, our GTN surrogate model empowers an automatic circuit design framework to discover circuits of comparable quality to those identified through high-fidelity simulation while reducing the time required by up to 98.2%. With curriculum learning, the enhanced GTN achieves a 51% improvement for performance prediction of large-scale circuits compared to the GTN model without this strategy. These advancements establish GTN as a scalable framework for automated analog circuit design across varying circuit complexity levels.

Lu, Haoshu [New Jersey Institute of Technology (NJ↗

Structure and mechanism of Staphylococcus aureus oleate hydratase (OhyA)

Flavin adenine dinucleotide (FAD)-dependent bacterial oleate hydratases (OhyAs) catalyze the addition of water to isolated fatty acid carbon–carbon double bonds. Staphylococcus aureus uses OhyA to counteract the host innate immune response by inactivating antimicrobial unsaturated fatty acids. Mechanistic information explaining how OhyAs catalyze regiospecific and stereospecific hydration is required to understand their biological functions and the potential for engineering new products. In this study, we deduced the catalytic mechanism of OhyA from multiple structures of S. aureus OhyA in binary and ternary complexes with combinations of ligands along with biochemical analyses of relevant mutants. The substrate-free state shows Arg81 is the gatekeeper that controls fatty acid entrance to the active site. FAD binding engages the catalytic loop to simultaneously rotate Glu82 into its active conformation and Arg81 out of the hydrophobic substrate tunnel, allowing the fatty acid to rotate into the active site. FAD binding also dehydrates the active site, leaving a single water molecule connected to Glu82. This active site water is a hydronium ion based on the analysis of its hydrogen bond network in the OhyA•PEG400•FAD complex. We conclude that OhyA accelerates acid-catalyzed alkene hydration by positioning the fatty acid double bond to attack the active site hydronium ion, followed by the addition of water to the transient carbocation intermediate. Structural transitions within S. aureus OhyA channel oleate to the active site, curl oleate around the substrate water, and stabilize the hydroxylated product to inactivate antimicrobial fatty acids.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Experimental study of UV induced tensile properties deterioration and chemical aging of polyurea–POSS composites

Ultraviolet (UV) radiation present in natural sunlight degrades the chemical and mechanical properties of polymeric matrices in composites. Polyurea possess a unique set of chemical and mechanical properties due to its complex microstructure comprising of hard and soft segments (phases). The primary objective of the work presented here is to characterize the chemical and mechanical degradation of pure polyurea and polyurea - polyhedral oligomeric silsesquioxane (POSS) nanocomposites subjected to UV radiation exposure for 45 days. Control specimen in as-received condition (reference specimen) and UV-aged tensile test specimen are tested in the study. Analysis of Fourier transformed infrared (FTIR) spectra for UV-aged polyurea and polyurea-POSS composites reveals that the addition of POSS nanoparticles accelerate the deterioration of polyurea matrix by breaking the hard phase network and altering the mechanism of degradation for the soft phase present in polyurea microstructure. The deterioration of hard and soft segments cause a significant reduction in tensile properties of polyurea-POSS composites.

Materials Science↗

Magnetic Reconnection as the Driver of the Solar Wind

We present EUV solar observations showing evidence for omnipresent jetting activity driven by small-scale magnetic reconnection at the base of the solar corona. We argue that the physical mechanism that heats and drives the solar wind at its source is ubiquitous magnetic reconnection in the form of small-scale jetting activity (i.e., a.k.a. jetlets). This jetting activity, like the solar wind and the heating of the coronal plasma, are ubiquitous regardless of the solar cycle phase. Each event arises from small-scale reconnection of opposite polarity magnetic fields producing a short-lived jet of hot plasma and Alfv´en waves into the corona. The discrete nature of these jetlet events leads to intermittent outflows from the corona, which homogenize as they propagate away from the Sun and form the solar wind. This discovery establishes the importance of small-scale magnetic reconnection in solar and stellar atmospheres in understanding ubiquitous phenomena such as coronal heating and solar wind acceleration. Based on previous analyses linking the switchbacks to the magnetic network, we also argue that these new observations might provide the link between the magnetic activity at the base of the corona and the switchback solar wind phenomenon. These new observations need to be put in the bigger picture of the role of magnetic reconnection and the diverse form of jetting in the solar atmosphere.

magnetic reconnection↗

Scaling neural simulations in STACS

Abstract As modern neuroscience tools acquire more details about the brain, the need to move towards biological-scale neural simulations continues to grow. However, effective simulations at scale remain a challenge. Beyond just the tooling required to enable parallel execution, there is also the unique structure of the synaptic interconnectivity, which is globally sparse but has relatively high connection density and non-local interactions per neuron. There are also various practicalities to consider in high performance computing applications, such as the need for serializing neural networks to support potentially long-running simulations that require checkpoint-restart. Although acceleration on neuromorphic hardware is also a possibility, development in this space can be difficult as hardware support tends to vary between platforms and software support for larger scale models also tends to be limited. In this paper, we focus our attention on Simulation Tool for Asynchronous Cortical Streams (STACS), a spiking neural network simulator that leverages the Charm++ parallel programming framework, with the goal of supporting biological-scale simulations as well as interoperability between platforms. Central to these goals is the implementation of scalable data structures suitable for efficiently distributing a network across parallel partitions. Here, we discuss a straightforward extension of a parallel data format with a history of use in graph partitioners, which also serves as a portable intermediate representation for different neuromorphic backends. We perform scaling studies on the Summit supercomputer, examining the capabilities of STACS in terms of network build and storage, partitioning, and execution. We highlight how a suitably partitioned, spatially dependent synaptic structure introduces a communication workload well-suited to the multicast communication supported by Charm++. We evaluate the strong and weak scaling behavior for networks on the order of millions of neurons and billions of synapses, and show that STACS achieves competitive levels of parallel efficiency.

59 BASIC BIOLOGICAL SCIENCES↗

Community Choice Aggregation(CCA) Data Collection Webinar for Status and Trends in the Voluntary Market Report (2024 Data) [Slides]

We have subcontracted LEAN Energy US, to help us improve our CCA data collection effort for the Annual Voluntary Energy Markets Data Report. LEAN Energy US (Local Energy Aggregation Network) is a national 501(c)3 non-profit organization dedicated to accelerating the country's transition to clean and renewable power, supporting competition and customer choice in the energy sector, and maintaining affordable electricity rates. We work in partnership with a range of organizations to actively support the formation and operational success of Community Choice Aggregation (CCA) programs around the country. This webinar, hosted in partnership with LEAN Energy US, is intended to introduce their members to our data collection effort and encourage CCAs in their network to participate.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Open Science principles for accelerating trait-based science across the Tree of Life

Synthesizing trait observations and knowledge across the Tree of Life remains a grand challenge for biodiversity science. Species traits are widely used in ecological and evolutionary science, and new data and methods have proliferated rapidly. Yet accessing and integrating disparate data sources remains a considerable challenge, slowing progress toward a global synthesis to integrate trait data across organisms. Trait science needs a vision for achieving global integration across all organisms. In this perspective, we outline how the adoption of key Open Science principles—open data, open source and open methods—is transforming trait science, increasing transparency, democratizing access and accelerating global synthesis. To enhance widespread adoption of these principles, we introduce the Open Traits Network (OTN), a global, decentralized community welcoming all researchers and institutions pursuing the collaborative goal of standardizing and integrating trait data across organisms. We demonstrate how adherence to Open Science principles is key to the OTN community and outline five activities that can accelerate the synthesis of trait data across the Tree of Life, thereby facilitating rapid advances to address scientific inquiries and environmental issues. Lessons learned along the path to a global synthesis of trait data will provide a framework for addressing similarly complex data science and informatics challenges.

59 BASIC BIOLOGICAL SCIENCES↗

Large-Scale NASA Science Applications on the Columbia Supercluster

Columbia, NASA's newest 61 teraflops supercomputer that became operational late last year, is a highly integrated Altix cluster of 10,240 processors, and was named to honor the crew of the Space Shuttle lost in early 2003. Constructed in just four months, Columbia increased NASA's computing capability ten-fold, and revitalized the Agency's high-end computing efforts. Significant cutting-edge science and engineering simulations in the areas of space and Earth sciences, as well as aeronautics and space operations, are already occurring on this largest operational Linux supercomputer, demonstrating its capacity and capability to accelerate NASA's space exploration vision. The presentation will describe how an integrated environment consisting not only of next-generation systems, but also modeling and simulation, high-speed networking, parallel performance optimization, and advanced data analysis and visualization, is being used to reduce design cycle time, accelerate scientific discovery, conduct parametric analysis of multiple scenarios, and enhance safety during the life cycle of NASA missions. The talk will conclude by discussing how NAS partnered with various NASA centers, other government agencies, computer industry, and academia, to create a national resource in large-scale modeling and simulation.

Brooks, Walter↗

Even Higher-Level Synthesis: An Exploration of AI Hardware Accelerators using HLS4ML

With the rise of artificial intelligence, the popularization of deep learning, and a constantly evolving industry, the demand for flexible and efficient tools has never been greater. As algorithms grow more complex, their runtime and energy consumption increase exponentially. Customized hardware accelerators, long used for specific mathematical operations, remain essential for managing modern applications' computational and power demands. Hardware accelerators can speed up complex computations by orders of magnitude, but their manual design and verification processes are often challenging and time-consuming. High-Level Synthesis (HLS) provides a solution by transforming high-level algorithm descriptions, typically written in C++ or SystemC, into synthesizable RTL suitable for hardware implementation. This approach reduces development time for RTL engineers while offering flexibility beyond what traditional handwritten RTL can provide. We extended this capability to the machine-learning domain with the open-source framework hls4ml, which allows neural networks trained in Python frameworks like Tensorflow or PyTorch to be synthesized into efficient hardware representations for the traditional FPGA and ASIC flows. This breakthrough addresses the growing need for reduced design turnaround and easy verification of ML hardware accelerators with low latency and power efficiency constraints. During this tutorial, we will demonstrate how Python complements HLS by simplifying the ML design process, bridging the gap between software and hardware development. Attendees will explore how we translate neural networks modeled in Python into fixed-point C++ models suitable for HLS workflows. We will dive into strategies like Value-Range Analysis and Quantization-Aware Training, which optimize these designs for deployment and evaluate their accuracy, power consumption, and energy efficiency. To exemplify these concepts, experts from Fermilab will share their experiences applying this technology to high-energy physics experiments, where real-time, low-latency processing is critical. Over the years, Fermilab engineers have demonstrated how deep neural networks, optimized for hardware using hls4ml, can meet the stringent requirements of trigger systems at the CERN Large Hadron Collider. These systems rely on rapid decision-making to process immense data volumes while retaining only the most relevant events for further analysis. The application of hls4ml has also been extended to innovative technologies like smart pixel arrays. These smart pixels integrate ML inference capabilities directly into sensor devices, enabling localized data processing at the pixel level. This approach drastically reduces the need to transmit raw data to external processing units, significantly decreasing power consumption and latency. By embedding neural networks within the pixel architecture, the smart pixels can identify and prioritize relevant data in real time, providing a highly efficient solution for edge computing in scenarios such as particle detectors and imaging systems. Fermilab's work highlights the potential of hardware-accelerated ML in scenarios where both speed and power efficiency are mission-critical. Through this tutorial, attendees will gain valuable insights into the challenges and solutions of deploying ML in hardware. Understanding how HLS and hls4ml streamline the development of neural network-based hardware accelerators is fundamental for the industry's future. Participants will learn how these technologies are shaping the future of AI and scientific computing.

Di Guglielmo, Giuseppe [Fermilab]↗

ION Configuration Editor

The configuration of ION (Inter - planetary Overlay Network) network nodes is a manual task that is complex, time-consuming, and error-prone. This program seeks to accelerate this job and produce reliable configurations. The ION Configuration Editor is a model-based smart editor based on Eclipse Modeling Framework technology. An ION network designer uses this Eclipse-based GUI to construct a data model of the complete target network and then generate configurations. The data model is captured in an XML file. Intrinsic editor features aid in achieving model correctness, such as field fill-in, type-checking, lists of valid values, and suitable default values. Additionally, an explicit "validation" feature executes custom rules to catch more subtle model errors. A "survey" feature provides a set of reports providing an overview of the entire network, enabling a quick assessment of the model s completeness and correctness. The "configuration" feature produces the main final result, a complete set of ION configuration files (eight distinct file types) for each ION node in the network.

Borgen, Richard L.↗

Keeping LAMMPS cutting edge

Since its inception 30 years ago, LAMMPS has grown to be a world-class molecular dynamics code and a cornerstone of computational materials science research. This project aimed to keep LAMMPS at the forefront of molecular dynamics simulations by adapting LAMMPS to the latest developments in machine learning technology and hardware. Initially, the project set out to provide a unified implementation of active learning for efficient training data generation in LAMMPS, but the research trajectory pivoted to address more immediate and impactful opportunities. On the hardware side, recent record-breaking molecular dynamics simulations were developed on the Cerebras wafer-scale AI chip, and this project has developed an interface between LAMMPS and the hardware-specific molecular dynamics code to accelerate and simplify development and user adoption. On the software side, PyTorch’s Ahead-of-Time (AOT) compilation features promised increased performance for state-of-the-art equivariant neural network potentials, and this project laid the groundwork for their adoption in LAMMPS, resulting in a nearly 20x acceleration in extreme cases. Combined with a comprehensive benchmark study of LAMMPS across all current exascale systems, this project has reinforced LAMMPS’s role as a versatile, high-performance tool for current and future materials science applications.

36 MATERIALS SCIENCE↗

Southeast Regional CO 2 Utilization and Storage Acceleration Partnership (SECARB-USA): Initial Inventory of Non-Technical Challenges to CCUS Deployment

The “Southeast Regional CO 2 Utilization and Storage Acceleration Partnership” (SECARBUSA) project supports the U.S. Department of Energy (DOE) Office of Fossil Energy's (FE) mission to help the United States meet its need for secure, affordable, and environmentally sound fossil energy supplies by utilizing the advancements made by the current Regional Carbon Sequestration Partnership (RCSP) Initiative to continue to identify and address knowledge gaps. The primary project objective is to identify and address regional onshore storage and transport challenges facing commercial deployment of carbon dioxide (CO 2 ) capture, utilization, and storage (CCUS) technologies. The Research Partners and a selected industry network of experienced CCUS project developers and operators will coordinate their capabilities to accelerate CCUS deployment by achieving four primary research objectives: 1) address key technical challenges; 2) facilitate data collection, sharing and analysis; 3) assess transportation and distribution infrastructure; and 4) promote regional technology transfer and dissemination of knowledge. The SECARB-USA Region includes the states of Alabama, Arkansas, Florida, Georgia, Louisiana, Mississippi, North Carolina, South Carolina, Tennessee, and Virginia and portions of Kentucky, Missouri, Oklahoma, Texas, and West Virginia. Under subtask 5.2: Non-Technical Challenges to CCUS Deployment, the Southern States Energy Board (SSEB) will define and identify An Inventory of Non-Technical Challenges to CCUS Deployment. As an initial step, SSEB organized an Industry and Non-Governmental Organization (NGO) Working Group comprised of knowledgeable market participants to assist in the development of an initial list of non-technical challenges to CCUS development.

42 ENGINEERING↗