Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “network acceleration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Software Tools Ecosystem Project (STEP) Midyear Report CY2025

This document provides a technical project report for the first six months of 2025 for the Software Tools Ecosystem Project (STEP). The mission of STEP is to enable critical software tools to proactively adapt to emerging platform technologies (such as new accelerators, storage devices, network technologies, and smart devices) and emerging application use cases (such as advanced machine learning and workflow frameworks) so that they continue to meet the needs of scientific computing and provide a strong foundation for future Advanced Scientific Computing Research activities. Our challenges include the wide breadth of our stakeholders and rapidly evolving platform technology dependencies.

97 MATHEMATICS AND COMPUTING↗

Software Tools Ecosystem Project (STEP): CY2025 Annual Report

This document provides a technical project report for the Software Tools Ecosystem Project (STEP) during calendar year 2025. The mission of STEP is to enable critical software tools to proactively adapt to emerging platform technologies (such as new accelerators, storage devices, network technologies, and smart devices) and emerging application use cases (such as advanced machine learning and workflow frameworks) so that they continue to meet the needs of scientific computing and provide a strong foundation for future Advanced Scientific Computing Research activities.

97 MATHEMATICS AND COMPUTING↗

Addressing GPU memory limitations for Graph Neural Networks in High-Energy Physics applications

Introduction Reconstructing low-level particle tracks in neutrino physics can address some of the most fundamental questions about the universe. However, processing petabytes of raw data using deep learning techniques poses a challenging problem in the field of High Energy Physics (HEP). In the Exa.TrkX Project, an illustrative HEP application, preprocessed simulation data is fed into a state-of-art Graph Neural Network (GNN) model, accelerated by GPUs. However, limited GPU memory often leads to Out-of-Memory (OOM) exceptions during training, due to the large size of models and datasets. This problem is exacerbated when deploying models on High-Performance Computing (HPC) systems designed for large-scale applications. Methods We observe a high workload imbalance issue during GNN model training caused by the irregular sizes of input graph samples in HEP datasets, contributing to OOM exceptions. We aim to scale GNNs on HPC systems, by prioritizing workload balance in graph inputs while maintaining model accuracy. Our paper introduces diverse balancing strategies aimed at decreasing the maximum GPU memory footprint and avoiding the OOM exception, across various datasets. Results Our experiments showcase memory reduction of up to 32.14% compared to the baseline. We also demonstrate the proposed strategies can avoid OOM in application. Additionally, we create a distributed multi-GPU implementation using these samplers to demonstrate the scalability of these techniques on the HEP dataset. Discussion By assessing the performance of these strategies as data loading samplers across multiple datasets, we can gauge their effectiveness in both single-GPU and distributed environments. Our experiments, conducted on datasets of varying sizes and across multiple GPUs, broaden the applicability of our work to various GNN applications that handle input datasets with irregular graph sizes.

Lee, Claire Songhyun↗

Accelerating discrete dislocation dynamics simulations with graph neural networks

Discrete dislocation dynamics (DDD) is a widely employed computational method to study plasticity at the mesoscale that connects the motion of dislocation lines to the macroscopic response of crystalline materials. However, the computational cost of DDD simulations remains a bottleneck that limits its range of applicability. Here, we introduce a new DDD-GNN framework in which the expensive time-integration of dislocation motion is entirely substituted by a graph neural network (GNN) model trained on DDD trajectories. As a first application, we demonstrate the feasibility and potential of our method on a simple yet relevant model of a dislocation line gliding through an array of obstacles. We show that the DDD-GNN model is stable and reproduces very well unseen ground-truth DDD simulation responses for a range of straining rates and obstacle densities, without the need to explicitly compute nodal forces or dislocation mobilities during time-integration. Our approach opens new promising avenues to accelerate DDD simulations and to incorporate more complex dislocation motion behaviors.

36 MATERIALS SCIENCE↗

Upgrading Fermilab s Accelerator Control System with ACORN

The Fermilab Accelerator Complex is the largest national user facility in the Office of High Energy Physics (DOE/HEP) program and the only national user facility operating at Fermilab. Fermilab serves as the host to the Long Baseline Neutrino Facility/Deep Underground Neutrino Experiment (LBNF/DUNE), the laboratory’s flagship project for neutrino science that is under construction. LBNF/DUNE will be powered by megawatt beams from an upgraded accelerator, the Proton Improvement Plan II (PIP-II) that will replace the laboratory’s aging linear accelerator with a new one based on superconducting radio-frequency cavities. The Accelerator Controls Operations Research Network (ACORN) Project will support LBNF/DUNE and PIP-II by modernizing the accelerator control system. The project is at the conceptual design phase and looking to achieve Critical Decision 1 (CD-1) later this year. The scope and structure of the project will be presented, along with an overview of how that has changed in the past year. Current design and technology choices will be shared. Specific challenges facing the project will be addressed, along with current thinking on solutions.

Roehrig, Christian [Fermilab]↗

Graph learning for particle accelerator operations

Particle accelerators play a crucial role in scientific research, enabling the study of fundamental physics and materials science, as well as having important medical applications. This study proposes a novel graph learning approach to classify operational beamline configurations as good or bad. By considering the relationships among beamline elements, we transform data from components into a heterogeneous graph. We propose to learn from historical, unlabeled data via our self-supervised training strategy along with fine-tuning on a smaller, labeled dataset. Additionally, we extract a low-dimensional representation from each configuration that can be visualized in two dimensions. Leveraging our ability for classification, we map out regions of the low-dimensional latent space characterized by good and bad configurations, which in turn can provide valuable feedback to operators. This research demonstrates a paradigm shift in how complex, many-dimensional data from beamlines can be analyzed and leveraged for accelerator operations.

43 PARTICLE ACCELERATORS↗

Predicting beam transmission using 2-dimensional phase space projections of hadron accelerators

We present a method to compress the 2D transverse phase space projections from a hadron accelerator and use that information to predict the beam transmission. This method assumes that obtaining at least three projections of the 4D transverse phase space is possible and that an accurate simulation model is available for the beamline. Using a simulated model, we show that—a computer can train a convolutional autoencoder to reduce phase-space information which can later be used to predict the beam transmission. Finally, we argue that although using projections from a realistic nonlinear distribution produces less accurate results, the method still generalizes well.

43 PARTICLE ACCELERATORS↗

Accelerated computation of lattice thermal conductivity using neural network interatomic potentials

We report with the development of the density functional theory (DFT) and ever-increasing computational capacity, an accurate prediction of lattice thermal conductivity based on the Boltzmann transport theory becomes computationally feasible, contributing to a fundamental understanding of thermal conductivity as well as a choice of the optimal materials for specific applications. However, steep costs in evaluating third-order force constants limit the theoretical investigation to crystals with high symmetry and few atoms in the unit cell. Currently, machine learning potentials are garnering attention as a computationally efficient high-fidelity model of DFT, and several studies have demonstrated that the lattice thermal conductivity could be computed accurately via the machine learning potentials. However, test materials were mostly crystals with high symmetries, and the applicability of machine learning potentials to a wide range of materials has yet to be demonstrated. Furthermore, establishing a standard training set that provides consistent accuracy and computational efficiencies across a wide range of materials would be useful. To address these issues, herein we compute lattice thermal conductivities at 300 K using neural network interatomic potentials. As test materials, we select 25 materials with diverse symmetries and a wide range of lattice thermal conductivities between 10 -1 and 10 3 Wm -1 K -1 . Among various choices of training sets, we find that molecular dynamics trajectories at 50–700 K consistently provide results at par with DFT for the test materials. In contrast to pure DFT approaches, the computational cost in the present approach is uniform over the test materials, yielding a speed gain of 2–10 folds. When a smaller reduced training set is used, the relative efficiency increases by up to ~50 folds without sacrificing accuracy significantly. The current work will broaden the application scope of machine learning potentials by establishing a robust framework for accurately computing lattice thermal conductivity with machine learning potentials.

36 MATERIALS SCIENCE↗

Strategy for Modernizing a 40-Year-Old Accelerator Control System

Modernizing the Fermilab accelerator control system is essential to future operations of the laboratory's accelerator complex. The existing control system has evolved over four decades and uses hardware that is no longer available and software that uses obsolete frameworks. The Accelerator Controls Operations Research Network (ACORN) Project will modernize the control system and replace end-of-life power supplies to enable future accelerator complex operations with megawatt particle beams. The project team is evaluating three design concepts, and the future deployment of artificial intelligence capabilities for accelerator operations is an important consideration. An overview of the ACORN Project will be presented, including R&D used for evaluating the conceptual designs in the context of requirements for future accelerator operations.

43 PARTICLE ACCELERATORS↗

Modernizing Fermilab s Control Hardware

Modernizing the Fermilab accelerator control system is essential to future operations of the laboratory's accelerator complex. The existing control system has evolved over four decades and uses hardware that is no longer available. The Accelerator Controls Operations Research Network (ACORN) Project will modernize the control system and replace end-of-life power supplies to enable future accelerator complex operations with megawatt particle beams. The ACORN project is planning to replace Fermilab’s obsolete CAMAC crate-and-card controls hardware with modern MicroTCA hardware. There are over 2,000 CAMAC cards and over 250 CAMAC crates serving various functions in Fermilab’s control system. We will present an overview of the existing CAMAC hardware and the proposed MicroTCA replacement plan.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Modernizing Fermilab's Accelerator Control Hardware

Modernizing the Fermilab accelerator control system is essential to future operations of the laboratory's accelerator complex. The existing control system has evolved over four decades and uses hardware that is no longer available.The Accelerator Controls Operations Research Network (ACORN) Project will modernize the control system and replace end-of-life power supplies to enable future accelerator complex operations with megawatt particle beams. The ACORN project is planning to replace Fermilab s obsolete CAMAC crate-and-card controls hardware with modern MicroTCA hardware. There are over 2,000 CAMAC cards and over 250 CAMAC crates serving various functions in Fermilab s control system. We will review the existing CAMAC hardware and provide updates on the status of the ACORN project, which includes the conceptual design of the MicroTCA replacement hardware and software tools for supporting the installation of hundreds of MicroTCA crates.

43 PARTICLE ACCELERATORS↗

Introduction and Status of Fermilab's ACORN Project

Modernizing the Fermilab accelerator control system is essential to future operations of the laboratory's accelerator complex. The existing control system has evolved over four decades and uses hardware that is no longer available and software that uses obsolete frameworks. The Accelerator Controls Operations Research Network (ACORN) Project will modernize the control system and replace end-of-life power supplies to enable future accelerator complex operations with megawatt particle beams. An overview of the ACORN Project and a summary of recent research and development activities will be presented.

43 PARTICLE ACCELERATORS↗

Toward accelerating rare-earth metal extraction using equivariant neural networks

The separation of rare-earth metals, vital for numerous advanced technologies, is hampered by their similar chemical properties, making ligand discovery a significant challenge. Traditional experimental and quantum chemistry approaches for identifying effective ligands are often resource-intensive. We introduce a machine learning protocol based on an equivariant neural network, Allegro, for the rapid and accurate prediction of binding energies in rare-earth complexes. Key to this work is our newly curated dataset of rare-earth metal complexes—made publicly available to foster further research—systematically generated using the Architector program. This dataset distinctively features functionalized derivatives of proven rare-earth-chelating scaffolds, hydroxypyridinone (HOPO), catecholamide (CAM), and their thio-analogues, selected for their established efficacy in binding these elements. Trained on this valuable resource, our Allegro models demonstrate excellent performance, particularly when trained to directly predict DFT-level binding energies, yielding highly accurate results that closely correlate with theoretical calculations on a diverse test set. Furthermore, this strategy exhibited strong out-of-sample generalization, accurately predicting binding energies for an isomeric HOPO-derivative ligand not seen during training. By substantially reducing computational demands, this machine learning framework, alongside the provided dataset, represent powerful tools to accelerate the high-throughput screening and rational design of novel ligands for efficient rare-earth metal separation.

Gupta, Ankur K. [Lawrence Berkeley National Labora↗

An MLIR-based Compiler Flow for System-Level Design and Hardware Acceleration

The generation of custom hardware accelerators for applications implemented within high-level productive programming frameworks requires considerable manual effort. To automate this process, we introduce \sodaopt, a compiler tool that extends the MLIR infrastructure. \sodaopt automatically searches, outlines, tiles, and pre-optimizes relevant code regions to generate high-quality accelerators through high-level synthesis. \sodaopt can support any high-level programming framework and domain-specific language that interface with the MLIR infrastructure. By leveraging MLIR, \sodaopt solves compiler optimization problems with specialized abstractions. Backend synthesis tools connect to \sodaopt through progressive intermediate representation lowerings. \sodaopt interfaces to a design space exploration engine to identify the combination of compiler optimization passes and options that provides high-performance generated designs for different backends and targets. We demonstrate the practical applicability of the compilation flow by exploring the automatic generation of accelerators for deep neural networks operators outlined at arbitrary granularity and by combining outlining with tiling on large convolution layers. Experimental results with kernels from the PolyBench benchmark show that \sodaopt high-level optimizations improve execution delays of synthesized accelerators up to 60x. We also show that for the selected kernels, our solution outperforms the current of state-of-the art in more than 70% of the benchmarks and provides better average speedup in 55% of them.

Bohm Agostini, Nicolas↗

ThunderSecure: deploying real-time intrusion detection for 100G research networks by leveraging stream-based features and one-class classification network

Nowadays, data generated by large-scale scientific experiments are on the scale of petabytes per month. These data are transferred through dedicated high-bandwidth networks (40/100G) across distributed sites for processing, storage, and analysis. Like general purpose networks, research networks experience intrusions. However, monitoring anomalies in such high-speed network traffics is challenging given current cyber-infrastructure. Moreover, traditional network intrusion detection systems (NIDS) are signature based. However, anomaly patterns are difficult to define and that rulesets are often not updated frequently enough to reflect the changes of attack behaviors. We present ThunderSecure, a high-throughput, unsupervised learning-based intrusions detection system for 100G research networks. ThunderSecure implements an efficient packet processing and detection pipeline using multi-cores and GPUs. It extracts statistical and temporal features from real-time network data streams and feeds them to a one-class anomaly detection network. A baseline of normal distribution will be created based on the training observation. Testing traffic deviated from the learned profile will be marked as anomalies. We trained ThunderSecure on hundreds of billions of science data packets mirrored from two 100G network connections at Fermi National Accelerator Laboratory. The detection performance was evaluated on traffic captured from the same research network days and weeks after the training with different types of attack flows injected. Results show that ThunderSecure can recognize science data traffic captured long after the training and made nearly certain detection on the segment of the streams where anomalous flows were injected.

100G research network↗

Bridging the Gap Between Modern UX Design and Particle Accelerator Control Room Interfaces

Accelerator control systems often represent relatively complex and safety-sensitive human-machine interfaces within process control industries. These systems are technically robust and reflect the cumulative integration of solutions built and adapted across decades. One of the regular, unfortunate casualties of provisional accelerator control system updates is their human-system interfaces (HSIs) which often lag behind modern usability and design standards. An additional challenge is that although there is a multitude of established human factors (HF), and user experience (UX) principles for everyday digital applications, there are very few (if any) established principles for complex and safety-critical applications for an accelerator. This paper argues for the importance of established HF and UX principles (herein referred to as human-centered design principles) into the development of accelerator HSIs, emphasizing the need for clarity, consistency, responsiveness, and cognitive accessibility. Drawing from HF/UX best practices and human-centered design, this paper discusses how these approaches can enhance operator performance, reduce human error, and improve accelerator personnel collaboration. Case studies from Accelerator Control Operations Research Network (ACORN) at Fermilab are explored to demonstrate how interfaces built with human-centered design principles can scale with system complexity while remaining intuitive and efficient for diverse user roles including operators, machine experts, and engineers. By bridging the gap between traditional control system design and modern human-centered design methods, this paper provides a roadmap for evolving accelerator HSIs into more usable, maintainable, and effective tools.

Hill, Rachael [Idaho Natl. Lab.]↗