Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “network acceleration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25

Learning-Accelerated ADMM for Distributed DC Optimal Power Flow

We suggest a novel data-driven method to accelerate the convergence of Alternating Direction Method of Multipliers (ADMM) for solving distributed DC optimal power flow (DC-OPF) where lines are shared between independent network partitions. Using previous observations of ADMM trajectories for a given system under varying load, the method trains a recurrent neural network (RNN) to predict the converged values of dual and consensus variables. Given a new realization of system load, a small number of initial ADMM iterations is taken as input to infer the converged values and directly inject them into the iteration. We empirically demonstrate that the online injection of these values into the ADMM iteration accelerates convergence by a significant factor for partitioned 14-, 118-and 2848-bus test systems under differing load scenarios. The proposed method has several advantages: it maintains the security of private decision variables inherent in consensus ADMM; inference is fast and so may be used in online settings; RNN-generated predictions can dramatically improve time to convergence but, by construction, can never result in infeasible ADMM subproblems; it can be easily integrated into existing software implementations. While we focus on the ADMM formulation of distributed DC-OPF in this paper, the ideas presented are naturally extended to other distributed optimization problems.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Learning-Accelerated ADMM for Distributed DC Optimal Power Flow

We propose a novel data-driven method to accelerate the convergence of Alternating Direction Method of Multipliers (ADMM) for solving distributed DC optimal power flow (DC-OPF) where lines are shared between independent network partitions. Using previous observations of ADMM trajectories for a given system under varying load, the method trains a recurrent neural network (RNN) to predict the converged values of dual and consensus variables. Given a new realization of system load, a small number of initial ADMM iterations is taken as input to infer the converged values and directly inject them into the iteration. We empirically demonstrate that the online injection of these values into the ADMM iteration accelerates convergence by a significant factor for partitioned 14-, 118- and 2848-bus test systems under differing load scenarios. The proposed method has several advantages: it maintains the security of private decision variables inherent in consensus ADMM; inference is fast and so may be used in online settings; RNN-generated predictions can dramatically improve time to convergence but, by construction, can never result in infeasible ADMM subproblems; it can be easily integrated into existing software implementations. While we focus on the ADMM formulation of distributed DC-OPF in this paper, the ideas presented are naturally extended to other distributed optimization problems.

alternating direction method of multipliers↗

Modeling Freight Traffic Demand and Highway Networks for Hydrogen Fueling Station Planning: A Case Study of U.S. Interstate 75 Corridor

The use of hydrogen as an alternative transportation fuel has gained much interest in recent years. Hydrogen can be utilized in electric vehicles equipped with hydrogen powertrains (including hydrogen internal combustion engines or fuel cells). Given that most of the freight in the U.S. is transported via diesel trucks, transition to hydrogen fuel would help achieve significant environmental benefits as well as accelerate the decarbonization of the freight transportation sector. This paper presents the methodology and results of a case study on modeling freight traffic demand and highway networks based on publicly available data for the Interstate 75 freight corridor. The purpose of this study is to prepare input traffic and network data that can support the planning of a hydrogen fueling station infrastructure. In particular, the data can be used for siting and characterizing an optimized framework of hydrogen fueling stations from candidate diesel stations along the Interstate 75 corridor. The methodologies developed and presented in this paper may be readily expanded and applied to any transport corridor given the data availability. This paper is the first in a series that will build out a comprehensive model to optimize a consolidated national hydrogen refueling infrastructure eco-system targeted at commercial vehicles.

Uddin, Majbah↗

Energy Materials Chemistry Integrating Theory, Experiment and Data Science (Final Report)

The Energy Materials Chemistry Integrating Theory, Experiment and Data Science (EM-CITED) project is a multidisciplinary research effort focused on accelerating discovery of scientific knowledge via incorporation of data science and artificial intelligence in materials chemistry research. The project aims to advance materials chemistry-aware data science to unify theory and experiment knowledge streams. The work resulted in foundational AI frameworks for materials chemistry – Deep Reasoning Networks (DRNets), Hierarchical Correlation Learning for Multi-property Prediction (H-CLMP), and Material-to-Spectrum (Mat2Spec) prediction – as well as a host of strategies for accelerated scientific discoveries through principled incorporation of data science in computational and experimental research.

36 MATERIALS SCIENCE↗

Observations related to the acceleration, injection, and interplanetary propagation of energetic protons during the solar cosmic ray event on February 16, 1984

This paper presents an analysis of data collected by the worldwide network of neutron monitors and from IMP-8 cosmic-ray telescopes, as well as by particle detectors on the GOES 5 and 6 and ICE satellites, on the solar cosmic ray event that took place on February 16, 1984. Using these data, the intensity-time (IT) profiles, the anisotropy-time profiles, the energy spectra, and the pitch angle distributions of the solar protons near earth were deduced. It was found that the solar protons propagated essentially scatter-free from the sun to the earth. The solar protons had easy access to the IMF lines to earth; the time from the onset to maximum intensity and the shape of the IT profiles at earth as a function of energy could be explained by the diffusion of the flare protons near the acceleration region. The energy spectrum of the solar flare protons injected into the undisturbed IMF at the sun was changing with time in both amplitude and shape. The observations suggest a shock acceleration process.

Debrunner, H.↗

SMALE: Enhancing Scalability of Machine Learning Algorithms on Extreme-Scale Computing Platforms

Deployment and execution of machine learning tasks on extreme-scale computing platforms face several significant technical challenges: 1) High computing cost incurred by dense networks – The computing workload of deep networks with densely-connected topology increases rapidly with the network size, imposing a non-scalable computing model of extreme-scale computing platforms; 2) Non-optimized workload distribution – Many advanced deep learning algorithms, e.g., sparsification and irregular net-work topology, produce very unbalanced workload distribution on extreme-scale computing platforms. The computation efficiency is greatly hindered by the incurred data and computation redundancies as well as long tails of the node with extensive workload; 3) Constraints in data movement and I/O bottle-neck – Inter-node data movement in extreme-scale computing platforms are associated with high energy and latency costs, and subject to the constraints of I/O bandwidth; and 4) Generalization of algorithm realization and acceleration on computing platforms – The large varieties of machine learning algorithms and structures of extreme-scale computing platforms make the derivation of a generalized algorithm realization and acceleration method very challenging, which, however, is the requirement by domain scientists and interested users. We call the above challenges Smale’s Problems in Machine Learning and Understanding for High-Performance Computing Scientific Discovery. The objective of our three-year research project is to develop a holistic innovation set at structure, assembly, and acceleration layers of machine learning algorithms to address the above challenges in algorithm deployment and execution. Three tasks are particularly performed, including: At the algorithm structure level, we investigate the techniques that can structurally sparsify on the topology of deep networks for computing workload reduction. We also study clustering and pruning techniques that can optimize the workload distributions over the extreme-scale computing platforms; At the algorithm assembly level, we derive a unified learning framework for unsupervised transfer learning and dynamic growing capabilities. Novel training methods are also exploited to enhance the training efficiency of the proposed framework; At the algorithm acceleration level, we will develop a series of techniques that can accelerate the computation of sparse matrix operations, which are one of the core executions in deep learning and optimize memory access of the concerned platforms. Our proposed techniques attack the fundamental problems in machine learning algorithms running on extreme-scale computing platforms by vertically integrating the solutions at three closely entangled layers, paving the long-term scaling path of machine learning applications under DOE context. Three tasks corresponding to the above respective research orientations are performed during the three-year project period with our collaborators at ORNL. The outcome of the proposed project is anticipated to form a holistic solution set of novel algorithms and network topologies, efficient training techniques, and fast acceleration methods to promote the computing scalability of the machine learning applications of particular interest to DOE.

97 MATHEMATICS AND COMPUTING↗

Liquid-phase mega-electron-volt ultrafast electron diffraction

The conversion of light into usable chemical and mechanical energy is pivotal to several biological and chemical processes, many of which occur in solution. To understand the structure–function relationships mediating these processes, a technique with high spatial and temporal resolutions is required. Here, we report on the design and commissioning of a liquid-phase mega-electron-volt (MeV) ultrafast electron diffraction instrument for the study of structural dynamics in solution. Limitations posed by the shallow penetration depth of electrons and the resulting information loss due to multiple scattering and the technical challenge of delivering liquids to vacuum were overcome through the use of MeV electrons and a gas-accelerated thin liquid sheet jet. To demonstrate the capabilities of this instrument, the structure of water and its network were resolved up to the 3rd hydration shell with a spatial resolution of 0.6 Å; preliminary time-resolved experiments demonstrated a temporal resolution of 200 fs.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Advancing Fusion Research and Development at TAE Technologies Through INFUSE Program

The U.S. Department of Energy’s Innovation Network for Fusion Energy (INFUSE) program serves as a crucial catalyst by fostering public-private partnership that accelerates technological innovation for fusion energy research and development (R&D) in the private sector. Further, this article provides a comprehensive overview of the technical goals and accomplishments of projects awarded to TAE Technologies through the INFUSE program since 2019. We offer high-level perspectives on how these projects have contributed to fusion energy R&D, and we address key challenges encountered during these collaborations.

field-reversed configuration↗

FLASH: FPGA-Accelerated Smart Switches with GCN Case Study

Some communication switches, e.g., the Mellanox SHArP and those in the IBM BlueGene clusters, are augmented to process packets at the application level with fixed-function collectives. This approach, however, lacks flexibility, which limits their applicability in diverse and dynamic workloads. Recently, a new type of programmable packet processor, which uses high-level languages, e.g., P4, has emerged as possible candidates. P4-based switches, however, fall short in certain applications, including machine learning, where capabilities not currently supported by P4 are needed. These include more complex calculation, such as sparse computation and fused multiply-accumulate, data-intensive floating point operations, data reuse, and significant memory. The problem addressed here is that such a switch augmentation needs to support: a large amount of state, significant flexible compute capability, and ease of programming, all while maintaining full functionality, including ensuring high throughput, and demonstrating utility. In this work, we propose a programmable look-aside-type accelerator that can be embedded into, or attached to, existing communication switch pipelines and that is capable of processing packets at line-rate. The proposed in-switch accelerator is based on mixing an ISA (subset of RISC-V instructions) with dataflow graphs (found in CGRAs). To augment performance, vector instructions are also supported. To facilitate usability, we have developed a complete toolchain to compile user-provided C/C++ codes to appropriate back-end instructions for configuring the accelerator. While this approach is flexible enough to support various workloads, in this paper, we consider Graph Convolutional Networks (GCNs) as a case study. Experimental results show that this approach considerably improves the performance of distributed GCN applications.

Haghi, Pouya↗

Cooperative Research and Development Agreement With Georgetown University Report: National Institutes of Health, National Center for Advancing Translational Sciences Clinical and Translational Science Award

Oak Ridge National Laboratory (ORNL) is participating with Georgetown University (GU) as a subrecipient in response to the National Institutes of Health (NIH), National Center for Advancing Translational Sciences (NCATS) Clinical and Translational Science Award (CTSA) (U54 Clinical Trial Optional) funding opportunity announcement. This Cooperative Research and Development Agreement is put in to place to facilitate the development and implementation of clinical interventions that demonstrably improve human health is currently a complex, recursive, and inefficient process that leads to delays of years or decades before discoveries in biomedical research result in health benefits for patients and communities. NCATS conducts and supports research in the science of translation, to discover the mechanistic and operational principles of the intervention development and dissemination process, thereby providing the scientific foundation for improvements in translational efficiency that will accelerate the realization of interventions that improve human health. Under NCATS’ leadership, the CTSA Program supports a national network of medical research institutions called hubs. GU is the lead institution in one of the NIH hubs that was created as a result of a previous NIH CTSA. The missions of the GU have historically included the advancement of health through research in the clinical and biomedical sciences, the education of future leaders in medical and nursing practice and academia, and the provision of compassionate and scientifically competent patient care and service to the Washington, DC community and the nation. GU is the lead institution for the Georgetown-Howard Universities Center for Clinical and Translational Science (GHUCCTS), a multi-institutional partnership of medical research institutions forged from a desire to promote clinical research and translational science. Through multiple collaborations among these institutions, GHUCCTS is transforming clinical research and translational science in order to bring new scientific advances to health care. Oak Ridge National Laboratory is the Department of Energy's (DOE) largest science and energy laboratory. Managed since April 2000 by a partnership of the University of Tennessee and Battelle, ORNL was established in 1943 as a part of the secret Manhattan Project to pioneer a method for producing and separating plutonium. During the 1950s and 1960s, ORNL became an international center for the study of nuclear energy and related research in the physical and life sciences. With the creation of DOE in the 1970s, ORNL's mission broadened to include a variety of energy technologies and strategies. Today the laboratory supports the nation with a peacetime science and technology mission that is just as important as, but very different from, its role during the Manhattan Project. ORNL is home to the world's premier center for high performance supercomputing to enable scientific discovery. ORNL has extensive expertise in various areas of computer science that are uniquely situated to support GU. Additionally, ORNL’s leading computational user facilities present a unique opportunity to leverage the largest scale machines for open science in support of the stated mission of the NCATS CTSA. ORNL's partnership with GU will offer unparalleled opportunity in data analytics, deep-learning, artificial intelligence, and urban dynamics.

59 BASIC BIOLOGICAL SCIENCES↗

Federated IRI Science Testbed (FIRST): A Concept Note

The Department of Energy’s (DOE’s) vision for an Integrated Research Infrastructure (IRI) is to empower researchers to smoothly and securely meld the DOE’s world-class user facilities and research infrastructure in novel ways in order to radically accelerate discovery and innovation. Performant IRI arises through the continuous interoperability of research workflows with compute, storage, and networking infrastructure, fulfilling researchers’ quests to gain insight from observational and experimental data. Decades of successful research, pilot projects, and demonstrations point to the extraordinary promise of IRI but also indicate the intertwined technological, policy, and sociological hurdles it presents. Creating, developing, and stewarding the conditions for seamless interoperability of DOE research infrastructure, with clear value propositions to stakeholders to opt into an IRI ecosystem, will be the next big step. Governance, funding, and resource allocation are beyond the scope of this document: it seeks to provide a high-level view of potential benefits, focus areas, and the working groups whose formation would further define the testbed’s design, activities, and goals.

97 MATHEMATICS AND COMPUTING↗

LANSCE Science Overview [Slides]

LANSCE is primarily focused on three science questions for NNSA and contributes to several other programs. We will primarily consider the NNSA questions: (1) How can we advance our understanding of dynamic material behavior using focused experiments? (2) What tools are needed to address Advanced Manufacturing and Aging? (3) How do we constrain the nuclear reaction networks involved in weapons?

36 MATERIALS SCIENCE↗

FracML: A Machine Learning Based Tool to Quantify Reservoir Scale Fracture Network for CO2 Storage

Poster on “FRACML: A Machine Learning Based Tool to Quantify Reservoir Scale Fracture Network for CO2 Storage” for the CCUS 2025 conference held in Houston, Texas March 3-5, 2025. The accurate characterization of subsurface fracture networks is essential for the secure operation of carbon capture, utilization, and storage (CCUS) projects. A thorough understanding of the spatial distribution of subsurface faults and fractures is crucial for predicting CO2 plume evolution and minimizing risks such as potential leakage into overlying formations or induced seismicity. In this context, robust fracture network quantification plays a pivotal role in reservoir management, providing the data necessary to fine-tune operational parameters, and ensure the environmental and economic viability of CCUS projects. As part of the U.S. Department of Energy’s SMART (Science-informed Machine Learning for Accelerating Real-time Decisions in Subsurface Applications) initiative, we focused on the development and application of a machine learning-based tool (FRACML) designed to quantify and map fracture networks using real-world (non-synthetic) data from an active CO2 injection site. Our objective is to demonstrate the utility of this tool in improving operational efficiency and safety across CCUS sites.

artifical intelligence / machine learning (AI/ML)↗

wa-hls4ml and lui-gnn: A benchmark and GNN-based surrogate model for hls4ml resource and latency estimation

As machine learning (ML) increasingly serves as a tool for addressing real-time challenges in scientific applications, the development of advanced tooling has significantly reduced the time required to iterate on various designs. These advancements have solved major obstacles, but also exposed new challenges. For example, processes that were not previously considered bottlenecks, such as model synthesis, are now becoming limiting factors in the rapid iteration of designs. To reduce these emerging constraints, multiple efforts are being launched toward designing an ML-based surrogate model that estimates resource usage of synthesized accelerator architectures. This model would reduce the design iteration time, especially when designing within a set of given hardware constraints. This approach shows considerable potential, but as it stands, the effort is early and would benefit from coordination and standardization to assist future work as it emerges. We introduce wa-hls4ml, a benchmark for ML accelerator resource and latency estimation, and its corresponding initial dataset of more than 100,000 fully connected neural networks, all synthesized using hls4ml and targeting Xilinx FPGAs. In addition to the resource utilization and latency data provided, the dataset includes generated artifacts and log files for many of the synthesized neural networks, in order to support future research in ML-based code generation. The benchmark evaluates the performance of resource and latency predictors against several common ML model architectures, primarily originating from scientific domains, as exemplar models, as well as the average performance across a subset of the dataset. We measure the performance of a given predictor model through multiple metrics, including $R^2$ score and SMAPE on regression tasks, as well as inference time to further characterize the estimator under test. Additionally, we introduce the latency/utilization inference graph neural network (lui-gnn), a surrogate model that uses a graph neural network to represent input architectures in the form of a directed graph. This graph representation allows for a diverse set of model architectures to all be effectively handled by a surrogate model. We present the architecture and performance of the model, as evaluated by the new proposed benchmark, including SMAPE, $R^2$ score, and inference times, and find that lui-gnn generally predicts latency and utilization for the 75\% quantile within several percent of the synthesized resources on the synthetic test dataset, indicating that this approach of estimating resource and latency via a surrogate models has promise and warrants further research.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Using EMG to anticipate head motion for virtual-environment applications

In virtual environment (VE) applications, where virtual objects are presented in a see-through head-mounted display, virtual images must be continuously stabilized in space in response to user's head motion. Time delays in head-motion compensation cause virtual objects to "swim" around instead of being stable in space which results in misalignment errors when overlaying virtual and real objects. Visual update delays are a critical technical obstacle for implementing head-mounted displays in applications such as battlefield simulation/training, telerobotics, and telemedicine. Head motion is currently measurable by a head-mounted 6-degrees-of-freedom inertial measurement unit. However, even given this information, overall VE-system latencies cannot be reduced under about 25 ms. We present a novel approach to eliminating latencies, which is premised on the fact that myoelectric signals from a muscle precede its exertion of force, thereby limb or head acceleration. We thus suggest utilizing neck-muscles' myoelectric signals to anticipate head motion. We trained a neural network to map such signals onto equivalent time-advanced inertial outputs. The resulting network can achieve time advances of up to 70 ms.

Clinical Trial↗

The Source of Alfven Waves That Heat the Solar Corona

We suggest a source for high-frequency Alfven waves invoked in coronal heating and acceleration of the solar wind. The source is associated with small-scale magnetic loops in the chromospheric network.

Alfven Waves Heat Solar Corona solar wind solar wi↗

Development of Machine Learning Algorithms to Segment and Study Images of Astromaterial Samples

Introduction: Micrometer-scale chemical analyses of chondritic meteorites and mission-returned asteroid samples can reveal details of the physical and chemical processes operating in the early solar system, including processes that gave rise to planets, moons, and minor bodies. These primitive astromaterials are comprised of chondrules, calcium- and aluminum-rich inclusions (CAI), and many other silicates, oxides, metals, sulfides, and fine-grained materials. The chemical and mineralogical complexity of these samples, vast populations of different components, and heterogeneity across mm to km scales, all limit our understanding of the origin and evolution of these materials. Here, we describe recent efforts to use machine learning techniques to automate the segmentation of chemical maps of chondritic meteorites, designed to aid studies of asteroid samples returned by spacecraft. By automating the task of segmentation it will become possible to rapidly analyze and interpret the sizes, shapes, mineralogy, chemistry, and other properties of every chondrule, calcium- and aluminum-rich inclusion (CAI) and other clast within and between asteroid samples. Sample return missions significantly accelerate and heighten the need to develop such new data analysis techniques, and associated data repositories. Techniques: Neural networks require abundant training data, i.e. images which have been segmented by a human user. We have manually segmented data available from previous petrologic and chemical work at NASA Johnson Space Center and the American Museum of Natural History [1-4]. These data were derived from energy- and wavelength-dispersive X-ray spectroscopy (EDS, WDS) mapping of samples from many chondrite groups. The Deeplabv3+ [5] neural network architecture was trained on human-labeled masks and used to create machine-labeled masks. Several different algorithms were investigated, with inputs ranging from common RGB image formats through to hyperspectral datasets, with raw data comprising greyscale maps of Mg, Ca, and Al, with or without Si, Fe, Ti for both EDS and WDS data, and extending to other elements in EDS only. Each greyscale image was paired with a binary mask for each labelled particle type. Results: The trained algorithms can segment (Fig 1), classify, and measure the dimensions of thousands of particles in chemical maps of a standard 1-inch round petrographic section in seconds to minutes, rather than many hours needed by a human. Accuracy of the algorithms varied from chondrite to chondrite and across particle types. Further results and details of the algorithms will be presented at the workshop. Future directions: Machine learning has the potential to revolutionize our understanding of complex particle populations contained within primitive astromaterial, with segmentation being a critical first step. Example applications include better understanding of particle transport, nebular reservoirs, parent body accretion, and a deeper understanding of the relationships between particle populations and bulk rock elemental and isotopic compositions. In addition to benefits that machine learning can bring to individual researchers, building a community data repository of thousands to millions of particles across hundreds of samples will open up many other possibilities. For example, with a large enough dataset it will be possible to search for exceptionally closely matching particles across disparate samples. Such a capability would enable a single CAI from OSIRISREx or Hayabusa/II samples to be matched to chondritic CAIs that exhibit near-identical size, texture, and mineralogy, down to the level of similar core phenocrysts, zonation, and rim sequences. Such comparative analyses will help to disentangle precursor chemistry, chronology, gas/dust reservoirs during heating, and accretion. Such an endeavor would be impossible without machine learning and a large community data repository of astromaterial chemical/mineralogic maps.

Machine Learning↗