Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “network acceleration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24

Sensitivity of He Flames in X-Ray Bursts to Nuclear Physics

Through the use of axisymmetric 2D hydrodynamic simulations, we further investigate laterally propagating flames in X-ray bursts (XRBs). Our aim is to understand the sensitivity of a propagating helium flame to different nuclear physics. Using the Castro simulation code, we confirm the phenomenon of enhanced energy generation shortly after a flame is established by adding 12 C(p, γ) 13 N(α, p) 16 O to the network, in agreement with the past literature. This sudden outburst of energy leads to a short accelerating phase, causing a drastic alteration in the overall dynamics of the flame in XRBs. Furthermore, we investigate the influence of different plasma screening routines on the propagation of the XRB flame. We finally examine the performance of simplified spectral deferred correction, a novel approach to hydrodynamics and reaction coupling incorporated in Castro, as an alternative to operator splitting.

79 ASTRONOMY AND ASTROPHYSICS↗

Space Technology Mission Directorate Game Changing Development Program FY2015 Annual Program Review: Advanced Manufacturing Technology

The Advance Manufacturing Technology (AMT) Project supports multiple activities within the Administration's National Manufacturing Initiative. A key component of the Initiative is the Advanced Manufacturing National Program Office (AMNPO), which includes participation from all federal agencies involved in U.S. manufacturing. In support of the AMNPO the AMT Project supports building and Growing the National Network for Manufacturing Innovation through a public-private partnership designed to help the industrial community accelerate manufacturing innovation. Integration with other projects/programs and partnerships: STMD (Space Technology Mission Directorate), HEOMD, other Centers; Industry, Academia; OGA's (e.g., DOD, DOE, DOC, USDA, NASA, NSF); Office of Science and Technology Policy, NIST Advanced Manufacturing Program Office; Generate insight within NASA and cross-agency for technology development priorities and investments. Technology Infusion Plan: PC; Potential customer infusion (TDM, HEOMD, SMD, OGA, Industry); Leverage; Collaborate with other Agencies, Industry and Academia; NASA roadmap. Initiatives include: Advanced Near Net Shape Technology Integrally Stiffened Cylinder Process Development (launch vehicles, sounding rockets); Materials Genome; Low Cost Upper Stage-Class Propulsion; Additive Construction with Mobile Emplacement (ACME); National Center for Advanced Manufacturing.

Vickers, John↗

SPARTA: High-Level Synthesis of Parallel Multi-Threaded Accelerators

This article presents a methodology for the Synthesis of PARallel multi-Threaded Accelerators (SPARTA) from OpenMP annotated C/C++ specifications. SPARTA extends an open-source HLS tool, enabling the generation of accelerators that provide latency tolerance for irregular memory accesses through multithreading, support fine-grained memory-level parallelism through a hot-potato deflection-based network-on-chip (NoC), support synchronization constructs, and can instantiate memory-side caches. Our approach is based on a custom runtime OpenMP library, providing flexibility and extensibility. Experimental results show high scalability when synthesizing irregular graph kernels. The accelerators generated with our approach are, on average, 2.29x faster than state-of-the-art HLS methodologies.

Design automation↗

Decentralized Schemes with Overlap for Solving Graph-Structured Optimization Problems

We present a new algorithmic paradigm for the decentralized solution of graph-structured optimization problems that arise in the estimation and control of network systems. A key and novel design concept of the proposed approach is that it uses overlapping subdomains to promote and accelerate convergence. We show that the algorithm converges if the size of the overlap is sufficiently large and that the convergence rate improves exponentially with the size of the overlap. The proposed approach provides a bridge between fully decentralized and centralized architectures and is flexible in that it enables the implementation of asynchronous schemes, handling of constraints, and balancing of computing, communication, and data privacy needs. The proposed scheme is tested in an estimation problem for a 9241-node power network and we show that it outperforms the alternating direction method of multipliers.

asynchronous↗

Toward fusion plasma scenario planning for NSTX-U using machine-learning-accelerated models

One of the most promising devices for realizing power production through nuclear fusion is the tokamak. To maximize performance, it is preferable that tokamak reactors achieve advanced operating scenarios characterized by good plasma confinement, improved magnetohydrodynamic (MHD) stability, and a largely non-inductively driven plasma current. Such scenarios could enable steady-state reactor operation with high \emph{fusion gain} --- the ratio of produced fusion power to the external power provided through the plasma boundary. Precise and robust control of the evolution of the plasma boundary shape as well as the spatial distribution of the plasma current, density, temperature, and rotation will be essential to achieving and maintaining such scenarios. The complexity of the evolution of tokamak plasmas, arising due to nonlinearities and coupling between various parameters, motivates the use of model-based control algorithms that can account for the system dynamics. In this work, a learning-based accelerated model trained on data from the National Spherical Torus Experiment Upgrade (NSTX-U) is employed to develop planning and control strategies for regulating the density and temperature profile evolution around desired trajectories. The proposed model combines empirical scaling laws developed across multiple devices with neural networks trained on empirical data from NSTX-U and a database of first-principles-based computationally intensive simulations. The reduced execution time of the accelerated model will enable practical application of optimization algorithms and reinforcement learning approaches for scenario planning and control development. An initial demonstration of applying optimization approaches to the learning-based model is presented, including a strategy for mitigating the effect of leaving the finite validity range of the accelerated model. The approach shows promise for actuator planning between experiments and in real-time.

machine learning↗

Four-dimensional phase-space reconstruction of flat and magnetized beams using neural networks and differentiable simulations

Beams with cross-plane coupling or extreme asymmetries between the two transverse phase spaces are often encountered in particle accelerators. Flat beams with large transverse-emittance ratios are critical for future linear colliders. Similarly, magnetized beams with significant cross-plane coupling are expected to enhance the performance of electron cooling in hadron beams. Preparing these beams requires precise control and characterization of the four-dimensional transverse phase space. In this study, we employ generative phase-space reconstruction techniques to rapidly characterize magnetized and flat-beam phase-space distributions using a conventional quadrupole-scan method. The reconstruction technique is experimentally demonstrated on an electron beam produced at the Argonne Wakefield Accelerator and successfully benchmarked against conventional diagnostics techniques. Specifically, we show that predicted beam parameters from the reconstructed phase-space distributions (e.g., as magnetization and flat-beam emittances) are in excellent agreement with those measured from the conventional diagnostic methods. Published by the American Physical Society 2024

43 PARTICLE ACCELERATORS↗

Tiny sensor-transmitter can withstand extreme acceleration, gives digital output

A self-pulsing oscillator transmits a pulsed signal. The time between pulses and the frequency are controlled by two networks. Variations in the component values in each of the two networks, due to environmental changes, appear as changes in frequency and time between pulses in the transmitted signal. Such a sensor is used to measure physical magnitudes.

Mossino, R. L.↗

Enabling a Larger Deep Space Mission Suite: A Deep Space Network Queuing Antenna for Demand Access

The advent of deep space small spacecraft, as exemplified by the Mars Cubesat One (MarCO), Lunar Trailblazer, Janus, the Escape and Plasma Acceleration and Dynamics Explorers (EscaPADE), and the thirteen Artemis 1 missions, opens the possibility that a much larger number of deep space spacecraft may be launched over the next 10 years and beyond. While scientifically exciting, the prospect of a (much) larger mission suite raises significant challenges for the current approach to ground stations and mission operations. We have been investigating an integrated approach for ground stations and missions operations to enable new modes of operation while maintaining the capabilities of the current operational techniques. This integrated approach is built around three core capabilities: (1) A queuing antenna that enables monitoring the status of a much larger number of spacecraft, and allows spacecraft to transmit requests for telemetry with NASA’s Deep Space Network (DSN); (2) a flexible scheduling system that expands the current DSN scheduling services to enable allocating time on DSN antennas in near real-time; and (3) a cloud-based ground data system that can be spun up and down according to how tracks are assigned by the flexible scheduling system. We shall show that an 18 meter DSN queuing antenna equipped with cyrogenic receivers would enable use of the DSN Demand Access Service for small spacecraft throughout the inner Solar System, thus providing service to a large mission suite. We first discuss the architecture of the queuing antenna and its supporting systems, including, for instance, the service required to generate the schedule for the queueing antenna (which dictates how it slews to monitor multiple spacecraft in a day of operations). Next, we describe the signaling scheme used to encode a request, which is inherited from the already operational DSN Beacon Tone Service, and describe two alternative ways to detect the incoming tone at the ground station, one based on maximum likelihood estimation (MLE), and another one based on Fast-Fourier Transfer (FFT) processing. We then use these results to estimate the maximum range at which a request can be reliably detected as a function of the spacecraft and ground station communication capabilities. Finally, the last part of this part of this paper briefly describes the prototyping effort undertaken at Morehead State University (MSU) and JPL to demonstrate the viability of this new DSN demand access. In particular, we describe the suite of tests conducted using MSU’s 21 meter ground station to validate its use a queuing antenna.

Mattle, Emily↗

Neural network methods for radiation detectors and imaging

Recent advances in image data proccesing through deep learning allow for new optimization and performance-enhancement schemes for radiation detectors and imaging hardware. This enables radiation experiments, which includes photon sciences in synchrotron and X-ray free electron lasers as a subclass, through data-endowed artificial intelligence. We give an overview of data generation at photon sources, deep learning-based methods for image processing tasks, and hardware solutions for deep learning acceleration. Most existing deep learning approaches are trained offline, typically using large amounts of computational resources. However, once trained, DNNs can achieve fast inference speeds and can be deployed to edge devices. A new trend is edge computing with less energy consumption (hundreds of watts or less) and real-time analysis potential. While popularly used for edge computing, electronic-based hardware accelerators ranging from general purpose processors such as central processing units (CPUs) to application-specific integrated circuits (ASICs) are constantly reaching performance limits in latency, energy consumption, and other physical constraints. These limits give rise to next-generation analog neuromorhpic hardware platforms, such as optical neural networks (ONNs), for high parallel, low latency, and low energy computing to boost deep learning acceleration (LA-UR-23-32395).

edge computing↗

Neural network-based model of galaxy power spectrum: fast full-shape galaxy power spectrum analysis

ABSTRACT We present a neural network-based emulator for the galaxy redshift-space power spectrum that enables several orders of magnitude acceleration in the galaxy clustering parameter inference, while preserving 3$\sigma$ accuracy better than 0.5 per cent up to $k_{\mathrm{max}}$ = 0.25 $\, h\text{Mpc}^{-1}$ within Lambda-cold dark matter ($\Lambda$CDM) and around 0.5 per cent $w_0$–$w_a$CDM. Our surrogate model only emulates the galaxy bias-invariant terms of one-loop perturbation theory predictions, these terms are then combined analytically with galaxy bias terms, counter-terms, and stochastic terms in order to obtain the non-linear redshift-space galaxy power spectrum. This allows us to avoid any galaxy bias prescription in the training of the emulator, which makes it more flexible. Moreover, we include the redshift $z \in [0,1.4]$ in the training which further avoids the need for re-training the emulator. We showcase the performance of the emulator in recovering the cosmological parameters of $\Lambda$CDM by analysing the suite of 25 AbacusSummit simulations that mimic the Dark Energy Spectroscopic Instrument luminous red galaxies at $z=0.5$ and 0.8, together as the emission line galaxies at $z=0.8$. We obtain similar performance in all cases, demonstrating the reliability of the emulator for any galaxy sample at any redshift in $0 \lt z \lt 1.4$. We will make our emulator public at github repository.

Trusov, Svyatoslav (ORCID:0000000224146720)↗

JANUS: Resilient and Adaptive Data Transmission for Enabling Timely and Efficient Cross-Facility Scientific Workflows

In modern science, the growing complexity of large-scale scientific projects has led to an increasing reliance on cross-facility scientific workflows, where resources and expertise from multiple institutions and geographic locations are leveraged to accelerate scientific discovery. These workflows often require transmitting huge amounts of scientific data through wide-area networks. Although high-speed networks like ESnet and transfer services such as Globus have improved data mobility, several challenges remain. The sheer volume of data can overwhelm network bandwidth, widely used transport protocols such as TCP suffer from inefficiencies due to retransmissions triggered by packet loss, and existing fault-tolerance mechanisms like erasure coding introduce substantial overhead. In this paper, we propose Janus, a resilient and adaptable data transmission approach designed for cross-facility scientific workflows. Unlike traditional TCP-based methods, Janus leverages UDP, integrates erasure coding for fault tolerance, and combines it with error-bounded lossy compression to reduce overhead. This novel design allows users to balance data transmission time and accuracy, optimizing transfer performance based on specific scientific requirements. Additionally, Janus dynamically adjusts erasure coding parameters in response to real-time network conditions, ensuring efficient data transfers even in fluctuating environments. We develop optimization models for determining ideal configurations and implement adaptive data transfer protocols to enhance reliability. Through extensive simulations and real-network experiments, we demonstrate that Janus significantly improves transfer efficiency while maintaining data fidelity.

Esaulov, Vladislav [Georgia State University, Atla↗

Deep reinforcement learning for dynamic control of fuel injection timing in multi-pulse compression ignition engines

Conventional compression-ignition (CI) engines have long offered high thermal efficiencies and torque across a wide range of loads, but often require extensive exhaust gas treatment that decreases efficiency to meet ever-increasing emissions regulations. One strategy to decrease emissions is to split the fuel injection into a series of smaller injections. In this paper, we explore a new way of discovering optimal control strategies for the next generation of CI engines using deep reinforcement learning (DRL). We outline a DRL procedure to maximize the weighted reward of engine work while minimizing end-of-cycle NO x emissions. Through the procedure outlined in this paper, we show that the DRL agent is able to reduce NO x emissions threefold while only decreasing network by 2%. We demonstrate the use of transfer learning (TL) across hierarchies of physical models to accelerate the learning process, making this approach feasible for a range of control problems within this space. This paper presents a framework and demonstration for using DRL to design control systems in technology areas such as multi-pulse engine control where a hierarchy of models combined with multi-objective rewards are used for optimal operation.

33 ADVANCED PROPULSION SYSTEMS↗

Toward a Seamless Integration of Computing, Experimental, and Observational Science Facilities: A Blueprint to Accelerate Discovery

The Department of Energy, Office of Science operates world-leading facilities for experimental, observational, and computational science. DOE supercomputing facilities will reach performance at the scale of ExaFLOPs in the coming years, enabling new vistas of scale and precision for large scale simulations and data analysis. Experimental scientific facilities are undergoing similar upgrades that will lead to higher data rates and correspondingly larger computational demands, and will increase the need for near-real-time processing and resilient support for more complex workflows. A transformation of science is underway, with workloads at supercomputing facilities increasingly driven by this explosion of data from instruments and experimental facilities, as well as the accelerating use of Artificial Intelligence (AI) as a tool for scientific discovery. A seamless integration of computing, networking, instruments, and experimental facilities is required to support these emerging workloads and open up a new frontier of U.S. leadership in scientific discovery. We propose to accomplish this by providing frictionless access to the ASCR supercomputing facilities. We describe our vision of combining the power of ASCR supercomputers and networking infrastructure into an integrated scalable fabric, available to end user scientists via interfaces that aim to automate and simplify access to high performance computing systems. This will enable unprecedented computational science capabilities for experimental and observational facilities, and will create new opportunities to combine large simulations and modeling with experimental facility data analysis. This blueprint for creating an integrated network of computational and experimental facilities will provide an enriched discovery environment and open doors for new scientific communities to access the DOE’s world-leading computing and networking capabilities.

97 MATHEMATICS AND COMPUTING↗

Development and Commercialization of Heavy-Duty Battery Electric Trucks Under Diverse Climate Conditions (DTNA EMG Innovation eCascadia 2.0 Final Technical Report)

During this project DTNA advanced heavy-duty transportation technologies to produce the eCascadia 2.0, a fully commercialized Class 7/8 electric tractor with the range and durability to meet the needs of 70% of freight movement in the United States. The eCascadia2.0 provides flexibility, operational performance, efficiency, and maintenance cost savings to the end user. The most direct outcome of this project is a commercially viable zero-emission heavy duty option with a 250-mile daily range and sufficient payload for regional haul duty cycles. The larger outcome is an acceleration of the market transformation away from petroleum-based fuels. DTNA’s comprehensive sales team, dealer network, customer network and maintenance and support teams will ensure that this is not a standalone zero-emission truck, but rather one that can be produced, marketed, and operated at scale to realize vast greenhouse gas and criteria pollutant emissions reductions, while also bringing zero emission vehicles closer to cost parity.

33 ADVANCED PROPULSION SYSTEMS↗

The 1989 NASA-ASEE Summer Faculty Fellowship Program in Aeronautics and Research

The 1989 NASA-ASEE Summer Faculty Fellowship Program at the Goddard Space Flight Center was conducted during 5 Jun. 1989 to 11 Aug. 1989. The research projects were previously assigned. Work summaries are presented for the following topics: optical properties data base; particle acceleration; satellite imagery; telemetry workstation; spectroscopy; image processing; stellar spectra; optical radar; robotics; atmospheric composition; semiconductors computer networks; remote sensing; software engineering; solar flares; and glaciers.

Boroson, Harold R.↗

Reduced‐Order Modeling of Energetic Materials Using Physics‐Aware Recurrent Convolutional Neural Networks in a Latent Space (LatentPARC)

Physics-aware deep learning (PADL) has gained popularity for use in spatiotemporal dynamics simulations, such as those in computational modeling of energetic materials (EM). We show that the challenge PADL methods face while learning complex field evolution problems can be simplified and accelerated by decoupling it into two tasks: learning complex geometric features in evolving fields and modeling dynamics over these features in a lower-dimensional feature space. We build upon our previous work on physics-aware recurrent convolutional neural networks (PARC). PARC embeds knowledge of underlying physics into its neural network architecture for more robust and accurate prediction of evolving physical fields. PARC was shown to effectively learn complex nonlinear features such as the formation of hotspots and coupled shock fronts in various initiation scenarios of EMs, as a function of microstructures, serving effectively as a microstructure-aware burn model. Here, we further accelerate PARC and reduce its computational cost by projecting the original dynamics onto a lower-dimensional invariant manifold, or “latent space.” The projected latent representation encodes the complex geometry of evolving fields (e.g., temperature and pressure) in a set of data-driven features. The reduced dimension of this latent space allows us to learn the dynamics during the initiation of EM with a lighter and more efficient model. We observe a significant decrease in training and inference time while maintaining results comparable to PARC at inference. This work takes steps towards enabling rapid prediction of EM thermomechanics at larger scales and characterization of EM structure–property–performance linkages at a full application scale.

Mathematics and Computing↗

DPM: A deep learning PDE augmentation method with application to large-eddy simulation

A framework is introduced that leverages known physics to reduce overfitting in machine learning for scientific applications. The partial differential equation (PDE) that expresses the physics is augmented with a neural network that uses available data to learn a description of the corresponding unknown or unrepresented physics. Training within this combined system corrects for missing, unknown, or erroneously represented physics, including discretization errors associated with the PDE's numerical solution. For optimization of the network within the PDE, an adjoint PDE is solved to provide high-dimensional gradients, and a stochastic adjoint method (SAM) further accelerates training. Additionally, the approach is demonstrated for large-eddy simulation (LES) of turbulence. High-fidelity direct numerical simulations (DNS) of decaying isotropic turbulence provide the training data used to learn sub-filter-scale closures for the filtered Navier–Stokes equations. Out-of-sample comparisons show that the deep learning PDE method outperforms widely-used models, even for filter sizes so large that they become qualitatively incorrect. It also significantly outperforms the same neural network when a priori trained based on simple data mismatch, not accounting for the full PDE. Measures of discretization errors, which are well-known to be consequential in LES, point to the importance of the unified training formulation's design, which without modification corrects for them. For comparable accuracy, simulation runtime is significantly reduced. A relaxation of the typical discrete enforcement of the divergence-free constraint in the solver is also successful, instead allowing the DPM to approximately enforce incompressibility physics. Since the training loss function is not restricted to correspond directly to the closure to be learned, training can incorporate diverse data, including experimental data.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗