Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel time integration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

An Integrated High-performance Computing and Digital Real-time Simulation Testbed to Benchmark Closed-loop Load Shedding Algorithms in Power Systems

An integrated testbed using digital real-time simulator (DRTS) and a high-performance computing (HPC) cluster is presented here to compare speed and performance of computational schemes to mitigate time-critical issues in electric power systems. The first approach in this testbed validation is taken by running a set of closed-loop load shedding algorithms to compare and contrast two paradigms of arresting cascading failure propagation. Two algorithms involve solving DC and AC power flow model-based optimization problems to compute load shedding at different buses, while a model-based stochastic search using parallel computing provides a viable alternative. The algorithms are implemented in the DRTS-HPC testbed for the IEEE 14-bus benchmark transmission system. As a proof of the concept, simulation results are presented for implementation of closed-loop load-shedding algorithms for cascading failures in the DRTS-HPC testbed

24 POWER TRANSMISSION AND DISTRIBUTION↗

Simulating single-particle dynamics in magnetized plasmas: The RMF code

The RMF (Rotating Magnetic Field) code is designed to calculate the motion of a charged particle in a given electromagnetic field. It integrates Hamilton’s equations in cylindrical coordinates using an adaptive predictor-corrector double-precision variable-coefficient ordinary differential equation solver for speed and accuracy. RMF has multiple capabilities for the field. Particle motion is initialized by specifying the position and velocity vectors. Here, the six-dimensional state vector and derived quantities are saved as functions of time. A post-processing graphics code, XDRAW, is used on the stored output to plot up to 12 windows of any two quantities using different colors to denote successive time intervals. Multiple cases of RMF may be run in parallel and perform data mining on the results. Recent features are a synthetic diagnostic for simulating the observations of charge-exchange-neutral energy distributions and RF grids to explore a Fermi acceleration parallel to static magnetic fields.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Real-Time GPU-Accelerated OFDR With an Integrated Auxiliary Interferometer

A GPU-accelerated optical frequency domain reflectometry (OFDR) system with an improved integrated auxiliary interferometer is proposed. Unlike conventional approaches that require separate auxiliary interferometers and multiple detection channels, the proposed OFDR system embeds this functionality directly into the signal via an intentional beat component. This enables self-calibration of laser nonlinearity while maintaining a cost-effective hardware configuration. Building on this simplified configuration, the system leverages GPU acceleration with an NVIDIA RTX 4070 Ti to achieve real-time performance, delivering high-throughput signal processing for continuous OFDR interrogation. The signal processing pipeline comprises signal capture, resampling for nonlinearity compensation, and frequency shift computation, all optimized for parallel execution. Hardware benchmarking demonstrates substantial acceleration over CPU implementations, achieving up to a 45× speedup for resampling and frequency shift computations and enabling processing latencies below 30 ms. Thermal response validation is conducted under two complementary scenarios: localized heating using a water bath and cryogenic-temperature conditions using liquid nitrogen. Under localized heating, the system achieves an accuracy of 0.249 °C with a thermal sensitivity of 5.971 GHz/°C, while cryogenic-temperature validation demonstrates a frequency shift response with a sensitivity of 2.383 GHz/°C and an accuracy of 2.04 °C. The high acceleration of the proposed GPU-accelerated OFDR system and its accuracy are achieved by exploiting CUDA-based stride indexing, enabling efficient parallel segmentation and processing of large datasets without additional memory copies. The benchmarking results confirm the robustness, accuracy, and deployability of the proposed OFDR system across a wide temperature range, establishing it as a practical platform for real-time distributed fiber sensing in structurally dynamic environments.

Harb, Salah [Lawrence Berkeley National Laboratory↗

Femtojoule optical nonlinearity for deep learning with incoherent illumination

Optical neural networks (ONNs) are a promising computational alternative for deep learning due to their inherent massive parallelism for linear operations. However, the development of energy-efficient and highly parallel optical nonlinearities, a critical component in ONNs, remains an outstanding challenge. Here, we introduce a nonlinear optical microdevice array (NOMA) compatible with incoherent illumination by integrating the liquid crystal cell with silicon photodiodes at the single-pixel level. We fabricate NOMA with more than half a million pixels, each functioning as an optical analog of the rectified linear unit at ultralow switching energy down to 100 femtojoules per pixel. With NOMA, we demonstrate an optical multilayer neural network. Our work holds promise for large-scale and low-power deep ONNs, computer vision, and real-time optical image processing.

36 MATERIALS SCIENCE↗

One-shot omnidirectional pressure integration through matrix inversion

In this work, we present a method to perform 2D and 3D omnidirectional pressure integration from velocity measurements with a single-iteration matrix inversion approach. This work builds upon our previous work, where the rotating parallel ray approach was extended to the limit of infinite rays by taking continuous projection integrals of the ray paths and recasting the problem as an iterative matrix inversion problem. This iterative matrix equation is now 'fast-forwarded' to the 'infinity' iteration, leading to a different matrix equation that can be solved in a single step, thereby presenting the same computational complexity as the Poisson equation. We observe computational speedups of ~10 6 when compared to brute-force omnidirectional integration methods, enabling the treatment of grids of ~10 9 points and potentially even larger in a desktop setup at the time of publication. Further examination of the boundary conditions of our one-shot method shows that omnidirectional pressure integration implements a boundary condition where the boundary points are treated as interior points to the extent that information is available. Finally, we show how the method can be extended from the regular grids typical of particle image velocimetry to the unstructured meshes characteristic of particle tracking velocimetry data.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

New insights on divertor parallel flows,E × B drifts, and fluctuations from in situ, two-dimensional probe measurement in the Tokamak à Configuration Variable

Abstract In situ , two-dimensional (2D) Langmuir probe measurements across a large part of the TCV outer divertor are reported in L-mode discharges with and without divertor baffles. This provides detailed insights into time averaged profiles, particle fluxes, and fluctuation behavior in different divertor regimes. The presence of the baffles is shown to substantially increase the divertor neutral pressure for a given upstream density and to facilitate the access to detachment, an effect that increases with plasma current. The detailed, 2D probe measurements allow for a divertor particle balance, including ion flux contributions from parallel flows and E × B drifts. The poloidal flux contribution from the latter is often comparable or even larger than the former, and the divertor parallel flow direction reverses in some conditions, pointing away from the target. In most conditions, the integrated particle flux at the outer target can be predominantly ascribed to ionization along the outer divertor leg, consistent with a closed-box approximation of the divertor. The exception is a strongly detached divertor, achieved here only with baffles, where the total poloidal ion flux even decreases towards the outer target, indicative of significant plasma recombination. The most striking observation from relative density fluctuation measurements along the outer divertor leg is the transition from poloidally uniform fluctuation levels in attached conditions to fluctuations strongly peaking near the X-point when approaching detachment.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Post-puff SOL broadening on MAST-U under high-recycling conditions: evidence consistent with cross-field transport changes

Transient broadening of the scrape-off layer (SOL) density profile can modify main-chamber first-wall particle fluxes and divertor loading, yet its control parameters remain debated between divertor-regime transitions, neutral dynamics and changes in cross-field transport. We investigate fueling-driven SOL density-profile evolution and post-fueling relaxation on MAST-Upgrade (MAST-U) in ohmic L-mode, using two otherwise similar double-null Conventional Divertor discharges (I p = 450 kA, B T = 0.33 T) with identical 50 ms low-field side gas puffs; the only intentional difference is the puff start time. Upstream Thomson scattering shows that both discharges develop a transient far-SOL density profile modification, expressed as an increased SOL-width metric λ n e and a far-SOL enhancement consistent with a shoulder-like signature. In the earlier-puff case, the SOL-width metric remains elevated after puff termination and the post-puff decay time scales are systematically longer across the analysed radii. Outer-target Langmuir probes indicate high-recycling conditions during the analysed window (few-eV T e with radially peaked j sat,∥ and q ∥ profiles, without signatures of deep detachment). The outer-divertor collisionality proxy Λ div is elevated in both cases and does not discriminate between the different post-puff persistence. The target-integrated ion flux proxy ∫ J sat dA evolves nearly identically in both discharges when aligned to the puff start. Taken together, these observations suggest that the late/post-puff upstream SOL evolution is not set by parallel exhaust to the outer targetalone and is more consistent with upstream cross-field redistribution, with a possible role for plasma–neutral coupling and fueling geometry.

MAST-U↗

Motion Planning Algorithms for Safety and Quantum Computing Efficiency

Motion planning remains a fundamental problem in robotics. Sampling-based algorithms use randomization to allow efficient solutions to this complex problem. As mobile robots and autonomous vehicles become more prevalent in everyday life, motion planning must be applied to increasingly challenging scenarios. Safety has become a paramount concern in motion planning for ensuring robotic applications enrich human lives. To date, many motion planning techniques to increase safety in the face of uncertain and dynamic environments have been developed. This dissertation first addresses distributional safety of Rapidly-Exploring Random Trees (RRT) through our algorithm W-Safe RRT. To acknowledge distributional uncertainty and poor modeling, W-Safe RRT uses the Wasserstein metric to provide a probabilistic bound on the distributional distance between a robot and obstacles. Human-interpretable environmental agent classification allows online safety margin adaptation. We propose and analyze an integrating region method for online classification that increases actor labeling accuracy based on behavioral feature values when compared to state of the art methods. The method performs class assignments based on local maximum likelihood in a created behavioral feature-space, allowing a notion of classification uncertainty. Model-based methods with safety guarantees can quickly become computationally in tractable, especially with multiple agents, higher dimensions, and plentiful unknowns. Sampling based algorithms have been parallelized for computation with multi-core computers and GPUs. We consider the use of quantum algorithms and computers for sampling-based motion planning for the first time. Quantum computing performs operations on superpositions of states and can solve certain problems much more efficiently than classical computers, but introduces previously unseen challenges. With Quantum-RRT, we recast the motion planning problem into a database-search structure and use Quantum Amplitude Amplification to find reachable states in the database with a quadratic performance increase over classical methods. We address two error sources with this method: quantum measurement and quantum oracle errors. We then extend this method to Parallel Quantum-RRT, which uses a manager-worker architecture with multiple parallel quantum workers to increase database search efficiency. We compare algorithm architectures and characterize probabilities of multiple workers finding solutions. Lastly, we test in simulation the quantum algorithms against classical versions in a wide variety of scenarios, concluding that a similar parallelization improvement is to be found in the quantum case as was found in the parallelization of classical RRT.

97 MATHEMATICS AND COMPUTING↗

rustpix

rustpix is a high-performance, open-source Rust library with first-class Python bindings (via PyO3) for processing pixel-detector data in neutron imaging. It targets time-stamping detectors such as Timepix3 (TPX3) at ORNL's Spallation Neutron Source (VENUS beamline), where each detected neutron deposits charge across a cluster of pixels within a very high-rate event stream (96M+ hits/sec). rustpix parses TPX3 event data in parallel using memory-mapped I/O, offers four interchangeable clustering algorithms (ABS adjacency-based search, DBSCAN, graph/union-find connected components, and a parallel grid method), and extracts weighted, super-resolved centroids to produce neutron-event lists. A streaming architecture lets it process files larger than available memory. rustpix is distributed as a pip-installable Python package (with NumPy integration), Rust crates, a command-line tool, and an interactive GUI; it writes HDF5, Apache Arrow, and CSV; and it is designed to extend to TPX4 and other detector types. Released as open-source under the MIT License.

Zhang, Chen [Oak Ridge National Laboratory (ORNL),↗

Three practical workflow schedulers for easy maximum parallelism

Runtime scheduling and workflow systems are an increasingly popular algorithmic component in HPC because they allow full system utilization with relaxed synchronization requirements. There are so many special-purpose tools for task scheduling, one might wonder why more are needed. Use cases seen on the Summit supercomputer needed better integration with MPI and greater flexibility in job launch configurations. Preparation, execution, and analysis of computational chemistry simulations at the scale of tens of thousands of processors revealed three distinct workflow patterns. A separate job scheduler was implemented for each one using extremely simple and robust designs: file-based, task-list based, and bulk-synchronous. Comparing to existing methods shows unique benefits of this work, including simplicity of design, suitability for HPC centers, short startup time, and well-understood per-task overhead. All three new tools have been shown to scale to full utilization of Summit, and have been made publicly available with tests and documentation. This work presents a complete characterization of the minimum effective task granularity for efficient scheduler usage scenarios. Here, these schedulers have the same bottlenecks, and hence similar task granularities as those reported for existing tools following comparable paradigms.

97 MATHEMATICS AND COMPUTING↗

Extreme-scale EV charging infrastructure planning for last-mile delivery using high-performance parallel computing

Here, this paper addresses stochastic charger location and allocation problems under queue congestion for last-mile delivery using electric vehicles (EVs). The objective is to decide where to open charging stations and how many chargers of each type to install, subject to budgetary and waiting-time constraints. We formulate the problem as a mixed-integer non-linear program, where each station-charger pair is modeled as a multiserver queue with stochastic arrivals and service times to capture the notion of waiting in fleet operations. The model is extremely large, with billions of variables and constraints for a typical metropolitan area; even loading the model in solver memory is difficult, let alone solving it. To address this challenge, we develop a Lagrangian-based dual decomposition framework that decomposes the problem by station and leverages parallelization on high-performance computing systems, where the subproblems are solved by using a cutting plane method and their solutions are collected at the master level. We also develop a three-step rounding heuristic to transform the fractional subproblem solutions into feasible integral solutions. Computational experiments on data from the Chicago metropolitan area with hundreds of thousands of households and thousands of candidate stations show that our approach produces high-quality solutions in cases where existing exact methods cannot even load the model in memory. We also analyze various policy scenarios, demonstrating that combining existing depots with newly built stations under multiagency collaboration substantially reduces costs and congestion. These findings offer a scalable and efficient framework for developing sustainable large-scale EV charging networks.

Capacity allocation↗

Virtual Self-Excited Induction Generator-Based Grid-Forming Inverter Control for Robust Voltage Regulation Under Nonideal Loading

This paper presents a generator-inspired control methodology for grid-forming (GFM) inverters that deliberately emulates a self-excited induction generator so that the inverter can hold its voltage and frequency under difficult loading and severe terminal disturbances across wide voltage and frequency ranges. The design integrates a Lyapunov energy function-based inner loop to provide high bandwidth and strong disturbance rejection, and it complements this with a passivity-based argument that furnishes a coherent large-signal stability guarantee beyond small-signal limits. Analytical insights are developed via the Krylov-Bogoliubov-Mitropolsky averaging method, which reveals an intrinsic resistive droop characteristic; these closed-form relations both explain the observed dynamics and yield simple, decentralized tuning rules. The methodology is validated on a controller-hardware-in-the-loop platform and exercised in real time across balanced, unbalanced, and nonlinear loads, as well as during parallel operation. Across these scenarios, the inverter maintains balanced three-phase voltages, limits harmonic content, settles quickly with well-damped transients, and remains resilient when multiple units operate in parallel. The contributions are a self-excited-machine-inspired GFM controller with enhanced dynamic performance and robustness, a single stability rationale grounded in passivity, closed-form expressions that guide tuning, and comprehensive hardware-in-the-loop validations demonstrating effectiveness and superiority under challenging operating conditions.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Large-Scale Welding Process Simulation by GPU Parallelized Computing

The computational design of industrially relevant welded structures is extremely time consuming due to coupled physics and high nonlinearity. Previously, most welding distortion and residual stress simulations have been limited to small coupons and reduced order (from three-dimensional [3D] to two-dimensional [2D]), or inherent strain approximations were used for large structures. In this current study, an explicit finite element code based on a graphics processing unit was utilized to perform 3D transient thermomechanical simulation of structural components during welding. Laser brazing of aluminum alloy panels as representative of automotive manufacturing scenarios was simulated to predict out-of-plane distortion under different clamping conditions. The predicted deformation pattern and magnitude were validated by laser scanning data of physical assemblies. In addition, the code was used to investigate residual stresses developed during multipass arc welding of a nuclear industry pressurizer surge nozzle and subsequent welding repair where a 3D simulation was necessary. Taking the experimental data as reference, the 3D model predicted better residual stress distribution than a typical 2D asymmetrical model. Stress evolution in welding repair was also presented and discussed in this study. Furthermore, the efficient numerical model made it feasible to use integrated computational welding engineering to simulate welding processes for large-scale structures.

97 MATHEMATICS AND COMPUTING↗

Computer Science Research Needs for Parallel Discrete Event Simulation (PDES)

Historically, scientific computing efforts have demonstrated the clear need for, and effective use of, supercomputing with traditional time-stepped simulations. Nevertheless, there are several areas in the mission spaces of the U.S. Department of Energy and other agencies waiting to tap advanced computing research using a different, discrete event style of modeling, simulation, and analysis. These span a wide spectrum of applications including energy grid resilience, urban planning and policy, transportation science, building technologies, emergency response and planning, environmental impact analysis, computational epidemiology, Internet communications, cyber security, and cyber-physical systems, to name only a few. Even within traditional scientific applications, the role of discrete event modes of execution is increasing in the form of new event-based mathematical solvers such as quantized state integration methods and discrete-continuous hybrid system solvers. Co-design of advanced supercomputing hardware systems is another area that exploits discrete event simulation at its core for effective analyses. Complex systems, entity behaviors and interconnections play a significant role in all these applications, which are mapped to large-scale models with discrete event formulations.

97 MATHEMATICS AND COMPUTING↗

Integrated hydrogeophysical modelling and data assimilation for geoelectrical leak detection

Time-lapse electrical resistivity tomography (ERT) measurements provide indirect observations of hydrological processes in the Earth's shallow subsurface at high spatial and temporal resolution. ERT has been used in the past decades to detect leaks and monitor the evolution of associated contaminant plumes. Specifically, inverted resistivity images allow visualization of the dynamic changes in the structure of the plume. However, existing methods do not allow the direct estimation of leak parameters (e.g. leak rate, location, etc.) and their uncertainties. We propose an ensemble-based data assimilation framework that evaluates proposed hydrological models against observed time-lapse ERT measurements without directly inverting for the resistivities. Each proposed hydrological model is run through the parallel coupled hydro-geophysical simulation code PFLOTRAN-E4D to obtain simulated ERT measurements. The ensemble of model proposals is then updated using an iterative ensemble smoother. In this paper, we demonstrate the proposed framework on synthetic and field ERT data from controlled tracer injection experiments. Our results show that the approach allows joint identification of contaminant source location, initial release time, and solute loading from the cross-borehole time-lapse ERT data, alongside with an assessment of uncertainties in these estimates. We demonstrate a reduction in site-wide uncertainty by comparing the prior and posterior plume mass discharges at a selected image plane. This framework is particularly attractive to sites that have previously undergone extensive geological investigation (e.g., nuclear sites). It is well suited to complement ERT imaging and we discuss practical issues in its application to field problems.

58 GEOSCIENCES↗

Scalable simulation of coupled adsorption and transport of methane in confined complex porous media with density preconditioning

The growing significance of shales and tight formations in the transition to less carbon-intensive and clean energy drives the research endeavor to understand the physics of gas flow within these systems. However, shales are composed of massively heterogeneous physical and chemical features. Most nano-sized pores connect to millimeter-scale fractures, leading to multiscale transport. These nano-scale pore throats demonstrate non-classical flow behavior, such as non-negligible slip velocities and adsorbed gas layers at the boundary. As a result, classical computational fluid dynamics models do not capture the physics. In this work, we develop a coupling scheme for the multiple-relaxation-time (MRT) lattice Boltzmann (LB) method that integrates the Peng-Robinson equation of state into a pseudo-potential interaction model to capture the physics of methane flow in irregular networks of channels that represent nano-scale porous media. We use atomistic simulations to calibrate and validate our model in slit nano-channels. We propose a preconditioning scheme to initialize the coupled transport and adsorption simulation of methane in complex porous media. The results of this implementation of LB agree with Direct Simulation Monte Carlo (DSMC) and Molecular Dynamics (MD) simulations. We then scale up the LB implementation through vectorization and indirect addressing. We parallelize it using Message Passing Interface (MPI) and OpenMP frameworks to simulate transport and adsorption in complex media with a million lattices. Additionally, we analyze the differences between coupled and transport-only simulations in two case studies and show that considering phase behavior, i.e., adsorption, can significantly change the flow behavior. This work constitutes an important step towards bridging the gap between molecular flow and system-scale behavior of complex disordered porous media.

42 ENGINEERING↗

Benchmarking of massively parallel phase-field codes for directional solidification

We present a detailed benchmark comparing two state-of-the-art phase-field implementations for simulating alloy solidification under experimentally relevant conditions. The study investigates the directional solidification of Al-3wt%Cu under high-velocity solidification conditions and SCN-0.46wt% camphor under microgravity conditions from National Aeronautics and Space Administration (NASA) DECLIC-DSI-R experiments. Both codes, one employing finite-difference discretization with uniform mesh and GPU-acceleration (GPU-PF) and the other one employing finite-element discretization with adaptive-mesh and CPU-parallelization (PRISMS-PF), solve the same quantitative phase-field formulation that incorporates an anti-trapping current for the solidification of dilute alloys. We evaluate the predictions of each code for dendritic morphology, primary spacing, and tip dynamics in both 2D and 3D, as well as their numerical convergence and computational performance. While existing benchmark problems have primarily focused on simplified or small-scale simulations, they do not reflect the computational and modeling challenges posed by employing experimentally relevant time and length scales. Our results provide a practical framework for assessing phase-field code performance as well as validating and facilitating their application in integrated computational materials engineering (ICME) workflows that require integration with realistic experimental data.

36 MATERIALS SCIENCE↗

Graph-Learning-Assisted State and Event Tracking for Solar-Penetrated Power Grids with Heterogeneous Data Sources

Unlike transmission systems, distribution systems do not typically contain sufficient metering to enable real-time state estimation. The lack of sufficient real-time measurements prohibits accurate and timely monitoring of the state of distribution systems. As a result, control and optimal operation of distribution systems, especially those containing large numbers of renewable generation units are not possible without proper data and information about the current state of the system. The main motivation of this project is to address this shortcoming by developing an approach which provides “predicted” real-time measurements so that they can be used to execute a distribution system state estimator. Thus, the objective of the project is to make the distribution systems fully observable, such that the hosting capacity for solar generation can be accurately estimated, and unnecessary solar curtailments can be avoided. In order to accomplish this goal, the project investigated the use of a grid-model-informed machine learning (ML) tool which integrates heterogeneous data streams obtained from AMI meters, SCADA as well as PMU measurements and created synchronous measurement snapshots for the state estimator (SE); and developed a hybrid robust SE which provides not only accurate state estimates but also real-time feedback for the ML model refinement.

14 SOLAR ENERGY↗