Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel time integration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Examination of Semi-Analytical Solution Methods in the Coarse Operator of Parareal Algorithm for Power System Simulation

With continuing advances in high-performance parallel computing platforms, parallel algorithms have become powerful tools for development of faster than real-time power system dynamic simulations. In particular, it has been demonstrated in recent years that parallel-in-time (Parareal) algorithms have the potential to achieve such an ambitious goal. Here, the selection of a fast and reasonably accurate coarse operator of the Parareal algorithm is crucial for its effective utilization and performance. This paper examines semi-analytical solution (SAS) methods as the coarse operators of the Parareal algorithm and explores performance of the SAS methods to the standard numerical time integration methods. Two promising time-power series-based SAS methods were considered; Adomian decomposition method and Homotopy analysis method with a windowing approach for improving the convergence. Numerical performance case studies on 10-generator 39-bus system and 327-generator 2383-bus system were performed for these coarse operators over different disturbances, evaluating the number of Parareal iterations, computational time, and stability of convergence. All the coarse operators tested with different scenarios have converged to the same corresponding true solution (if they are convergent) and the SAS methods provide comparable computational speed, while having more stable convergence to the true solution in many cases.

97 MATHEMATICS AND COMPUTING↗

Practical Implementation of GPU-based Computing at the Grid Edge for Resilience Scenarios

This paper presents a practical implementation of GPU-accelerated computing at the grid edge to enhance power system resilience through next-generation smart meters. Advanced Metering Infrastructure (AMI) systems rely predominantly on centralized processing architectures, which limit real-time response capabilities during grid disturbances. This work proposes the integration of GPU-enabled computational platforms directly within smart meter to enable local execution support for power system analytics, fault detection algorithms, and optimization routines. The proposed framework uses the Julia programming language to leverage highperformance parallel computing capabilities while maintaining code portability and development efficiency. We use two experimental scenarios to benchmark the computational feasibility of this approach: sparse linear system solutions representative of power flow analyses, and multi-stage production cost simulations incorporating unit commitment and economic dispatch operations. Results demonstrate that computationally intensive power system algorithms, such as those supporting resilience scenario calculations, can be effectively executed at the distribution edge using commercially available embedded GPU hardware. Keywords—GPU acceleration, edge computing, smart meters, grid resilience, AMI, resilience.

De Souza, Reubun [School of Electrical Engineering↗

PhytoOracle: Scalable, modular phenomics data processing pipelines

As phenomics data volume and dimensionality increase due to advancements in sensor technology, there is an urgent need to develop and implement scalable data processing pipelines. Current phenomics data processing pipelines lack modularity, extensibility, and processing distribution across sensor modalities and phenotyping platforms. To address these challenges, we developed PhytoOracle (PO), a suite of modular, scalable pipelines for processing large volumes of field phenomics RGB, thermal, PSII chlorophyll fluorescence 2D images, and 3D point clouds. PhytoOracle aims to ( i ) improve data processing efficiency; ( ii ) provide an extensible, reproducible computing framework; and ( iii ) enable data fusion of multi-modal phenomics data. PhytoOracle integrates open-source distributed computing frameworks for parallel processing on high-performance computing, cloud, and local computing environments. Each pipeline component is available as a standalone container, providing transferability, extensibility, and reproducibility. The PO pipeline extracts and associates individual plant traits across sensor modalities and collection time points, representing a unique multi-system approach to addressing the genotype-phenotype gap. To date, PO supports lettuce and sorghum phenotypic trait extraction, with a goal of widening the range of supported species in the future. At the maximum number of cores tested in this study (1,024 cores), PO processing times were: 235 minutes for 9,270 RGB images (140.7 GB), 235 minutes for 9,270 thermal images (5.4 GB), and 13 minutes for 39,678 PSII images (86.2 GB). These processing times represent end-to-end processing, from raw data to fully processed numerical phenotypic trait data. Repeatability values of 0.39-0.95 (bounding area), 0.81-0.95 (axis-aligned bounding volume), 0.79-0.94 (oriented bounding volume), 0.83-0.95 (plant height), and 0.81-0.95 (number of points) were observed in Field Scanalyzer data. We also show the ability of PO to process drone data with a repeatability of 0.55-0.95 (bounding area).

59 BASIC BIOLOGICAL SCIENCES↗

Jacobian-free Newton–Krylov method for the simulation of non-thermal plasma discharges with high-order time integration and physics-based preconditioning

A preconditioning framework for the numerical simulation of non-thermal streamer discharges is developed using the Jacobian-free Newton-Krylov (JFNK) method. A reduced plasma fluid model is considered, consisting of electrons, one positive ion, one negative ion, and the electrostatic potential. Here, the plasma kinetics model includes ionization, electron-ion recombination, electron attachment, electron detachment, and ion-ion recombination. The governing equations are made dimensionless, discretized in space with finite differences, and integrated in time with a fully implicit method based on high-order backward differentiation formulas. The preconditioning framework is based on a linearized form of the governing equations and physics-based operator splitting. The efficiency of the preconditioning strategy is assessed through two test cases: streamer propagation between parallel plates and an axisymmetric pin-to-pin discharge. The fully implicit approach overcomes traditional restrictions in the time step size due to processes such as electron drift, electron diffusion, and dielectric relaxation. Excellent performance is observed through relevant statistics of the JFNK solver, although the number of linear iterations increases for the pin-to-pin discharge when nonlinear numerical boundary conditions are imposed at the electrodes. Performance studies show scalability with O(100-1000) processors for O(10M) unknowns with ample room for optimization.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

The Peridigm Meshfree Peridynamics Code

Abstract Peridigm is a meshfree peridynamics code written in C++ for use on large-scale parallel computers. It was originally developed at Sandia National Laboratories and is currently managed as an open-source, community driven software project. Its primary features include bond-based, state-based, and non-ordinary state-based constitutive models, bond failure laws, contact, and support for explicit and implicit time integration. To date, Peridigm has been used primarily by methods developers focused on solid mechanics and material failure. Peridigm utilizes foundational software components from Sandia’s Trilinos project and was designed for extensibility. This paper provides an overview of the solution methods implemented in Peridigm , a discussion of its software infrastructure, and demonstrates the use of Peridigm for the solution of several example problems.

97 MATHEMATICS AND COMPUTING↗

xesn: Echo state networks powered by Xarray and Dask

Xesn is a Python package that allows scientists to easily design Echo State Networks (ESNs) for forecasting problems. ESNs are a Recurrent Neural Network architecture introduced by Jaeger (2001) that are part of a class of techniques termed Reservoir Computing. One defining characteristic of these techniques is that all internal weights are determined by a handful of global, scalar parameters, thereby avoiding problems during backpropagation and reducing training time significantly. Because this architecture is conceptually simple, many scientists implement ESNs from scratch, leading to questions about computational performance. Xesn offers a straightforward, standard implementation of ESNs that operates efficiently on CPU and GPU hardware. The package leverages optimization tools to automate the parameter selection process, so that scientists can reduce the time finding a good architecture and focus on using ESNs for their domain application. Importantly, the package flexibly handles forecasting tasks for out-of-core, multi-dimensional datasets, eliminating the need to write parallel programming code. Xesn was initially developed to handle the problem of forecasting weather dynamics, and so it integrates naturally with Python packages that have become familiar to weather and climate scientists such as Xarray (Hoyer & Hamman, 2017). However, the software is ultimately general enough to be utilized in other domains where ESNs have been useful, such as in signal processing (Jaeger & Haas, 2004).

97 MATHEMATICS AND COMPUTING↗

Optimization of the moderators in the STS preliminary design

This report details the results for an optimization of the dimensions of the moderators in the preliminary design of the Spallation Neutron Source Second Target Station (STS). This study uses the optimization algorithms of Dakota and an unstructured mesh model for the moderators in MCNP. More details on the unstructured mesh model and the automated mesh generation can be found in [3]. Parallel to this effort, the same moderator geometries have been optimized using a constructive solid geometry (CSG) MCNP model. More details on this model and its results can be found in [4]. Three optimal designs are selected for each moderator: one that is optimized for maximum peak brightness, one for maximum time-integrated brightness, and one for a combination of peak and time-integrated brightness. The backbone of the optimization work flow is provided by Dakota. For each set of design parameters requested by Dakota, a new solid geometry is automatically built in Creo and SpaceClaim, and subsequently exported to Attila4MC to generate an unstructured mesh geometry for MCNP. After the MCNP calculation is finished, the objective function (e.g., brightness metric) is returned to Dakota. After the new design has been evaluated, a result-file is written, and Dakota proposes the next set of design parameters to be evaluated. The loop continues until a specified convergence criterion has been met. The design parameters of the cylindrical (upper) moderator include the hydrogen radius, the premoderator thickness (top, bottom, radial), the beryllium radius and the horizontal position of the moderator. The crucial design choice is the hydrogen radius. A radius of 62 mm is shown to provide the maximum time-integrated brightness. The maximum peak brightness occurs with a radius of 40 mm. A combined (middle) design, which balances peak and time-integrated brightnesses, is obtained with a hydrogen radius of 50 mm. The premoderator thicknesses and the beryllium radius are slightly larger in the design optimized for time-integrated brightness than in the design optimized for peak brightness. The sensitivity to these two parameters is relatively small close to the optimal configurations. The hydrogen vessel and vacuum vessel wall thicknesses are dependent on the radius of the liquid hydrogen due to structural integrity requirements. The increased wall thicknesses for larger vessels significantly penalize the time-integrated brightness, with the maximum obtainable value reduced by more than 10% relative to earlier studies which used fixed vessel wall thicknesses. The impact of the variable wall thicknesses is much less for the peak brightness and combined brightness designs. The design parameters of the tube (lower) moderator selected for the optimization are the tube length, the annular premoderator thickness, the beryllium radius and the horizontal position of the moderator. The tube length is the crucial parameter and is chosen large (210 mm) and small (125 mm) in the designs optimized for time-integrated and peak brightness respectively. A combined optimal design has a tube length of 170 mm. The premoderator thickness and the beryllium radius are chosen larger in the design optimized for time-integrated brightness.

42 ENGINEERING↗

Neural Networks for Nuclear Reactions in MAESTROeX

We demonstrate the use of neural networks to accelerate the reaction steps in the MAESTROeX stellar hydrodynamics code. A traditional MAESTROeX simulation uses a stiff ODE integrator for the reactions; here, we employ a ResNet architecture and describe details relating to the architecture, training, and validation of our networks. Our customized approach includes options for the form of the loss functions, a demonstration that the use of parallel neural networks leads to increased accuracy, and a description of a perturbational approach in the training step that robustifies the model. We test our approach on millimeter-scale flames using a single-step, 3-isotope network describing the first stages of carbon fusion occurring in Type Ia supernovae. We train the neural networks using simulation data from a standard MAESTROeX simulation, and show that the resulting model can be effectively applied to different flame configurations. This work lays the groundwork for more complex networks, and iterative time-integration strategies that can leverage the efficiency of the neural networks.

79 ASTRONOMY AND ASTROPHYSICS↗

Effects of magnetic field assisted heat treatment on the microstructure and mechanical properties of Fe-0.63 %C alloy

This study investigates the influence of an applied magnetic field on the microstructural evolution and mechanical properties of hypoeutectoid steels subjected to heat treatment. Tensile tests and microstructural analysis were performed on samples processed under varying magnetic field strengths (0 T, 5 T, and 9 T) and different austenitization incubation times. The results indicate that the application of a magnetic field alters the fraction of proeutectoid ferrite phase without changing the cooling rates and heat treatment process. Additionally, pearlite microstructural features such as lamellar spacing and misorientation angles exhibit variations under different field strengths. While the pearlite nodule diameter remains largely unaffected, an increase in percentage elongation and strength is observed in the 5 T treated sample, attributed to changes in microstructural features with the magnetic field. Additionally, the percentage elongation is reduced in the samples heat treated with reduced austenitization incubation times. The study further demonstrates that low-angle misorientations increased in the samples taken parallel to the magnetic field direction, influencing the mechanical response. These findings suggest that applying a magnetic field during heat treatment provides an additional driving force for phase transformations, offering a manufacturing process for tailoring microstructures and optimizing mechanical properties. Moreover, integrating magnetic fields in heat treatment processes has potential benefits in energy efficiency.

High magnetic field↗

HPC-enabled computation of demand models at scale

The purpose of this project is to examine the energy impact of urban-scale traffic for the Los Angeles Basin by developing and implementing a scalable traffic assignment model. An energy optimization function will be posed and when integrated into the optimization code for travel assignment it can be mathematically proven to converge. The energy optimization function can then be compared to the typical travel time optimization that is traditionally used in traffic assignment models. The analysis will begin with static traffic assignment models with the routing for all origin and destinations computed in parallel on high performance computing facilities. Convergence of the numerical methods rely on the solution of convex programs (or extensions of these). This step will mostly consist of demonstrating the ability to parallelize the Frank Wolfe algorithm on various platforms.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

HPC4Mobilty w/ UCB

The purpose of this project is to examine the energy impact of urban-scale traffic for the Los Angeles Basin by developing and implementing a scalable traffic assignment model. An energy optimization function will be posed and when integrated into the optimization code for travel assignment it can be mathematically proven to converge. The energy optimization function can then be compared to the typical travel time optimization that is traditionally used in traffic assignment models. The analysis will begin with static traffic assignment models with the routing for all origin and destinations computed in parallel on high performance computing facilities. Convergence of the numerical methods rely on the solution of convex programs (or extensions of these). This step will mostly consist of demonstrating the ability to parallelize the Frank Wolfe algorithm on various platforms. This work will contribute to LBNL’s efforts to develop new processes, analytical tools, program designs, and business models to advance the state of the art in next-generation sustainable transportation solutions.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Insights From Dayflow: A Historical Streamflow Reanalysis Dataset for the Conterminous United States

Abstract Reconstructed historical streamflow time series can supplement limited streamflow gauge observations. However, there are common challenges of typical modeling approaches: process‐based hydrologic models can be data/computation‐intensive, and statistics‐based models can be region/stream‐specific. Here we present a nationally scalable modeling framework integrating the simulated runoff from the Variable Infiltration Capacity (VIC) model with the Routing Application for Parallel computatIon of Discharge (RAPID) routing model leveraging high‐performance computing. We demonstrate an efficient method of assimilating streamflow at US Geological Survey (USGS) streamflow monitoring sites using a simple hierarchical approach in the VIC‐RAPID framework. The result is a reconstructed 36‐year (1980–2015) daily and monthly streamflow dataset (Dayflow) at ∼2.7 million NHDPlusV2 stream reaches in the conterminous US (CONUS). We perform a comprehensive evaluation at 7,526 USGS sites and characterize their error statistics. The results demonstrate that 49% of the USGS sites demonstrate Kling–Gupta Efficiency (KGE) > 0.5 and 58% of the sites show percentage bias within ±20% for the daily naturalized streamflow. Streamflow data assimilation across CONUS shows an overall improvement over naturalized streamflow, notably in the western semiarid‐to‐arid regions. Comparison to other national and global streamflow reanalysis datasets such as the National Water Model and Global Reach‐scale A priori Discharge Estimates for SWOT demonstrates improved KGE, reduced bias, and directions for Dayflow improvements. Investigations of error statistics with key hydrologic, hydroclimatic, and geomorphologic basin characteristics reveal region‐specific patterns which may help improve future framework applications. Overall, Dayflow may enable a better understanding of hydrologic conditions in a changing environment, especially in locations currently not represented by streamflow monitoring networks.

54 ENVIRONMENTAL SCIENCES↗

Task Parallelism to Optimize Performance of Environmental Modeling Software

Climate modeling is an integral part of environmental research, from studying rare phenomena to predicting future climate trends. The need for more accurate models is only growing, but as climate modeling capabilities advance, existing workflows require optimization to recoup performance. A solution comes in the form of task parallelism, a novel programming capability that provides an opportunity for optimization at execution time by allowing tasks to be executed in parallel, reducing runtime significantly. Using Parsl, an intuitive and scalable parallel scripting library for Python, we implement task parallelism within support software to aid in the continuous advancement of climate modeling technology.

54 ENVIRONMENTAL SCIENCES↗

Rapid measurement of RH-dependent aerosol hygroscopic growth using a humidity-controlled fast integrated mobility spectrometer (HFIMS)

Abstract. The ability of aerosol particles to uptake water (hygroscopic growth) is an important determinant of aerosol optical properties and radiative effects. Aerosol hygroscopic growth is traditionally measured by humidified tandem differential mobility analyzers (HTDMA), in which size-selected dry particles are exposed to elevated relative humidity (RH), and the size distribution of humidified particles is subsequently measured using a scanning mobility particle sizer. As a scanning mobility particle sizer can measure only one particle size at a time, HTDMA measurements are time consuming, and ambient measurements are often limited to a single RH level. Pinterich et al. (2017b) showed that fast measurements of aerosol hygroscopic growth are possible using a humidity-controlled fast integrated mobility spectrometer (HFIMS). In HFIMS, the size distribution of humidified particles is rapidly captured by a water-based fast integrated mobility spectrometer (WFIMS), leading to a factor of ∼10 increase in measurement time resolution. In this study we present a prototype HFIMS that extends fast hygroscopic growth measurements to a wide range of atmospherically relevant RH values, allowing for more comprehensive characterizations of aerosol hygroscopic growth. A dual-channel humidifier consisting of two humidity conditioners in parallel is employed such that aerosol RH can be quickly stepped among different RH levels by sampling from alternating conditioners. The measurement sequence is also optimized to minimize the transition time between different particle sizes. The HFIMS is capable of measuring aerosol hygroscopic growth of six particle diameters under five RH levels ranging from 20 % to 85 % (30 separate measurements) every 25 min. The performance of this HFIMS is characterized and validated using laboratory-generated ammonium sulfate aerosol standards. Measurements of ambient aerosols are shown to demonstrate the capability of HFIMS to capture the rapid evolution of aerosol hygroscopic growth and its dependence on both size and RH.

54 ENVIRONMENTAL SCIENCES↗

DESIGN & DEVELOPMENT OF A SECURE SELF-LEVELING WIRELESS RECHARGING PLATFORM FOR AN AERIAL DRONE ON AN UNMANNED SURFACE VESSEL

The Design and Development of an automated recharging station for an aerial drone, onboard a small, unmanned surface vessel, is described. Drones require a landing surface that is level within five degrees of the surrounding terrain for repeated reliable landing and takeoff. System constraints and at-sea application necessitate a compact, lightweight, and secure solution. A passive self-leveling platform and an accompanying automated parallel-pusher drone restraint mechanism have been designed and fabricated to aid in achieving a level landing surface and holding the drone in place while it charges. The self-leveling mechanism has been analyzed and subjected to initial laboratory tests. The testing of the drone restraint mechanism to verify its weight capacity and closing time, and the integration of the platform with a custom conductive contact wireless charging pad are identified as future work. The resulting cohesive unit will be tested for performance optimization and implementation onboard the unmanned surface vehicle.

McKinney, Adriana↗

Testing and Analysis of Grid Forming Inverter Control for Achieving Resilient and Economic Operation of an Islanded Microgrid

This investigation examines the feasibility of operating a battery energy storage system (BESS) in parallel with synchronous generation by using grid forming (GFM) control in order to achieve frequency control objectives while mitigating increases to operating costs in the context of an islanded microgrid. The BESS GFM control system, which is based on conventional droop techniques, is modeled along with the overall microgrid using the Real Time Digital Simulator (RTDS) to allow for integration of genset controller hardware. A series of simulations are performed to test the voltage and frequency regulation capability of the BESS control system when the primary frequency regulating genset is tripped offline. The results of the simulations suggest that the GFM control scheme will successfully maintain frequency and voltage stability, which will enable operation without a back-up genset while not compromising the microgrid resiliency to contingencies.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

JWST Noise Floor. I. Random Error Sources in JWST NIRCam Time Series

James Webb Space Telescope (JWST) transmission and emission spectra will provide invaluable glimpses of transiting exoplanet atmospheres, including possible biosignatures. This promising science from JWST, however, will require exquisite precision and understanding of systematic errors that can impact the time series of planets crossing in front of and behind their host stars. Here, we provide estimates of the random noise sources affecting JWST Near-Infrared Camera (NIRCam) time-series data on the integration-to-integration level. We find that 1/f noise can limit the precision of grism time series for two groups (230–1000 ppm depending on the extraction method and extraction parameters), but will average down like the square root of N frames/reads. The current NIRCam grism time-series mode is especially affected by 1/f noise because its GRISMR dispersion direction is parallel to the detector fast-read direction, but could be alleviated in the GRISMC direction. Care should be taken to include as many frames as possible per visit to reduce this 1/f noise source: thus, we recommend the smallest detector subarray sizes one can tolerate, four output channels, and readout modes that minimize the number of skipped frames (RAPID or BRIGHT2). We also describe a covariance-weighting scheme that can significantly lower the contributions from 1/f noise as compared to sum extraction. We evaluate the noise introduced by pre-amplifier offsets, random telegraph noise, and high dark current resistor capacitor (RC) pixels and find that these are correctable below 10 ppm once background subtraction and pixel masking are performed. We explore systematic error sources in a companion paper.

79 ASTRONOMY AND ASTROPHYSICS↗

Multirate partitioned Runge–Kutta methods for coupled Navier–Stokes equations

Earth system models are complex integrated models of atmosphere, ocean, sea ice, and land surface. Coupling the components can be a significant challenge due to the difference in physics, temporal, and spatial scales. Further, this study explores multirate partitioned Runge-Kutta methods for the fluid-fluid interaction problem and demonstrates its parallel performance by using the PETSc library. We consider compressible Navier-Stokes equations with gravity coupled through a rigid-lid interface. Our large-scale numerical experiments reveal that multirate partitioned Runge-Kutta coupling schemes (1) can conserve total mass; (2) have second-order accuracy in time; and (3) provide favorable strong- and weak-scaling performance on modern computing architectures. We also show that the speedup factors of multirate partitioned Runge-Kutta methods match theoretical expectations over their base (single-rate) method.

54 ENVIRONMENTAL SCIENCES↗