Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel time integration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

CaTS: Integration of Geant4 and Opticks

CaTS [6]is an advanced example that is part of Geant4 since version 11.0. It demonstrates the use of Opticks to offload the simulation of optical photons to GPUs. Opticks interfaces with the Geant4 toolkit to collect all the necessary information to generate and trace optical photons, re-implements the optical physics processes to be run on the GPU, and automatically translates the Geant4 geometry into a GPU appropriate format. To trace the photons, Opticks uses NVIDIA OptiX®. In this report, we describe CaTS and the integration of Opticks with Geant4. We demonstrate that the generation and tracing of optical photons represents an ideal application to be offloaded to GPUs, fully utilizing the high degree of available parallelism. In a typical liquid argon TPC simulation, a speedup of several hundred times is observed compared to an equivalent simulation using single threaded Geant4.

Wenzel, Hans↗

User Interface Developed for Controls/CFD Interdisciplinary Research

The NASA Lewis Research Center, in conjunction with the University of Akron, is developing analytical methods and software tools to create a cross-discipline "bridge" between controls and computational fluid dynamics (CFD) technologies. Traditionally, the controls analyst has used simulations based on large lumping techniques to generate low-order linear models convenient for designing propulsion system controls. For complex, high-speed vehicles such as the High Speed Civil Transport (HSCT), simulations based on CFD methods are required to capture the relevant flow physics. The use of CFD should also help reduce the development time and costs associated with experimentally tuning the control system. The initial application for this research is the High Speed Civil Transport inlet control problem. A major aspect of this research is the development of a controls/CFD interface for non-CFD experts, to facilitate the interactive operation of CFD simulations and the extraction of reduced-order, time-accurate models from CFD results. A distributed computing approach for implementing the interface is being explored. Software being developed as part of the Integrated CFD and Experiments (ICE) project provides the basis for the operating environment, including run-time displays and information (data base) management. Message-passing software is used to communicate between the ICE system and the CFD simulation, which can reside on distributed, parallel computing systems. Initially, the one-dimensional Large-Perturbation Inlet (LAPIN) code is being used to simulate a High Speed Civil Transport type inlet. LAPIN can model real supersonic inlet features, including bleeds, bypasses, and variable geometry, such as translating or variable-ramp-angle centerbodies. Work is in progress to use parallel versions of the multidimensional NPARC code.

Source record↗

7.2 kV Three-Port SiC Single-Stage Current-Source Solid-State Transformer With 90 kV Lightning Protection

This article proposes a multiport modular single-stage current-source solid-state transformer (SST) for applications like photovoltaic, energy storage integration, electric vehicle fast charging, data center, etc. The 7.2 kV 50 kVA current-source SST consists of five input-series output-parallel modules, each based on 3.3 kV SiC reverse-blocking MOSFET-plus-diode modules. The proposed SST has some unique features. First, compared to the voltage-source or matrix converter-based SSTs, the current-source SST has a unique advantage of single-stage isolated AC/DC or AC/AC conversion with an inductive DC link, but no medium-voltage (MV) AC experiments have been reported. This article for the first time demonstrates MV AC current-source SST up to 7.5 kV peak. Second, the multiport SST has a buffer port for active power decoupling (APD) or energy storage integration. The double-line-frequency power ripple from single-phase AC grid normally results in a large capacitor size in MV SSTs. The APD scheme is proposed in MV applications for the first time to enable a reduced DC link and the electrolytic capacitor-less SST with high reliability. Third, as a direct grid-connected converter without line-frequency transformer, insulation and protection are critical. A medium-frequency transformer design passes 55 kV basic-insulation level (BIL) and 60 kV high potential dielectrics withstand test with only 0.09% leakage inductance. Importantly, a lightning protection scheme is presented to protect the SST itself from 90 kV BIL impulse. Fourth, the proposed current-source SST topology is a modular soft-switching solid-state transformer (M-S4T) with full-range zero-voltage switching and controlled dv/dt for low electromagnetic interference. Furthermore, these concepts are verified in a three-port M-S4T prototype with forced oil cooling under single-module, stacked-module, steady-state, and dynamic operations.

14 SOLAR ENERGY↗

Implementing machine learning methods on QICK hardware for qubit readout & control

Quantum readout and control is a fundamental aspect of quantum computing that requires accurate measurement of qubit states. Errors emerge in all stages, from initialization to readout, and identifying errors in post-processing necessitates resource-intensive statistical analysis. In our work, we use a lightweight fully-connected neural network (NN) to classify states of a transmon system with no prior processing. Our NN accelerator yields higher fidelities (92%) than the classical matched filter method (84%). By exploiting the natural parallelism of NNs and their placement near the source of data on field-programmable gate arrays (FPGAs), we can achieve ultra-low latency on the Quantum Instrumentation Control Kit (QICK). Integrating machine learning methods on QICK opens several pathways for efficient real-time processing of quantum circuits.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Peak holding circuit for extremely narrow pulses

An improved pulse stretching circuit comprising: a high speed wide-band amplifier connected in a fast charge integrator configuration; a holding circuit including a capacitor connected in parallel with a discharging network which employs a resistor and an FET; and an output buffer amplifier. Input pulses of very short duration are applied to the integrator charging the capacitor to a value proportional to the input pulse amplitude. After a predetermined period of time, conventional circuitry generates a dump pulse which is applied to the gate of the FET making a low resistance path to ground which discharges the capacitor. When the dump pulse terminates, the circuit is ready to accept another pulse to be stretched. The very short input pulses are thus stretched in width so that they may be analyzed by conventional pulse height analyzers.

Oneill, R. W.↗

Route planning in a four-dimensional environment

Robots must be able to function in the real world. The real world involves processes and agents that move independently of the actions of the robot, sometimes in an unpredictable manner. A real-time integrated route planning and spatial representation system for planning routes through dynamic domains is presented. The system will find the safest most efficient route through space-time as described by a set of user defined evaluation functions. Because the route planning algorthims is highly parallel and can run on an SIMD machine in O(p) time (p is the length of a path), the system will find real-time paths through unpredictable domains when used in an incremental mode. Spatial representation, an SIMD algorithm for route planning in a dynamic domain, and results from an implementation on a traditional computer architecture are discussed.

Slack, M. G.↗

Element-topology-independent preconditioners for parallel finite element computations

A family of preconditioners for the solution of finite element equations are presented, which are element-topology independent and thus can be applicable to element order-free parallel computations. A key feature of the present preconditioners is the repeated use of element connectivity matrices and their left and right inverses. The properties and performance of the present preconditioners are demonstrated via beam and two-dimensional finite element matrices for implicit time integration computations.

Park, K. C.↗

New Tools for Automating Arcjet Sample Recession Tracking and Analysis

Arcjet Computer Vision (arcjetCV) has been significantly upgraded to enhance accuracy and performance in tracking material recession and shock-material standoff in test videos. These improvements include integrating new machine learning models, developing a specialized edge detection class, and incorporating a more comprehensive training dataset. These upgrades have refined the software’s ability to automate time-resolved recession tracking, making it more precise and reliable for analyzing complex physical processes. In parallel, a new tool called STARscan (Spatial Targeting and Alignment Rig for Scanning) is being developed to capture detailed 3D surface data before and after testing. By comparing these pre- and post-test scans with arcjetCV’s automated video analysis results, users can achieve a more comprehensive assessment of material recession. This method enables cross-validation of results, improving confidence in the analysis of tested materials. The expanded capabilities of arcjetCV have been successfully demonstrated on videos from various facilities, including the NASA Ames arcjets, UIUC’s PlasmatronX, and the VKI Plasmatron. It has been adopted as a new standard for in-situ recession tracking by the Mars Sample Return Project and Orion. ArcjetCV’s improved efficiency and accuracy are critical for reducing testing uncertainties and validating heatshield material performance under extreme conditions. The software’s user-friendly graphical interface ensures ease of use, enabling seamless processing and precise analysis of arcjet videos, providing deeper insights into material behavior in hypersonic environments. ArcjetCV is now available on both PyPI and Conda, allowing easy installation via "pip install arcjetCV" or through the Conda package manager, ensuring broad accessibility and streamlined deployment for users across various platforms.

Ablation↗

Revisiting the Soyuz-1 Parachute Failure in the Context of Safety in the Modern Era

The Soyuz‑1 accident remains one of the most consequential parachute related failures in human spaceflight history and provides enduring lessons for modern Entry, Descent, and Landing (EDL) system design. Occurring during the height of the Cold War and the Space Race, the mission unfolded under extraordinary political and schedule pressure as the Soviet Union sought to maintain its early leadership in space achievements following the death of chief designer Sergei Korolev. Despite unresolved propulsion, electrical, and parachute system deficiencies, Soyuz‑1 proceeded to launch and immediately encountered critical inflight anomalies, including a failed solar panel deployment, attitude control issues, and communication dropouts. Upon reentry, a malfunction in the parachute system, driven by a primary main canopy that failed to deploy, and subsequent entanglement of the reserve main canopy with the primary drogue parachute, resulted in insufficient deceleration and the fatal crash of cosmonaut Vladimir Komarov. Subsequent investigations revealed deep rooted cultural and organizational issues within the Soviet space program, including inadequate testing, suppression of dissent, undocumented last minute design changes, and the absence of integrated parachute system verification. More than 200 design flaws were identified after the accident, and firsthand accounts, including those from Yuri Gagarin, highlighted widespread concern prior to launch. Over time, the Soviet program implemented substantial reforms: systematic design corrections, rigorous process documentation, and an extensive series of drop tests that ultimately transformed the Soyuz system into one of the world’s most reliable human-rated return vehicles. This paper examines the technical architecture of the Soyuz‑1 parachute system, reconstructs the likely deployment sequence and failure mechanism, and analyzes the cultural contributors that shaped the accident. The study draws parallels to modern spacecraft parachute development, emphasizing the critical importance of integrated system testing, transparent engineering culture, and continuous hardware surveillance. These lessons remain directly relevant to today’s NASA and Commercial Crew Programs (CCP), where the Government continues to refine its understanding of aggregate risk and strengthen overall astronaut safety in the face of increasingly complex parachute systems.

Aaron L Morris↗

Computer Science Research Needs for Parallel Discrete Event Simulation (PDES)

Historically, scientific computing efforts have demonstrated the clear need for, and effective use of, supercomputing with traditional time-stepped simulations. Nevertheless, there are several areas in the mission spaces of the U.S. Department of Energy and other agencies waiting to tap advanced computing research using a different, discrete event style of modeling, simulation, and analysis. These span a wide spectrum of applications including energy grid resilience, urban planning and policy, transportation science, building technologies, emergency response and planning, environmental impact analysis, computational epidemiology, Internet communications, cyber security, and cyber-physical systems, to name only a few. Even within traditional scientific applications, the role of discrete event modes of execution is increasing in the form of new event-based mathematical solvers such as quantized state integration methods and discrete-continuous hybrid system solvers. Co-design of advanced supercomputing hardware systems is another area that exploits discrete event simulation at its core for effective analyses. Complex systems, entity behaviors and interconnections play a significant role in all these applications, which are mapped to large-scale models with discrete event formulations. To make advancements in all the aforementioned scientific areas, many technical aspects need to be more thoroughly studied and deeply understood in parallel discrete event simulation (PDES). The unique dynamics inherent in a discrete event modeling approach, by their very nature, intersect and influence the entire stack of the computing system, including (a) the unique nature of the instruction sets exercised in PDES workloads without a predominance of high-precision floating point operations, (b) virtual time-constrained multi-threaded execution of many logical processes per processor, (c) extremely variable and difficult to predict network traffic characteristics, (d) interfaces and inter-dependencies with machine learning and artificial intelligence codes at higher software layers, and (e) highly challenging load balancing needs, especially in effectively accounting for accelerated/extremely heterogeneous computing in current and future high-performance computing systems. Efficient and accurate parallel execution of PDES workloads is also dominated by challenges in dealing with their asynchronous concurrency fundamentally present at the model level. Conservative synchronization, optimistic/speculative synchronization, and their hybrid schemes open new questions in fundamental computer science with respect to reversibility of computation and prediction (lookahead) of behaviors inherent within model codes. On the implementation front, there are relatively few scalable, general-purpose parallel discrete event simulators in the world, and even fewer have been studied on emerging hardware platforms. To enable scientific advances using PDES, the research needs in computer science must also be pursued and met in the intersection of the algorithmic and hardware-aware aspects of scalable PDES engines. This report is aimed at capturing a computer science-oriented view of this important area of research in PDES, presenting a sample of important applications with their inherent discrete event technology elements. Needs are outlined in core areas of parallel discrete event research as well as cross-cutting directions in computer science research that positively impact scientific advancements across several important application areas. A selection of priority research opportunities in advanced computing for PDES is identified to serve as reference for key research topics and their order of importance for scientific advancements.

97 MATHEMATICS AND COMPUTING↗

Model-based, in-situ, non-destructive qualification and certification of parts made by autonomous additive manufacturing

To address the significant productivity challenges associated with the qualification and certification (Q&C) tasks of additively manufactured (AM) parts, which have traditionally relied on rigorous post‐build inspection and testing, we propose an integrated framework that combines model‐based qualification and certification (MBQ&C) with autonomous additive manufacturing (AAM). MBQ&C employs high‐fidelity predictive models, developed within the Integrated Computational Materials Engineering (ICME) paradigm, to simulate process–structure–property–performance relationships for assessing a part’s fitness for use. Since predictive models are commonly machine learning (ML)-based or reduced-order surrogates of validated physics models, they run efficiently, enabling timely inference. In parallel, the self-driving AAM utilises ML-based adaptive, closed‐loop control strategies to avoid, mitigate, or repair defects and anomalies during fabrication, thereby increasing the likelihood of producing acceptable parts. A key feature of the combined AAM-MBQ&C framework is that predictive models explicitly incorporate defects or anomalies that persist after the build, using instance-specific data captured via in-situ sensing. This customisation enables a build‐specific assessment of fitness for use, rather than relying on nominal or generic parameters. Such individualised evaluation provides a robust basis for Q&C-related acceptance decisions relating to each build. Additionally, the rapid solution capabilities of ML or reduced-order models enable the determination of a part’s suitability for service shortly after build completion. As the framework matures, it has the potential to substantially reduce reliance on conventional point‐design approaches—such as time‐consuming post‐build computed tomography scanning and costly destructive testing. Thus, the AAM-MBQ&C framework represents a transformative, scalable strategy for quality assurance of AM components, as parts produced within a stable, validated, and certified envelope can be certified with reduced testing. Key benefits include: (1) significant gains in Q&C productivity through efficient, model-centric assessment; (2) performance-based classification of defects into critical and non-critical categories; (3) the ability to predict potential deviations in the performance of parts affected by real-time, adaptive process control interventions relative to those produced under a certified process, and (4) the enabling of virtual Q&C for service environments that are difficult, hazardous, or impractical to access or reproduce experimentally. Collectively, these capabilities strengthen the business case for AM, particularly for high‐consequence and mission‐critical applications. Finally, although this work focuses on powder-based AM, the proposed techniques could be extended to AM processes employing alternative feedstock forms.

Gunasegaram, Dayalan↗

Examination of Semi-Analytical Solution Methods in the Coarse Operator of Parareal Algorithm for Power System Simulation

With continuing advances in high-performance parallel computing platforms, parallel algorithms have become powerful tools for development of faster than real-time power system dynamic simulations. In particular, it has been demonstrated in recent years that parallel-in-time (Parareal) algorithms have the potential to achieve such an ambitious goal. Here, the selection of a fast and reasonably accurate coarse operator of the Parareal algorithm is crucial for its effective utilization and performance. This paper examines semi-analytical solution (SAS) methods as the coarse operators of the Parareal algorithm and explores performance of the SAS methods to the standard numerical time integration methods. Two promising time-power series-based SAS methods were considered; Adomian decomposition method and Homotopy analysis method with a windowing approach for improving the convergence. Numerical performance case studies on 10-generator 39-bus system and 327-generator 2383-bus system were performed for these coarse operators over different disturbances, evaluating the number of Parareal iterations, computational time, and stability of convergence. All the coarse operators tested with different scenarios have converged to the same corresponding true solution (if they are convergent) and the SAS methods provide comparable computational speed, while having more stable convergence to the true solution in many cases.

97 MATHEMATICS AND COMPUTING↗

Practical Implementation of GPU-based Computing at the Grid Edge for Resilience Scenarios

This paper presents a practical implementation of GPU-accelerated computing at the grid edge to enhance power system resilience through next-generation smart meters. Advanced Metering Infrastructure (AMI) systems rely predominantly on centralized processing architectures, which limit real-time response capabilities during grid disturbances. This work proposes the integration of GPU-enabled computational platforms directly within smart meter to enable local execution support for power system analytics, fault detection algorithms, and optimization routines. The proposed framework uses the Julia programming language to leverage highperformance parallel computing capabilities while maintaining code portability and development efficiency. We use two experimental scenarios to benchmark the computational feasibility of this approach: sparse linear system solutions representative of power flow analyses, and multi-stage production cost simulations incorporating unit commitment and economic dispatch operations. Results demonstrate that computationally intensive power system algorithms, such as those supporting resilience scenario calculations, can be effectively executed at the distribution edge using commercially available embedded GPU hardware. Keywords—GPU acceleration, edge computing, smart meters, grid resilience, AMI, resilience.

De Souza, Reubun [School of Electrical Engineering↗

ANALYSIS OF THE MSL/MEDLI ENTRY DATA WITH COUPLED CFD AND MATERIAL RESPONSE.

The Mars Science Laboratory (MSL) was protected during its atmospheric entry by an instrumented heat-shield using NASA's Phenolic Impregnated Carbon Ablator (PICA) material. PICA is a lightweight carbon fiber/polymeric resin material that offers out-standing performances for protecting probes during planetary entry. The Mars Entry Descent and Landing Instrument (MEDLI) suite on MSL offers unique in-flight validation data for models of material response and atmospheric entry. MEDLI recorded, among other things, time-resolved in-depth temperature data of PICA using thermocouple sensors assembled in the MEDLI Integrated Sensor Plugs (MISP). The objective of this work is to showcase and analyze the coupling between the material response and the aerothermal environment. As shown in Figure 1, the workflow is divided into the following steps. First, the aerothermal properties are computed in the Data Parallel Line Relaxation (DPLR) code [3] and used with the Nonequilibrium air radiation (NEQAIR) program [8] to compute radiative heating. Second, the thermal response inside the material is computed in the Porous material Analysis Toolbox based on Open-FOAM (PATO) using a fixed blowing correction parameter. Third, the pyrolysis gases computed in PATO are used as inputs to a blowing boundary condition within DPLR. Fourth, the new environment properties from DPLR are used in NEQAIR to provide an updated solution, then both the updated aerothermal environment and radiative heating are used in PATO without blowing correction. The third and fourth steps are then repeated until convergence in surface temperature is obtained. Convergence in the radiative heating is generally achieved before surface temperature, at which point the radiative heating is no longer updated. Char mass loss rates are forced to zero to produce a non-receding surface condition. For early time points in the trajectory, where flow around the MSL aeroshell is rarefied, the Direct Simulation Monte Carlo (DSMC) code, SPARTA, is used to compute the aerothermal environment. Iteration between PATO and SPARTA is not performed due to the computational cost of DSMC simulations. Preliminary results of the coupling between PATO and DPLR for the MSL heatshield atmospheric entry model are presented in Figures 2-4 at 65 seconds after entry interface. Figure 2 shows the surface temperature results from an uncoupled simulation in PATO with the blowing correction parameter applied (left) along with the coupled surface temperature after iteration (right). Figure 3 shows the surface temperature along the centerline from windward to leeward for easier comparison. Figure 4 shows the coupled and uncoupled pyrolysis gas blowing rate. Mars 2020 used a similar heatshield consisting of PICA for thermal protection during entry, descent, and landing. In preparation for Mars 2020 post-flight analysis, the predictive material response capability is benchmarked against flight data from MEDLI. This work represents an important milestone toward the development of validated predictive capabilities for designing thermal protection systems for planetary probes.

Mars Science Laboratory↗

CFD Research, Parallel Computation and Aerodynamic Optimization

During the last five years, CFD has matured substantially. Pure CFD research remains to be done, but much of the focus has shifted to integration of CFD into the design process. The work under these cooperative agreements reflects this trend. The recent work, and work which is planned, is designed to enhance the competitiveness of the US aerospace industry. CFD and optimization approaches are being developed and tested, so that the industry can better choose which methods to adopt in their design processes. The range of computer architectures has been dramatically broadened, as the assumption that only huge vector supercomputers could be useful has faded. Today, researchers and industry can trade off time, cost, and availability, choosing vector supercomputers, scalable parallel architectures, networked workstations, or heterogenous combinations of these to complete required computations efficiently.

Ryan, James S.↗

PhytoOracle: Scalable, modular phenomics data processing pipelines

As phenomics data volume and dimensionality increase due to advancements in sensor technology, there is an urgent need to develop and implement scalable data processing pipelines. Current phenomics data processing pipelines lack modularity, extensibility, and processing distribution across sensor modalities and phenotyping platforms. To address these challenges, we developed PhytoOracle (PO), a suite of modular, scalable pipelines for processing large volumes of field phenomics RGB, thermal, PSII chlorophyll fluorescence 2D images, and 3D point clouds. PhytoOracle aims to ( i ) improve data processing efficiency; ( ii ) provide an extensible, reproducible computing framework; and ( iii ) enable data fusion of multi-modal phenomics data. PhytoOracle integrates open-source distributed computing frameworks for parallel processing on high-performance computing, cloud, and local computing environments. Each pipeline component is available as a standalone container, providing transferability, extensibility, and reproducibility. The PO pipeline extracts and associates individual plant traits across sensor modalities and collection time points, representing a unique multi-system approach to addressing the genotype-phenotype gap. To date, PO supports lettuce and sorghum phenotypic trait extraction, with a goal of widening the range of supported species in the future. At the maximum number of cores tested in this study (1,024 cores), PO processing times were: 235 minutes for 9,270 RGB images (140.7 GB), 235 minutes for 9,270 thermal images (5.4 GB), and 13 minutes for 39,678 PSII images (86.2 GB). These processing times represent end-to-end processing, from raw data to fully processed numerical phenotypic trait data. Repeatability values of 0.39-0.95 (bounding area), 0.81-0.95 (axis-aligned bounding volume), 0.79-0.94 (oriented bounding volume), 0.83-0.95 (plant height), and 0.81-0.95 (number of points) were observed in Field Scanalyzer data. We also show the ability of PO to process drone data with a repeatability of 0.55-0.95 (bounding area).

59 BASIC BIOLOGICAL SCIENCES↗

Application of integration algorithms in a parallel processing environment for the simulation of jet engines

The application of Predictor corrector integration algorithms developed for the digital parallel processing environment are investigated. The algorithms are implemented and evaluated through the use of a software simulator which provides an approximate representation of the parallel processing hardware. Test cases which focus on the use of the algorithms are presented and a specific application using a linear model of a turbofan engine is considered. Results are presented showing the effects of integration step size and the number of processors on simulation accuracy. Real time performance, interprocessor communication, and algorithm startup are also discussed.

Krosel, S. M.↗

Multidisciplinary propulsion simulation using NPSS

The current status of the Numerical Propulsion System Simulation (NPSS) program, a cooperative effort of NASA, industry, and universities to reduce the cost and time of advanced technology propulsion system development, is reviewed. The technologies required for this program include (1) interdisciplinary analysis to couple the relevant disciplines, such as aerodynamics, structures, heat transfer, combustion, acoustics, controls, and materials; (2) integrated systems analysis; (3) a high-performance computing platform, including massively parallel processing; and (4) a simulation environment providing a user-friendly interface. Several research efforts to develop these technologies are discussed.

Claus, Russell W.↗