Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel time integration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

In-depth analysis on parallel processing patterns for high-performance Dataframes

The Data Science domain has expanded monumentally in both research and industry communities during the past decade, predominantly owing to the Big Data revolution. Artificial Intelligence (AI) and Machine Learning (ML) are bringing more complexities to data engineering applications, which are now integrated into data processing pipelines to process terabytes of data. Typically, a significant amount of time is spent on data preprocessing in these pipelines, and hence improving its efficiency directly impacts the overall pipeline performance. The community has recently embraced the concept of Dataframes as the de-facto data structure for data representation and manipulation. However, the most widely used serial Dataframes today (R, pandas) experience performance limitations while working on even moderately large data sets. We believe that there is plenty of room for improvement by taking a look at this problem from a high-performance computing point of view. In a prior publication, we presented a set of parallel processing patterns for distributed dataframe operators and the reference runtime implementation, Cylon. In this paper, we are expanding on the initial concept by introducing a cost model for evaluating the said patterns. Furthermore, we evaluate the performance of Cylon on the ORNL Summit supercomputer.

97 MATHEMATICS AND COMPUTING↗

Stochastic evaluation of fourth-order many-body perturbation energies

A scalable, stochastic algorithm evaluating the fourth-order many-body perturbation (MP4) correction to energy is proposed. Three hundred Goldstone diagrams representing the MP4 correction are computer generated and then converted into algebraic formulas expressed in terms of Green’s functions in real space and imaginary time. They are evaluated by the direct (i.e., non-Markov, non-Metropolis) Monte Carlo (MC) integration accelerated by the redundant-walker and control-variate algorithms. The resulting MC-MP4 method is efficiently parallelized and is shown to display O(n 5.3 ) size-dependence of cost, which is nearly two ranks lower than the O(n 7 ) dependence of the deterministic MP4 algorithm. Furthermore, it evaluates the MP4/aug-cc-pVDZ energy for benzene, naphthalene, phenanthrene, and corannulene with the statistical uncertainty of 10 mE h (1.1% of the total basis-set correlation energy), 38 mE h (2.6%), 110 mE h (5.5%), and 280 mE h (9.0%), respectively, after about 10 9 MC steps.

74 ATOMIC AND MOLECULAR PHYSICS↗

Determining the tilt of the Raman laser beam using an optical method for atom gravimeters

The tilt of a Raman laser beam is a major systematic error in precision gravity measurement using atom interferometry. The conventional approach to evaluating this tilt error involves modulating the direction of the Raman laser beam and conducting time-consuming gravity measurements to identify the error minimum. In this work, we demonstrate a method to expediently determine the tilt of the Raman laser beam by transforming the tilt angle measurement into characterization of parallelism, which integrates the optical method of aligning the laser direction, commonly used in freely falling corner-cube gravimeters, into an atom gravimeter. A position-sensing detector (PSD) is utilized to quantitatively characterize the parallelism between the test beam and the reference beam, thus measuring the tilt precisely and rapidly. After carefully positioning the PSD and calibrating the relationship between the distance measured by the PSD and the tilt angle measured by the tiltmeter, we achieved a statistical uncertainty of less than 30 µrad in the tilt measurement. Furthermore, we compared the results obtained through this optical method with those from the conventional tilt modulation method for gravity measurement. The comparison validates that our optical method can achieve tilt determination with an accuracy level of better than 200 µrad, corresponding to a systematic error of 20 µGal in g measurement. This work has practical implications for real-world applications of atom gravimeters.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Dynamic Response of a Semiactive Suspension System with Hysteretic Nonlinear Energy Sink Based on Random Excitation by means of Computer Simulation

This paper aims to investigate the property and behavior of the hysteretic nonlinear energy sink (HNES) coupled to a half vehicle system which is a nine-degree-of-freedom, nonlinear, and semiactive suspension system in order to improve the ride comfort and increase the stability in shock mitigation by using the computer simulation method. The HNES model is a semiactive suspension device, which comprises the famous Bouc–Wen (B-W) model employed to describe the force produced by both the purely hysteretic spring and linear elastic spring of potentially negative stiffness connected in parallel, for the half vehicle system. Nine nonlinear motion equations of the half vehicle system are derived in terms of the seven displacements and the two dimensionless hysteretic variables, which are integrated numerically by employing the direct time integration method for studying both the variables of vertical displacements, velocities, accelerations, chassis pitch angle, and the ride comfort and driver safety, respectively, based on the bump and random road inputs of the pseudoexcitation method as excitation signal. Simulation results show that, compared with the HNES model and the magnetorheological (MR) model coupled to the half vehicle system, the ride comfort and stability have been evidently improved. A successful validation process has been performed, which indicated that both the ride comfort and driver safety properties of the HNES model coupled to half vehicle significantly improved.

Chen, Hui↗

Sub-kW Class Hall-Effect Thruster Power Processing Unit for Wide Output Range Applications

The National Aeronautics and Space Administration (NASA) Small Spacecraft Electric Propulsion (SSEP) project is maturing high-propellant throughput sub-kilowatt Hall-effect thruster technologies to enable small spacecraft deep space science and exploration missions with high delta-v requirements. In support of this effort, development of a power processing unit (PPU) capable of providing discharge power of up to 1 kW continues to be pursued at the NASA Glenn Research Center (GRC). Previous reported work included a successful integrated test of a scalable, modular breadboard discharge power supply with the NASA-H64M laboratory model Hall-effect thruster and presentation of notional designs for the various auxiliary power supplies needed for thruster operation. Since that time, auxiliary power supply designs have been completed and fabricated, with the cathode heater and keeper power supplies being successfully tested with a hollow cathode assembly (HCA) in the NASA GRC Vacuum Facility 56 (VF-56). The desire for a lower mass, higher efficiency, and more versatile PPU to maximize performance of power and mass-limited small spacecraft has led to the exploration of a discharge power supply based on a series-parallel (LCC) resonant topology. This topology has enabled the discharge power supply to operate over a wider output range at switching frequencies 4-5 times higher than previous design iterations. Simulation models of the topology have been developed and a breadboard of the topology has been fabricated and evaluated on both resistive loads and an integrated Hall thruster test. This paper will present collected performance and integrated test data from both the fabricated auxiliary and resonant discharge power supplies. Advantages of the resonant converter architecture over more traditional pulse-width modulated (PWM) techniques in Hall-effect thruster discharge power supply applications will also be described.

electric propulsion↗

Sub-kW Class Hall-Effect Thruster Power Processing Unit for Wide Output Range Applications

The National Aeronautics and Space Administration (NASA) Small Spacecraft Electric Propulsion (SSEP) project is maturing high-propellant throughput sub-kilowatt Hall-effect thruster technologies to enable small spacecraft deep space science and exploration missions with high delta-v requirements. In support of this effort, development of a power processing unit (PPU) capable of providing discharge power of up to 1 kW continues to be pursued at the NASA Glenn Research Center (GRC). Previous reported work included a successful integrated test of a scalable, modular breadboard discharge power supply with the NASA-H64M laboratory model Hall-effect thruster and presentation of notional designs for the various auxiliary power supplies needed for thruster operation. Since that time, auxiliary power supply designs have been completed and fabricated, with the cathode heater and keeper power supplies being successfully tested with a hollow cathode assembly (HCA) in the NASA GRC Vacuum Facility 56 (VF-56). The desire for a lower mass, higher efficiency, and more versatile PPU to maximize performance of power and mass-limited small spacecraft has led to the exploration of a discharge power supply based on a series-parallel (LCC) resonant topology. This topology has enabled the discharge power supply to operate over a wider output range at switching frequencies 4-5 times higher than previous design iterations. Simulation models of the topology have been developed and a breadboard of the topology has been fabricated and evaluated on both resistive loads and an integrated Hall thruster test. This paper will present collected performance and integrated test data from both the fabricated auxiliary and resonant discharge power supplies. Advantages of the resonant converter architecture over more traditional pulse-width modulated (PWM) techniques in Hall-effect thruster discharge power supply applications will also be described.

electric propulsion↗

New Features of the NEQAIR Radiation Code

The longest-lived code for predicting shock layer radiation, NEQAIR, is now in its 5th decade of service. Substantial changes to the code have been made over the previous decade, the most recent report of which was at the 5th Workshop on Radiation in High Temperature Gases in 2014, for the version referred to as NEQAIR14. This paper will review some of the improvements made to the NEQAIR code since then, which is now at v15.2. Some of these features are discussed briefly below. NEQAIR15 and subsequent versions have enabled parallel evaluation of multiple lines of sight. This is accomplished by utilizing the HDF5 file format and placing multiple lines into a single file, LOS.h5, which is used for both input and output. This approach enables straightforward parallel execution both over the number of lines of sight and the number of points per line. For large problems, runtime reduces linearly with the number of nodes deployed since each line is processed independently by a subset of MPI ranks. Three applications of the multi-line solver are discussed. The first has to do with performing loosely coupled radiation-flowfield solutions. In this case the computed absorption and emission coefficients are used to evaluate the total energy absorbed or emitted at each point, allowing evaluation of the volumetric source term in the flowfield. The second computation is for obtaining heat flux from nonuniform flows, which require integration over spherical co-ordinates. These are of particular interest for evaluating radiation on the vehicle backshell. This 3D option improves the angular integration scheme and allows adaptive line selection that together reduce the number of lines required by about an order of magnitude. The final application is for remote observation, which is essentially the 3D integration problem over a small solid angle. For all three of these computations, data can be stored in the HDF5 file which allows a NEQAIR run to be restarted when it times out, or to add atmospheric absorption or instrument scan functions. An additional level of parallelism is enabled in NEQAIR15.2 using GPU routines. The GPU parallelism has realized up to 8x speed-up when running on a single core but diminishes as CPU parallelism is increased. For running multi-line simulations, it may be easier to reserve a large number of CPU nodes than to obtain the number of GPU nodes required for similar performance. A GUI, known as NEQTPY, allows for reading and creating input files, running NEQAIR, and displaying results. A significant feature of NEQTPY is the ability to perform spectral fits to data. The fits can operate on a single line spectrum (radiance vs. wavelength) or a 3D input file with multiple columns of data. Other new features include improved constants, additional species, more detailed non-Boltzmann modelling, advanced user controls, the ability to read and calculate spectra from HITRAN datafiles, photodissociation and photoionization cross-sections. A “fast” automatic grid option may reduce the size and time of spectral calculations while still maintaining good accuracy for total heat flux.

Brett A Cruden↗

Deep Generative Models that Solve PDEs: Distributed Computing for Training Large Data-Free Models

Recent progress in scientific machine learning (SciML) has opened up the possibility of training novel neural network architectures that solve complex partial differential equations (PDEs). Several (nearly data free) approaches have been recently reported that successfully solve PDEs, with examples including deep feed forward networks, generative networks, and deep encoder-decoder networks. However, practical adoption of these approaches is limited by the difficulty in training these models, especially to make predictions at large output resolutions (≥1024×1024). Here we report on a software framework for data parallel distributed deep learning that resolves the twin challenges of training these large SciML models - training in reasonable time as well as distributing the storage requirements. Our framework provides several out of the box functionality including (a) loss integrity independent of number of processes, (b) synchronized batch normalization, and (c) distributed higher-order optimization methods. We show excellent scalability of this framework on both cloud as well as HPC clusters, and report on the interplay between bandwidth, network topology and bare metal vs cloud. We deploy this approach to train generative models of sizes hitherto not possible, showing that neural PDE solvers can be viably trained for practical applications. We also demonstrate that distributed higher-order optimization methods are 2-3× faster than stochastic gradient-based methods and provide minimal convergence drift with higher batch-size.

PDEs↗

Parallel decomposition methods for the solution of electromagnetic scattering problems

This paper contains a overview of the methods used in decomposing solutions to scattering problems onto coarse-grained parallel processors. Initially, a short summary of relevant computer architecture is presented as background to the subsequent discussion. After the introduction of a programming model for problem decomposition, specific decompositions of finite difference time domain, finite element, and integral equation solutions to Maxwell's equations are presented. The paper concludes with an outline of possible software-assisted decomposition methods and a summary.

Cwik, Tom↗

Spectral Analysis of Integrated Pressures on Patches with Unsteady Pressure-Sensitive Paint Measurements

The technique of Pressure-Sensitive Paint (PSP) is commonly used in the aerospace industry to measure surface pressures on the model of launch vehicles and airplanes in the wind tunnel test. Recent research has demonstrated that Unsteady Pressure-Sensitive Paint (uPSP) can be an essential tool for the assessment of the unsteady, aerodynamic phenomena. The work described in this paper is a part of NASA’s development of a new state-of-the-art uPSP capability in production wind tunnels. This paper describes the spectral analysis of integrated pressures on patches of the scale model of the Space Launch System (SLS) Block 1 cargo vehicle with the uPSP measurements, which were collected in the Ascent Transient Aerodynamics Test (ATAT) with the Unitary Plan Wind Tunnel 11-by-11-foot Transonic Wind Tunnel in September 2019 at NASA Ames Research Center. The patches are defined with x station values, indicating the position along length of the SLS vehicle, and azimuth angles of the scale model. For each patch, the polygons are determined from the surface cells of the grid of the model, clipped with the edges of the patch, and each of the polygons is divided into triangles. The integrated pressure of the patch is determined as the ratio of the sum of the forces on the triangles over the sum of the areas of the triangles. For each run of the test, the Cross Power Spectral Density (CPSD) and magnitude squared coherence are computed from the time series of the integrated pressures on the patches. The pressure integration is coded in C++ and the spectral analysis is coded in MATLAB. The results were generated with the execution of the compiled C++ and MATLAB codes in parallel on the NASA Pleiades supercomputer. The results of pressure integration and spectral analysis are presented in this paper, and the data consistency of the test is also demonstrated. Funding for this research was provided by the NASA Aerosciences Evaluation and Test Capabilities Project.

Pressure-Sensitive Paint↗

Controls Status of Fermilab's PIP-II Project

The Fermilab Proton Improvement Project II (PIP-II) is building a new Super Conducting Linear Accelerator (SCL) accelerating protons to 800 MeV for injection into the rest of the FNAL beam complex. Key progress since the last status report given at ICALEPCS includes the adoption of modern DevOps practices with continuous integration and GitOps-based deployments, commissioning of EPICS-based systems at the Cryomodule Test Facility, and integration of a Virtual Accelerator framework for application development ahead of installation. In parallel, web-based applications using Dart and Flutter have matured, providing secure, unified access to both EPICS and legacy ACNET data. Data acquisition and timing systems have also evolved. This paper presents the current state of controls, emphasizing these recent developments and outlining upcoming milestones as PIP-II approaches commissioning of its cryoplant in 2026 and the Warm Front End in 2027.

Crisp, D. B. [Fermilab]↗

Optimization and Experimental Validation of Annular Finned PCM-HX for a Domestic Hot Water Heater Application

The load profile for domestic water heating is time-dependent and can result in high energy demand during peak operating times. Shifting this peak load can have significant environmental and economic impacts. Phase change material (PCM)-based thermal energy storage (TES) is a potentially useful technology for peak load shifting in domestic hot water (DHW) applications thanks to its high latent heat and energy density. In this study, an annular finned-tube PCM-HX design concept was optimized for a load-shifting TES unit to meet the Department of Energy standard for a medium-usage DHW heater using a resistance-capacitance model (RCM) integrated with a Multi-Objective Genetic Algorithm. The optimized design comprised 70 identical annular finned-tube PCM-HX units connected in parallel and utilizing RT62HC as the PCM. A single PCM-HX unit was prototyped and tested in a vertically oriented setup with upward heat transfer fluid (HTF) flow. The hot water supply time was defined based on a cutoff temperature of 51.7°C. The as-designed mass flow rate (1.5 g/s) was tested to assess the performance of the prototyped PCM-HX unit for RCM validation. For the experimental investigation, RTD sensor bundles measured HTF temperature at the PCM-HX inlet and outlet, and a Coriolis flow meter accurately measured the HTF mass flow rate. The simulated discharging power underpredicted the experimental result by about 12%, and the simulated hot water supply time underpredicted the experimental result by approximately 13% for the as-designed mass flow rate (1.5 g/s). The average deviation of the hot water supply temperature between the experimental and RCM results during the complete PCM solidification process was 1.3 K for the as-designed mass flow rate. The overall good agreement between the experimental and RCM results provides confidence that computationally efficient models such as RCM can be utilized for design optimization of PCM-HXs.

42 ENGINEERING↗

Performance of an Optimized Eta Model Code on the Cray T3E and a Network of PCs

In the year 2001, NASA will launch the satellite TRIANA that will be the first Earth observing mission to provide a continuous, full disk view of the sunlit Earth. As a part of the HPCC Program at NASA GSFC, we have started a project whose objectives are to develop and implement a 3D cloud data assimilation system, by combining TRIANA measurements with model simulation, and to produce accurate statistics of global cloud coverage as an important element of the Earth's climate. For simulation of the atmosphere within this project we are using the NCEP/NOAA operational Eta model. In order to compare TRIANA and the Eta model data on approximately the same grid without significant downscaling, the Eta model will be integrated at a resolution of about 15 km. The integration domain (from -70 to +70 deg in latitude and 150 deg in longitude) will cover most of the sunlit Earth disc and will continuously rotate around the globe following TRIANA. The cloud data assimilation is supposed to run and produce 3D clouds on a near real-time basis. Such a numerical setup and integration design is very ambitious and computationally demanding. Thus, though the Eta model code has been very carefully developed and its computational efficiency has been systematically polished during the years of operational implementation at NCEP, the current MPI version may still have problems with memory and efficiency for the TRIANA simulations. Within this work, we optimize a parallel version of the Eta model code on a Cray T3E and a network of PCs (theHIVE) in order to improve its overall efficiency. Our optimization procedure consists of introducing dynamically allocated arrays to reduce the size of static memory, and optimizing on a single processor by splitting loops to limit the number of streams. All the presented results are derived using an integration domain centered at the equator, with a size of 60 x 60 deg, and with horizontal resolutions of 1/2 and 1/3 deg, respectively. In accompanying charts we report the elapsed time, the speedup and the Mflops as a function of the number of processors for the non-optimized version of the code on the T3E and theHIVE. The large amount of communication required for model integration explains its poor performance on theHIVE. Our initial implementation of the dynamic memory allocation has contributed to about 12% reduction of memory but has introduced a 3% overhead in computing time. This overhead was removed by performing loop splitting in some of the high demanding subroutines. When the Eta code is fully optimized in order to meet the memory requirement for TRIANA simulations, a non-negligeable overhead may appear that may seriously affect the efficiency of the code. To alleviate this problem, we are considering implementation of a new algorithm for the horizontal advection that is computationally less expensive, and also a new approach for marching in time.

Kouatchou, Jules↗

Enabling Modular Autonomous Feedback‐Loops in Materials Science through Hierarchical Experimental Laboratory Automation and Orchestration

Abstract Materials acceleration platforms (MAPs) operate on the paradigm of integrating combinatorial synthesis, high‐throughput characterization, automatic analysis, and machine learning. Within a MAP, one or multiple autonomous feedback loops may aim to optimize materials for certain functional properties or to generate new insights. The scope of a given experiment campaign is defined by the range of experiment and analysis actions that are integrated into the experiment framework. Herein, the authors present a method for integrating many actions within a hierarchical experimental laboratory automation and orchestration (HELAO) framework. They demonstrate the capability of orchestrating distributed research instruments that can incorporate data from experiments, simulations, and databases. HELAO interfaces laboratory hardware and software distributed across several computers and operating systems for executing experiments, data analysis, provenance tracking, and autonomous planning. Parallelization is an effective approach for accelerating knowledge generation provided that multiple instruments can be effectively coordinated, which the authors demonstrate with parallel electrochemistry experiments orchestrated by HELAO. Efficient implementation of autonomous research strategies requires device sharing, asynchronous multithreading, and full integration of data management in experimental orchestration, which to the best of the authors’ knowledge, is demonstrated for the first time herein.

36 MATERIALS SCIENCE↗

Supervised enhancer prediction with epigenetic pattern recognition and targeted validation

Enhancers are important non-coding elements, but they have traditionally been hard to characterize experimentally. The development of massively parallel assays allows the characterization of large numbers of enhancers for the first time. Here, we developed a framework using Drosophila STARR-seq to create shape-matching filters based on meta-profiles of epigenetic features. We integrated these features with supervised machine-learning algorithms to predict enhancers. We further demonstrated that our model could be transferred to predict enhancers in mammals. We comprehensively validated the predictions using a combination of in vivo and in vitro approaches, involving transgenic assays in mice and transduction-based reporter assays in human cell lines (153 enhancers in total). The results confirmed that our model can accurately predict enhancers in different species without re-parameterization. Finally, we examined the transcription factor binding patterns at predicted enhancers versus promoters. Here, we demonstrated that these patterns enable the construction of a secondary model that effectively distinguishes enhancers and promoters.

59 BASIC BIOLOGICAL SCIENCES↗

Application of high-performance computing to numerical simulation of human movement

We have examined the feasibility of using massively-parallel and vector-processing supercomputers to solve large-scale optimization problems for human movement. Specifically, we compared the computational expense of determining the optimal controls for the single support phase of gait using a conventional serial machine (SGI Iris 4D25), a MIMD parallel machine (Intel iPSC/860), and a parallel-vector-processing machine (Cray Y-MP 8/864). With the human body modeled as a 14 degree-of-freedom linkage actuated by 46 musculotendinous units, computation of the optimal controls for gait could take up to 3 months of CPU time on the Iris. Both the Cray and the Intel are able to reduce this time to practical levels. The optimal solution for gait can be found with about 77 hours of CPU on the Cray and with about 88 hours of CPU on the Intel. Although the overall speeds of the Cray and the Intel were found to be similar, the unique capabilities of each machine are better suited to different portions of the computational algorithm used. The Intel was best suited to computing the derivatives of the performance criterion and the constraints whereas the Cray was best suited to parameter optimization of the controls. These results suggest that the ideal computer architecture for solving very large-scale optimal control problems is a hybrid system in which a vector-processing machine is integrated into the communication network of a MIMD parallel machine.

NASA Discipline Musculoskeletal↗

Development of tailorable advanced blanket insulation for advanced space transportation systems

Two items of Tailorable Advanced Blanket Insulation (TABI) for Advanced Space Transportation Systems were produced. The first consisted of flat panels made from integrally woven, 3-D fluted core having parallel fabric faces and connecting ribs of Nicalon silicon carbide yarns. The triangular cross section of the flutes were filled with mandrels of processed Q-Fiber Felt. Forty panels were prepared with only minimal problems, mostly resulting from the unavailability of insulation with the proper density. Rigidizing the fluted fabric prior to inserting the insulation reduced the production time. The procedures for producing the fabric, insulation mandrels, and TABI panels are described. The second item was an effort to determine the feasibility of producing contoured TABI shapes from gores cut from flat, insulated fluted core panels. Two gores of integrally woven fluted core and single ply fabric (ICAS) were insulated and joined into a large spherical shape employing a tadpole insulator at the mating edges. The fluted core segment of each ICAS consisted of an Astroquartz face fabric and Nicalon face and rib fabrics, while the single ply fabric segment was Nicalon. Further development will be required. The success of fabricating this assembly indicates that this concept may be feasible for certain types of space insulation requirements. The procedures developed for weaving the ICAS, joining the gores, and coating certain areas of the fabrics are presented.

Calamito, Dominic P.↗