Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel time integration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20

Large-Scale Numerical Simulations of Human Motion

This paper examines the feasibility of using massively-parallel and vector-processing supercomputers to solve large-scale optimal control problems for human movement. Specifically, we compare the computational expense of determining the optimal controls for the single support phase of walking using a conventional serial machine (a Silicon Graphics Personal Iris 4D25 workstation), a MIMD parallel machine (an Intel iPSC/860 comprising 128 processors), and a parallel-vector-processing machine (a Cray Y-MP 8/864). With the human body modeled as a 14 degree-of-freedom linkage actuated by 46 musculotendinous units, computation of the optimal controls for walking could take up to 3 months of CPU time on the Iris. Both the Cray Y-MP and the Intel iPSC/860 are able to reduce this time to practical levels. The optimal control solution for walking can be found with about 77 hours of CPU time on the Cray, and with about 88 hours of CPU time on the Intel. Although the overall speeds of the Cray and the Intel were found to be similar, the unique capabilities of each machine are best suited to different parts of the optimal control algorithm used. The Intel performed best in the calculation of the derivatives of the performance criterion and the constraints. In contrast, the Cray performed best during parameter optimization of the controls. These results suggest that the ideal computer architecture for solving very large-scale optimal control problems is a hybrid system in which a vector-processing machine is integrated into the communication network of a MIMD parallel machine.

Anderson, Frank C.↗

Modelling parallel programs and multiprocessor architectures with AXE

AXE, An Experimental Environment for Parallel Systems, was designed to model and simulate for parallel systems at the process level. It provides an integrated environment for specifying computation models, multiprocessor architectures, data collection, and performance visualization. AXE is being used at NASA-Ames for developing resource management strategies, parallel problem formulation, multiprocessor architectures, and operating system issues related to the High Performance Computing and Communications Program. AXE's simple, structured user-interface enables the user to model parallel programs and machines precisely and efficiently. Its quick turn-around time keeps the user interested and productive. AXE models multicomputers. The user may easily modify various architectural parameters including the number of sites, connection topologies, and overhead for operating system activities. Parallel computations in AXE are represented as collections of autonomous computing objects known as players. Their use and behavior is described. Performance data of the multiprocessor model can be observed on a color screen. These include CPU and message routing bottlenecks, and the dynamic status of the software.

Yan, Jerry C.↗

ParFlow Sand Tank: A tool for groundwater exploration

The ParFlow Sand Tank model is an open source application designed to allow users to interactively simulate and visualize groundwater movement through the subsurface. The app is designed for both research and education; teaching hydrogeology concepts and making it easy explore and run sophisticated groundwater simulations. Our goal is to support increased accessibility and usability of research grade hydrology tools for research and teaching. The Sand Tank application simulates groundwater and surface water fluxes as well as contaminant transport in real time using the integrated physical hydrology model ParFlow (Kollet & Maxwell, 2006; Maxwell & Miller, 2005; Osei-Kuffuor et al., 2014) and the particle tracking code EcoSlim (Maxwell et al., 2019). ParFlow is a numerical hydrology model that simulates spatially distributed groundwater and surface water flow. It is a well established research tool with more than 90 publications documenting its development use to advance our understanding of groundwater dynamics and groundwater surface water interactions from the hillslope to the continental scale e.g. (Condon et al., 2020; Condon & Maxwell, 2019; Maxwell & Condon, 2016). It is designed for efficient parallel computation and has been run on many platforms spanning from laptops to supercomputers. However, one of the challenges of ParFlow is that it requires significant training and hydrologic expertise to develop simulations. The Sand Tank application makes this model accessible to anyone for education and exploration. Our application uses ParFlow for its simulation backend and ParaView for the data loading and processing. The communication infrastructure relies on the ParaViewWeb framework. We use model templates deployed in Docker images to setup the Sand Tank framework. Users can build the application locally or interact with it through our web deployment. When interacting with a template users can interactively change model parameters like subsurface processes or pump/inject water into the subsurface and watch the system respond to their changes in real time as the simulation runs. Additionally, our template setup will allow more advanced users to build custom templates of increasing complexity for both research and educational purposes.

54 ENVIRONMENTAL SCIENCES↗

An evidential reasoning extension to quantitative model-based failure diagnosis

The detection and diagnosis of failures in physical systems characterized by continuous-time operation are studied. A quantitative diagnostic methodology has been developed that utilizes the mathematical model of the physical system. On the basis of the latter, diagnostic models are derived each of which comprises a set of orthogonal parity equations. To improve the robustness of the algorithm, several models may be used in parallel, providing potentially incomplete and/or conflicting inferences. Dempster's rule of combination is used to integrate evidence from the different models. The basic probability measures are assigned utilizing quantitative information extracted from the mathematical model and from online computation performed therewith.

Gertler, Janos J.↗

Smart-Pixel Array Processors Based on Optimal Cellular Neural Networks for Space Sensor Applications

A smart-pixel cellular neural network (CNN) with hardware annealing capability, digitally programmable synaptic weights, and multisensor parallel interface has been under development for advanced space sensor applications. The smart-pixel CNN architecture is a programmable multi-dimensional array of optoelectronic neurons which are locally connected with their local neurons and associated active-pixel sensors. Integration of the neuroprocessor in each processor node of a scalable multiprocessor system offers orders-of-magnitude computing performance enhancements for on-board real-time intelligent multisensor processing and control tasks of advanced small satellites. The smart-pixel CNN operation theory, architecture, design and implementation, and system applications are investigated in detail. The VLSI (Very Large Scale Integration) implementation feasibility was illustrated by a prototype smart-pixel 5x5 neuroprocessor array chip of active dimensions 1380 micron x 746 micron in a 2-micron CMOS technology.

Fang, Wai-Chi↗

Microfluidic Devices for Studying Biomolecular Interactions

Microfluidic devices for monitoring biomolecular interactions have been invented. These devices are basically highly miniaturized liquid-chromatography columns. They are intended to be prototypes of miniature analytical devices of the laboratory on a chip type that could be fabricated rapidly and inexpensively and that, because of their small sizes, would yield analytical results from very small amounts of expensive analytes (typically, proteins). Other advantages to be gained by this scaling down of liquid-chromatography columns may include increases in resolution and speed, decreases in the consumption of reagents, and the possibility of performing multiple simultaneous and highly integrated analyses by use of multiple devices of this type, each possibly containing multiple parallel analytical microchannels. The principle of operation is the same as that of a macroscopic liquid-chromatography column: The column is a channel packed with particles, upon which are immobilized molecules of the protein of interest (or one of the proteins of interest if there are more than one). Starting at a known time, a solution or suspension containing molecules of the protein or other substance of interest is pumped into the channel at its inlet. The liquid emerging from the outlet of the channel is monitored to detect the molecules of the dissolved or suspended substance(s). The time that it takes these molecules to flow from the inlet to the outlet is a measure of the degree of interaction between the immobilized and the dissolved or suspended molecules. Depending on the precise natures of the molecules, this measure can be used for diverse purposes: examples include screening for solution conditions that favor crystallization of proteins, screening for interactions between drugs and proteins, and determining the functions of biomolecules.

Wilson, Wilbur W.↗

Active Learning for Metamaterial Optimization on HPC and QC Integrated Systems

Active learning algorithms, integrating machine learning, quantum computing and optics simulation in an iterative loop, offer a promising approach to optimizing metamaterials. However, these algorithms can face difficulties in optimizing highly complex structures due to computational limitations. High-performance computing (HPC) and quantum computing (QC) integrated systems can address these issues by enabling parallel computing. In this study, we develop an active learning algorithm working on HPC-QC integrated systems. We evaluate the performance of optimization processes within active learning (i.e., training a machine learning model, problem-solving with quantum computing, and evaluating optical properties through wave-optics simulation) for highly complex metamaterial cases. Our results showcase that utilizing multiple cores on the integrated system can significantly reduce computational time, thereby enhancing the efficiency of optimization processes. Therefore, we expect that leveraging HPC-QC integrated systems helps effectively tackle large-scale optimization challenges in general.

Kim, Seongmin↗

Computer architecture for efficient algorithmic executions in real-time systems: New technology for avionics systems and advanced space vehicles

Improvements and advances in the development of computer architecture now provide innovative technology for the recasting of traditional sequential solutions into high-performance, low-cost, parallel system to increase system performance. Research conducted in development of specialized computer architecture for the algorithmic execution of an avionics system, guidance and control problem in real time is described. A comprehensive treatment of both the hardware and software structures of a customized computer which performs real-time computation of guidance commands with updated estimates of target motion and time-to-go is presented. An optimal, real-time allocation algorithm was developed which maps the algorithmic tasks onto the processing elements. This allocation is based on the critical path analysis. The final stage is the design and development of the hardware structures suitable for the efficient execution of the allocated task graph. The processing element is designed for rapid execution of the allocated tasks. Fault tolerance is a key feature of the overall architecture. Parallel numerical integration techniques, tasks definitions, and allocation algorithms are discussed. The parallel implementation is analytically verified and the experimental results are presented. The design of the data-driven computer architecture, customized for the execution of the particular algorithm, is discussed.

Carroll, Chester C.↗

Massively Parallel Dantzig-Wolfe Decomposition Applied to Traffic Flow Scheduling

Optimal scheduling of air traffic over the entire National Airspace System is a computationally difficult task. To speed computation, Dantzig-Wolfe decomposition is applied to a known linear integer programming approach for assigning delays to flights. The optimization model is proven to have the block-angular structure necessary for Dantzig-Wolfe decomposition. The subproblems for this decomposition are solved in parallel via independent computation threads. Experimental evidence suggests that as the number of subproblems/threads increases (and their respective sizes decrease), the solution quality, convergence, and runtime improve. A demonstration of this is provided by using one flight per subproblem, which is the finest possible decomposition. This results in thousands of subproblems and associated computation threads. This massively parallel approach is compared to one with few threads and to standard (non-decomposed) approaches in terms of solution quality and runtime. Since this method generally provides a non-integral (relaxed) solution to the original optimization problem, two heuristics are developed to generate an integral solution. Dantzig-Wolfe followed by these heuristics can provide a near-optimal (sometimes optimal) solution to the original problem hundreds of times faster than standard (non-decomposed) approaches. In addition, when massive decomposition is employed, the solution is shown to be more likely integral, which obviates the need for an integerization step. These results indicate that nationwide, real-time, high fidelity, optimal traffic flow scheduling is achievable for (at least) 3 hour planning horizons.

Rios, Joseph Lucio↗

New laser materials for laser diode pumping

The potential advantages of laser diode pumped solid state lasers are many with high overall efficiency being the most important. In order to realize these advantages, the solid state laser material needs to be optimized for diode laser pumping and for the particular application. In the case of the Nd laser, materials with a longer upper level radiative lifetime are desirable. This is because the laser diode is fundamentally a cw source, and to obtain high energy storage, a long integration time is necessary. Fluoride crystals are investigated as host materials for the Nd laser and also for IR laser transitions in other rare earths, such as the 2 micron Ho laser and the 3 micron Er laser. The approach is to investigate both known crystals, such as BaY2F8, as well as new crystals such as NaYF8. Emphasis is on the growth and spectroscopy of BaY2F8. These two efforts are parallel efforts. The growth effort is aimed at establishing conditions for obtaining large, high quality boules for laser samples. This requires numerous experimental growth runs; however, from these runs, samples suitable for spectroscopy become available.

Jenssen, H. P.↗

America's Next Great Ship: Space Launch System Core Stage Transitioning from Design to Manufacturing

The Space Launch System (SLS) Program is essential to achieving the Nation's and NASA's goal of human exploration and scientific investigation of the solar system. As a multi-element program with emphasis on safety, affordability, and sustainability, SLS is becoming America's next great ship of exploration. The SLS Core Stage includes avionics, main propulsion system, pressure vessels, thrust vector control, and structures. Boeing manufactures and assembles the SLS core stage at the Michoud Assembly Facility (MAF) in New Orleans, LA, a historical production center for Saturn V and Space Shuttle programs. As the transition from design to manufacturing progresses, the importance of a well-executed manufacturing, assembly, and operation (MA&O) plan is crucial to meeting performance objectives. Boeing employs classic techniques such as critical path analysis and facility requirements definition as well as innovative approaches such as Constraint Based Scheduling (CBS) and Cirtical Chain Project Management (CCPM) theory to provide a comprehensive suite of project management tools to manage the health of the baseline plan on both a macro (overall project) and micro level (factory areas). These tools coordinate data from multiple business systems and provide a robust network to support Material & Capacity Requirements Planning (MRP/CRP) and priorities. Coupled with these tools and a highly skilled workforce, Boeing is orchestrating the parallel buildup of five major sub assemblies throughout the factory. Boeing and NASA are transforming MAF to host state of the art processes, equipment and tooling, the most prominent of which is the Vertical Assembly Center (VAC), the largest weld tool in the world. In concert, a global supply chain is delivering a range of structural elements and component parts necessary to enable an on-time delivery of the integrated Core Stage. SLS is on plan to launch humanity into the next phase of space exploration.

Birkenstock, Benjamin↗

Path Planning: Differential Dynamic Programming and Model Predictive Path Integral Control on VTOL Aircraft

This paper explores two optimal control approaches, widely used in robotics, to establish their viability as real-time trajectory planners for vehicle configurations envisioned for the emerging aviation sector of Urban Air Mobility (UAM). Differential Dynamic Programming (DDP) enables planning over highly nonlinear dynamics using second-order approximations along a nominal trajectory, and displays quadratic convergence to a local solution. Model Predictive Path Integral (MPPI) is a stochastic sampling-based algorithm that can optimize for general cost criteria, including potentially highly nonlinear formulations, and supports parallel computation through the use of modern GPU hardware. In this work, DDP and MPPI were implemented using model predictive control (MPC), and the results indicate they are able to successfully transition the aircraft over different flight envelopes and generate trajectories unique to UAM vehicles.

Differential Dynamic Programming↗

Analysis and design of a high power, digitally-controlled spacecraft power system

The progress to date on the analysis and design of a high power, digitally controlled spacecraft power system is described. Several battery discharger topologies were compared for use in the space platform application. Updated information has been provided on the battery voltage specification. Initially it was thought to be in the 30 to 40 V range. It is now specified to be 53 V to 84 V. This eliminated the tapped-boost and the current-fed auto-transformer converters from consideration. After consultations with NASA, it was decided to trade-off the following topologies: (1) boost converter; (2) multi-module, multi-phase boost converter; and (3) voltage-fed push-pull with auto-transformer. A non-linear design optimization software tool was employed to facilitate an objective comparison. Non-linear design optimization insures that the best design of each topology is compared. The results indicate that a four-module, boost converter with each module operating 90 degrees out of phase is the optimum converter for the space platform. Large-signal and small-signal models were generated for the shunt, charger, discharger, battery, and the mode controller. The models were first tested individually according to the space platform power system specifications supplied by NASA. The effect of battery voltage imbalance on parallel dischargers was investigated with respect to dc and small-signal responses. Similarly, the effects of paralleling dischargers and chargers were also investigated. A solar array and shunt model was included in these simulations. A model for the bus mode controller (power control unit) was also developed to interface the Orbital replacement Unit (ORU) model to the platform power system. Small signal models were used to generate the bus impedance plots in the various operating modes. The large signal models were integrated into a system model, and time domain simulations were performed to verify bus regulation during mode transitions. Some changes have subsequently been incorporated into the models. The changes include the use of a four module boost discharger, and a new model for the mode controller, which includes the effects of saturation. The new simulations for the boost discharger show the improvement in bus ripple that can be achieved by phase-shifted operation of each of the boost modules.

Lee, F. C.↗

Supervised autonomous rendezvous and docking system technology evaluation

Technology for manned space flight is mature and has an extensive history of the use of man-in-the-loop rendezvous and docking, but there is no history of automated rendezvous and docking. Sensors exist that can operate in the space environment. The Shuttle radar can be used for ranges down to 30 meters, Japan and France are developing laser rangers, and considerable work is going on in the U.S. However, there is a need to validate a flight qualified sensor for the range of 30 meters to contact. The number of targets and illumination patterns should be minimized to reduce operation constraints with one or more sensors integrated into a robust system for autonomous operation. To achieve system redundancy, it is worthwhile to follow a parallel development of qualifying and extending the range of the 0-12 meter MSFC sensor and to simultaneously qualify the 0-30(+) meter JPL laser ranging system as an additional sensor with overlapping capabilities. Such an approach offers a redundant sensor suite for autonomous rendezvous and docking. The development should include the optimization of integrated sensory systems, packaging, mission envelopes, and computer image processing to mimic brain perception and real-time response. The benefits of the Global Positioning System in providing real-time positioning data of high accuracy must be incorporated into the design. The use of GPS-derived attitude data should be investigated further and validated.

Marzwell, Neville I.↗

BISON Robustness and Performance Improvements

BISON is a modern finite-element based nuclear fuel performance code that has been under development at the Idaho National Laboratory (USA) since 2009 [1]. The code is applicable to both steady and transient fuel behavior and can be used to analyze 1D (spherically symmetric), 2D (axisymmetric and generalized plane strain) or 3D geometries. BISON is the fuel performance code used within CASL for LWR fuel under both normal operating and accident conditions. BISON is built using the INL Multiphysics ObjectOriented Simulation Environment, or MOOSE [2, 3]. MOOSE is a massively parallel, finite element-based framework to solve systems of coupled non-linear partial differential equations using the Jacobian-Free Newton Krylov (JFNK) method [4]. This enables investigation of computationally large problems, for example a full stack of discrete pellets in a LWR fuel rod, or every rod in a full reactor core. MOOSE supports the use of complex two and three-dimensional meshes and uses implicit time integration, important for the widely varied time scale in nuclear fuel simulation. An object-oriented architecture is employed which greatly minimizes the programming effort required to add new material and behavioral models. The flexibility of the implicit and fully coupled multiphysics approach comes with a need for constructing suitable approximations for the Jacobian matrix of the coupled system used for either preconditioning a Krylov solve or in a direct Newton solve. Preconditioning options for Bison problems need to be revisited with new preconditioning methods becoming available.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Nonlinear simulations of GAEs in NSTX-U

A set of nonlinear simulations has been performed in order to study the nonlinear evolution of unstable global Alfvén eigenmodes in the National Spherical Torus Experiment-Upgrade (NSTX-U). Results of the single toroidal mode number, n, simulations are compared with a full nonlinear simulation (all toroidal harmonics included). In single-n simulations, the conservation of two integrals of motion of a particle in a cyclotron resonance with a monochromatic wave is demonstrated, resulting in a one-dimensional evolution of the particle distribution in (E,μ,pϕ) phase-space. Nonlinear simulations (both single-n and full nonlinear) show a significant redistribution of the resonant fast ions, especially in the pitch parameter. Thus, the changes in the resonant particle's parallel and perpendicular energies can be several times larger than the total particle energy change, with only a small fraction transferring into the excitation of the mode itself. This implies that even a relatively small amplitude mode can significantly modify the beam distribution in the resonant region. For the NSTX-U case considered, the single-n simulation results are close to full nonlinear simulation only for the most unstable mode, in which case the saturation amplitudes and changes in the fast ion distribution are comparable. In contrast, peak amplitudes of subdominant modes in all-n simulations are smaller by a factor of 3–10 compared to single-n runs due to the flattening of the beam ion distribution by the fastest growing mode.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Summer Internship Report: ARA2 Benchmarking

Over the past decade, the RISC-V Instruction Set Architecture (ISA) has emerged as a significant player in both academic and industrial processor design due to its open-source nature, modular extension system, and versatility across domains ranging from microcontrollers to high-performance computing (HPC). One of its most important recent advancements is the RISC-V Vector Extension (RVV), which enables explicit data-level parallelism through vector registers and vectorized instructions. Unlike traditional SIMD (Single Instruction, Multiple Data) architectures that fix vector lengths at design time, RVV uses the concept of VLEN (vector register length) as a hardware-independent parameter and allows software to adapt dynamically to the available vector width. This flexible approach ensures portability across implementations while enabling scalable performance. The ARA2 core is a parameterizable RISC-V vector processor developed at the Integrated Systems Lab at ETH Zürich and the University of Bologna. Designed as a tightly-coupled accelerator to a scalar RISC-V core, ARA2 implements the RVV 1.0 specification and offers tunable architectural parameters such as the number of vector lanes, VLEN, and cache sizes.

97 MATHEMATICS AND COMPUTING↗

Receptivity of Hypersonic Boundary Layers Due to Acoustic Disturbances over Blunt Cone

The transition process induced by the interaction of acoustic disturbances in the free-stream with boundary layers over a 5-degree straight cone and a wedge with blunt tips is numerically investigated at a free-stream Mach number of 6.0. To compute the shock and the interaction of shock with the instability waves the Navier-Stokes equations are solved in axisymmetric coordinates. The governing equations are solved using the 5th -order accurate weighted essentially non-oscillatory (WENO) scheme for space discretization and using third-order total-variation-diminishing (TVD) Runge-Kutta scheme for time integration. After the mean flow field is computed, acoustic disturbances are introduced at the outer boundary of the computational domain and unsteady simulations are performed. Generation and evolution of instability waves and the receptivity of boundary layer to slow and fast acoustic waves are investigated. The mean flow data are compared with the experimental results. The results show that the instability waves are generated near the leading edge and the non-parallel effects are stronger near the nose region for the flow over the cone than that over a wedge. It is also found that the boundary layer is much more receptive to slow acoustic wave (by almost a factor of 67) as compared to the fast wave.

Kara, K.↗