Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel machines”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32

Turbine Electrified Energy Management for Single Aisle Aircraft

Electrified aircraft propulsion technology is being developed to reduce the environmental impacts of the aviation industry. This is prompting the exploration of potential uses and benefits of hybrid systems in which electric powertrains are integrated with more traditional gas turbine propulsion systems. Turbine Electrified Energy Management (TEEM) is an energy management approach for hybrid-electric architectures in which electric machines are connected to the turbofan shafts and used to suppress the off-design operation naturally associated with engine transients. This reduces the need to maintain a large amount of compressor operability margin, thus allowing further exploration of the engine design space. In this study, a 19,000 lbf engine within a parallel hybrid propulsion system is considered along with a 30,000 lbf standalone engine. Data from prior TEEM applications are used to approximate the electric machine sizing required to achieve operability benefits. The TEEM controller is shown to improve operability during transients through the reduction of stall margin undershoots and the decrease of transient variations in component performance maps by over 29%.

Controls↗

Turbine Electrified Energy Management for Single Aisle Aircraft

Electrified aircraft propulsion technology is being developed to reduce the environmental impacts of the aviation industry. This is prompting the exploration of potential uses and benefits of hybrid systems in which electric powertrains are integrated with more traditional gas turbine propulsion systems. Turbine Electrified Energy Management (TEEM) is an energy management approach for hybrid-electric architectures in which electric machines are connected to the turbofan shafts and used to suppress the off-design operation naturally associated with engine transients. This reduces the need to maintain a large amount of compressor operability margin, thus allowing further exploration of the engine design space. In this study, a 19,000 lbf engine within a parallel hybrid propulsion system is considered along with a 30,000 lbf standalone engine. Data from prior TEEM applications are used to approximate the electric machine sizing required to achieve operability benefits. The TEEM controller is shown to improve operability during transients through the reduction of stall margin undershoots and the decrease of transient variations in component performance maps by over 29%.

EAP↗

Data-driven closure modeling for hypersonic turbulent flows

The Reynolds-averaged Navier–Stokes (RANS) equations remain a workhorse technology for simulating compressible fluid flows of practical interest. Due to model-form errors, however, RANS models can yield erroneous predictions that preclude their use on mission-critical problems. This report summarizes work performed from FY22-FY24 focused on improving RANS models for hypersonic flows using data-driven modeling and scientific machine learning. In this work we: 1. Investigate the current capabilities of RANS models in Sandia’s parallel aerodynamics and re-entry code (SPARC) for hypersonic flows with a focus on shock boundary layer interactions (SBLIs), 2. Assess several established corrections that exist in the literature aimed at improving predictions for SBLIs, 3. Develop improved models for the Reynolds stress tensor using tensor-basis neural networks, 4. Develop a neural-network-based variable turbulent Prandtl number model to reduce errors in wall heating in SBLIs. 5. Begin future investigations including employing the LIFE framework to improve wall heating predictions in SBLIs as well as the ensemble Kalman filter. We find that current RANS models in SPARC are deficient for complex SBLI flows. In particular, no current model jointly predicts wall heat flux, wall shear stress, and wall pressure with reasonable accuracy. Existing corrections help, but do not alleviate this issue altogether. The development of improved models for the Reynolds stress tensor via tensor-basis neural networks results in more predictive RANS models across a suite of low-speed and high-speed cases. For hypersonic boundary layers, the inclusion of the wall-normal Reynolds stress via TBNNs has an appreciable impact on the wall-normal momentum balance and wall quantities. However, we find that improvements to the Reynolds stress tensor do not address the over-prediction in wall heat flux in SBLIs. We find that a neural-network-based variable turbulent Prandtl number model systematically and substantially improves wall heating predictions for a range of SBLI cases.

97 MATHEMATICS AND COMPUTING↗

Wave scheduling - Decentralized scheduling of task forces in multicomputers

Decentralized operating systems that control large multicomputers need techniques to schedule competing parallel programs called task forces. Wave scheduling is a probabilistic technique that uses a hierarchical distributed virtual machine to schedule task forces by recursively subdividing and issuing wavefront-like commands to processing elements capable of executing individual tasks. Wave scheduling is highly resistant to processing element failures because it uses many distributed schedulers that dynamically assign scheduling responsibilities among themselves. The scheduling technique is trivially extensible as more processing elements join the host multicomputer. A simple model of scheduling cost is used by every scheduler node to distribute scheduling activity and minimize wasted processing capacity by using perceived workload to vary decentralized scheduling rules. At low to moderate levels of network activity, wave scheduling is only slightly less efficient than a central scheduler in its ability to direct processing elements to accomplish useful work.

Van Tilborg, A. M.↗

Mapping robust parallel multigrid algorithms to scalable memory architectures

The convergence rate of standard multigrid algorithms degenerates on problems with stretched grids or anisotropic operators. The usual cure for this is the use of line or plane relaxation. However, multigrid algorithms based on line and plane relaxation have limited and awkward parallelism and are quite difficult to map effectively to highly parallel architectures. Newer multigrid algorithms that overcome anisotropy through the use of multiple coarse grids rather than relaxation are better suited to massively parallel architectures because they require only simple point-relaxation smoothers. In this paper, we look at the parallel implementation of a V-cycle multiple semicoarsened grid (MSG) algorithm on distributed-memory architectures such as the Intel iPSC/860 and Paragon computers. The MSG algorithms provide two levels of parallelism: parallelism within the relaxation or interpolation on each grid and across the grids on each multigrid level. Both levels of parallelism must be exploited to map these algorithms effectively to parallel architectures. This paper describes a mapping of an MSG algorithm to distributed-memory architectures that demonstrates how both levels of parallelism can be exploited. The result is a robust and effective multigrid algorithm for distributed-memory machines.

Overman, Andrea↗

Performance of a plasma fluid code on the Intel parallel computers

One approach to improving the real-time efficiency of plasma turbulence calculations is to use a parallel algorithm. A parallel algorithm for plasma turbulence calculations was tested on the Intel iPSC/860 hypercube and the Touchtone Delta machine. Using the 128 processors of the Intel iPSC/860 hypercube, a factor of 5 improvement over a single-processor CRAY-2 is obtained. For the Touchtone Delta machine, the corresponding improvement factor is 16. For plasma edge turbulence calculations, an extrapolation of the present results to the Intel (sigma) machine gives an improvement factor close to 64 over the single-processor CRAY-2.

Lynch, V. E.↗

Computing for the DUNE Long-Baseline Neutrino Oscillation Experiment

This is a talk given at Computers in High Energy Physics in Adelaide, South Australia, Australia in November 2019. It is partially intended to explain the context of DUNE Computing for computing specialists. The DUNE collaboration consists of over 180 institutions from 33 countries. The experiment is in preparation now with commissioning of the first 10kT fiducial volume Liquid Argon TPC expected over the period 2025-2028 and a long data taking run with 4 modules expected from 2029 and beyond. An active prototyping program is already in place with a short test beam run with a 700T, 15,360 channel prototype of single-phase readout at the neutrino platform at CERN in late 2018 and tests of a similar sized dual-phase detector scheduled for mid-2019. The 2018 test beam run was a valuable live test of our computing model. The detector produced raw data at rates of up to ~2GB/s. These data were stored at full rate on tape at CERN and Fermilab and replicated at sites in the UK and Czech Republic. In total 1.2 PB of raw data from beam and cosmic triggers were produced and reconstructed during the six week test beam run. Baseline predictions for the full DUNE detector data, starting in the late 2020's are 30-60 PB of raw data per year. In contrast to traditional HEP computational problems, DUNE's Liquid Argon TPC data consist of simple but very large (many GB) 2D data objects which share many characteristics with astrophysical images. This presents opportunities to use advances in machine learning and pattern recognition as a frontier user of High Performance Computing facilities capable of massively parallel processing.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Status of DUNE Offline Computing

We summarize the status of Deep Underground Neutrino Experiment (DUNE) Offline Software and Computing program. We describe plans for the computing infrastructure needed to acquire, catalog, reconstruct, simulate and analyze the data from the DUNE experiment and its prototypes in pursuit of the experiment's physics goals of precision measurements of neutrino oscillation parameters, detection of astrophysical neutrinos, measurement of neutrino interaction properties and searches for physics beyond the Standard Model. In contrast to traditional HEP computational problems, DUNE's Liquid Argon Time Projection Chamber data consist of simple but very large (many GB) data objects which share many characteristics with astrophysical images. We have successfully reconstructed and simulated data from 4% prototype detector runs at CERN. The data volume from the full DUNE detector, when it starts commissioning late in this decade will present memory management challenges in conventional processing but significant opportunities to use advances in machine learning and pattern recognition as a frontier user of High Performance Computing facilities capable of massively parallel processing. Our goal is to develop infrastructure resources that are flexible and accessible enough to support creative software solutions as HEP computing evolves.

43 PARTICLE ACCELERATORS↗

Overcoming the field-of-view to diameter trade-off in microendoscopy via computational optrode-array microscopy

High-resolution microscopy of deep tissue with large field-of-view (FOV) is critical for elucidating organization of cellular structures in plant biology. Microscopy with an implanted probe offers an effective solution. However, there exists a fundamental trade-off between the FOV and probe diameter arising from aberrations inherent in conventional imaging optics (typically, FOV < 30% of diameter). Here, we demonstrate the use of microfabricated non-imaging probes (optrodes) that when combined with a trained machine-learning algorithm is able to achieve FOV of 1x to 5x the probe diameter. Further increase in FOV is achieved by using multiple optrodes in parallel. With a 1 × 2 optrode array, we demonstrate imaging of fluorescent beads (including 30 FPS video), stained plant stem sections and stained living stems. Our demonstration lays the foundation for fast, high-resolution microscopy with large FOV in deep tissue via microfabricated non-imaging probes and advanced machine learning.

59 BASIC BIOLOGICAL SCIENCES↗

Automated Generation of Message-Passing Programs: An Evaluation Using CAPTools

Scientists at NASA Ames Research Center have been developing computational aeroscience applications on highly parallel architectures over the past ten years. During that same time period, a steady transition of hardware and system software also occurred, forcing us to expend great efforts into migrating and re-coding our applications. As applications and machine architectures become increasingly complex, the cost and time required for this process will become prohibitive. In this paper, we present the first set of results in our evaluation of interactive parallelization tools. In particular, we evaluate CAPTool's ability to parallelize computational aeroscience applications. CAPTools was tested on serial versions of the NAS Parallel Benchmarks and ARC3D, a computational fluid dynamics application, on two platforms: the SGI Origin 2000 and the Cray T3E. This evaluation includes performance, amount of user interaction required, limitations and portability. Based on these results, a discussion on the feasibility of computer aided parallelization of aerospace applications is presented along with suggestions for future work.

Hribar, Michelle R.↗

Automated Generation of Message-Passing Programs: An Evaluation Using CAPTools

Scientists at NASA Ames Research Center have been developing computational aeroscience applications on highly parallel architectures over the past ten years. During that same time period, a steady transition of hardware and system software also occurred, forcing us to expend great efforts into migrating and re-coding our applications. As applications and machine architectures become increasingly complex, the cost and time required for this process will become prohibitive. In this paper, we present the first set of results in our evaluation of interactive parallelization tools. In particular, we evaluate CAPTool's ability to parallelize computational aeroscience applications. CAPTools was tested on serial versions of the NAS Parallel Benchmarks and ARC3D, a computational fluid dynamics application, on two platforms: the SGI Origin 2000 and the Cray T3E. This evaluation includes performance, amount of user interaction required, limitations and portability. Based on these results, a discussion on the feasibility of computer aided parallelization of aerospace applications is presented along with suggestions for future work.

Hribar, Michelle R.↗

File-access characteristics of parallel scientific workloads

Phenomenal improvements in the computational performance of multiprocessors have not been matched by comparable gains in I/O system performance. This imbalance has resulted in I/O becoming a significant bottleneck for many scientific applications. One key to overcoming this bottleneck is improving the performance of parallel file systems. The design of a high-performance parallel file system requires a comprehensive understanding of the expected workload. Unfortunately, until recently, no general workload studies of parallel file systems have been conducted. The goal of the CHARISMA project was to remedy this problem by characterizing the behavior of several production workloads, on different machines, at the level of individual reads and writes. The first set of results from the CHARISMA project describe the workloads observed on an Intel iPSC/860 and a Thinking Machines CM-5. This paper is intended to compare and contrast these two workloads for an understanding of their essential similarities and differences, isolating common trends and platform-dependent variances. Using this comparison, we are able to gain more insight into the general principles that should guide parallel file-system design.

Nieuwejaar, Nils↗

Ensemble models for circuit topology estimation, fault detection and classification in distribution systems

This paper presents a methodology for simultaneous fault detection, classification, and topology estimation for adaptive protection of distribution systems. The methodology estimates the probability of the occurrence of each one of these events by using a hybrid structure that combines three sub-systems, a convolutional neural network for topology estimation, a fault detection based on predictive residual analysis, and a standard support vector machine with probabilistic output for fault classification. The input to all these sub-systems is the local voltage and current measurements. A convolutional neural network uses these local measurements in the form of sequential data to extract features and estimate the topology conditions. The fault detector is constructed with a Bayesian stage (a multitask Gaussian process) that computes a predictive distribution (assumed to be Gaussian) of the residuals using the input. Since the distribution is known, these residuals can be transformed into a Standard distribution, whose values are then introduced into a one-class support vector machine. The structure allows using a one-class support vector machine without parameter cross-validation, so the fault detector is fully unsupervised. Finally, a support vector machine uses the input to perform the classification of the fault types. All three sub-systems can work in a parallel setup for both performance and computation efficiency. In conclusion, we test all three sub-systems included in the structure on a modified IEEE123 bus system, and we compare and evaluate the results with standard approaches.

24 POWER TRANSMISSION AND DISTRIBUTION↗

On-Board AC Charging Topology Integrated with Electric Vehicle Motor Drive System

On-board AC charging is a convenient and widely adopted method for recharging electric vehicles (EVs) directly from standard alternating current (AC) power sources. This paper presents a novel topology for AC charging of EVs that utilizes EV 3-phase electric machine windings as the input inductors, thus eliminating the requirement for bulky grid interfacing inductors and resulting in a compact and cost-effective integrated motor drive and charger system. The proposed approach leverages the motor windings and parallel operating half-bridge inverter during the charging process by interconnecting the inverter phases with the motor windings in a mechanically interleaved and electrically paralleled manner. The implementation of this unique and innovative idea, achieved through precise control and arrangement of the motor winding as a series inductor, successfully eliminates the possibility of unintended motion of the electric machine during the charging process.

ADVANCED PROPULSION SYSTEMS↗

Real-time optimization of multi-cell industrial evaporative cooling towers using machine learning and particle swarm optimization

Existing electrical generating stations must operate with greater flexibility due to increasing renewable energy penetration on the electrical grid, and many coal-fired power stations have transitioned away from baseload operation to load-following operation to aid in grid stability. In cases where multiple independently controlled cooling tower cells are used in parallel for the cooling purposes of such stations, there is an opportunity to increase plant efficiency through data-driven optimization across their full load ranges. This work presents a novel application of real-time optimization using machine learning and particle swarm optimization on a multi-cell induced-draft cooling tower servicing a coal-fired power station under variable load. This is the first work to demonstrate simultaneous optimization of a multi-cell cooling tower, in addition to using machine learning for closed-loop control on a cooling tower. A novel control configuration is presented that ensures original control logic is not adversely affected and that the overall plant process is not disrupted using only existing hardware and operational data. To verify this methodology, the 12 independent cooling tower cells are simulated in parallel using historic operating data to demonstrate the effectiveness of real-time optimization compared to current practice. An artificial neural network is trained to predict overall cooling tower power consumption using only operational data and ambient conditions with an R2 value of greater than 0.96. The real-time optimization using particle swarm yields 6.7% annual energy usage savings compared to current practices, although the extent of the real-time savings varies greatly with both plant load and environmental conditions. This is particularly significant for a variable load situation because frequent ramping typically results in reduced overall efficiency. Furthermore, this proposed AI-based solution presents an opportunity to improve the overall heat rate of a load-following coal-fired power plant without the need to perform extensive first-principles modeling or add additional hardware to the cooling tower, resulting in more resources conserved and less overall emissions per unit of electricity generated.

42 ENGINEERING↗

Prediction of DIII-D Pedestal Structure From Externally Controllable Parameters

The sharp increase of pressure at the edge of a high confinement mode (H-mode) plasma, the pedestal, strongly impacts overall plasma performance. Predicting the pedestal is a necessity to control and optimize tokamak operations. Here, an experimental data-driven machine learning (ML) approach is presented that predicts the pedestal heights and widths of electron density (n e ) and electron temperature (T e ) profiles as well as the separatrix ne from externally controllable parameters such as the plasma shape, heating method and power, and gas puff rate and integrated gas puff. The OMFIT framework was used with DIII-D data to efficiently, robustly, and automatically build a database of pedestal parameters to train machine learning models. Database creation was enabled by the search engine tool for DIII-D data, TokSearch, which parallelizes data fetching, enabling fast searches through basic signals of thousands of DIII-D shots and selection of relevant time intervals. Principal Component Analysis (PCA) separated the database into three clusters that represent classes of plasma shapes that are regularly used in DIII-D. The most important parameters for setting the pedestal structure were plasma current (I p ), toroidal magnetic field (B Φ ), neutral beam heating power (P NBI ) and shaping quantities. The Deep Jointly Informed Neural Networks (DJINN) algorithm was applied to identify suitable neural network (NN) architectures that appropriately capture the features of the pedestal database. Separate NNs were implemented for each pedestal parameter, and ensembling methods were used to improve the prediction accuracy and allowed estimation of the prediction uncertainty. The pedestal predictions of the test dataset lie within the measurement uncertainties of the pedestal parameters. The NN outperformed simple Linear Regression (LR) analysis, indicating non-linear dependencies in the pedestal structure. The presented achievements illustrate a promising path for future research, using feature extraction to infer experimental trends and thereby improve pedestal models as well as deploying NN for a fast pedestal prediction in DIII-D scenario development.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Neuromorphic Graph Algorithms: Extracting Longest Shortest Paths and Minimum Spanning Trees

Neuromorphic computing is poised to become a promising computing paradigm in the post Moore's law era due to its extremely low power usage and inherent parallelism. Traditionally speaking, a majority of the use cases for neuromorphic systems have been in the field of machine learning. In order to expand their usability, it is imperative that neuromorphic systems be used for non-machine learning tasks as well. The structural aspects of neuromorphic systems (i.e., neurons and synapses) are similar to those of graphs (i.e., nodes and edges), However, it is not obvious how graph algorithms would translate to their neuromorphic counterparts. In this work, we propose a preprocessing technique that introduces fractional offsets on the synaptic delays of neuromorphic graphs in order to break ties. This technique, in turn, enables two graph algorithms: longest shortest path extraction and minimum spanning trees.

Kay, Bill↗

MLtool++ package for machine learning and its applications to materials data

We are developing Mltool++ package of software programs for machine learning (ML). Given the MLtool Python code, we create a faster C++ code with the potential for parallelization. We have extracted materials data from the literature. One dataset contains melting temperatures of stoichiometric 1:1 metallic compounds XZ, composed by elements X={Al, Ti, V, Cr, Zr, Nb, Mo, Hf, Ta, W} and Z={Co, Ni, Cu, Rh, Pd, Ag, Ir, Pt, Au}, and another contains solid-solid symmetry-breaking phase transition temperatures. We studied dependences of temperatures on composition, found several correlations, and parametrized them by analytical functions. Mltool++ package is generic and applicable to any tabulated numeric data.

Pierce M. Pettit↗