Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel machines”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38

Iterated Gauss-Seidel GMRES

The GMRES algorithm of Saad and Schultz [SIAM J. Sci. Stat. Comput., 7 (1986), pp. 856-869] is an iterative method for approximately solving linear systems Ax = b, with initial guess x0 and residual r0 = b Ax0. The algorithm employs the Arnoldi process to generate the Krylov basis vectors (the columns of Vk ). It is well known that this process can be viewed as a QR factorization of the matrix Bk = [r0, AVk] at each iteration. Despite an O (..epsilon..)..kappa.. (Bk ) loss of orthogonality, for unit roundoff ..epsilon..and condition number ..kappa.. , the modified Gram-Schmidt formulation was shown to be backward stable in the seminal paper by Paige et al. [SIAM J. Matrix Anal.Appl., 28 (2006), pp. 264-284]. We present an iterated Gauss-Seidel formulation of the GMRES algorithm (IGS-GMRES) based on the ideas of Ruhe [Linear Algebra Appl., 52 (1983), pp. 591-601] and Swirydowicz et al. [Numer. Linear Algebra Appl., 28 (2020), pp. 1-20]. IGS-GMRES maintains orthogonality to the level O (..epsilon..)..kappa.. (Bk ) or O (..epsilon..), depending on the choice of one or two iterations; for two Gauss-Seidel iterations, the computed Krylov basis vectors remain orthogonal to working accuracy and the smallest singular value of Vk remains close to one. The resulting GMRES method is thus backward stable. We show that IGS-GMRES can be implemented with only a single synchronization point per iteration, making it relevant to large-scale parallel computing environments. We also demonstrate that, unlike MGS-GMRES, in IGS-GMRES the relative Arnoldi residual corresponding to the computed approximate solution no longer stagnates above machine precision even for highly nonnormal systems.

Arnoldi-QR↗

Advanced Algal Biofoundries for the Production of Polyurethane Precursors

The primary goal of the BEEPs project was to develop a process that could accelerate the development of algae as bioproduction platforms, from initial chemical product concept to an economically viable market supply. Under this program we elected to develop strains of algae that could generate polyurethane precursors, while simultaneously developing basic genetic tools to enable improved algal production systems. This program was specifically designed to incorporate National Laboratories as a means to utilize the expertise and facilities for new bio-production platforms. To that end, we designed a program to collaborate with the Agile BioFoundry at Lawrence Berkeley National Laboratory (LBNL) and computational platforms at Pacific Northwest National Laboratory (PNNL). In addition to these National Laboratory partners, we also had academic partners from UC Davis and Georgia Tech, as well as commercial partners Algenesis Materials and BASF. To achieve these goals, we initially focused on developing the genetic tools and high throughput screening technologies necessary to generate and assess production of polymer precursors (succinic acid) in algae and cyanobacteria, including advanced promoters and biosensors. In parallel, we computationally identified potential production bottlenecks and then used the developed genetic tools to increase production rates and yields. Constant feedback of data was used in conjunction with machine learning, high-throughput cell sorting, and synthetic biology, for additional targeted metabolic engineering. Multiple rounds of tool design, building, testing, and learning were supplied to partners at PNNL and LBNL to develop new models and tools that could expedite bioproduction platform development and increase yield performance. We targeted, and achieved, the FOA requirement yield metric of 20 g/L, as a milestone and deliverable from at least one of our engineered strains for the production of succinic acid.

09 BIOMASS FUELS↗

Non-Intrusive Parallel-in-Time Solvers for Partial Differential Equations (Final Report)

Many time-dependent problems and simulations are often modeled using Partial Differential Equations. Traditional modeling approaches that use sequential time-stepping are reaching a bottleneck in optimizing efficiency. The Center of Applied Science and Computing at Lawrence Livermore National Laboratory extensively works on parallelizing these algorithms to leverage the increasing computational power from the growing number of processors in computer hardware. In particular, they aim to design non-intrusive algorithms that can generalize to a variety of problems and sizes without requiring additional information from or modifications on the original problems. Multigrid Reduction in Time (MGRIT) is a parallel-in-time algorithm that is designed to be non-intrusive. This project focuses on increasing the efficiency of MGRIT by approximating the coarse-grid operator using machine learning approaches as a means to find the most non-intrusive, or general, solution.

97 MATHEMATICS AND COMPUTING↗

SPARTAN (Scalable Probabilistic Application Reconfigurable Tensor Autonomous Network)

The technical founder of Ludwig Computing Inc has been competitively selected for support by Cyclotron Road, a U.S. Department of Energy (DOE) Advanced Manufacturing Office (AMO) Lab-Embedded Entrepreneurship Program (LEEP) through an approved merit review process. Ludwig Computing Inc, supported by the U.S. Department of Energy's Advanced Manufacturing Office through the Cyclotron Road program, has investigated the advantages of probabilistic computing for real-world compute-intensive applications. This research adds to the understanding of alternative computing paradigms by exploring a unique hardware-software co-design that integrates quantum computing methods with nature-inspired problem-solving techniques. The project's focus on areas such as combinatorial optimization, graph analytics, and machine learning demonstrates the potential for significant advancements in computational efficiency and performance. By harnessing natural randomness to streamline large circuits into fewer devices, Ludwig's approach enables massive parallelism, potentially offering higher throughput, speed, and energy efficiency compared to conventional hardware solutions. This work benefits the public by paving the way for more efficient computing solutions that could address complex real-world problems while potentially reducing energy consumption in data-intensive industries.

97 MATHEMATICS AND COMPUTING↗

Application of parallel distributed processing to space based systems

The concept of using Parallel Distributed Processing (PDP) to enhance automated experiment monitoring and control is explored. Recent very large scale integration (VLSI) advances have made such applications an achievable goal. The PDP machine has demonstrated the ability to automatically organize stored information, handle unfamiliar and contradictory input data and perform the actions necessary. The PDP machine has demonstrated that it can perform inference and knowledge operations with greater speed and flexibility and at lower cost than traditional architectures. In applications where the rule set governing an expert system's decisions is difficult to formulate, PDP can be used to extract rules by associating the information an expert receives with the actions taken.

Macdonald, J. R.↗

Lunar Simulant Deposition Technique for Dust Tolerance Studies

A renewed interest in lunar exploration has spawned an array of development efforts for lunar surface assets. These systems depend on the reliable operation of mechanisms and components that may be susceptible to performance degradations or failure due to dust. The Uniform Dust Deposition System was developed at the NASA Glenn Research Center to provide repeatable, uniform, and automated deposition of simulants on surfaces of interest for dust mitigation testing. The system is capable of depositing simulants on test articles up to 60 cm in diameter and 15 cm high in a dry air environment with less than 1 percent relative humidity while keeping users safe from aerosolized dust. The automation of the system allows for high testing throughput while not sacrificing test quality and allows the user to reduce data in parallel. The additional development of a simulant preparation technique complements the repeatability of the deposition physics during testing. The system includes an imaging subsystem that leverages the power of machine learning to count simulant particles and measure their size, thereby allowing for accurate predictions of surface deposition densities (coefficient of determination R^(2) = 0.93) from images alone. The coverage of dust on a surface was shown to be uniform (coefficient of variation CV < 0.11), allowing developers to accurately evaluate the performance of their technology with a prescribed amount of lunar simulant, information that can be used to develop and refine models. The accuracy of the system is currently less than desired for a single deposition run, with a standard deviation (SD) ranging from 18 to 24 mg, or 0.839 to 1.184 mg/sq. cm , for a 5-cm-diameter area. However, the accuracy can be improved by performing multiple deposition runs to build dust to a desired level. Testing has shown that a SD of 0.2 to 0.6 mg, or 0.076 to 0.227 mg/sq. cm, can be achieved for a 5-cm-diameter area using this technique.

Stephen Gerdts↗

Failure Behavior and Control-Based Mitigation for a Parallel Hybrid Propulsion System

NASA is pursuing research to advance Electrified Aircraft Propulsion (EAP) technologies that address fuel burn and emission reduction goals. EAP brings the potential for improved performance over the state of the art. However, for these systems to be practical and certifiable, they need to possess adequate robustness to adverse conditions including a variety of system failures that are not applicable to conventional turbofans today. Numerous EAP concepts interface gas turbine engines with an electrical power system that includes electric machines and sometimes electrical energy storage. The expansion of the powertrain increases the probability of encountering a failure and introduces new failure modes. Failures within the electrical power system may also impact the gas turbine engine(s) to which the electrical powertrain is coupled. This effort investigates failures originating in the electrical power system and their impact on the parallel hybrid propulsion system. Reversionary control strategies are also demonstrated to reduce the impact of the failures. Failure mitigation strategies were devised and employed in simulation. Various failure scenarios were simulated including those occurring during steady state operation, transients, and takeoff and landing scenarios. The timing of the failure and delay in failure identification and activation of mitigation strategies are noteworthy variables in the study. While the system remained stable throughout all failure scenarios, delays in failure identification could result in undesirable conditions such as increased operating temperatures and reduced stall margin. The results demonstrate successful mitigation of failures through reversionary control modes and help to generate confidence in the robustness of the conceptual parallel hybrid propulsion system.

Failure behavior↗

Failure Behavior and Control Based Mitigation for a Parallel Hybrid Propulsion System

NASA is pursuing research to advance Electrified Aircraft Propulsion (EAP) technologies that address fuel burn and emission reduction goals. EAP brings the potential for improved performance over the state of the art. However, for these systems to be practical and certifiable, they need to possess adequate robustness to adverse conditions including a variety of system failures that are not applicable to conventional turbofans today. Numerous EAP concepts interface gas turbine engines with an electrical power system that includes electric machines and sometimes electrical energy storage. The expansion of the powertrain increases the probability of encountering a failure and introduces new failure modes. Failures within the electrical power system may also impact the gas turbine engine(s) to which the electrical powertrain is coupled. This effort investigates failures originating in the electrical power system and their impact on the parallel hybrid propulsion system. Reversionary control strategies are also demonstrated to reduce the impact of the failures. Failure mitigation strategies were devised and employed in simulation. Various failure scenarios were simulated including those occurring during steady state operation, transients, and takeoff and landing scenarios. The timing of the failure and delay in failure identification and activation of mitigation strategies are noteworthy variables in the study. While the system remained stable throughout all failure scenarios, delays in failure identification could result in undesirable conditions such as increased operating temperatures and reduced stall margin. The results demonstrate successful mitigation of failures through reversionary control modes and help to generate confidence in the robustness of the conceptual parallel hybrid propulsion system.

Failure behavior↗

Part-scale microstructure prediction for laser powder bed fusion Ti-6Al-4V using a hybrid mechanistic and machine learning model

Laser powder bed fusion (LPBF) Ti-6Al-4V is widely studied for use in structural applications in aerospace and medical industries, but mechanical anisotropy and microstructural inhomogeneity prohibits its wider adoption. Although successful microstructure prediction models have been developed, a remaining challenge is their limited integration across length/time scales and validation by experimental studies. Here, this work proposes a physics-augmented machine learning surrogate model to unite predictions of LPBF temperature, β phase morphology and texture, and α/α’ formation into a single framework that is calibrated and validated with experiments. First, a phase field (PF) model of the martensitic β→α’ transformation is developed and calibrated using data from in-situ synchrotron cyclic heating/cooling studies quantifying the variation of α phase fraction with time. In parallel, an established finite difference-Monte Carlo (FDMC) model predicts the part-scale temperature profile and β grain formation during solidification. A dataset is developed using LPBF cyclic temperature descriptors from the FDMC model as inputs and corresponding α/α’ phase fraction and width from the PF model as outputs. Five machine learning (ML) regression models are tested and optimized, having mean absolute error in testing ≤ 4 %, and the k-nearest neighbors (KNN) model is selected as the best performing. The KNN model is called at the nodal level during post-processing of the FDMC model to replace and downscale the response of the PF model. The combined agility and accuracy of the hybrid FDMC-ML model enables part-scale microstructure predictions that can be further used for property predictions to accelerate AM process optimization.

36 MATERIALS SCIENCE↗

DEEP Solar: Data DrivEn Modeling and Analytics for Enhanced System Layer ImPlementation

Realizing the SETO 2030 mission of reducing solar energy costs to 3-5 c/kWh will require innovative enabling research on effective, cost-efficient integration of local PV within distribution systems. However, the intermittent and variable nature of PVs compels operators to impose conservative hosting capacity constraints. Given the extremely high variability of (intermittent and unpredictable) solar energy generation, relaxing the capacity constraints (which are currently around 15%) and achieving 100% or greater integration of renewables will require a fundamental transformation of the power grid via the utilization of exponentially larger amounts of AMI enabled fine-grained data. To address the challenges in increasing the penetration of renewable energy based DERs, this project envisions an Enhanced System Layer (ESL) at the distribution network level that is reliable, cost-effective and scalable to millions of Distributed Energy Resources (DERs)/devices. This includes developing: 1) Transformative and highly scalable machine learning based predictive analytics tools that plug into distribution system planning and provide real-time situational awareness at the distribution level for short and long-term operational planning. The tools will be built using novel data-driven energy models of millions of active nodes with AMI, 2) Adaptive stochastic analysis and optimization algorithms for real-time grid operations, 3) Dynamic Scenario Analysis using parallel Cloudenabled implementations with < 1 minute computational cycle times.

14 SOLAR ENERGY↗

Delamination-informed lifecycle decisions: A dielectric and machine learning framework for composite sorting and recycling

Composite materials are widely used in aerospace, marine, and automotive sectors due to their high strength-to-weight ratio and durability. However, their long-term reliability can be compromised by damage accumulation. Specifically, delamination initiation serves as a precursor to structural failure, which is often difficult to detect during damage inspection. Identifying and sorting delamination initiation in samples not only increases operational safety while providing critical information for end-of-life decisions, which influences both the service life extension value and the efficiency of fiber extraction during recycling. This research addresses two challenges: (1) developing a nondestructive, ex-situ framework to sort composite materials based on damage severity, particularly delamination, and (2) understanding how damage in composites influences resin removal during pyrolysis. Both experimental work and finite element analysis were performed to predict critical stress levels that are associated with delamination onset. Based on these results, three loading levels 50 %, 75 %, and 90 % of maximum stress, were selected for controlled experiments, generating composite samples with varying extents of damage for machine learning model training. Microscopic imaging of these samples confirmed the damage progression from matrix cracking to delamination, validating the computational predictions. We explored supervised machine learning using dielectric measurements to classify damage states. Preliminary results show an artificial neural network can identify early delamination which is a potential precursor to failure, with 94.44 % accuracy on our dataset. A parallel investigation into the effect of damage severity on pyrolysis recycling showed that heavily delaminated samples required significantly less energy for comparable matrix removal than undamaged samples.

dielectric variables↗

IMS Rapid Response 2024 Summary Report: A Machine Learning Potential for the Periodic Table

Stockpile stewardship and nuclear waste remediation are inherently chemically complex, involving practically the full diversity of the periodic table, but existing methods are too expensive or not functional for a large diversity of atom types. Overall, the field of machine learning interatomic potentials (MLIPs) has advanced dramatically in 2024 with large high-accuracy datasets existing for bulk, surface, and organic chemical systems and new online leaderboards for diverse chemistry. To participate in, and bring LANL interests into this ecosystem, here, we have built upon existing technologies created by LANL to create a framework capable of creating machine learning interatomic potentials (MLIPs) for over 90 atom types. Our results have created a massively diverse coordination complex training dataset more than 3 times the size of existing datasets, parallelized MLIP training over multiple GPUs, enabling the training of an MLIP spanning the periodic table at 20 times the speed of prior training on 32 GPUs. These advances are substantial towards creation on foundational MLIPs for LANL-specific application areas.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

User Manual - HydraGNN v5.0: Distributed Implementation of Multi-Tasking Graph Neural Networks

This document serves as the user manual for HydraGNN v5.0, a scalable graph neural network (GNN) architecture for simultaneous prediction of multiple target properties using multi-task learning (MTL). This version of HydraGNN has been developed primarily to support the development, training, and deployment of predictive graph-based deep learning (DL) models for atomistic materials modeling. HydraGNN is templated over 13 message-passing policies, including invariant models (GIN, PNA, PNAPlus, GAT, MFC, CGCNN, SAGE, SchNet, DimeNet) and equivariant models (EGNN, PNAEq, PAINN, MACE), and supports distributed training via distributed data parallelism (DDP), DeepSpeed, and Fully Sharded Data Parallelism (FSDP) on leadership-class supercomputers. Although HydraGNN can be applied to problems beyond atomistic materials modeling, its current use is confined to homogeneous graphs. Additional capabilities include machine-learned interatomic potentials with energy-conserving forces, General, Powerful, and Scalable Graph Transformer (GraphGPS) global attention, periodic boundary conditions, hyperparameter optimization, mixed-precision training, and uncertainty quantification.

97 MATHEMATICS AND COMPUTING↗

From atomistic models to machine learning: Predictive design of nanocarbons under extreme conditions

The formation of technologically valuable nanocarbon structures under extreme conditions, such as those produced during high-explosive detonations, remains poorly understood but holds significant potential for the development of controlled synthesis pathways. While detonation shockwaves provide the high-pressure, high-temperature environment required for nanodiamond formation, subsequent cooling and decompression dictate whether the diamond phase is preserved or transformed into other nanocarbon structures. Here, in this study, we employ GPU-accelerated reactive molecular dynamics (ReaxFF) simulations to investigate the graphitization and structural remodeling of detonation nanodiamond under nonlinear quench and pressure-release trajectories. We further investigate how the initial nanodiamond morphology; cuboctahedral, octahedral, or hexagonal prism influences the resulting transformation products. Evolution of nanostructure, allotrope (via simulated x-ray diffraction), carbon hybridization, and ring statistics are tracked during a two-stage quench from 5000 K to 60 GPa. Rapid cooling combined with slow decompression optimizes cubic diamond retention, whereas slow cooling with rapid pressure release promotes surface-to-core graphitization, producing concentric sp 2 -hybridized layers and hollowed inner shells. Octahedral nanodiamonds evolve into carbon nano-onions, initially forming bucky diamonds that progressively transform into fully sp 2 -hybridized structures, while hexagonal prisms preferentially form parallel-stacked graphite layers resembling carbon dots. Transient hexagonal diamond (lonsdaleite) emerges as an interfacial phase, suggesting potential reversibility in the shock-induced graphite-to-diamond transformation pathway transformation route. To extend predictive capabilities, we trained machine learning (ML) regressors on over 10 5 node-hours of molecular dynamics (MD) trajectories. A multilayer perceptron (MLP) model reliably predicts the number of graphitized layers from temperature–pressure trajectories with a coefficient of determination (R 2 ) exceeding 0.90. This high predictive fidelity enables efficient, high-throughput mapping of the synthesis parameter space for optimized graphitization outcomes. Collectively, morphological control combined with optimized quench–decompression conditions promote the selective synthesis of nanocarbon allotropes. This work establishes a data-driven framework for the rational, a priori design of carbon nanomaterials for applications in energy storage, sensing, and biomedicine.

Detonation nanodiamond remodeling↗

Scalable parallel communications

Coarse-grain parallelism in networking (that is, the use of multiple protocol processors running replicated software sending over several physical channels) can be used to provide gigabit communications for a single application. Since parallel network performance is highly dependent on real issues such as hardware properties (e.g., memory speeds and cache hit rates), operating system overhead (e.g., interrupt handling), and protocol performance (e.g., effect of timeouts), we have performed detailed simulations studies of both a bus-based multiprocessor workstation node (based on the Sun Galaxy MP multiprocessor) and a distributed-memory parallel computer node (based on the Touchstone DELTA) to evaluate the behavior of coarse-grain parallelism. Our results indicate: (1) coarse-grain parallelism can deliver multiple 100 Mbps with currently available hardware platforms and existing networking protocols (such as Transmission Control Protocol/Internet Protocol (TCP/IP) and parallel Fiber Distributed Data Interface (FDDI) rings); (2) scale-up is near linear in n, the number of protocol processors, and channels (for small n and up to a few hundred Mbps); and (3) since these results are based on existing hardware without specialized devices (except perhaps for some simple modifications of the FDDI boards), this is a low cost solution to providing multiple 100 Mbps on current machines. In addition, from both the performance analysis and the properties of these architectures, we conclude: (1) multiple processors providing identical services and the use of space division multiplexing for the physical channels can provide better reliability than monolithic approaches (it also provides graceful degradation and low-cost load balancing); (2) coarse-grain parallelism supports running several transport protocols in parallel to provide different types of service (for example, one TCP handles small messages for many users, other TCP's running in parallel provide high bandwidth service to a single application); and (3) coarse grain parallelism will be able to incorporate many future improvements from related work (e.g., reduced data movement, fast TCP, fine-grain parallelism) also with near linear speed-ups.

Maly, K.↗

Generation of gear tooth surfaces by application of CNC machines

This study will demonstrate the importance of application of computer numerically controlled (CNC) machines in generation of gear tooth surfaces with new topology. This topology decreases gear vibration and will extend the gear capacity and service life. A preliminary investigation by a tooth contact analysis (TCA) program has shown that gear tooth surfaces in line contact (for instance, involute helical gears with parallel axes, worm gear drives with cylindrical worms, etc.) are very sensitive to angular errors of misalignment that cause edge contact and an unfavorable shape of transmission errors and vibration. The new topology of gear tooth surfaces is based on the localization of bearing contact, and the synthesis of a predesigned parabolic function of transmission errors that is able to absorb a piecewise linear function of transmission errors caused by gear misalignment. The report will describe the following topics: description of kinematics of CNC machines with six degrees of freedom that can be applied for generation of gear tooth surfaces with new topology. A new method for grinding of gear tooth surfaces by a cone surface or surface of revolution based on application of CNC machines is described. This method provides an optimal approximation of the ground surface to the given one. This method is especially beneficial when undeveloped ruled surfaces are to be ground. Execution of motions of the CNC machine is also described. The solution to this problem can be applied as well for the transfer of machine tool settings from a conventional generator to the CNC machine. The developed theory required the derivation of a modified equation of meshing based on application of the concept of space curves, space curves represented on surfaces, geodesic curvature, surface torsion, etc. Condensed information on these topics of differential geometry is provided as well.

Litvin, F. L.↗

Reaction Mechanism Generator v3.0: Advances in Automatic Mechanism Generation

In chemical kinetics research, kinetic models containing hundreds of species and tens of thousands of elementary reactions are commonly used to understand and predict the behavior of reactive chemical systems. Reaction Mechanism Generator (RMG) is a software suite developed to automatically generate such models by incorporating and extrapolating from a database of known thermochemical and kinetic parameters. Here, we present the recent version 3 release of RMG and highlight improvements since the previously published description of RMG v1.0. Most notably, RMG can now generate heterogeneous catalysis models in addition to the previously available gas- and liquid-phase capabilities. For model analysis, new methods for local and global uncertainty analysis have been implemented to supplement first-order sensitivity analysis. The RMG database of thermochemical and kinetic parameters has been significantly expanded to cover more types of chemistry. The present release includes parallelization for faster model generation and a new molecule isomorphism approach to improve computational performance. RMG has also been updated to use Python 3, ensuring compatibility with the latest cheminformatics and machine learning packages. Overall, RMG v3.0 includes many changes which improve the accuracy of the generated chemical mechanisms and allow for exploration of a wider range of chemical systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Sparse Tensor Benchmark Suite for CPUs and GPUs

Tensor computations present significant performance chal- lenges that impact a wide spectrum of applications ranging from machine learning, healthcare analytics, social network analysis, data mining to quantum chemistry and signal processing. Efforts to improve the perfor- mance of tensor computations include exploring data layout, execution scheduling, and parallelism in common tensor kernels. This work presents a benchmark suite for arbitrary-order sparse tensor kernels using state- of-the-art tensor formats: coordinate (COO) and hierarchical coordinate (HiCOO) on CPUs and GPUs. It presents a set of reference tensor kernel implementations that are compatible with real-world tensors and power law tensors extended from synthetic graph generation techniques. We also propose Roofline performance models for these kernels to provide insights of computer platforms from sparse tensor view. This benchmark suite along with the synthetic tensor generator is publicly available.

Li, Jiajia↗