Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel machines”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37

Grid-Forming and Grid-Following Inverter Comparison of Droop Response

With the increase in penetration of inverter-based resources (IBRs) in the electrical power system, the ability of these devices to provide grid support to the system has become a necessity. With standards previously developed for the interconnection requirements of grid-following inverters (GFLI) (most commonly photovoltaic inverters), it has been well documented how these inverters “should” respond to changes in voltage and frequency. However, with other IBRs such as grid-forming inverters (GFMIs) (used for energy storage systems, standalone systems, and as uninterruptable power supplies) these requirements are either: not yet documented, or require a more in deep analysis. With the increased interest in microgrids, GFMIs that can be paralleled onto a distribution system have become desired. With the proper control schemes, a GFMI can help maintain grid stability through fast response compared to rotating machines. This paper will present an experimental comparison of commercially available GFMI and GFLI ' responses to voltage and frequency deviation, as well as the GFMI operating as a standalone system and subjected to various changes in loads.

Grid Support, Inverter, Droop Control, Volt-VAR, F↗

Draft ASME Code Case to qualify L-PBF 316H material for Section III, Division 5 applications

This report documents the AMMT program’s development and submission of a draft ASME Code Case to qualify Laser Powder Bed Fusion (L PBF) Type 316H stainless steel for Section III, Divi-sion 5 Class A and SM high temperature nuclear applications. It summarizes the technical basis, the comprehensive high temperature mechanical test database assembled between 2023–2026, and the proposed code language and qualification framework submitted to ASME. The work was co-ordinated across multiple national laboratories and leverages prior ASME efforts to integrate additive manufacturing into the Boiler & Pressure Vessel Code. The body of the report describes the experimental database and analysis supporting the Code Case: tensile, creep, fatigue, creep fatigue, and thermal aging tests collected from multiple additive manufacturing sites, machine types, and powder lots, with material processed by a solution anneal heat treatment. The dataset — including both full size and subsized specimens and tests oriented parallel and perpendicular to build direction — shows limited tensile anisotropy, tensile properties comparable to wrought 316H, creep strength within the scatter of wrought material, but markedly reduced creep ductility above about 650 °C associated with rapid σ phase formation in L PBF microstructures. The draft Code Case itself prescribes a staged qualification model (manufacturing process qualification, component qualification, and per build witness testing), treats L PBF components as equivalent to Type 316 weld metal for design and inspection, and requires mechanical, chemical, and metallographic controls tied to ASTM/ISO 52946. Key acceptance criteria include tensile tests within a 90% prediction interval of the AMMT dataset, a creep fatigue screening test adapted from ASME Section III, Division 5, Subsection HB, HBB 2800 but with the cycle acceptance reduced to 100 for L PBF material, and double volumetric inspection of production components. The report concludes that the present data support treating L PBF 316H as analogous to conventional fusion weld metal for Division 5 design and inspection, while highlighting important caveats: the σ phase driven loss of creep ductility above ~650 °C, preliminary indications of enhanced creep fatigue sensitivity in some lots, and remaining gaps in long term aging and additional cyclic testing. Recommended next actions include completing outstanding cyclic and long duration creep/aging tests on the solution annealed condition, supporting inclusion of the 316H chemistry and heat treatment in ASTM/ISO 52946, and continuing engagement with ASME and NRC during balloting and review to enable industry adoption.

Messner, Mark C. (ORCID:0000000200404385)↗

Pattern recognition and control in manipulation

A new approach to the use of sensors in manipulator or robot control is discussed. The concept addresses the problem of contact or near-contact type of recognition of three-dimensional forms of objects by proprioceptive and/or exteroceptive sensors integrated with the terminal device. This recognition of object shapes both enhances and simplifies the automation of object handling. Several examples have been worked out for the 'Belgrade hand' and for a parallel jaw terminal device, both equipped with proprioceptive (position) and exteroceptive (proximity) sensors. The control applications are discussed in the framework of a multilevel man-machine system control. The control applications create interesting new issues which, in turn, invite novel theoretical considerations. An important issue is the problem of stability in control when the control is referenced to patterns.

Bejczy, A. K.↗

Efficient Parallel Sparse Symmetric Tucker Decomposition for High-Order Tensors

Tensor based methods are receiving renewed attention in recent years due to their prevalence in diverse real-world applications. There is considerable literature on tensor representations and algorithms for tensor decompositions, both for dense and sparse tensors. Many applications in hypergraph analytics, machine learning, psychometry, and signal processing result in tensors that are both sparse and symmetric, making it an important class for further study. Similar to the critical Tensor Times Matrix chain operation (TTMc) in general sparse tensors, the Sparse Symmetric Tensor Times Same Matrix chain (S3TTMc) operation is compute and memory intensive due to high tensor order and the associated factorial explosion in the number of non-zeros. In this work, we present a novel compressed storage format CSS for sparse symmetric tensors, along with an efficient parallel algorithm for the S3TTMc operation. We theoretically establish that S3TTMc on CSS achieves a better memory versus run-time trade-off compared to state-of-the-art implementations. We demonstrate experimental findings that confirm these results and achieve up to 2.9× speedup on synthetic and real datasets.

Shivakumar, Shruti↗

An updated LLVM-based quantum research compiler with further OpenQASM support

Abstract Quantum computing is a rapidly growing field with the potential to change how we solve previously intractable problems. Emerging hardware is approaching a complexity that requires increasingly sophisticated programming and control. Scaffold is an older quantum programming language that was originally designed for resource estimation for far-future, large quantum machines, and ScaffCC is the corresponding LLVM-based compiler. For the first time, we provide a full and complete overview of the language itself, the compiler as well as its pass structure. While previous works Abhari et al (2015 Parallel Comput. 45 2–17), Abhari et al (2012 Scaffold: quantum programming language https://cs.princeton.edu/research/techreps/TR-934-12 ), have piecemeal descriptions of different portions of this toolchain, we provide a more full and complete description in this paper. We also introduce updates to ScaffCC including conditional measurement and multidimensional qubit arrays designed to keep in step with modern quantum assembly languages, as well as an alternate toolchain targeted at maintaining correctness and low resource count for noisy-intermediate scale quantum (NISQ) machines, and compatibility with current versions of LLVM and Clang. Our goal is to provide the research community with a functional LLVM framework for quantum program analysis, optimization, and generation of executable code.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Quantum-inspired tempering for ground state approximation using artificial neural networks

A large body of work has demonstrated that parameterized artificial neural networks (ANNs) can efficiently describe ground states of numerous interesting quantum many-body Hamiltonians. However, the standard variational algorithms used to update or train the ANN parameters can get trapped in local minima, especially for frustrated systems and even if the representation is sufficiently expressive. We propose a parallel tempering method that facilitates escape from such local minima. This methods involves training multiple ANNs independently, with each simulation governed by a Hamiltonian with a different "driver" strength, in analogy to quantum parallel tempering, and it incorporates an update step into the training that allows for the exchange of neighboring ANN configurations. We study instances from two classes of Hamiltonians to demonstrate the utility of our approach using Restricted Boltzmann Machines as our parameterized ANN. The first instance is based on a permutation-invariant Hamiltonian whose landscape stymies the standard training algorithm by drawing it increasingly to a false local minimum. The second instance is four hydrogen atoms arranged in a rectangle, which is an instance of the second quantized electronic structure Hamiltonian discretized using Gaussian basis functions. We study this problem in a minimal basis set, which exhibits false minima that can trap the standard variational algorithm despite the problem’s small size. We show that augmenting the training with quantum parallel tempering becomes useful to finding good approximations to the ground states of these problem instances.

Albash, Tameem↗

Optimizing Management of Persistent Data Structures in High-Performance Analytics

Large-scale data analytics workflows ingest massive input data into various data structures, including graphs and key-value datastores. These data structures undergo multiple transformations and computations and are typically reused in incremental and iterative analytics workflows. Persisting in-memory views of these data structures enables reusing them beyond the scope of a single program run while avoiding repetitive raw data ingestion overheads. Memory-mapped I/O enables persisting in-memory data structures without data serialization and deserialization overheads. However, memory-mapped I/O lacks the key feature of persisting consistent snapshots of these data structures for incremental ingestion and processing. The obstacles to efficient virtual memory snapshots using memory-mapped I/O include background writebacks outside the application’s control, and the significantly high storage footprint of such snapshots. To address these limitations, we present Privateer, a memory and storage management tool that enables storage-efficient virtual memory snapshotting while also optimizing snapshot I/O performance. Here, we integrated Privateer into Metall, a state-of-the-art persistent memory allocator for C++, and the Lightning Memory-Mapped Database (LMDB), a widely-used key-value datastore in data analytics and machine learning. Privateer optimized application performance by 1.22× when storing data structure snapshots to node-local storage, and up to 16.7× when storing snapshots to a parallel file system. Privateer also optimizes storage efficiency of incremental data structure snapshots by up to 11× using data deduplication and compression.

Computer science↗

A GPU-Accelerated Population Generation, Sorting, and Mutation Kernel for an Optimization-Based Causal Inference Model

We develop a GPU-accelerated machine learning generative adversarial network model that can be used with observational data for the purpose of constructing causal inferences. The theoretical basis of our machine learning model is novel and is conceptualized to be operable and scalable for high performance computing platforms. Our GPU-accelerated code enables large-scale parallelization of the computation within a common and accessible computing environment. This will expand the reach of our model and empower research in new substantive domains while maintaining the underlying theoretical properties.

Cho, Wendy K. Tam↗

EJFAT: Towards Intelligent Compute Destination Load Balancing

To handle increased data flow, Jefferson Lab (JLab) is partnering with ESnet for development of an AI/ML directed compute work Load Balancer (LB) of UDP streamed data. The LB is FPGA based featuring dynamically configurable, low latency and high throughput destination address switching. The LB provides integration of edge and core computing to support JLab experimental programs, the Electron-Ion Collider, as well as data centers of the future. In the ESnet/JLab FPGA Accelerated Transport (EJFAT) initiative, the function of the LB Data Plane (DP) is to redirect data streams to selectable (but unknown to sender) destination hosts based on current worload and within that host to destination ports as a function of sub- stream id. This effects hierarchical scaling, first across compute machines for processing over a series of events and second, across ports so different data source sub-streams may be assigned to different processors for further parallelization. The LB Control Plane (CP) programs the DP using compute farm telemetry to direct and balance workloads across a compute cluster as the operating conditions require. While Proportional/Integrative/Derivative (PID) controllers are often seen in similar applications, here we investigate the feasibility of a Reinforcement Learning (RL) based schedule manager running in the CP to provide dynamic updates to the DP scheduling policy.

Lawrence, David↗

Inventory estimation on the massively parallel processor

This paper describes algorithms for efficiently computing inventory estimates from satellite based images. The algorithms incorporate a one dimensional feature extraction which optimizes the pairwise sum of Fisher distances. Biases are eliminated with a premultiplication by the inverse of the analytically derived error matrix. The technique is demonstrated with a numerical example using statistics obtained from an actual Landsat scene. Attention was given to implementation of the Massively Parallel processor (MPP). A timing analysis demonstrates that the inventory estimation can be performed an order of magnitude faster on the MPP than on a conventional serial machine.

Argentiero, P. D.↗

Concurrent processing in nonlinear structural stability

A concurrent processing algorithm is developed for materially nonlinear stability analysis of imperfect columns with biaxial partial rotational end restraints. The algorithm for solving the governing nonlinear ordinary differential equations is implemented on a multiprocessor computer called the 'Finite Element Machine', developed at the NASA Langley Research Center. Numerical results are obtained on up to nine concurrent processors. A substantial computational gain is achieved in using the parallel processing approach.

Darbhamulla, S. P.↗

Highly parallel computation

Highly parallel computing architectures are the only means to achieve the computation rates demanded by advanced scientific problems. A decade of research has demonstrated the feasibility of such machines and current research focuses on which architectures designated as multiple instruction multiple datastream (MIMD) and single instruction multiple datastream (SIMD) have produced the best results to date; neither shows a decisive advantage for most near-homogeneous scientific problems. For scientific problems with many dissimilar parts, more speculative architectures such as neural networks or data flow may be needed.

Denning, Peter J.↗

Dynamic remapping of parallel computations with varying resource demands

A large class of computational problems is characterized by frequent synchronization, and computational requirements which change as a function of time. When such a problem must be solved on a message passing multiprocessor machine, the combination of these characteristics lead to system performance which decreases in time. Performance can be improved with periodic redistribution of computational load; however, redistribution can exact a sometimes large delay cost. We study the issue of deciding when to invoke a global load remapping mechanism. Such a decision policy must effectively weigh the costs of remapping against the performance benefits. We treat this problem by constructing two analytic models which exhibit stochastically decreasing performance. One model is quite tractable; we are able to describe the optimal remapping algorithm, and the optimal decision policy governing when to invoke that algorithm. However, computational complexity prohibits the use of the optimal remapping decision policy. We then study the performance of a general remapping policy on both analytic models. This policy attempts to minimize a statistic W(n) which measures the system degradation (including the cost of remapping) per computation step over a period of n steps. We show that as a function of time, the expected value of W(n) has at most one minimum, and that when this minimum exists it defines the optimal fixed-interval remapping policy. Our decision policy appeals to this result by remapping when it estimates that W(n) is minimized. Our performance data suggests that this policy effectively finds the natural frequency of remapping. We also use the analytic models to express the relationship between performance and remapping cost, number of processors, and the computation's stochastic activity.

Nicol, D. M.↗

Nanosecond machine learning regression with deep boosted decision trees in FPGA for high energy physics

We present a novel application of the machine learning / artificial intelligence method called boosted decision trees to estimate physical quantities on field programmable gate arrays (FPGA). The software package fwXmachina features a new architecture called parallel decision paths that allows for deep decision trees with arbitrary number of input variables. It also features a new optimization scheme to use different numbers of bits for each input variable, which produces optimal physics results and ultraefficient FPGA resource utilization. Problems in high energy physics of proton collisions at the Large Hadron Collider (LHC) are considered. Estimation of missing transverse momentum (E T miss ) at the first level trigger system at the High Luminosity LHC (HL-LHC) experiments, with a simplified detector modeled by Delphes, is used to benchmark and characterize the firmware performance. The firmware implementation with a maximum depth of up to 10 using eight input variables of 16-bit precision gives a latency value of $\mathcal{O}$(10) ns, independent of the clock speed, and $\mathcal{O}$(0.1)% of the available FPGA resources without using digital signal processors.

Instruments & Instrumentation↗

Implementation of an ADI method on parallel computers

In this paper the implementation of an ADI method for solving the diffusion equation on three parallel/vector computers is discussed. The computers were chosen so as to encompass a variety of architectures. They are the MPP, an SIMD machine with 16-Kbit serial processors; Flex/32, an MIMD machine with 20 processors; and Cray/2, an MIMD machine with four vector processors. The Gaussian elimination algorithm is used to solve a set of tridiagonal systems on the Flex/32 and Cray/2 while the cyclic elimination algorithm is used to solve these systems on the MPP. The implementation of the method is discussed in relation to these architectures and measures of the performance on each machine are given. Simple performance models are used to describe the performance. These models highlight the bottlenecks and limiting factors for this algorithm on these architectures. Finally conclusions are presented.

Fatoohi, Raad A.↗

Towards real-time simulation of large space structures: Stabilization of fluid/thermal/structure interactions and implementation on high performance supercomputers

Within the Center for Space Construction, the SIMSTRUC project's objectives center around the development of simulation tools for the realistic analysis of large space structures. The word 'tools' is the broad sense; it designates mathematical models, finite element/finite difference formulations, computational algorithms, implementations on advanced computer architectures, and visualization capabilities. The results of our activities during the first year within the SIMSTRUC project are reported. On the modeling side, an alternative approach to fluid/thermal/structure interaction analysis that is a departure from the 'loosely coupled' and 'unified' approaches that are being currently practiced are described. The advantages of our approach both in terms of accuracy and computational efficiency were demonstrated. On the computational side, a software architecture for parallel/vector and massively parallel supercomputers that speeds up finite element and finite difference computations by several orders of magnitude is presented. As an example, the simulation of the deployment of a space structure that used to require over six hours of a workstation using a conventional finite element software, now runs on a multiprocessor using a parallel computation strategy in less than three seconds. In order to promote the physical understanding of the simulation behavior, a real-time visualization capability on the Connection Machine, which allows the analyst to watch the graphical animation of the results at the same time these are generated, was also developed. It is believed that by combining efficient analytical formulations with the state-of-the-art high performance computer implementations and superfast visualization capabilities, SIMSTRUC is moving fast towards the real-time simulation of large space structures. The designers as well as the researchers will certainly benefit from this technology.

Farhat, C.↗

Microsystem Cooler Development

A patented microsystem Stirling cooler is under development with potential application to electronics, sensors, optical and radio frequency (RF) systems, microarrays, and other microsystems. The microsystem Stirling cooler is most suited to volume-limited applications that require cooling below the ambient or sink temperature. Primary components of the planar device include: two diaphragm actuators that replace the pistons found in traditional-scale Stirling machines; and a micro-regenerator that stores and releases thermal energy to the working gas during the Stirling cycle. The use of diaphragms eliminates frictional losses and bypass leakage concerns associated with pistons, while permitting reversal of the hot and cold sides of the device during operation to allow precise temperature control. Three candidate microregenerators were custom fabricated for initial evaluation: two constructed of porous ceramic, and one made of multiple layers of nickel and photoresist in an offset grating pattern. An additional regenerator was prepared with a random stainless steel fiber matrix commonly used in existing Stirling machines for comparison to the custom fabricated regenerators. The candidate regenerators were tested in a piezoelectric-actuated test apparatus designed to simulate the Stirling refrigeration cycle. In parallel with the regenerator testing, electrostatically-driven comb-drive diaphragm actuators for the prototype device have been designed for deep reactive ion etching (DRIE) fabrication.

Moran, Matthew E.↗

Accelerating Scientific Workflows on HPC Platforms with In Situ Processing

Scientific workflows drive most modern large-scale science breakthroughs by allowing scientists to define their computations as a set of jobs executed in a given order based on their data dependencies. Workflow management systems (WMSs) have become key to automating scientific workflows-executing computational jobs and orchestrating data transfers between those jobs running on complex high-performance computing (HPC) platforms. Traditionally, WMSs use files to communicate between jobs: a job writes out files that are read by other jobs. However, HPC machines face a growing gap between their storage and compute capabilities. To address that concern, the scientific community has adopted a new approach called in situ, which bypasses costly parallel filesystem I/O operations with faster in-memory or in-network communications. When using in situ approaches, communication and computations can be interleaved. In this work, we leverage the Decaf in situ dataflow framework to accelerate task-based scientific workflows managed by the Pegasus WMS, by replacing file communications with faster MPI messaging. We propose a new execution engine that uses Decaf to manage communications within a sub-workflow (i.e., set of jobs) to optimize inter-job communications. We consider two workflows in this study: (i) a synthetic workflow that benchmarks and compares file- and MPI-based communication; and (ii) a realistic bioinformatics workflow that computes mu-tational overlaps in the human genome. Experiments show that in situ communication can improve the bioinformatics workflow execution time by 22% to 30% compared with file communication. Our results motivate further opportunities and challenges for bridging traditional WMSs with in situ frameworks.

Decaf↗