Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel machines”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Prediction of DIII-D Pedestal Structure From Externally Controllable Parameters

The sharp increase of pressure at the edge of a high confinement mode (H-mode) plasma, the pedestal, strongly impacts overall plasma performance. Predicting the pedestal is a necessity to control and optimize tokamak operations. Here, an experimental data-driven machine learning (ML) approach is presented that predicts the pedestal heights and widths of electron density (n e ) and electron temperature (T e ) profiles as well as the separatrix ne from externally controllable parameters such as the plasma shape, heating method and power, and gas puff rate and integrated gas puff. The OMFIT framework was used with DIII-D data to efficiently, robustly, and automatically build a database of pedestal parameters to train machine learning models. Database creation was enabled by the search engine tool for DIII-D data, TokSearch, which parallelizes data fetching, enabling fast searches through basic signals of thousands of DIII-D shots and selection of relevant time intervals. Principal Component Analysis (PCA) separated the database into three clusters that represent classes of plasma shapes that are regularly used in DIII-D. The most important parameters for setting the pedestal structure were plasma current (I p ), toroidal magnetic field (B Φ ), neutral beam heating power (P NBI ) and shaping quantities. The Deep Jointly Informed Neural Networks (DJINN) algorithm was applied to identify suitable neural network (NN) architectures that appropriately capture the features of the pedestal database. Separate NNs were implemented for each pedestal parameter, and ensembling methods were used to improve the prediction accuracy and allowed estimation of the prediction uncertainty. The pedestal predictions of the test dataset lie within the measurement uncertainties of the pedestal parameters. The NN outperformed simple Linear Regression (LR) analysis, indicating non-linear dependencies in the pedestal structure. The presented achievements illustrate a promising path for future research, using feature extraction to infer experimental trends and thereby improve pedestal models as well as deploying NN for a fast pedestal prediction in DIII-D scenario development.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Deep learning for electron and scanning probe microscopy: From materials design to atomic fabrication

Machine learning and artificial intelligence (ML/AI) are rapidly becoming an indispensable part of physics research, with applications ranging from theory and materials prediction to high-throughput data analysis. In parallel, the recent successes in applying ML/AI methods for autonomous systems from robotics through self-driving cars to organic and inorganic synthesis are generating enthusiasm for the potential of these techniques to enable automated and autonomous experiment in imaging. Here, we discuss recent progress in application of machine learning methods in scanning transmission electron microscopy and scanning probe microscopy, from applications such as data compression and exploratory data analysis to physics learning to atomic fabrication.

36 MATERIALS SCIENCE↗

Sage: parallel semi-asymmtric graph algorithms for NVRAMs

Non-volatile main memory (NVRAM) technologies provide an attractive set of features for large-scale graph analytics, including byte-addressability, low idle power, and improved memory-density. NVRAM systems today have an order of magnitude more NVRAM than traditional memory (DRAM). NVRAM systems could therefore potentially allow very large graph problems to be solved on a single machine, at a modest cost. However, a significant challenge in achieving high performance is in accounting for the fact that NVRAM writes can be much more expensive than NVRAM reads. In this paper, we propose an approach to parallel graph analytics using the Parallel Semi-Asymmetric Model (PSAM), in which the graph is stored as a read-only data structure (in NVRAM), and the amount of mutable memory is kept proportional to the number of vertices. Similar to the popular semi-external and semi-streaming models for graph analytics, the PSAM approach assumes that the vertices of the graph fit in a fast read-write memory (DRAM), but the edges do not. In NVRAM systems, our approach eliminates writes to the NVRAM, among other benefits. To experimentally study this new setting, we develop Sage, a parallel semi-asymmetric graph engine with which we implement provably-efficient (and often work-optimal) PSAM algorithms for over a dozen fundamental graph problems. We experimentally study Sage using a 48--core machine on the largest publicly-available real-world graph (the Hyperlink Web graph with over 3.5 billion vertices and 128 billion edges) equipped with Optane DC Persistent Memory, and show that Sage outperforms the fastest prior systems designed for NVRAM. Importantly, we also show that Sage nearly matches the fastest prior systems running solely in DRAM, by effectively hiding the costs of repeatedly accessing NVRAM versus DRAM.

97 MATHEMATICS AND COMPUTING↗

Studying CPU and memory utilization of applications on Fujitsu A64FX and Nvidia Grace Superchip

ARM-based manycore CPU architectures are well-positioned to provide the rising memory throughput requirements of modern data intensive scientific applications in High Performance Computing (HPC). The Fujitsu A64FX CPU platform is based on the ARM v8.2A architecture, and is the processor of the flagship Japanese supercomputer - "Fugaku", which was previously ranked as the #1 supercomputer in the world according to the Top500 list. The Nvidia Grace superchip features 144 Neoverse V2 cores based on the ARMv9 architecture with 4x128b SVE2, providing exceptional computational power. The chip supports up to 480GB of memory, making it ideal for AI, machine learning, and scientific computing workloads. In this paper, we conduct a thorough performance exploration of a variety of parallel bandwidth-sensitive benchmarks and applications compiled with the native Fujitsu compiler on a Fugaku A64FX compute node and ARM (LLVM) Compiler on an NVIDIA Grace superchip compute node, engaging all the computational cores per cluster using OpenMP multithreading (assuming the cores can drive the available bandwidth). Our ultimate goals are to study the resource utilization of scientific applications and benchmarks on A64FX and Grace superchip, considering graph application scenarios ( GAP Benchmark suite) and eleven appli- cation proxies from the Rodinia heterogeneous benchmark suite (considering domains such as Data Mining, Bioinformatics, Fluid Dynamics, Pattern Recognition, etc.). Through exhaustive performance monitoring, we quantify the resource utilization of diverse OpenMP-based HPC applications on both the Fujitsu A64FX and the Nvidia Grace Superchip platforms.

benchmarking, Performance Analysis, High performan↗

Machine Learning for Distributed Acoustic Sensing data (MLDAS) v1.0.1

MLDAS is a Python-written package for exploratory data analysis and deep learning training on Distributed Acoustic Sensing data. The machine learning tools are powered by the PyTorch library and designed to work efficiently on large scale datasets using parallel computing. Various SLURM scripts as well as a tutorial have also been made available to allow geophysicists to quickly and easily implement the available tools in their analysis workflow on supercomputer facilities.

Dumont, Vincent↗

Dynamic Ride-Matching for Large-Scale Transportation Systems

Efficient dynamic ride-matching (DRM) in large-scale transportation systems is a key driver in transport simulations to yield answers to challenging problems. Although the DRM problem is simple to solve, it quickly becomes a computationally challenging problem in large-scale transportation system simulations. Therefore, this study thoroughly examines the DRM problem dynamics and proposes an optimization-based solution framework to solve the problem efficiently. To benefit from parallel computing and reduce computational times, the problem’s network is divided into clusters utilizing a commonly used unsupervised machine learning algorithm along with a linear programming model. Then, these sub-problems are solved using another linear program to finalize the ride-matching. At the clustering level, the framework allows users adjusting cluster sizes to balance the trade-off between the computational time savings and the solution quality deviation. A case study in the Chicago Metropolitan Area, U.S., illustrates that the framework can reduce the average computational time by 58% at the cost of increasing the average pick up time by 26% compared with a system optimum, that is, non-clustered, approach. Another case study in a relatively small city, Bloomington, Illinois, U.S., shows that the framework provides quite similar results to the system-optimum approach in approximately 62% less computational time.

33 ADVANCED PROPULSION SYSTEMS↗

Prediction of DIII-D Pedestal Structure from Externally Controllable Parameters

The sharp increase of pressure at the edge of a high confinement mode (H-mode) plasma, the pedestal, strongly impacts overall plasma performance. Predicting the pedestal is a necessity to control and optimize tokamak operations. An experimental data-driven machine learning (ML) approach is presented that predicts the pedestal heights and widths of electron density (ne) and electron temperature (Te) profiles as well as the separatrix ne from externally controllable parameters such as the plasma shape, heating method and power, and gas puff rate and integrated gas puff. The OMFIT framework was used with DIII-D data to efficiently, robustly, and automatically build a database of pedestal parameters to train machine learning models. Database creation was enabled by the search engine tool for DIII-D data, TokSearch, which parallelizes data fetching, enabling fast searches through basic signals of thousands of DIII-D shots and selection of relevant time intervals. Principal Component Analysis (PCA) separated the database into three clusters that represent classes of plasma shapes that are regularly used in DIII-D. The most important parameters for setting the pedestal structure were plasma current (Ip), toroidal magnetic field (Bφ), neutral beam heating power (PNBI) and shaping quantities. The Deep Jointly Informed Neural Networks (DJINN) algorithm was applied to identify suitable neural network (NN) architectures that appropriately capture the features of the pedestal database. Separate NNs were implemented for each pedestal parameter, and ensembling methods were used to improve the prediction accuracy and allowed estimation of the prediction uncertainty. The pedestal predictions of the test dataset lie within the measurement uncertainties of the pedestal parameters. The NN outperformed simple Linear Regression (LR) analysis, indicating non-linear dependencies in the pedestal structure. The presented achievements illustrate a promising path for future research, using feature extraction to infer experimental trends and thereby improve pedestal models as well as deploying NN for a fast pedestal prediction in DIII-D scenario development.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Active Learning for Metamaterial Optimization on HPC and QC Integrated Systems

Active learning algorithms, integrating machine learning, quantum computing and optics simulation in an iterative loop, offer a promising approach to optimizing metamaterials. However, these algorithms can face difficulties in optimizing highly complex structures due to computational limitations. High-performance computing (HPC) and quantum computing (QC) integrated systems can address these issues by enabling parallel computing. In this study, we develop an active learning algorithm working on HPC-QC integrated systems. We evaluate the performance of optimization processes within active learning (i.e., training a machine learning model, problem-solving with quantum computing, and evaluating optical properties through wave-optics simulation) for highly complex metamaterial cases. Our results showcase that utilizing multiple cores on the integrated system can significantly reduce computational time, thereby enhancing the efficiency of optimization processes. Therefore, we expect that leveraging HPC-QC integrated systems helps effectively tackle large-scale optimization challenges in general.

Kim, Seongmin↗

Model-based, in-situ, non-destructive qualification and certification of parts made by autonomous additive manufacturing

To address the significant productivity challenges associated with the qualification and certification (Q&C) tasks of additively manufactured (AM) parts, which have traditionally relied on rigorous post‐build inspection and testing, we propose an integrated framework that combines model‐based qualification and certification (MBQ&C) with autonomous additive manufacturing (AAM). MBQ&C employs high‐fidelity predictive models, developed within the Integrated Computational Materials Engineering (ICME) paradigm, to simulate process–structure–property–performance relationships for assessing a part’s fitness for use. Since predictive models are commonly machine learning (ML)-based or reduced-order surrogates of validated physics models, they run efficiently, enabling timely inference. In parallel, the self-driving AAM utilises ML-based adaptive, closed‐loop control strategies to avoid, mitigate, or repair defects and anomalies during fabrication, thereby increasing the likelihood of producing acceptable parts. A key feature of the combined AAM-MBQ&C framework is that predictive models explicitly incorporate defects or anomalies that persist after the build, using instance-specific data captured via in-situ sensing. This customisation enables a build‐specific assessment of fitness for use, rather than relying on nominal or generic parameters. Such individualised evaluation provides a robust basis for Q&C-related acceptance decisions relating to each build. Additionally, the rapid solution capabilities of ML or reduced-order models enable the determination of a part’s suitability for service shortly after build completion. As the framework matures, it has the potential to substantially reduce reliance on conventional point‐design approaches—such as time‐consuming post‐build computed tomography scanning and costly destructive testing. Thus, the AAM-MBQ&C framework represents a transformative, scalable strategy for quality assurance of AM components, as parts produced within a stable, validated, and certified envelope can be certified with reduced testing. Key benefits include: (1) significant gains in Q&C productivity through efficient, model-centric assessment; (2) performance-based classification of defects into critical and non-critical categories; (3) the ability to predict potential deviations in the performance of parts affected by real-time, adaptive process control interventions relative to those produced under a certified process, and (4) the enabling of virtual Q&C for service environments that are difficult, hazardous, or impractical to access or reproduce experimentally. Collectively, these capabilities strengthen the business case for AM, particularly for high‐consequence and mission‐critical applications. Finally, although this work focuses on powder-based AM, the proposed techniques could be extended to AM processes employing alternative feedstock forms.

Gunasegaram, Dayalan↗

Neuromorphic Graph Algorithms: Cycle Detection, Odd Cycle Detection, and Max Flow

Neuromorphic computing is poised to become a promising computing paradigm in the post Moore’s law era due to its extremely low power usage and inherent parallelism. Spiking neural networks are the traditional use case for neuromorphic systems, and have proven to be highly effective at machine learning tasks such as control problems. More recently, neuromorphic systems have been applied outside of the arena of machine learning, primarily in the field of graph algorithms. Neuromorphic systems have been shown to perform graph algorithms faster and with lower power consumption than their traditional (GPU/CPU) counterparts, and are hence an attractive option for a co-processing unit in future high performance computing systems, where graph algorithms play a critical role. In this paper, we present a neuromorphic implementation of cycle detection, odd cycle detection, and the Ford-Fulkerson max-flow algorithm. We further evaluate the performance of these implementations using the NEST neuromorphic simulator by using spike counts and simulation time as proxies for energy consumption and run time. In addition to gains inherent in neuromorphic systems, we show that within the neuromorphic implementations early stopping criteria can be implemented to further improve performance.

Kay, Bill↗

Combined Environments: Driving multiple loads on Z

Pulsed power drivers such as the Z generator of Sandia National Laboratories typically deliver high current (>20MA) to single experiments. This project is intended to develop and assess ways to simultaneously drive multiple targets on a single pulsed power driver (specifically a neutron and an x-ray producing target driven in a single experiment). The combined x-ray/neutron environment produced will then be used to investigate potential synergistic effects in integrated circuits. A pre-requisite for being able to design and study multiple targets on Z is first adapting simulation tools to be able to model them effectively. This will enable us to assess the tradeoffs between the different ways multiple targets can be combined, and to better understand how existing and future pulsed power machines can be used to generate combined testing environments. This report is limited to documenting the initial development of a parallel load modeling capability that is presently being applied to design experiments to produce combined neutron/x-ray environments on Z.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Understanding performance variability in standard and pipelined parallel Krylov solvers

In this work, we collect data from runs of Krylov subspace methods and pipelined Krylov algorithms in an effort to understand and model the impact of machine noise and other sources of variability on performance. We find large variability of Krylov iterations between compute nodes for standard methods that is reduced in pipelined algorithms, directly supporting conjecture, as well as large variation between statistical distributions of runtimes across iterations. Based on these results, we improve upon a previously introduced nondeterministic performance model by allowing iterations to fluctuate over time. We present our data from runs of various Krylov algorithms across multiple platforms as well as our updated non-stationary model that provides good agreement with observations. We also suggest how it can be used as a predictive tool.

97 MATHEMATICS AND COMPUTING↗

Parthenon—a performance portable block-structured adaptive mesh refinement framework

On the path to exascale the landscape of computer device architectures and corresponding programming models has become much more diverse. While various low-level performance portable programming models are available, support at the application level lacks behind. To address this issue, we present the performance portable block-structured adaptive mesh refinement (AMR) framework Parthenon, derived from the well-tested and widely used Athena++ astrophysical magnetohydrodynamics code, but generalized to serve as the foundation for a variety of downstream multi-physics codes. Parthenon adopts the Kokkos programming model, and provides various levels of abstractions from multidimensional variables, to packages defining and separating components, to launching of parallel compute kernels. Parthenon allocates all data in device memory to reduce data movement, supports the logical packing of variables and mesh blocks to reduce kernel launch overhead, and employs one-sided, asynchronous MPI calls to reduce communication overhead in multi-node simulations. Using a hydrodynamics miniapp, we demonstrate weak and strong scaling on various architectures including AMD and NVIDIA GPUs, Intel and AMD x86 CPUs, IBM Power9 CPUs, as well as Fujitsu A64FX CPUs. At the largest scale on Frontier (the first TOP500 exascale machine), the miniapp reaches a total of 1.7 × 10 13 zone-cycles/s on 9216 nodes (73,728 logical GPUs) at [Formula: see text] weak scaling parallel efficiency (starting from a single node). In combination with being an open, collaborative project, this makes Parthenon an ideal framework to target exascale simulations in which the downstream developers can focus on their specific application rather than on the complexity of handling massively-parallel, device-accelerated AMR.

97 MATHEMATICS AND COMPUTING↗

In-depth analysis on parallel processing patterns for high-performance Dataframes

The Data Science domain has expanded monumentally in both research and industry communities during the past decade, predominantly owing to the Big Data revolution. Artificial Intelligence (AI) and Machine Learning (ML) are bringing more complexities to data engineering applications, which are now integrated into data processing pipelines to process terabytes of data. Typically, a significant amount of time is spent on data preprocessing in these pipelines, and hence improving its efficiency directly impacts the overall pipeline performance. The community has recently embraced the concept of Dataframes as the de-facto data structure for data representation and manipulation. However, the most widely used serial Dataframes today (R, pandas) experience performance limitations while working on even moderately large data sets. We believe that there is plenty of room for improvement by taking a look at this problem from a high-performance computing point of view. In a prior publication, we presented a set of parallel processing patterns for distributed dataframe operators and the reference runtime implementation, Cylon. In this paper, we are expanding on the initial concept by introducing a cost model for evaluating the said patterns. Furthermore, we evaluate the performance of Cylon on the ORNL Summit supercomputer.

97 MATHEMATICS AND COMPUTING↗

Machine learning-based optimization of air-cooled heat sinks

Machine learning-based models using Artificial Neural Network (ANN) and greedy search algorithm are used to optimize air-cooled parallel plate-finned heat sinks (PPFHSs) subjected to laminar flow over an extensive range of design parameters. Here, the thermal and hydraulic performances of PPFHSs are represented by heat transfer coefficient (h) and pressure drop (ΔP), respectively. Optimization objectives for PPFHS designs can vary from industry to industry depending on their design priorities. The present study proposes a novel and generalized optimization method that defines practical optimization objectives and provides an accurate optimization process to design effective PPFHSs for a wide range of industrial applications with different design requirements. Three optimization objectives are presented in this study: (i) the largest h PΔ, (ii) the largest h within a specified maximum allowed flow rate, and (iii) the lowest weight that maximizes h for operation within the maximum allowed flow rate. While the shortcoming of the first objective is demonstrated, the other two objectives are found to be suitable for designing effective heat sinks (HSs) across different applications. Results suggest a promising trend from the third objective to develop HSs with ~ 37-68% lower weight, 80-85% reduced ΔP, and negligible penalty in h compared with optimized HSs obtained from the second objective. However, since the third objective leads to HSs with thinner fins, structural analysis should be performed to ensure reliable operation of the HSs.

42 ENGINEERING↗

Exploring model complexity in machine learned potentials for simulated properties

Abstract Machine learning (ML) enables the development of interatomic potentials with the accuracy of first principles methods while retaining the speed and parallel efficiency of empirical potentials. While ML potentials traditionally use atom-centered descriptors as inputs, different models such as linear regression and neural networks map descriptors to atomic energies and forces. This begs the question: what is the improvement in accuracy due to model complexity irrespective of descriptors? We curate three datasets to investigate this question in terms of ab initio energy and force errors: (1) solid and liquid silicon, (2) gallium nitride, and (3) the superionic conductor Li $$_{10}$$ 10 Ge(PS $$_{6}$$ 6 ) $$_{2}$$ 2 (LGPS). We further investigate how these errors affect simulated properties and verify if the improvement in fitting errors corresponds to measurable improvement in property prediction. By assessing different models, we observe correlations between fitting quantity (e.g. atomic force) error and simulated property error with respect to ab initio values. Graphical abstract

Rohskopf, A. (ORCID:0000000227128296)↗

Implementing machine learning methods on QICK hardware for qubit readout & control

Quantum readout and control is a fundamental aspect of quantum computing that requires accurate measurement of qubit states. Errors emerge in all stages, from initialization to readout, and identifying errors in post-processing necessitates resource-intensive statistical analysis. In our work, we use a lightweight fully-connected neural network (NN) to classify states of a transmon system with no prior processing. Our NN accelerator yields higher fidelities (92%) than the classical matched filter method (84%). By exploiting the natural parallelism of NNs and their placement near the source of data on field-programmable gate arrays (FPGAs), we can achieve ultra-low latency on the Quantum Instrumentation Control Kit (QICK). Integrating machine learning methods on QICK opens several pathways for efficient real-time processing of quantum circuits.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Spotlight: efficient automated global optimization in rietveld analysis of diffraction data

Performing reliable Rietveld analysis on tens or hundreds of powder diffraction datasets from parametric or time-resolved experiments often poses a bottleneck in extracting meaningful results from the data. While automated analysis of data has recently been demonstrated, high temperature annealing studies, during which phase transformations occur and lattice parameters may change due to repartitioning of elements, are prime examples where automation by a simple phase identification from a database of room temperature structures or automation by sequential refinements is likely to fail. To enable reliable, efficient, automated Rietveld analysis, we present a Python package named Spotlight , building on established Rietveld packages such as MAUD, GSAS , or GSAS-II , which extends the refinement of best fit parameters to a global optimization using an ensemble of optimizers leveraging hierarchical parallel execution on high-performance computing clusters. Spotlight further enables the efficient design of refinement plans through the iterative automated machine-learning of a surrogate for the refinement on which the global optimizations are performed until results from the surrogate converge to the response surface data. We demonstrate Spotlight with the analysis of uranium molybdenum and Ti–6Al–4V datasets, as well as in two open-source tutorials analyzing aluminium oxide and lead sulphate.

36 MATERIALS SCIENCE↗