Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel machines”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Estimating the randomness of quantum circuit ensembles up to 50 qubits

Random quantum circuits have been utilized in the contexts of quantum supremacy demonstrations, variational quantum algorithms for chemistry and machine learning, and blackhole information. The ability of random circuits to approximate any random unitaries has consequences on their complexity, expressibility, and trainability. To study this property of random circuits, we develop numerical protocols for estimating the frame potential, the distance between a given ensemble and the exact randomness. Our tensor-network-based algorithm has polynomial complexity for shallow circuits and is high-performing using CPU and GPU parallelism. We study 1. local and parallel random circuits to verify the linear growth in complexity as stated by the Brown–Susskind conjecture, and; 2. hardware-efficient ansätze to shed light on its expressibility and the barren plateau problem in the context of variational algorithms. Our work shows that large-scale tensor network simulations could provide important hints toward open problems in quantum information science.

97 MATHEMATICS AND COMPUTING↗

Mitigating Inter-Job Interference via Process-Level Quality-of-Service

Jobs on most high-performance computing (HPC) systems share the network with other concurrently executing jobs. Network sharing leads to contention that can severely degrade performance. Here we investigate the use of Quality of Service (QoS) mechanisms to reduce the negative impacts of network contention. QoS allows users to manage resource sharing between network flows and to provide bandwidth guarantees to specific flows. Our results show that careful use of QoS reduces the impact of network contention for specific jobs, resulting in up to a 40% performance improvement. In some cases, it completely eliminates the impact of contention. It achieves these improvements with limited negative impact to other jobs; any job that experiences performance loss typically degrades less than 5%, and often much less. Our approach can help ensure that HPC machines maintain high levels of throughput as per-node compute power continues to increase faster than network bandwidth.

97 MATHEMATICS AND COMPUTING↗

A semi-automated algorithm for designing stellarator divertor and limiter plates and application to HSX

We present a semi-automated algorithm for designing three-dimensional divertor or limiter plates targeting low heat loads. The algorithm designs the plates in two stages: firstly, the parallel heat flux distribution is caught on vertically-inclined plates at one or several toroidal locations. Secondly, the power per unit area is reduced by stretching, tilting and bending the plates toroidally. Heat transport is modelled using the EMC3-Lite code, which uses an anisotropic diffusion model. We apply this scheme to HSX, a medium-sized stellarator located at the University of Wisconsin–Madison. Starting from the current machine with an extended vessel wall, we construct plates which are able to effectively catch and spread the heat for three different magnetic configurations. The scheme has a computational cost in the order of tens of CPU-minutes, making it a powerful tool for semi-automated plasma-facing component design in three-dimensional environments.

anisotropic diffusion↗

Track reconstruction as a service for collider physics

Optimizing charged-particle track reconstruction algorithms is crucial for efficient event reconstruction in Large Hadron Collider (LHC) experiments due to their significant computational demands. Existing track reconstruction algorithms have been adapted to run on massively parallel coprocessors, such as graphics processing units (GPUs), to reduce processing time. Nevertheless, challenges remain in fully harnessing the computational capacity of coprocessors in a scalable and non-disruptive manner. This paper proposes an inference-as-a-service approach for particle tracking in high energy physics experiments. To evaluate the efficacy of this approach, two distinct tracking algorithms are tested: Patatrack, a rule-based algorithm, and Exa.TrkX, a machine learning-based algorithm. The as-a-service implementations show enhanced GPU utilization and can process requests from multiple CPU cores concurrently without increasing per-request latency. The impact of data transfer is minimal and insignificant compared to running on local coprocessors. This approach greatly improves the computational efficiency of charged particle tracking, providing a solution to the computing challenges anticipated in the High-Luminosity LHC era.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

AI ATAC 1: An Evaluation of Prominent Commercial Malware Detectors

This work presents an evaluation of six prominent commercial endpoint malware detectors, a network malware detector, and a file-conviction algorithm from a cyber technology vendor. The evaluation was administered as the first of the Artificial I ntelligence Applications t o Autonomous Cybersecurity (AI ATAC) prize challenges, funded by / completed in service of the US Navy. The experiment employed 100K files (50/50% benign/malicious) with a stratified distribution of file types, including ~1K zero-day program executables (increasing experiment size two orders of magnitude over previous work). We present an evaluation process of delivering a file to a fresh virtual machine donning the detection technology, waiting 90s to allow static detection, then executing the file and waiting another period for dynamic detection; this allows greater fidelity in the observational data than previous experiments, in particular, resource and time-to-detection statistics. To execute all 800K trials (100K files × 8 tools), a software framework is designed to choreograph the experiment into an automated, time-synced, and reproducible workflow with substantial parallelization. Software with base classes for this framework are provided. A cost-benefit model was configured to integrate the tools’ detection statistics into a comparable quantity by simulating costs of use. This provides a ranking methodology for cyber competitions and a lens for reasoning about the varied statistical results. The results provide insights on state of commercial malware detection.

Bridges, Robert↗

A non-cooperative meta-modeling game for automated third-party calibrating, validating and falsifying constitutive laws with parallelized adversarial attacks

The evaluation of constitutive models, especially for high-risk and high-regret engineering applications, requires efficient and rigorous third-party calibration, validation and falsification. While there are numerous efforts to develop paradigms and standard procedures to validate models, difficulties may arise due to the sequential, manual, and often biased nature of the commonly adopted calibration and validation processes, thus slowing down data collections, hampering the progress towards discovering new physics, increasing expenses and possibly leading to misinterpretations of the credibility and application ranges of proposed models. This work attempts to introduce concepts from game theory and machine learning techniques to overcome many of these existing difficulties. Here, we introduce an automated meta-modeling game where two competing AI agents systematically generate experimental data to calibrate a given constitutive model and to explore its weakness such that the experiment design and model robustness can be improved through competitions. The two agents automatically search for the Nash equilibrium of the meta-modeling game in an adversarial reinforcement learning framework without human intervention. In particular, a protagonist agent seeks to find the more effective ways to generate data for model calibrations, while an adversary agent tries to find the most devastating test scenarios that expose the weaknesses of the constitutive model calibrated by the protagonist. By capturing all possible design options of the laboratory experiments into a single decision tree, we recast the design of experiments as a game of combinatorial moves that can be resolved through deep reinforcement learning by the two competing players. Our adversarial framework emulates idealized scientific collaborations and competitions among researchers to achieve a better understanding of the application range of the learned material laws and prevent misinterpretations caused by conventional AI-based third-party validation. Numerical examples are given to demonstrate the wide applicability of the proposed meta-modeling game with adversarial attacks on both human-crafted constitutive models and machine learning models.

97 MATHEMATICS AND COMPUTING↗

MBX V1.2: Accelerating Data-Driven Many-Body Molecular Dynamics Simulations

The MBX software provides an advanced platform for molecular dynamics simulations, leveraging state-of-the-art MB-pol and MB-nrg data-driven many-body potential energy functions. Developed over the past decade, these potential energy functions integrate physics-based and machine-learned many-body terms trained on electronic structure data calculated at the "gold standard" coupled-cluster level of theory. Recent advancements in MBX have focused on optimizing its performance, resulting in the release of MBX v1.2. While the inherently many-body nature of MB-pol and MB-nrg ensures high accuracy, it poses computational challenges. MBX v1.2 addresses these challenges with significant performance improvements, including enhanced parallelism that fully harnesses the power of modern multicore CPUs. In conclusion, these advancements enable simulations on nanosecond time scales for condensed-phase systems, significantly expanding the scope of high-accuracy, predictive simulations of complex molecular systems powered by data-driven many-body potential energy functions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Performance Analysis and Optimal Node-aware Communication for Enlarged Conjugate Gradient Methods

Krylov methods are a key way of solving large sparse linear systems of equations but suffer from poor strong scalability on distributed memory machines. Furthermore, this is due to high synchronization costs from large numbers of collective communication calls alongside a low computational workload. Enlarged Krylov methods address this issue by decreasing the total iterations to convergence, an artifact of splitting the initial residual and resulting in operations on block vectors. In this article, we present a performance study of an enlarged Krylov method, Enlarged Conjugate Gradients (ECG), noting the impact of block vectors on parallel performance at scale. Most notably, we observe the increased overhead of point-to-point communication as a result of denser messages in the sparse matrix-block vector multiplication kernel. Additionally, we present models to analyze expected performance of ECG, as well as motivate design decisions. Most importantly, we introduce a new point-to-point communication approach based on node-aware communication techniques that increases efficiency of the method at scale.

97 MATHEMATICS AND COMPUTING↗

Optimizing transmit field inhomogeneity of parallel RF transmit design in 7T MRI using deep learning

Ultrahigh field (UHF) Magnetic Resonance Imaging (MRI) provides a higher signal-to-noise ratio and, thereby, higher spatial resolution. However, UHF MRI introduces challenges such as transmit radiofrequency (RF) field (B+1) inhomogeneities, leading to uneven flip angles and image intensity anomalies. These issues can significantly degrade imaging quality and its medical applications. This study addresses B+1 field homogeneity through a novel deep learning-based strategy. Traditional methods like Magnitude Least Squares (MLS) optimization have been effective but are time-consuming and dependent on the patient’s presence. Recent machine learning approaches, such as RF Shim Prediction by Iteratively Projected Ridge Regression and deep learning frameworks, have shown promise but face limitations like extensive training times and oversimplified architectures. We propose a two-step deep learning strategy. First, we obtain the desired reference RF shimming weights from multi-channel B+1 fields using random-initialized Adaptive Moment Estimation. Then, we employ Residual Networks (ResNets) to train a model that maps B+1 fields to target RF shimming outputs. Our approach does not rely on pre-calculated reference optimizations for the testing process and efficiently learns residual functions. Comparative studies with traditional MLS optimization demonstrate our method’s advantages in terms of speed and accuracy. The proposed strategy achieves a faster and more efficient RF shimming design, significantly improving imaging quality at UHF. This advancement holds potential for broader applications in medical imaging and diagnostics.

Lu, Zhengyi [Vanderbilt University]↗

Phasor-Measurement-Unit-Based Data Analytics Using Digital Twin and PhasorAnalytics Software

A major objective of this project was to apply GE’s commercial machine learning and data analytics toolsets to large-scale, real-world, anonymized Phasor Measurement Unit (PMU) datasets in order to extract signatures, correlated and/or causal factors, and precursor patterns associated with significant power system phenomena. The project had a particular emphasis on extraction of insights relevant to asset health monitoring, real-time load modeling and cybersecurity monitoring. Additionally, the team was directed to undertake a comprehensive data quality analysis for the provided datasets and encouraged to estimate the ‘machine-learning readiness’ of the datasets by documenting any major obstacles to the application of commercial machine learning algorithms. To accomplish the aforementioned objectives, the project team’s work centered around the identification of key event signatures and application of the identified event signatures for event detection and event classification. The industry-validated, semi-supervised machine learning strategy employed for event signature identification involved several major tasks, including data-preprocessing, generation of an overabundance of features, normal data identification, normality modeling, and event signature identification through a methodical, quantitative ranking of features in order of relevance to each studied event type. Throughout the project, data quality issues and mitigation techniques were investigated. In this report, insights are provided regarding the readiness of the provided synchrophasor datasets for application of machine learning and data analytics. The methodologies employed for this technical strategy are summarized in this report. With regards to data preprocessing and feature generation, the provided Training and Test Datasets were ingested into GE’s big data environment. Subsequently, the team applied bad data cleansing and data imputation scripts, event detection scripts, and application programming interfaces (APIs) to the datasets for convenient data access. The project team completed development and validation of dozens of physics-based, statistics-based and transformation-based feature functions used for the extraction of over 60 synchrophasor features. Using a new parallel feature generation technology developed on this project, over 60 features have been rapidly generated for the full two years’ worth of Training and Test Dataset data associated with both the Eastern and Western interconnects. Even accommodating for temporal down-sampling inherent to the feature extraction procedure, this parallel feature generation activity resulted in a massive feature set with a storage requirement approximately equal to that of the raw training dataset itself. With regards to normal data identification and normality modeling, a normality model was built using the feature data extracted from the Training Dataset and iteratively refined subsequent to incremental adjustments and expansions of the Training Dataset feature data. With respect to event characterization and signature identification, an event signature identification pipeline was developed and used in conjunction with the normality model to identify over 15 event signatures for key event categories within the Training Dataset. The identified event signatures were used to characterize hundreds of key events in terms of relative severity, duration, and location of the event. An investigation was undertaken to identify correlated and causal factors involved in transformer events. A separate investigation into temporal trends in ring-down analysis results was undertaken to determine possible associations between system dynamics and various other factors such as loading, season or year. To validate the identified event signatures, additional work was undertaken to develop signature-based anomaly detection and classification tools suitable for convenient application to the synchrophasor datasets. The anomaly detection and classification tools, suitable for online application, were then applied to the entirety of the Eastern Interconnect Training and Test Datasets. Performance of the event detection and classification tools was evaluated upon receipt of the Test Dataset event logs (i.e., the labels for events contained in the Test Dataset), and promising results were obtained despite several challenges (documented herein) associated with application of supervised or semi-supervised machine learning methods to large-scale, anonymized datasets. Finally, the detection and classification tools were used to detect, classify, and characterize thousands of new events not included in the original event logs provided by the DOE within both the Training and Test Datasets.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Fokker-Planck simulations of fast ion ICRF and electron EC heating in a mirror plasma using CQL3D-m

The CQL3D-m continuum bounce-average Fokker-Planck code is adapted for magnetic mirror plasmas [1] and is now routinely used in no-free-parameter classical integrated modeling of mirror devices [2, 3]. In the present effort, we report on two RF methods of plasma heating in mirror machine. The fast ions (FI) are heated by Fast waves at 2nd-4th harmonic, where FIs originate from neutral beam injection at 45 degrees to the magnetic field. The scenario shows an efficient ion heating near the FI bouncing point. The electrons are heated by X-mode launched from the high magnetic field side towards the resonance. Different from the tokamak applications, CQL3D-m provides an evolving self-consistent ambipolar parallel electric field, which determines the shape of the loss cone and hence an accurate confinement time of both ions and electrons. Also, it includes a description of ion and electron sources and sinks (related to charge exchange and impact ionization) which are updated at every time step. CQL3D-m utilizes a fully nonlinear Coulomb collision operator that is important for the significantly non-Maxwellian ion distributions typically established in mirror plasmas.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Device for controlling additive manufacturing machinery

A computing device for controlling the operation of an additive manufacturing machine comprises a memory element and a processing element. The memory element is configured to store a three-dimensional model of a part to be manufactured, wherein the three-dimensional model defines a plurality of cross sections of the part. The processing element is in communication with the memory element. The processing element is configured to receive the three-dimensional model, determine a plurality of paths, each path including a plurality of parallel lines, determine a radiation beam power for each line, such that the radiation beam power varies non-linearly according to a length of the line, and determine a radiation beam scan speed for each line, such that the radiation beam scan speed is a function of a temperature of a material used to manufacture the part, the length of the line, and the radiation beam power for the line.

Barr, Christian↗

Device for controlling additive manufacturing machinery

A computing device for controlling the operation of an additive manufacturing machine comprises a memory element and a processing element. The memory element is configured to store a three-dimensional model of a part to be manufactured, wherein the three-dimensional model defines a plurality of cross sections of the part. The processing element is in communication with the memory element. The processing element is configured to receive the three-dimensional model, determine a plurality of paths, each path including a plurality of parallel lines, determine a radiation beam power for each line, such that the radiation beam power varies non-linearly according to a length of the line, and determine a radiation beam scan speed for each line, such that the radiation beam scan speed is a function of a temperature of a material used to manufacture the part, the length of the line, and the radiation beam power for the line.

Barr, Christian↗

Anticipating gelation and vitrification with medium amplitude parallel superposition (MAPS) rheology and artificial neural networks

Abstract Anticipating qualitative changes in the rheological response of complex fluids (e.g., a gelation or vitrification transition) is an important capability for processing operations that utilize such materials in real-world environments. One class of complex fluids that exhibits distinct rheological states are soft glassy materials such as colloidal gels and clay dispersions, which can be well characterized by the soft glassy rheology (SGR) model. We first solve the model equations for the time-dependent, weakly nonlinear response of the SGR model. With this analytical solution, we show that the weak nonlinearities measured via medium amplitude parallel superposition (MAPS) rheology can be used to anticipate the rheological aging transitions in the linear response of soft glassy materials. This is a rheological version of a technique called structural health monitoring used widely in civil and aerospace engineering. We design and train artificial neural networks (ANNs) that are capable of quickly inferring the parameters of the SGR model from the results of sequential MAPS experiments. The combination of these data-rich experiments and machine learning tools to provide a surrogate for computationally expensive viscoelastic constitutive equations allows for rapid experimental characterization of the rheological state of soft glassy materials. We apply this technique to an aging dispersion of Laponite ® clay particles approaching the gel point and demonstrate that a trained ANN can provide real-time detection of transitions in the nonlinear response well in advance of incipient changes in the linear viscoelastic response of the system.

Lennon, Kyle R. (ORCID:0000000212515461)↗

Theory-based scaling laws of near and far scrape-off layer widths in single-null L-mode discharges

Abstract Theory-based scaling laws of the near and far scrape-off layer (SOL) widths are analytically derived for L-mode diverted tokamak discharges by using a two-fluid model. The near SOL pressure and density decay lengths are obtained by leveraging a balance among the power source, perpendicular turbulent transport across the separatrix, and parallel losses at the vessel wall, while the far SOL pressure and density decay lengths are derived by using a model of intermittent transport mediated by filaments. The analytical estimates of the pressure decay length in the near SOL is then compared to the results of three-dimensional, flux-driven, global, two-fluid turbulence simulations of L-mode diverted tokamak plasmas, and validated against experimental measurements taken from an experimental multi-machine database of divertor heat flux profiles, showing in both cases a very good agreement. Analogously, the theoretical scaling law for the pressure decay length in the far SOL is compared to simulation results and to experimental measurements in TCV L-mode discharges, pointing out the need of a large multi-machine database for the far SOL decay lengths.

Physics↗

Predictive turbulence-driven flux model of scrape-off layer widths across confinement regimes in tokamaks

Reliable scrape-off layer (SOL) profile decay lengths predictions are needed to design and operate future tokamaks. The present manuscript describes a new model based on turbulent transport that is able to predict SOL widths for both L-mode and H-mode plasmas. The model is based upon the sheared-spectral filament paradigm (Peret et al (WEST Team) 2022 Phys. Plasmas 29 072306), however, incorporating the effects of thermal transport in order to calculate the parallel heat fluxes. The effects of magnetic shear and ExB shear on the cross-field transport are crucial to explain the shorter SOL decay lengths found in H-mode. The model is validated against a database of thousands of DIII-D L-mode and H-mode SOL profiles. We also calculate SOL decay length predictions in terms of plasma and engineer control parameters, which are in agreement with the multi-machine empirical H-mode scaling (Eich et al (ASDEX Upgrade Team and JET EFDA Contributors) 2013 Nucl. Fusion 53 093031), however, with an additional device geometry dependence. ITER SOL width predictions by the model are 3 times higher than the empirical scaling.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Enhanced accuracy through ensembling of randomly initialized auto-regressive models for dynamical systems

Computational mechanics simulations using traditional finite element methods (FEM) require prohibitively expensive computational resources for real-time engineering applications, design optimization, and digital twin implementations. While machine learning (ML) surrogate models offer significant computational speedups, autoregressive ML models for time-dependent mechanical systems suffer from error accumulation that compromises long-term prediction reliability - a critical concern for engineering applications where accuracy over extended time horizons is essential for safety and performance assessments. Here, we propose a deep ensemble framework specifically designed to address this challenge in computational mechanics applications, where multiple ML surrogate models with random weight initializations are trained in parallel and their predictions aggregated during inference. This approach leverages statistical diversity to maximize information gain from a fixed set of training data and to mitigate error propagation, while maintaining the computational efficiency that makes ML surrogates attractive for engineering practice. We validate the framework on three representative problems spanning critical areas of computational mechanics: stress field evolution in heterogeneous microstructures under complex loading (relevant to advanced materials design and composite analysis), planetary-scale shallow water dynamics (applicable to environmental and geotechnical engineering), and Gray-Scott reaction-diffusion systems (relevant to mass transport and chemical process engineering). Across all test cases, the ensemble approach demonstrates consistent error reduction of 15-33% compared to individual models. The codes for this work are available on GitHub (https://github.com/Graham-Brady-Research-Group/AutoregressiveEnsemble_SpatioTemporal_Evolution).

autoregressive prediction↗

Development, construction and tests of the Mu2e electromagnetic calorimeter mechanical structures

The “muon-to-electron conversion” (Mu2e) experiment at Fermilab will search for the charged lepton flavour violating neutrino-less coherent conversion of a muon into an electron in the field of an aluminum nucleus. The observation of this process would be the unambiguous evidence of the existence of physics beyond the standard model. Mu2e detectors comprise a straw-tracker, an electromagnetic calorimeter and an external veto for cosmic rays. In particular, the calorimeter provides excellent electron identification, a fast calorimetric online trigger, and complementary information to aid pattern recognition and track reconstruction. The detector has been designed as a state-of-the-art crystal calorimeter and employs 1348 pure Cesium Iodide (CsI) crystals readout by UV-extended silicon photosensors and fast front-end and digitization electronics. A design consisting of two identical annular matrices (named “disks”) positioned at the relative distance of 70 cm downstream the aluminum target along the muon beamline satisfies the Mu2e physics requirements. The hostile Mu2e operational conditions, in terms of radiation levels (total expected ionizing dose of 12 krad and a neutron fluence of 5 × 10$^{10}$ n/cm$^{2}$ @ 1 MeV$_{eq}$ (Si)/y), magnetic field intensity (1 T) and vacuum level (10$^{-4}$ Torr) have posed tight constraints on scintillating materials, sensors, electronics and on the design of the detector mechanical structures and material choice. The support structure of each 674 crystal matrix is composed of an aluminum hollow ring and parts made of open-cell vacuum-compatible carbon fiber. The photosensors and front-end electronics for the readout of each crystal are inserted in a machined copper holder and make a unique mechanical unit. The resulting 674 mechanical units are supported by a machined plate of vacuum-compatible plastic material. The plate also integrates the cooling system made of a network of copper lines flowing a low temperature radiation-hard fluid and placed in thermal contact with the copper holders to constitute a low resistance thermal bridge. The data acquisition electronics are hosted in aluminum custom crates positioned on the external lateral surface of the disks. The crates also integrate the electronics cooling system as lines running in parallel to the front-end system. In this paper we report on the calorimeter mechanical structure design, the mechanical and thermal simulations that have determined the design technological choices, and the status of component production, quality assurance tests and plans for assembly at Fermilab.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗