Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel machines”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Progress and R/D challenges for FCC-ee SRF

The FCC-ee machines present a huge challenge for the RF systems, which need to be adapted to very diverse beam conditions going from moderate energy and high current for the Z machine to high energy and low beam current for the ttbar. This inverse scaling results naturally from a fixed budget for the synchrotron radiation, which the SRF cavities need to compensate. A global solution was elaborated for the FCC Conceptual Design Report (Abada in Eur Phys J Spec Top 228):261–623, 2019), and is referred here as the baseline. Recently, further studies have led to a new optimized baseline, still based on traditional elliptical cavities. In parallel, a novel concept, named the Slotted Waveguide ELLiptical (SWELL), was proposed with the potential of greatly simplified logistics and reduced costs. Under several aspects, all these changes call for enhanced performance of the RF systems. A vigorous R&D program has therefore continued since the publication of the CDR, with the aim of pushing the performance and demonstrating the feasibility of a more advanced baseline and, more recently, of the SWELL option. The progress and challenges of this ambitious program were presented in the dedicated SRF sessions at FCC week 2022 and are summarized in this paper.

47 OTHER INSTRUMENTATION↗

Serial2Parallel

In the era of machine learning, we often need to run the same code/script many times with little or no variations (e.g., performance evaluation, data preprocessing, data generation, hyperparameter tuning, etc.). It is not a problem when you just need to do that a few times, but when the number of repetitions becomes very large, it can be a daunting task. The code “Serial2Parallel” provides an easy way for users to be able to run many numbers of any serial code/scripts in a parallel manner across multiple nodes in an message passing interface (MPI) cluster. The code includes the server program that deals with task pool management and client program that processes task. The server gets the tasks ready and waits for clients' connections. The client code pulls tasks from the server and processes them. The client code will run in parallel.

Sangkeun, MattLee↗

Development, construction and qualification tests of the mechanical structures of the electromagnetic calorimeter of the Mu2e experiment at Fermilab

The “muon-to-electron conversion” (Mu2e) experiment at Fermilab will search for the Charged Lepton Flavour Violating neutrino-less coherent conversion of a muon into an electron in the field of an aluminum nucleus. The observation of this process would be the unambiguous evidence of physics beyond the Standard Model. Mu2e detectors comprise a straw-tracker, an electromagnetic calorimeter and an external veto for cosmic rays. The calorimeter provides excellent electron identification, complementary information to aid pattern recognition and track reconstruction, and a fast calorimetric online trigger. The detector has been designed as a state-of-the-art crystal calorimeter and employs 1340 pure Cesium Iodide (CsI) crystals readout by UV-extended silicon photosensors and fast front-end and digitization electronics. A design consisting of two identical annular matrices (named “disks”) positioned at the relative distance of 70 cm downstream the aluminum target along the muon beamline satisfies the Mu2e physics requirements.The hostile Mu2e operational conditions, in terms of radiation levels (total ionizing dose of 12 krad and a neutron fluence of 5x1010 n/cm2 @ 1 MeVeq (Si)/y), magnetic field intensity (1 T) and vacuum level (10$^{-4}$ Torr) have posed tight constraints on the design of the detector mechanical structures and materials choice. The support structure of the two 670 crystal matrices employs two aluminum hollow rings and parts made of open-cell vacuum-compatible carbon fiber. The photosensors and service front-end electronics for each crystal are assembled in a unique mechanical unit inserted in a machined copper holder. The 670 units are supported by a machined plate made of vacuum-compatible plastic material. The plate also integrates the cooling system made of a network of copper lines flowing a low temperature radiation-hard fluid and placed in thermal contact with the copper holders to constitute a low resistance thermal bridge. The data acquisition electronics is hosted in aluminum custom crates positioned on the external lateral surface of the two disks. The crates also integrate the electronics cooling system as lines running in parallel to the front-end system.The constraints on the calorimeter mechanical structures design, the development from the conceptual design to the specifications of all the structural components, the status of components production and the components quality assurance tests are presented.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

DETECTING FIRE WITH MACHINE LEARNING-ENABLED VISUAL MONITORING FOR NUCLEAR POWER PLANT ENVIRONMENTS

Nuclear power plants are experiencing significant cost challenges to remain competitive with other energy-generation utilities. Unlike other industries, the cost of operation and maintenance activities is mostly attributed to workforce costs. To mitigate this, nuclear power plant stakeholders are increasingly interested in the development and deployment of machine learning methods to potentially automate or augment manually intensive tasks to reduce costs, especially for monitoring activities. One monitoring function that is visually demanding and that can occur frequently to meet the requirements of a fire protection program is visually monitoring an area for fire occurrence. Currently, fire watch activities consist of a worker physically stationed at a given location with the sole responsibility of observing a given area to ensure a fire is detected and mitigated promptly. This effort focused on the development and evaluation of a suitable deep convolutional neural network to classify individual video frames at a sub-second frequency for the occurrence of “fire” and “no fire” in varying industrial environments similar to nuclear power plants. It is believed that a trained neural network model could be integrated with existing facility video surveillance camera feeds to generate alerts when fire inferences occur in individual frames captured at sub-second temporal resolutions. Extensive effort was dedicated to identifying and curating suitable imagery training data representing varying environments and scene settings with and without flame features to maximize generalization in nuclear power plant environments. The data collection effort resulted in the aggregation of a large, labeled image library exceeding 12,000 images to support model training for diverse industrial environments. A deep neural network model incorporating parallel multi-scale capabilities was developed and trained to support accurate image-based detection of flame incidents of varying sizes and spectral feature properties within heterogeneous scenes. Analysis results show that the trained model can achieve high inference accuracy despite heterogeneous scene environments and components. Testing accuracy exceeded 95.0 percent with very low false positive and false negative inferences.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Optimizing inference of segmentation on high-resolution images in MLExchange

MLExchange is a machine learning (ML) operations platform providing web user-interfaces (UIs) for data visualization and analysis pipelines at synchrotron facilities. Among these UIs is the segmentation app which helps synchrotron users utilize ML algorithms to automatically segment high-resolution scientific images with minimal manual annotation effort. In this work, we share code optimizations that significantly speed up the segmentation inference workflow of large data in short time. By optimizing the sequence of CPU-GPU data transfers and introducing CPU parallelization to key operations, we improve the per-device, per-image frame computational efficiency and observe close to 3×$$\times$$ speedup over the original segmentation inference workflow run time when utilizing a single GPU. Further adaptations enabling multi-GPU inference yield more than 40×$$\times$$ speedup with 100 GPUs compared to the optimized single GPU inference workflow. This acceleration of the segmentation inference workflow will provide MLExchange users with easy access to segmentation results with little wait time.

Lu, Shizhao↗

Control And Optimization Modular Modeling Application For Nuclear Deployment

The purpose of the COMMAND code is to provide a flexible, scalable tool for use in developing, integrating, and testing the technologies necessary for achieving autonomous operations of advanced nuclear reactors. The code enables users to efficiently implement custom simulations and experiments by combining key methods from different software modules. These modules are focused on: modeling and simulation tools, such as nuclear simulation tools used for high-fidelity modeling (e.g., Reactor Excursion and Leak Analysis Program [RELAP5-3D] and Monte Carlo N-Particle [MCNP]); machine learning and optimization tools (e.g., anomaly detection and data-driven modeling techniques); advanced control in its digital, high-performance, and supervisory control forms (e.g., proportional integral derivative (PID) control and model predictive control (MPC); and integration with hardware through industrial communication protocols. To ensure flexibility and scalability, COMMAND was designed to be both modular—the software “pieces” all inherit from generic building blocks and can be combined and connected to create complicated simulations—and high performing—designed for parallel processing, enabling simulations and experiments to take advantage of multi-core computers, servers, and nodes. The code is written in the Python programming language due to the language's popularity, active community, and open-source and cross-platform nature. Maintaining consistency with other simulation tools used within the nuclear energy community, users implement simulations and experiments through text input files, which define components, parameters, connections, etc., through lines of text. Given that COMMAND is written in Python, these input files are native Python scripts, and so use the standard Python structure and formatting. This also enables users to take advantage of Python's extensive package library to develop custom capabilities for their specific use cases.

Faber, Jacob [Idaho National Laboratory (INL), Ida↗

Extended Physics-Informed Neural Networks (XPINNs): A Generalized Space-Time Domain Decomposition Based Deep Learning Framework for Nonlinear Partial Differential Equations

Here we propose a generalized space-time domain decomposition approach for the physics-informed neural networks (PINNs) to solve nonlinear partial differential equations (PDEs) on arbitrary complex-geometry domains. The proposed framework, named eXtended PINNs ( X P I N N s ), further pushes the boundaries of both PINNs as well as conservative PINNs (cPINNs), which is a recently proposed domain decomposition approach in the PINN framework tailored to conservation laws. Compared to PINN, the XPINN method has large representation and parallelization capacity due to the inherent property of deployment of multiple neural networks in the smaller subdomains. Unlike cPINN, XPINN can be extended to any type of PDEs. Moreover, the domain can be decomposed in any arbitrary way (in space and time), which is not possible in cPINN. Thus, XPINN offers both space and time parallelization, thereby reducing the training cost more effectively. In each subdomain, a separate neural network is employed with optimally selected hyperparameters, e.g., depth/width of the network, number and location of residual points, activation function, optimization method, etc. A deep network can be employed in a subdomain with complex solution, whereas a shallow neural network can be used in a subdomain with relatively simple and smooth solutions. We demonstrate the versatility of XPINN by solving both forward and inverse PDE problems, ranging from one-dimensional to three-dimensional problems, from time-dependent to time-independent problems, and from continuous to discontinuous problems, which clearly shows that the XPINN method is promising in many practical problems. The proposed XPINN method is the generalization of PINN and cPINN methods, both in terms of applicability as well as domain decomposition approach, which efficiently lends itself to parallelized computation. The XPINN code is available on h t t p s : / / g i t h u b . c o m / A m e y a J a g t a p / X P I N N s .

97 MATHEMATICS AND COMPUTING↗

Device for controlling additive manufacturing machinery

A computing device for controlling the operation of an additive manufacturing machine comprises a memory element and a processing element. The memory element is configured to store a three-dimensional model of a part to be manufactured, wherein the three-dimensional model defines a plurality of cross sections of the part. The processing element is in communication with the memory element. The processing element is configured to receive the three-dimensional model, determine a path across a surface of each cross section, wherein the path includes a plurality of parallel lines, calculate a power for a radiation beam to scan each of the lines, such that the power varies from line to line non-linearly according to a length of the line, and calculate a scan speed for the radiation beam for each of the lines, such that the scan speed varies line to line non-linearly according to the power of the radiation beam.

Barr, Christian↗

Deep Generative Models that Solve PDEs: Distributed Computing for Training Large Data-Free Models

Recent progress in scientific machine learning (SciML) has opened up the possibility of training novel neural network architectures that solve complex partial differential equations (PDEs). Several (nearly data free) approaches have been recently reported that successfully solve PDEs, with examples including deep feed forward networks, generative networks, and deep encoder-decoder networks. However, practical adoption of these approaches is limited by the difficulty in training these models, especially to make predictions at large output resolutions (≥1024×1024). Here we report on a software framework for data parallel distributed deep learning that resolves the twin challenges of training these large SciML models - training in reasonable time as well as distributing the storage requirements. Our framework provides several out of the box functionality including (a) loss integrity independent of number of processes, (b) synchronized batch normalization, and (c) distributed higher-order optimization methods. We show excellent scalability of this framework on both cloud as well as HPC clusters, and report on the interplay between bandwidth, network topology and bare metal vs cloud. We deploy this approach to train generative models of sizes hitherto not possible, showing that neural PDE solvers can be viably trained for practical applications. We also demonstrate that distributed higher-order optimization methods are 2-3× faster than stochastic gradient-based methods and provide minimal convergence drift with higher batch-size.

PDEs↗

Overview on Current Activities of Conduction-Cooled SRF Accelerators and their Applications

When Nb3Sn was reintroduced to the SRF community as an alternative to pure niobium, one key motivation has been to reduce the cryogenic requirements of new and existing accelerators by shifting from 2 K to 4 K operation. Meanwhile, a variety of implementations beyond research machines are being explored. The combination of Nb3Sn with conventional cryocoolers, enabling cryogen-free operation, has paved the way for the development of compact, standalone systems suitable for applications far beyond research, such as enhancing the durability of synthetics via crosslinking or sterilizing food and medical equipment, as well as environmental cleanup when it comes to decontaminating liquid and solid waste material. So, while fundamental R&D continues to refine Nb3Sn resonators, exploring improvements such as replacing the niobium substrate with copper, parallel research efforts are investigating how the increased beam power provided by SRF could expand the commercial use of electron beams. This presentation aims to deliver a comprehensive overview of ongoing research efforts to harness the benefits of SRF through Nb3Sn and conduction cooling.

Vennekate, John [Thomas Jefferson National Acceler↗

Considerations regarding the Use of Computer Vision Machine Learning in Safety-Related or Risk-Significant Applications in Nuclear Power Plants

With the advancements made to date in the field of artificial intelligence (AI), significant potential exists to utilize AI capabilities for nuclear power plant (NPP) applications. AI can replicate human decision making and it is usually faster and more accurate than humans. For implementations that impact critical NPP applications (e.g., safety-related or non-safety systems that potentially affect overall plant risk), a deeper safety analysis of the AI methods is necessary. AI applied to NPP operations could resemble the use of digital I&C (DI&C) because such applications involve digital computer hardware and custom-designed software that input plant data, execute complex software algorithms, and output the results to a system or licensed human operator to potentially provoke an action. For AI methods to be compliant with current safety requirements for DI&C, AI compatibility must be evaluated, and AI-related gaps may exist that prevent the prompt deployment of AI in NPPs. This effort aims to evaluate how example AI technologies align with the DI&C safety framework, and discusses how they could be analyzed, modeled, tested, and validated in a manner similar to typical DI&C technologies. Because AI is a broad field that encompasses areas such as machine learning (ML), natural language processing, and computer vision, this research focused on a subset of methods categorized as the computer vision ML (CVML) methods. This report explores two CVML use cases, gauge reading and fire watch, considered relevant to the DI&C standards, as they could play a safety-critical role. For the gauge reading use case, a CVML-enabled technology that can read gauges at oblique angles is utilized. For the fire watch use case, a CVML-enabled technology is utilized that migrates fire watch from a manual (human) approach to automated fire detection. These use cases are mainly intended to give context to the CVML system discussion. This effort assumes the worst-case scenario, with the CVML system being used to replace a safety-related or risk-significant system, thus requiring evaluation. Evaluating CVML against most of the relevant safety requirements for DI&C yielded several CVML-specific considerations due to the uniqueness of its characteristics in comparison with typical DI&C systems. For example, CVML models often employ commonly used (open-source) datasets, and it is not always possible to determine the level of overlap among open-source datasets. Therefore, the independence of the developed CVML models when demonstrating diversity is questionable, therefore creating vulnerability to common cause failure (CCF). The design verification process is also impacted since the data overlap could result in overestimation of the software validation and verification (V&V) performance results. Section 2 of this report evaluates a list of the identified CVML-specific characteristics and discusses the resulting considerations and potential solutions in the context of each referenced requirement. A summation is provided in Section 3. This report is not to be used as a guideline. It was developed to identify and consider issues in the implementation of ML technologies used to augment activities that may have a bearing on plant operation. The report draws parallels to the use of DI&C technologies, for which many standards are available to guide their use in nuclear plant operation. It considers the technologies and some of the potential implications of their use in safety-related applications but is not intended to address regulatory or licensing related issues.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Machine Learning-Driven Conservative-to-Primitive Conversion in Hybrid Piecewise Polytropic and Tabulated Equations of State

We present a novel machine learning (ML)-based method to accelerate conservative-to-primitive inversion, focusing on hybrid piecewise polytropic and tabulated equations of state. Traditional root-finding techniques are computationally expensive, particularly for large-scale relativistic hydrodynamics simulations. To address this, we employ feedforward neural networks (NNC2PS and NNC2PL), trained in PyTorch (2.0+) and optimized for GPU inference using NVIDIA TensorRT (8.4.1), achieving significant speedups with minimal accuracy loss. The NNC2PS model achieves 𝐿 1 and 𝐿 ∞ errors of 4.54 × 10 −7 and 3.44 × 10−6, respectively, while the NNC2PL model exhibits even lower error values. TensorRT optimization with mixed-precision deployment substantially accelerates performance compared to traditional root-finding methods. Specifically, the mixed-precision TensorRT engine for NNC2PS achieves inference speeds approximately 400 times faster than a traditional single-threaded CPU implementation for a dataset size of 1,000,000 points. Ideal parallelization across an entire compute node in the Delta supercomputer (dual AMD 64-core 2.45 GHz Milan processors and 8 NVIDIA A100 GPUs with 40 GB HBM2 RAM and NVLink) predicts a 25-fold speedup for TensorRT over an optimally parallelized numerical method when processing 8 million data points. Moreover, the ML method exhibits sub-linear scaling with increasing dataset sizes. We release the scientific software developed, enabling further validation and extension of our findings. By exploiting the underlying symmetries within the equation of state, these findings highlight the potential of ML, combined with GPU optimization and model quantization, to accelerate conservative-to-primitive inversion in relativistic hydrodynamics simulations.

conservative-to-primitive conversion↗

Magnetic electron collimation in three-dimensional semi-metals

While electrons moving perpendicular to a magnetic field are confined to cyclotron orbits, they can move freely parallel to the field. This simple fact leads to complex current flow in clean, low carrier density semi-metals, such as long-ranged current jets forming along the magnetic field when currents pass through point-like constrictions. Occurring accidentally at imperfect current injection contacts, the phenomenon of "current jetting" plagues the research of longitudinal magneto-resistance, which is particularly important in topological conductors. Here we demonstrate the controlled generation of tightly focused electron beams in a new class of micro-devices machined from crystals of the Dirac semi-metal Cd 3 As 2 . The current beams can be guided by tilting a magnetic field and their range tuned by the field strength. Finite element simulations quantitatively capture the voltage induced at faraway contacts when the beams are steered towards them, supporting the picture of controlled electron jets. These experiments demonstrate direct control over the highly non-local signal propagation unique to 3D semi-metals in the current jetting regime, and may lead to applications akin to electron optics in free space.

36 MATERIALS SCIENCE↗

Inference-Engine v0.1.0

Given a pre-trained neural network, Inference-Engine performs maps network inputs to outputs by executing the forward pass through the provided network. Although the predominant programming language for machine-learning is Python, most high-performance computing (HPC) applications are written in Fortran, C, or C++. Inference-Engine aims to support HPC programs and is written in Fortran, a language with a large feature set supporting interoperability with C. This software exposes concurrency in a portable way by using standard language features that some modern Fortran compilers can exploit with various optimizations, including offloading computation to a Graphics Processing Unit (GPU). In particular, this software makes extensive use of Fortran's "do concurrent" parallel loop construct, implicitly parallel array statements, and pure procedures that can be invoked inside "do concurrent" blocks. Inference-Engine also supports dynamic choice of inference methods at runtime. Two current options include one method that uses Fortran's "dot_product" intrinsic function inside "do concurrent" blocks and another method that instead uses Fortran' "matmul" array intrinsic function. We plan to investigate automatic compiler offloading of "do concurrent" calculations to GPUs and compile-time substitution of optimized libraries such as the Basic Linear Algebra Library (BLAS) for "matmul" invocations. We also envision the potential for the choice of which method to use could happen at program launch based on in situ performance measurements on any given platform.

Rouson, Damian↗

Throughput Measurements and Profile Analysis of Cloud Networks

Cloud networks utilize virtual connections to connect virtual machines distributed across cloud sites. They are increasingly deployed due to flexible provisioning using software and cost-effectiveness in not requiring to build physical network infrastructure. However, their extensive virtualization makes it unclear how well the established practices of conventional networks translate to them. Here, we study throughput measurements over a Google Cloud network using a matching hardware emulated conventional network, which provide production and exploratory conditions, respectively. The measurements span connections representing local, cross-continental and around the Earth distances. We study the effects of parallel flows, congestion control algorithms and retransmissions on the network throughput profile expressed as a function of RTT. We compare the throughput profile of Google Cloud network with those of emulated network under various loss conditions, including those too disruptive or expensive in the former. Our analysis based on the concave-convex shape and utilization-concavity coefficients of throughput profiles indicates an overall agreement of performance between the two networks, thereby justifying the use of conventional network emulations to analyze cloud networks. In terms of practical use, our study establishes that BBR and BBRv2 alpha TCP achieve higher throughput compared to loss-based congestion control algorithms under most network configurations, especially, under losses at large RTT.

Phanekham, Derek [Southern Methodist Univ., Dallas↗

GASNet-EX RMA Communication Performance on Recent Supercomputing Systems

Partitioned Global Address Space (PGAS) programming models, typified by systems such as Unified Parallel C (UPC) and Fortran coarrays, expose one-sided Remote Memory Access (RMA) communication as a key building block for High Performance Computing (HPC) applications. Architectural trends in supercomputing make such programming models increasingly attractive, and newer, more sophisticated models such as UPC++, Legion and Chapel that rely upon similar communication paradigms are gaining popularity. GASNet-EX is a portable, open-source, high-performance communication library designed to efficiently support the networking requirements of PGAS runtime systems and other alternative models in emerging exascale machines. The library is an evolution of the popular GASNet communication system, building upon 20 years of lessons learned. We present microbenchmark results which demonstrate the RMA performance of GASNet-EX is competitive with MPI implementations on four recent, high-impact, production HPC systems. These results are an update relative to previously published results on older systems. The networks measured here are representative of hardware currently used in six of the top ten fastest supercomputers in the world, and all of the exascale systems on the U.S. DOE road map.

Hargrove, Paul H↗

Transforming microseismic clouds into near real-time visualization of the growing hydraulic fracture

SUMMARY Microseismic observations during unconventional reservoir stimulation are typically seen as a proxy for clusters of hydraulic fractures and the extent of the stimulated reservoir. Such straightforward interpretation is often misleading and fails to provide a physically reasonable image of the fracturing process. This paper demonstrates the application of a physics-based machine learning algorithm which enables a rapid and accurate fracture mapping from the microseismic data. Our training and validation data set relies on a history-matched geomechanical modelling workflow implemented in GEOS software for the Hydraulic Fracturing Test Site 1 (HFTS-1) project. For this study we augmented the simulated fracture growth through geostatistical modelling of induced seismicity, so that the synthetic microseismic catalogue matches the main statistical properties of the field observations. We formulated the problem of mapping the actual fracture in the clutter of events to parallel common video segmentation workflows: several past video frames (microseismic density snapshots) are passed through a deep convolutional network to classify whether a given voxel is associated with a fracture or intact rock. We found that for accurate fracture mapping, the network’s input and architecture must be augmented to incorporate the fluid injection parameters (pressure, rate, concentration of proppant, and location of the perforation within the cluster). The error rate for the network reached as little as 10 per cent of the fracture area, while a conventional microseismic interpretation approach yielded ∼300 per cent. Our approach also yields must faster predictions than conventional methods (minutes instead of weeks), and could enable engineers to make rapid decisions regarding engineering parameters (pumping rate, viscosity) in real time during stimulation.

58 GEOSCIENCES↗

Pedestal origin and extrapolation of high-density small edge-localised-modes peak parallel energy fluence in ITER and SPARC

Experimental analysis and simulations with the BOUT++ code show that small edge-localised modes (ELMs) in reactor-relevant high-density regimes originate in a region close to the separatrix and only marginally perturb the pedestal structure. The measured divertor peak parallel energy fluence (ε ∥,peak ) for a database of small ELM scenarios in DIII-D and ASDEX Upgrade can be reproduced, within 40 % accuracy on average, if an ad hoc modification of the Eich peak parallel ELM energy fluence model is applied to account for the small ELM pedestal birth location. This allows for first-order extrapolation of small-ELM divertor ε ∥,peak to ITER and SPARC, resulting in values that satisfy the nominal melting threshold of tungsten monoblocks of 12 MJ m −2 . The findings reported in this study, both via modelling and direct measurements, constitute a step forward in assessing small ELMs in high edge-collisionality scenarios as a viable plasma regime for the operation of next-generation fusion machines.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗