Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “concurrent computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Multiscale Concurrent Atomistic-Continuum (CAC) modeling of multicomponent alloys

We report strengthening in complex multicomponent systems such as solid solution alloys is controlled primarily by the dynamic interactions between dislocation lines and heterogeneously distributed solute species. Modeling of extended defect length scales in such multicomponent systems becomes prohibitively expensive, motivating the development of reduced order approaches. This work explores the application of the Concurrent Atomistic-Continuum (CAC) method to model dislocation mobility in random alloys at extended length scales. By employing recently developed average-atom interatomic potentials, the average “bulk” material response in coarse-grained regions interacts with true random solute species in the atomistic-scale domain. We demonstrate that spurious stresses in domain resolution transition regions are eliminated entirely due to the CAC formulation. Simultaneously, the key details of local stress fluctuation due to randomness in the dislocation core region are captured, and fluctuating stress smoothly decays to the long-range dislocation stress field response. Dislocation mobility calculations, for line lengths over 400 nm, are computed as a function of alloy composition in the model FeNiCr system and compared to full molecular dynamics (MD). The results capture the composition-dependent trends, while reducing degrees of freedom by nearly 40%. This approach can be readily extended to any system described by an EAM potential and facilitates the study of large-scale defect dynamics in complex solute environments to support computational alloy design.

36 MATERIALS SCIENCE↗

Layer-Parallel Training of Deep Residual Neural Networks

Residual neural networks (ResNets) are a promising class of deep neural networks that have shown excellent performance for a number of learning tasks, e.g., image classification and recognition. Mathematically, ResNet architectures can be interpreted as forward Euler discretizations of a nonlinear initial value problem whose time-dependent control variables represent the weights of the neural network. Hence, training a ResNet can be cast as an optimal control problem of the associated dynamical system. For similar time-dependent optimal control problems arising in engineering applications, parallel-in-time methods have shown notable improvements in scalability. This paper demonstrates the use of those techniques for efficient and effective training of ResNets. The proposed algorithms replace the classical (sequential) forward and backward propagation through the network layers with a parallel nonlinear multigrid iteration applied to the layer domain. This adds a new dimension of parallelism across layers that is attractive when training very deep networks. From this basic idea, we derive multiple layer-parallel methods. The most efficient version employs a simultaneous optimization approach where updates to the network parameters are based on inexact gradient information in order to speed up the training process. Finally, using numerical examples from supervised classification, we demonstrate that the new approach achieves a training performance similar to that of traditional methods, but enables layer-parallelism and thus provides speedup over layer-serial methods through greater concurrency.

97 MATHEMATICS AND COMPUTING↗

Advancing Multi-Hazard Risk and Safety Considerations for Aging Nuclear Facilities

While probabilistic risk assessment (PRA) of nuclear facilities is expected to include internal and external hazards for a risk-informed and performance-based design, the current state of practice treats each hazard independently. However, such an independent treatment of hazards may not account for the correlations between different hazards and their response of and damage to the structures, systems, and components (SSCs) in a plant resulting in underestimating the overall risk. This project proposes to advance the multi-hazard PRA of nuclear facilities to more adequately evaluate concurrent hazards and contribute to an increased safety of nuclear plants. A framework for multi-hazard PRA will be developed by identifying concurrent hazard events (both internal and external) and event sequences that include interdependencies through the response of SSCs. An example application of the multi-hazard PRA framework will be demonstrated by considering a generic pressurized water reactor (PWR) subjected to seismic and internal flooding hazards. Computational models for the response of components will be developed to generated multi-hazard fragility surfaces under seismic and flooding loads. A PRA model consisting of event and fault trees will also be developed to quantify the multi-hazard risk profile and compare it with the independent hazard risk profile. Overall, by advancing the multi-hazard PRA of nuclear facilities, this project enhances nuclear safety and reduces costs by mitigating unforeseen consequences caused by correlations between concurrent hazards.

97 - MATHEMATICS AND COMPUTING↗

VTK-m User's Guide (V.1.6)

High-performance computing relies on ever finer threading. Advances in processor technology include ever greater numbers of cores, hyperthreading, accelerators with integrated blocks of cores, and special vectorized instructions, all of which require more software parallelism to achieve peak performance. Traditional visualization solutions cannot support this extreme level of concurrency. Extreme scale systems require a new programming model and a fundamental change in how we design algorithms. To address these issues we created VTK-m: the visualization toolkit for multi-/many-core architectures.

97 MATHEMATICS AND COMPUTING↗

$\mathrm{RADICAL}$-Pilot and $\mathrm{PMIx}$/$\mathrm{PRRTE}$: Executing Heterogeneous Workloads at Large Scale on Partitioned $\mathrm{HPC}$ Resources

Execution of heterogeneous workflows on high-performance computing (HPC) platforms present unprecedented resource management and execution coordination challenges for runtime systems. Task heterogeneity increases the complexity of resource and execution management, limiting the scalability and efficiency of workflow execution. Re-source partitioning and distribution of tasks execution over portioned re-sources promises to address those problems but we lack an experimental evaluation of its performance at scale. Here this paper provides a performance evaluation of the Process Management Interface for Exascale (PMIx) and its reference implementation PRRTE on the leadership-class HPC plat-form Summit, when integrated into a pilot-based runtime system called RADICAL-Pilot. We partition resources across multiple PRRTE Distributed Virtual Machine (DVM) environments, responsible for launching tasks via the PMIx interface. We experimentally measure the work-load execution performance in terms of task scheduling/launching rate and distribution of DVM task placement times, DVM startup and termination overheads on the Summit leadership-class HPC platform. Integrated solution with PMIx/PRRTE enables using an abstracted, standardized set of interfaces for orchestrating the launch process, dynamic process management and monitoring capabilities. It extends scaling capabilities allowing to overcome a limitation of other launching mechanisms (e.g., JSM/LSF). Explored different DVM setup configurations provide insights on DVM performance and a layout to leverage it. Our experimental results show that heterogeneous workload of 65,500 tasks on 2048 nodes, and partitioned across 32 DVMs, runs steady with resource utilization not lower than 52%. While having less concurrently executed tasks resource utilization is able to reach up to 85%, based on results of heterogeneous workload of 8200 tasks on 256 nodes and 2 DVMs.

97 MATHEMATICS AND COMPUTING↗

Optimization and supervised machine learning methods for fitting numerical physics models without derivatives

Here, we address the calibration of a computationally expensive nuclear physics model for which derivative information with respect to the fit parameters is not readily available. Of particular interest is the performance of optimization-based training algorithms when dozens, rather than millions or more, of training data are available and when the expense of the model places limitations on the number of concurrent model evaluations that can be performed. As a case study, we consider the Fayans energy density functional model, which has characteristics similar to many model fitting and calibration problems in nuclear physics. We analyze hyperparameter tuning considerations and variability associated with stochastic optimization algorithms and illustrate considerations for tuning in different computational settings.

97 MATHEMATICS AND COMPUTING↗

Comparative modeling of the disregistry and Peierls stress for dissociated edge and screw dislocations in Al

Many elementary deformation processes in metals involve the motion of dislocations. The planes of glide and specific processes dislocations prefer depend heavily on their atomic core structures. Atomistic simulations are desirable for dislocation modeling but their application to even sub-micron scale problems is in general computationally costly. Accordingly, continuum-based approaches, such as the phase-field microelasticity, phase-field dislocation dynamics (PFDD), generalized Peierls–Nabarro (GPN) models, and the concurrent atomistic–continuum (CAC) method, have attracted increasing attention in the field of dislocation modeling because they well represent both short-range cores interactions and long-range stress fields of dislocations. To better understand their similarities and differences, it is useful to compare these methods in the context of benchmark simulations and predictions. In this paper, we apply the CAC method and different PFDD variants – one of them is equivalent to a GPN model – to simulate an extended (i.e., dissociated) dislocation in Al with initially pure edge or pure screw character in terms of the disregistry. CAC and discrete forms of PFDD are also employed to calculate the Peierls stress. By conducting comprehensive convergence studies, we quantify the dependence of these measures on time/grid resolution and simulation cell size. Several important but often overlooked differences between PFDD/GPN variants are clarified. In conclusion, our work sheds light on the advantages and limitations of each method, as well as the path towards enabling them to effectively model complex dislocation processes at larger length scales.

36 MATERIALS SCIENCE↗

Visible-Light Photoinitiation of (Meth)acrylate Polymerization with Autonomous Post-Conversion

Conversion plateaus rapidly in radical photopolymerizations (RPPs) following discontinuation of irradiation due to rapid termination of reactive radicals, which restricts the wider use of RPPs in applications that involve nonuniform light access including those with attenuated light transmission or irregular surfaces. Based on our recent report of a radical dark-curing photoinitiator (DCPI) that continues polymerization beyond the cessation of irradiation by enabling latent redox initiation with photo-released amine in the presence of a suitable oxidant, we developed a new DCPI with an absorption spectrum that extends well into the visible range. Our design process involved a series of computational investigations of candidate molecules, including a systematic study of substituents and their position-dependent effects on absorption characteristics, electronic transitions, and the photochemical mechanism and its associated energetics. Our quantum chemical computations identified the target compound 5,7-dimethoxy-6-bromo-3-aroylcoumarin-DMPT/BPh4 and predicted that it would facilitate the dark-curing mechanism by concurrent photo-radical generation and photo-induced release of an efficient redox reductant under visible irradiation. This reductant-tethered chromophore was then synthesized and optically characterized with UV–vis spectroscopy that revealed its strong visible-light absorption with a molar absorptivity of 5710 M–1 cm–1 at 405 nm and 50 M–1 cm–1 at 455 nm. We then demonstrated extensive dark-curing of >35% additional conversion over 25 min following brief activation of the shelf-stable one-part system by irradiation with a 455 nm LED that was ceased at 20% conversion. In contrast, shuttering irradiation of the control formulation at that same point resulted in immediate cessation of conversion, which plateaued at 20%. We determined a remarkable initiator efficiency of 2.82 that results from the additional redox-generated radicals with a 77% photo-reductant generation quantum yield. The combination of superior photo- and dark-curing efficiencies of this new visible DCPI is expected to open new application opportunities in RPP, especially those involving resins that are highly light attenuating, surfaces that possess irregular features that produce uneven irradiance, and production lines where continued dark-curing downstream of the light activation step enhances line efficiencies.

absorption spectrum↗

Three-dimensional magnetohydrodynamic modeling of auto-magnetizing liner implosions on the Z accelerator

Auto-magnetizing (AutoMag) liners are cylindrical tubes that employ helical current flow to produce strong internal axial magnetic fields prior to radial implosion on ~100 ns timescales. AutoMag liners have demonstrated strong uncompressed axial magnetic field production (>100 T) and remarkable implosion uniformity during experiments on the 20 MA Z accelerator. However, both axial field production and implosion morphology require further optimization to support the use of AutoMag targets in magnetized liner inertial fusion (MagLIF) experiments. Data from experiments studying the initiation and evolution of dielectric flashover in AutoMag targets on the Mykonos accelerator have enabled the advancement of magnetohydrodynamic (MHD) modeling protocols used to simulate AutoMag liner implosions. Implementing these protocols using ALEGRA has improved the comparison of simulations to radiographic data. Specifically, both the liner in-flight aspect ratio and the observed width of the encapsulant-filled helical gaps during implosion in ALEGRA simulations agree more closely with radiography data compared to previous GORGON simulations. Although simulations fail to precisely reproduce the measured internal axial magnetic field production, improved agreement with radiography data inspired the evaluation of potential design improvements with newly developed modeling protocols. Three-dimensional MHD simulation studies focused on improving AutoMag target designs, specifically seeking to optimize the axial magnetic field production and enhance the cylindrical implosion uniformity for MagLIF. Importantly, by eliminating the driver current prepulse and reducing the initial inter-helix gap widths in AutoMag liners, simulations indicate that the optimal 30–50 T range of precompressed axial magnetic field for MagLIF on Z can be accomplished concurrently with improved cylindrical implosion uniformity.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Using Apptainer in a Pilot-based Distributed Workload

GlideinWMS is a pilot and pressure-based workload manager for distributed scientific computing. Many experiments like CMS and Fermilab’s Neutrino experiments use it to provision elastic clusters for their analysis and simulations, split into close to a million concurrent jobs. Most user jobs require containers, and the pilots use Apptainer to set up the desired platform. For the pilots that run as regular batch jobs, Apptainer is safer, lighter, and easier to use than other containerization solutions. Many images used by the pilots are expanded SIF images distributed via the CernVM-FS: this combination is very efficient. At Fermilab, for example, we store on GitHub Dockerfiles that mimic the platform in the worker nodes of local clusters. GitHub workflows build and push the images to Docker Hub, and a service periodically pulls and converts them to the expanded SIF images in the CernVM-FS, so the scientists can find a familiar environment everywhere. Apptainer has also been used to run services inside the pilot jobs, like benchmarks that characterize the worker node being used, or a Triton Inference Server that allows sharing a GPU with all the jobs that run in parallel on a node.

Mambelli, Marco [Fermilab] (ORCID:0000000294892681↗

Enabling Capabilities and Resources: 2024 Principal Investigator Meeting Proceedings

As a major supporter of basic genome-enabled research, BER’s Biological Systems Science Division (BSSD) fosters scientific discovery by funding - fundamental biological research across disciplines in conjunction with enabling investigational tools and computational capabilities that include world-class user facilities. The overarching goal of BSSD is to provide the necessary fundamental science to understand, predict, manipulate, and design biological systems that underpin innovations for bioenergy and bioproduct production and enhance understanding of natural, DOE-relevant environmental processes (Biological Systems Science Division Strategic Plan, 2021). To accelerate the U.S. bioeconomy, BSSD pursues innovative science underpinning advances in sustainable biofuels and bioproducts and the development of next-generation technologies and computational resources for systems biology research. The 2024 BSSD Enabling Capabilities and Resources (ECR) Principal Investigator (PI) meeting brought together PIs across the BSSD ECR portfolio to confer on shared interests and opportunities. The meeting was held concurrently with the Genomic Science program (GSP) PI meeting to optimize collaboration on research to advance bioenergy and the bioeconomy. Rick Stevens of Argonne National Laboratory gave a keynote on How Generative Artificial Intelligence Can Impact Biological Research (see Keynote: How Generative Artificial Intelligence Can Impact Biological Research, this page). Plenary presentations included several joint sessions that illuminated the integration and understanding of the larger BSSD mission. GSP’s objective is to provide systems-level understanding of plants, microbes, and their communities through its Bioenergy Research, Biosystems Design, and Environmental Microbiome Research portfolios. The objective of the ECR portfolio is to support development of computational and instrumental platforms to advance fundamental GSP research—and BER more broadly— toward the overall goal of understanding the functional principles of living systems and their response to environmental challenges.

59 BASIC BIOLOGICAL SCIENCES↗

S4PST: Sustainability for Programming Systems and Tools: May Workshop Report

The US Department of Energy (DOE) Exascale Computing Project (ECP) has fostered and strengthened the use of modern software engineering practices for developing applications and libraries, and this effort has resulted in the coordinated and interoperable E4S1 and xSDK2 ecosystems. Although this approach is cost-effective, it relies on robust programming systems and tools (PST) as the underlying foundation for our HPC software. At present, our primary PST stack consists of traditional high-performance computing (HPC) languages, namely Fortran, C, C++, and the popular Python language for data analysis and AI workflows. These languages support various programming frameworks and run-time abstractions that enable parallelism and concurrency across multiple node architectures and thousands of nodes through a variety of interconnect systems. However, to accommodate users’ diverse needs, certain aspects of the HPC ecosystem are delegated to vendor-specific or third-party implementations that extend beyond a particular scientific domain. This broader scope results in a multitude of specifications and variations, which leads to a complex orchestration of many-ecosystems. Unfortunately, this complexity in the ecosystem imposes additional overhead costs on consumers during the latter stages of the development cycle. In addition to the software ecosystem challenge, the upcoming conclusion of the ECP by December 2023 has raised significant concerns within the HPC programming systems community, from both the economic and social perspectives. The ECP has implemented a management structure for software development and funding decisions across all ECP participants by following a conventional hierarchical and centralized approach. However, this structure has prompted certain considerations within the community, particularly in anticipation of the Software Sustainability initiative by the DOE’s Advanced Scientific Computing Research Program (ASCR). For the success of this new initiative, it is of utmost importance to secure consistent funding and foster close engagement with researchers and core developers of existing programming-system products. This collaboration is vital to maintaining the critical capabilities of the current software during the transition phase while proactively adapting to future technology and workforce trends. The community recognizes the significance of adapting to emerging trends and is aware of the inherent fragility of the HPC software ecosystem, particularly in relation to programming systems that cater to all users. The ability to adapt and evolve is essential to staying relevant and effectively addressing these technical, economic, and social challenges. The S4PST team, which represents one of the six ASCR Software Sustainability seedling projects, is dedicated to tackling these challenges through community-based approaches that go beyond the scope of the DOE. This involves collaboration between national laboratories with academia, non-DOE institutions, hardware and system vendors, and international partners. By fostering these partnerships, we aim to create a robust and sustainable HPC software ecosystem that can effectively meet the needs of the community. This new community effort, driven by the eight DOE labs, will take on the responsibility of guiding funding decisions for programming-systems development and maintenance with transparency and consistency across all decisions. Additionally, the team will offer common technical services to the programming systems community, irrespective of their funding situations, and facilitate community-wide incubation to proactively nurture the software ecosystem. By actively engaging with stakeholders and employing a collaborative approach, we can collectively shape the future of programming systems and ensure a robust and thriving HPC software landscape. On May 11–12, 2023, the S4PST team conducted its inaugural kick-off workshop at the Innovative Computing Laboratory (ICL) in the University of Tennessee, Knoxville, hosted by Hartwig Anzt. The workshop encompassed various sessions dedicated to presentations and discussions, with the aim of comprehending the team members’ perspectives on the vision of software sustainability. Additionally, the workshop aimed to identify the technical, economic, and social requirements for sustaining the programming-systems community in the field of HPC. This report provides a summary of the S4PST effort by highlighting five major thrust areas discussed during the workshop: (i) community, (ii) technical support, (iii) training and diversity, (iv) verification, validation and correctness, and (v) emerging technologies. It also encompasses an overview of the presentations and discussions held throughout the event, our views and potential synergies with other seedling efforts, along with the outcomes and key takeaways from our initial discussions.

97 MATHEMATICS AND COMPUTING↗

MassiveGNN: Efficient Training via Prefetching for Massively Connected Distributed Graphs

Graph Neural Networks (GNN) are indispensable in learning from graph-structured data, yet their rising computational costs, especially on massively connected graphs, pose significant challenges in terms of execution performance. To tackle this, distributed-memory solutions such as partitioning the graph to concurrently train multiple replicas of GNNs are in practice. However, approaches requiring a partitioned graph usually suffer from communication overhead and load imbalance, even under optimal partitioning and communication strategies due to irregularities in the neighborhood minibatch sampling. This paper proposes practical trade-offs for improving the sampling and communication overheads for representation learn- ing on distributed graphs (using popular GraphSAGE architecture) by developing a parameterized prefetch and eviction scheme on top of the state-of-the-art Amazon DistDGL distributed GNN framework, demonstrating about 15–40% improvement in end-to-end training performance on the NERSC Perlmutter supercomputer for various OGB datasets.

Machine Leanring, high performance comptuing, grap↗

On the use of a multigrid-reduction-in-time algorithm for multiscale convergence of turbulence simulations

Simulations of turbulent flow present challenges in terms of accuracy and affordability on modern highly-parallel computer architectures. A multigrid-reduction-in-time algorithm is used to provide a framework for separately evolving different scales of turbulence and for parallelizing the temporal domain, thereby increasing the concurrency. It is hypothesized that the space–time locality of the small scales of turbulence can be used to circumvent difficulties in applying temporal multigrid to flows dominated by inertial physics. For algorithms that fall well short of spectral accuracy (fourth-order is used in this work) attention must be paid to the accuracy of features on scales transferred between multigrid levels. Numerical experiments were performed using implicit large-eddy simulation. Results from applying the approach to an infinite-Reynolds number Taylor–Green flow and a double-shear flow at a Reynolds number of 11650 provide strong evidence that the approach has merit. The multigrid-reduction-in-time framework can be used to parallelize the temporal domain of a high-Reynolds-number turbulent flow and permit independent convergence of different scales. Establishing this foundation allows for future research in reducing the wall-clock time to solve turbulent flows while retaining the same accuracy as sequential solvers. In conclusion, current performance results from parallelizing the temporal domain are not competitive with those from sequential-in-time methods.

97 MATHEMATICS AND COMPUTING↗

Cation and anion topotactic transformations in cobaltite thin films leading to Ruddlesden-Popper phases.

Topotactic transformations involve structural changes between related crystal structures due to a loss or gain of material while retaining a crystallographic relationship. The perovskite oxide La0.7Sr0.3CoO3 (LSCO) is an ideal system for investigating phase transformations due to its high oxygen vacancy conductivity, relatively low oxygen vacancy formation energy, and strong coupling of the magnetic and electronic properties to the oxygen stoichiometry. While the transition between cobaltite perovskite and brownmillerite (BM) phases has been widely reported, further reduction beyond the BM phase lacks systematic studies. In this paper, we study the evolution of the physical properties of LSCO thin films upon exposure to highly reducing environments. We observe the rarely reported crystalline Ruddlesden-Popper phase, which involves the loss of both oxygen anions and cobalt cations upon annealing where the cobalt is found as isolated Co ions or Co nanoparticles. First-principles calculations confirm that the concurrent loss of oxygen and cobalt ions is thermodynamically possible through an intermediary BM phase. The strong correlation of the magnetic and electronic properties to the crystal structure highlights the potential of utilizing ion migration as a basis for emerging applications such as neuromorphic computing.

Chiu, I-Ting↗

DLHub: Simplifying publication, discovery, and use of machine learning models in science

Machine Learning (ML) has become a critical tool enabling new methods of analysis and driving deeper understanding of phenomena across scientific disciplines. There is a growing need for "learning systems" to support various phases in the ML lifecycle. While others have focused on supporting model development, training, and inference, few have focused on the unique challenges inherent in science, such as the need to publish and share models and to serve them on a range of available computing resources. In this paper, we present the Data and Learning Hub for science (DLHub), a learning system designed to support these use cases. Specifically, DLHub enables publication of models, with descriptive metadata, persistent identifiers, and flexible access control. It packages arbitrary models into portable servable containers, and enables low-latency, distributed serving of these models on heterogeneous compute resources. In this work, we show that DLHub supports low-latency model inference comparable to other model serving systems including TensorFlow Serving, SageMaker, and Clipper, and improved performance, by up to 95%, with batching and memoization enabled. We also show that DLHub can scale to concurrently serve models on 500 containers. Finally, we describe five case studies that highlight the use of DLHub for scientific applications.

97 MATHEMATICS AND COMPUTING↗