Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “extreme-scale”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

90 records · Page 5

Domain-decomposition nonlinear manifold reduced order model

This software combines nonlinear-manifold reduced order models (NM-ROMs) with domain decomposition (DD) techniques. NM-ROMs, which utilize a shallow, sparse autoencoder trained with full order model (FOM) snapshot data, approximate the FOM state on a nonlinear manifold. These models offer advantages over linear-subspace ROMs (LS-ROMs) particularly in scenarios with slowly decaying Kolmogorov n-width. However, the training of NM-ROMs involves a number of parameters that scale with the size of the FOM, and storing high-dimensional FOM snapshots can significantly increase the cost of ROM training for extreme-scale problems. To mitigate these costs, the software employs DD to partition the FOM into smaller subdomains, computes NM-ROMs for each, and then integrates these to form a global NM-ROM. This strategy offers multiple benefits: it enables parallel training of subdomain NM-ROMs, reduces the number of parameters needed, decreases the dimensional requirements of subdomain FOM training data, and allows for customization to the unique characteristics of each FOM subdomain. The use of a shallow, sparse autoencoder architecture in each subdomain NM-ROM facilitates the application of hyper-reduction (HR), simplifying the nonlinear complexities and enhancing computational speed. This software marks the inaugural application of NM-ROM combined with HR to a DD problem. It features an algebraic DD reformulation of the FOM, training of NM-ROMs with HR for each subdomain, and employs a sequential quadratic programming (SQP) solver for the evaluation of the coupled global NMROM. The effectiveness of the DD NM-ROM with HR is numerically demonstrated on the 2D steady-state Burgers' equation, showing an order of magnitude improvement in accuracy over the DD LS-ROM with HR.

Diaz, AlejandroN↗

Aero-structural design and optimization of 50 MW wind turbine with over 250-m blades

The quest for reduced levelized cost of energy has driven significant growth in wind turbine size; however, larger rotors face significant technical and logistical challenges. The largest published rotor design is 25 MW, and here we consider an even larger 50 MW design with blade length over 250 m. This paper shows that a 50 MW design is indeed possible from a detailed engineering perspective and presents a series of aero-structural blade designs, and critical assessment of technology pathways and challenges for extreme-scale rotors. The 50 MW rotor design begins with Monte Carlo simulations focused on optimizing carbon spar cap and root design. A baseline design resulted in a 250-m blade with mass of 502 tonnes. Subsequently, an aero-structural design and optimization were performed to reduce the blade mass/cost with more than 25% mass reduction and 30% cost reduction by determining optimal blade chord and airfoil thickness for best aero-structural performance.

Yao, Shulong↗

Globus service enhancements for exascale applications and facilities

Many extreme-scale applications require the movement of large quantities of data to, from, and among leadership computing facilities, as well as other scientific facilities and the home institutions of facility users. These applications, particularly when leadership computing facilities are involved, can touch upon edge cases (e.g., terabyte files) that had not been a focus of previous Globus optimization work, which had emphasized rather the movement of many smaller (megabyte to gigabyte) files. We report here on how automated client-driven chunking can be used to accelerate both the movement of large files and the integrity checking operations that have proven to be essential for large data transfers. In conclusion, we present detailed performance studies that provide insights into the benefits of these modifications in a range of file transfer scenarios.

97 MATHEMATICS AND COMPUTING↗

libEnsemble: A complete Python toolkit for dynamic ensembles of calculations

Almost all science and engineering applications eventually stop scaling: their runtime no longer decreases as available computational resources increase. Therefore, many applications will struggle to efficiently use emerging extreme-scale high-performance, parallel, and distributed systems. libEnsemble is a complete Python toolkit and workflow system for intelligently driving ensembles of experiments or simulations at massive scales. It enables and encourages multidisciplinary design, decision, and inference studies portably running on laptops, clusters, and supercomputers.

97 MATHEMATICS AND COMPUTING↗

ECP Software Technology Capability Assessment Report

The Exascale Computing Project (ECP) Software Technology (ST) Focus Area is responsible for developing critical software capabilities that will enable successful execution of ECP applications, and for providing key components of a productive and sustainable Exascale computing ecosystem that will position the US Department of Energy (DOE) and the broader high performance (HPC) community with a firm foundation for future extreme-scale computing capabilities. This ECP ST Capability Assessment Report (CAR) provides an overview and assessment of current ECP ST capabilities and activities, giving stakeholders and the broader HPC community information that can be used to assess ECP ST progress and plan their own efforts accordingly. ECP ST leaders commit to updating this document on regular basis (every six to 12 months).

97 MATHEMATICS AND COMPUTING↗

Summarizing the interoperabilities between xSDK members

Rapid, efficient production of high-quality, sustainable extreme-scale scientific applications is best accomplished using a rich ecosystem of state-of-the art reusable libraries, tools, lightweight frameworks, and defined software methodologies, developed by a community of scientists who are striving to identify, adapt, and adopt best practices in software engineering. The vision of the xSDK is to provide infrastructure for and interoperability of a collection of related and complementary software elements — developed by diverse, independent teams throughout the high-performance computing (HPC) community — that provide the building blocks, tools, models, processes, and related artifacts for rapid and efficient development of high-quality applications. A primary component, and challenge, of the xSDK is to improve interoperability among software libraries and domain components. This document summarizes the current and planned interoperabilities of the twenty-three xSDK member packages as of March 2021. Additionally, the status of the xSDK example codes that demonstrate and test the interoperabilities within the xSDK is provided.

97 MATHEMATICS AND COMPUTING↗

Science & Technology Review: The Road to Exascale Computing

At Lawrence Livermore National Laboratory, we focus on science and technology research to ensure our nation’s security. We also apply that expertise to solve other important national problems in energy, bioscience, and the environment. Science & Technology Review is published eight times a year to communicate, to a broad audience, the Laboratory’s scientific and technological accomplishments in fulfilling its primary missions. The publication’s goal is to help readers understand these accomplishments and appreciate their value to the individual citizen, the nation, and the world. The Department of Energy’s Exascale Computing Project (ECP) and Lawrence Livermore’s RADIUSS (Rapid Application Development via an Institutional Universal Software Stack) initiative benefit from strategically developed software tools. The front cover shows a simulation of advection under twisting rotation that uses high-order finite elements from Livermore’s Modular Finite Element Methods (MFEM) software library and GLVis visualization tool. On the back cover, the logo (also created with GLVis) for the MFEM project illustrates the curved mesh and sub-element resolution used in high-order simulations. MFEM and GLVis are key components of the ECP’s co-design Center for Efficient Exascale Discretizations (CEED) and RADIUSS. MFEM is also part of ECP’s Extreme-Scale Scientific Software Development Kit (xSDK).

97 MATHEMATICS AND COMPUTING↗

Integrated Research Infrastructure Architecture Blueprint Activity (Final Report 2023)

The complexity of scientific pursuits is increasing rapidly with aspects that require dynamic integration of experiment, observation, theory, modeling, simulation, visualization, machine learning (ML), artificial intelligence (AI), and analysis. Research projects across the Department of Energy (DOE) are increasingly data and compute intensive. Innovative research teams are accelerating the pace of discovery by using high-performance computational and data tools in their research workflows and leveraging multiple research infrastructures. Additionally, several recent high-level U.S. government reports underscore the necessity of a new advanced computing ecosystem for international competitiveness and national security. International competitors are moving forward with major research infrastructure integration efforts that seek to capture a competitive advantage in the global innovation race. Owing to its unparalleled constellation of world-class experimental and observational facilities and high-performance and extreme-scale computational, data, and networking infrastructure, DOE is positioned to be a global leader in this new era of integrated science. However, this new integration paradigm will demand continuing evolution to ensure the U.S. remains a global leader in research and innovation. The DOE Office of Science (SC) has seized on the strategic importance of integration and has adopted a vision for Integrated Research Infrastructure (IRI): To empower researchers to meld DOE’s world-class research tools, infrastructure, and user facilities seamlessly and securely in novel ways to radically accelerate discovery and innovation. To respond to the evolving computational requirements of research and the competitive international innovation landscape, experimental facilities could be connected with high performance computing resources for near real-time analysis, and resources should be provided for merging enormous and diverse data for AI/ML techniques and analysis.

97 MATHEMATICS AND COMPUTING↗

A Contextually-Aware Sensitivity Analysis to Guide the Design of Randomized Least Squares Solvers in Applications

Our work on the DOE-sponsored project “A Contextually-Aware Sensitivity Analysis to Guide the Design of Randomized Least Squares Solvers in Applications,” was an effort to address critical challenges in nu merical computing and its applications to optimization. The increasing demand for robust and scalable solutions to large-scale linear algebra problems has highlighted the limitations of traditional approaches, particularly in heterogeneous and extreme-scale computing environments. Randomized Numerical Linear Algebra (RandNLA) offers a promising framework to address these challenges, and this proposal builds on this foundation by introducing innovations in sensitivity analysis and computational adaptability.

97 MATHEMATICS AND COMPUTING↗

Rapid Optimization of Total Variation with Applications in Imaging, Additive Manufacturing, and Qualification

Total Variation optimization penalizes the gradient of a control variable or state. While this work focuses on image processing in particular, it has also found applications in inverse problems and topology optimization. In image processing, the goal is to maintain faithfulness to the original image while denoising and/or deblurring. Additionally, bilevel optimization over the spatially varying regularization weights can illuminate interfaces such as damage regions and other anomalies. We will address two fundamental challenges with TV-optimization: (i) the typical slow convergence of existing TV-optimization methods, and (ii) the selection of spatially varying TV parameters to promote interface detection. Additionally, we will apply such techniques to image data collected in additive manufacturing. In said context, stochasticity in build events induces flaws in the manufactured piece, compromising the integrity of said part. There is a critical need for in-situ monitoring to spot anomalies once they form, and in this setting we apply our total variation and hyperparameter solvers. We will develop a customized algorithm based on for extreme-scale TV-optimization that achieves super-linear or quadratic-convergence, a critical property for real-time, image-by-image analysis. A worst-case outcome is a preprocessing step that enhances image quality in-situ, specifically for out-of-focus and noisy images.

36 MATERIALS SCIENCE↗

Simulation Center for Runaway Electron Avoidance and Mitigation (SCREAM SciDAC) (Technical Final Report)

Runaway electrons can severely damage the plasma facing components on ITER during a major disruption and pose a major risk for tokamak fusion. It has been recognized that an adequate disruption mitigation system (DMS) is essential for the safe operation of ITER. The United States is responsible for the design and implementation of the disruption mitigation system on ITER, and in July 2016 the Simulation Center for Runaway Electron Avoidance and Mitigation (SCREAM) was launched by DOE, in a joint Fusion Energy Sciences (FES) and Advanced Scientific Computing Research (ASCR) collaboration. SCREAM was a comprehensive theory and simulation SciDAC center that provided physics guidance in the avoidance and mitigation of runaway electrons, and in tandem with domestic and international experiments, helped establish the qualitative and quantitative bases for safe operational scenarios and viable mitigation techniques. The SCREAM center assembled a national team of experts in runaway electron physics, tokamak disruptions, magnetohydrodynamic (MHD) simulation, and advanced algorithms and computing. The team combined advanced simulation and analysis capability facilitated by direct participation of ASCR SciDAC institutes with theoretical models and code development by FES scientists to focus on the runaway risk for ITER and tokamaks in general. The research scope was focussed on integrated simulations of kinetic runaway electrons, including MHD and fluid models of impurity transport, within a research plan guided by theory. The specific research tasks were (1) establish the fundamental physics of runaway generation, saturation, and dynamical evolution in a tokamak; (2) examine the critical path toward runaway avoidance; and (3) investigate the viability and effectiveness of the leading candidate schemes for runaway mitigation. In all three areas, members of the team carried out scoping studies that established the readiness for rapid and critical advances, especially in the deployment and further development of large-to extreme-scale simulation tools. Our multi-pronged computational approach included (1) relativistic Fokker-Planck solvers with discretization in phase space, (2) self-consistent particle-in-cell techniques, (3) particle-based Monte-Carlo, and (4) MHD-particle hybrid simulations. Cross-check between these different methods provided an additional means for verification and further bolstered the fidelity of our physics prediction. Validation against experimental results brings confidence to the predictive capability for ITER and frequently leads to new ideas for understanding and mitigating the thermal quench driven runaway electron phenomenon.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Tachyon: Intelligent Multi-Scale Modeling of Distributed Resilient Infrastructure and Workflows for Data Intensive HEP Analyses

The DOE High Energy Physics (HEP) program in Neutrino and Collider science drives data-intensive science and simulation on extreme-scale platforms. Modeling and optimizing the complex distributed components from experimental to leadership computing facilities are essential for HEP workflows to achieve required response times and resilience under various conditions. Tachyon proposes a framework for scalable modeling, simulation, and validation of key performance characteristics for the distributed infrastructure between FNAL and ALCF, along with associated HEP workflows.

Carothers, Chris [Rensselaer Poly.]↗

Tachyon: Intelligent Multi-Scale Modeling of Distributed Resilient Infrastructure and Workflows for Data Intensive HEP Analyses

The DOE High Energy Physics (HEP) program in Neutrino and Collider science drives data-intensive science and simulation on extreme-scale platforms. Modeling and optimizing the complex distributed components from experimental to leadership computing facilities are essential for HEP workflows to achieve required response times and resilience under various conditions. Tachyon proposes a framework for scalable modeling, simulation, and validation of key performance characteristics for the distributed infrastructure between FNAL and ALCF, along with associated HEP workflows.

Carothers, Chris [Rensselaer Poly.]↗

Characterization and Valuation of the Uncertainty of Calibrated Parameters in Microsimulation Decision Models

We evaluated the implications of different approaches to characterize the uncertainty of calibrated parameters of microsimulation decision models (DMs) and quantified the value of such uncertainty in decision making. We calibrated the natural history model of CRC to simulated epidemiological data with different degrees of uncertainty and obtained the joint posterior distribution of the parameters using a Bayesian approach. We conducted a probabilistic sensitivity analysis (PSA) on all the model parameters with different characterizations of the uncertainty of the calibrated parameters. We estimated the value of uncertainty of the various characterizations with a value of information analysis. We conducted all analyses using high-performance computing resources running the Extreme-scale Model Exploration with Swift (EMEWS) framework. The posterior distribution had a high correlation among some parameters. The parameters of the Weibull hazard function for the age of onset of adenomas had the highest posterior correlation of -0.958. When comparing full posterior distributions and the maximum-a-posteriori estimate of the calibrated parameters, there is little difference in the spread of the distribution of the CEA outcomes with a similar expected value of perfect information (EVPI) of $\$$653 and $\$$685, respectively, at a willingness-to-pay (WTP) threshold of $\$$66,000 per quality-adjusted life year (QALY). Ignoring correlation on the calibrated parameters’ posterior distribution produced the broadest distribution of CEA outcomes and the highest EVPI of $\$$809 at the same WTP threshold. Different characterizations of the uncertainty of calibrated parameters affect the expected value of eliminating parametric uncertainty on the CEA. Ignoring inherent correlation among calibrated parameters on a PSA overestimates the value of uncertainty.

97 MATHEMATICS AND COMPUTING↗

Strategies for Working Remotely: Responding to Pandemic-Driven Change.

In response to the COVID-19 pandemic, the Exascale Computing Project’s (ECP) Interoperable Design of Extreme-scale Application Software (IDEAS) productivity team launched the panel series Strategies for Working Remotely to facilitate informal, cross-organizational dialog in the absence of face-to-face meetings. In a time of pandemic, organizations increasingly need to reach across perceived boundaries to learn from each other, so that we can move beyond stand-alone silos to more connected multidisciplinary and multiorganizational configurations. The present paper argues that the unplanned transition to remote work, overuse of electronic communication, and need to unlearn habits associated with an overreliance on face-to-face, created unique opportunities to learn from the situation and accelerate cross-institutional cooperation and collaboration through online community dialog facilitated by informal panel discussions. Recommendations for facilitating online panel discussions to foster cross-organizational dialog are provided by applying the Simulation Experience Design Method.

Raybourn, Elaine M.↗

Adaptive Mesh Refinement Simulations for Turbulent Reacting Flow

With the increased availability of exascale computing hardware, detailed simulations of realistic devices can be performed at practically relevant time and length scales. Insights into the multiscale driving mechanisms in compressible reacting flow systems with complex geometry, such as combustors, can be used for design optimization and technology improvements. However, to effectively perform these simulations, advanced numerical algorithms must be used to maintain solution accuracy without incurring undue computational costs. PeleC, part of the Pele suite of codes, leverages block-structured adaptive mesh refinement (AMR) through the AMReX library to capture fine-scale flow features in compressible reacting flows. In this talk, we discuss recent improvements to the numerical algorithms, particularly in regard to describing flows at complex boundary structures, and PeleC's performance on exascale computing hardware. We will demonstrate that PeleC is well-suited for modern, extreme-scale, heterogenous compute platforms.

combustion↗

Parked aeroelastic field rotor response for a 20% scaled demonstrator of a 13‐MW downwind turbine

Abstract Aeroelastic parked testing of a unique downwind two‐bladed subscale rotor was completed to characterize the response of an extreme‐scale 13‐MW turbine in high‐wind parked conditions. A 20% geometric scaling was used resulting in scaled 20‐m‐long blades, whose structural and stiffness properties were designed using aeroelastic scaling to replicate the nondimensional structural aeroelastic deflections and dynamics that would occur for a lightweight, downwind 13‐MW rotor. The subscale rotor was mounted and field tested on the two‐bladed Controls Advanced Research Turbine (CART2) at the National Renewable Energy Laboratory's Flatiron Campus (NREL FC). The parked testing of these highly flexible blades included both pitch‐to‐run and pitch‐to‐feather configurations with the blades in the horizontal braked orientation. The collected experimental data includes the unsteady flapwise root bending moments and tip deflections as a function of inflow wind conditions. The bending moments are based on strain gauges located in the root section, whereas the tip deflections are captured by a video camera on the hub of the turbine pointed toward the tip of the blade. The experimental results are compared against computational predictions generated by FAST, a wind turbine simulation software, for the subscale and full‐scale models with consistent unsteady wind fields. FAST reasonably predicted the bending moments and deflections of the experimental data in terms of both the mean and standard deviations. These results demonstrate the efficacy of the first such aeroelastically scaled turbine test and demonstrate that a highly flexible lightweight downwind coned rotor can be designed to withstand extreme loads in parked conditions.

17 WIND ENERGY↗

The PetscSF Scalable Communication Layer

PetscSF, the communication component of the Portable, Extensible Toolkit for Scientific Computation (PETSc), is designed to provide PETSc's communication infrastructure suitable for exascale computers that utilize GPUs and other accelerators. PetscSF provides a simple application programming interface (API) for managing common communication patterns in scientific computations by using a star-forest graph representation. PetscSF supports several implementations based on MPI and NVSHMEM, whose selection is based on the characteristics of the application or the target architecture. An efficient and portable model for network and intra-node communication is essential for implementing large-scale applications. The Message Passing Interface, which has been the de facto standard for distributed memory systems, has developed into a large complex API that does not yet provide high performance on the emerging heterogeneous CPU-GPU-based exascale systems. Here, we discuss the design of PetscSF, how it can overcome some difficulties of working directly with MPI on GPUs, and we demonstrate its performance, scalability, and novel features.

97 MATHEMATICS AND COMPUTING↗