Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “HPC Utilization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

University Data Management Pilot Utilizing the Nuclear Research Data System

Background In 2022, the Office of Science and Technology Policy (OSTP) issued a memo that significantly reshaped the landscape of access to federally funded research. The memo mandated that all taxpayer-funded research be made available to the public without delay upon publication, without an embargo period, superseding the 2013 OSTP public access policy. This public access policy promotes transparency and the democratization of knowledge, ensuring that the fruits of scientific endeavors funded by federal agencies could be immediately accessed and built upon by scientists, educators, students, and the public at large. To implement the requirements of the OSTP guidance and DOE Public Access Plan, the Office of Nuclear Energy (NE) has implemented public access plan guidance and has identified several areas where better data management practices would further expand public access to important nuclear energy related scientific data, reports, and other technical products. Significant NE supported efforts are already underway for data management and public access to important nuclear energy related data.1 2 To address gaps in data management practices, and improve retention and accessibility of data, NE is actively exploring enhanced data management options utilizing its high-performance computing resources administered by its Nuclear Scientific User Facility Program. A newly piloted system, the Nuclear Research Data System (NRDS) acts as a portal for data collection and dissemination. Nuclear Energy University Program Research and Development Portfolio According to Web of Science, NEUP has produced 2,345 journal publication that have been cited more than 61,000 times3 and countless conference proceedings. These publications are publicly available through OSTI.gov and in the open literature. Additional scientific and technical products including project milestones that are not publications and NEUP project final reports are vetted through OSTI.gov and released once reviewed and approved by DOE. Since 2009, NEUP has awarded close to 1,000 different R&D projects in technical areas across the NE research programs. As of June 2023, 512 NEUP reports are publicly available on OSTI. The underlying data for projects is still held at universities, and data transfer, co-location, and dissemination has not occurred in a systematic way. NEUP data is currently accessible through myriad university-based data repositories, or through direct requests to PIs. The program identified this patchwork of repositories, or often lack of publicly available data, as a significant barrier to an organized, accessible, and comprehensive solution to sharing data with the larger nuclear energy community. Approach The goal of this pilot project is to establish a pathway to a consolidated long-term repository for NEUP project data. To accomplish this goal, the pilot strives to accomplish the following objectives: Establish data collection standards, including a standard set of required supplementary information to contextualize and support raw data files. Work with the HPC group collect and upload information and to modify the NRDS system, as needed, to support a standardized approach. Resolve potential barriers to successful roll out of an expanded data collection strategy, including modifying data management plan guidelines and establishing a document and data release process that accounts for potential intellectual property and/or export control concerns. Results Overall, the pilot was successful in collecting 8,982 raw and processes data files, 220 reports, 56 calibration files, and 5,931 other supplementary documents. Supplementary documents included experimental plans, methods, journal publications and conference proceedings, milestone reports, and final reports. Figure 2 shows the number of data sets and supplementary project information provided by each project. Projects has significantly different input, depending on experimental data produced and completeness of the datasets provided.

Data collection↗

Transportation Hub Infrastructure Expansion: Decision Support Under Uncertainty

The Athena project (www.athena-mobility.org) has worked to investigate the relationship between the Dallas-Fort Worth Airport (DFW) and the greater Dallas area in order to better understand and therefore better inform future decision-making regarding the critical infrastructure that influence mobility between the airport and the city. Through this work, infrastructure related to curbside pickup and drop-off, parking, public transit, and the road network congestion were identified as critical to the operation of the DFW transportation hub. The infrastructure analysis and expansion aspect of the Athena project is focused on the restructuring of the CTA curb as a hierarchical curb and the building or repurposing of parking infrastructure as the interplay between these two areas. Many sources of uncertainty exist that may impact future airport and transportation hub operations, such as passenger volume growth, population demographic changes over time, electric vehicle (EV) adoption rates, and autonomous vehicle (AV) adoption rates. Due to these sources of uncertainty, we have selected for our research a modeling framework that can capture various types of uncertainty and hedge against those uncertainties in the optimization process. We analyze road network and curb congestion, the rise of transportation networking companies, trends in parking usage, existing policies around this infrastructure, airport revenue streams, and other contributing factors to enable infrastructure decision making with less uncertainty. To accomplish this wholistic analysis, we have developed a novel multi-stage, multi-period stochastic optimization model which considers the airport's decisions from 2025-2045 under different possible future macro trajectories and day-to-day variations in operational conditions captured as "annual representation of operations" scenarios with respective probabilities. This model has also been designed to leverage the outputs of various efforts under the Athena project to create a combined decision framework for infrastructure decisions. These various efforts include the route optimization model, the ASPIRES simulation, the mode choice model, and the SUMO traffic simulation. Our computational experiments of this system at scale have resulted in a working version of our infrastructure model which enables the explicit representation and consideration of various sources of uncertainty in the decision process to enable robust, flexible decision-making. This model has been effectively run on NREL's HPC system, Eagle, with large numbers of stochastic scenarios and shows promise as a scalable tool for robust consideration of uncertainties in airport planning. We have tested our model using 30,240 operational circumstances in total, resulting in a problem with more 200 million variables. This model was solved in several different configurations, and a workflow to simulate the performance of the infrastructure model results was developed and deployed. In general, our results indicate that a combination of remote parking, remote curb infrastructure, and dynamic pricing can generate revenue, reduce emissions, accommodate emerging technologies such as AVs and EVs, and manage airport passenger growth over time. We note the success of the proposed strategy depends on the data collection and forecasting abilities of DFW. We have also seen that the AV adoption by TNCs might necessitate larger amounts of remote curb. The results of this work inform strategies for airport infrastructure decision making, as well as demonstrate the value of an adaptable model, but also indicate that there are avenues remaining where further research would be of value.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

US Department of Energy, Office of Science High Performance Computing Facility Operational Assessment 2019 Oak Ridge Leadership Computing Facility

Oak Ridge National Laboratory's (ORNL's) Leadership Computing Facility (OLCF) continues to surpass its operational target goals: supporting users; delivering fast, reliable computational ecosystems; creating innovative solutions for high performance computing (HPC) needs; and managing risks, safety, and security associated with operating some of the most powerful computers in the world. The results can be seen in the cutting-edge science conducted by users and the praise from the research community. Calendar year (CY) 2019 was a big year as OLCF staff ran five world-class resources (the leadershipclass computers Titan and Summit, the large analysis cluster called Eos, and the massive parallel filesystems called Atlas and Alpine)) and also began power and cooling upgrades for a 2021 exascale system called Frontier. While continuing exceptional operation of Titan, Eos, and Rhea, the OLCF released the Summit supercomputer for production on January 1, 2019. Summit debuted as the most capable and efficient system in its class and has been recognized as the most powerful system in the world for its performance on both the high performance linpack (HPL) and conjugate gradient (HPCG) benchmark applications since June 2018 according to TOP500. Summit represents the culmination of a multiyear effort between the OLCF, IBM, NVIDIA, and Mellanox to deliver a system that is unmatched for modeling, simulation, data analysis, and learning. To hit the ground running with science-ready applications on day one, application teams worked closely with the OLCF through the Center for Accelerated Application Readiness (CAAR) program for years in advance of the Summit deployment. CY 2019 was filled with outstanding results and accomplishments: a very high rating from users on overall satisfaction for the sixth year in a row; a tremendous amount of core-hours delivered to researchers from two leadership-class systems; and success in delivering on the allocation split of roughly 60%, 30%, and 10% of core-hours offered for the Innovative and Novel Computational Impact on Theory and Experiment (INCITE), Advanced Scientific Computing Research Leadership Computing Challenge (ALCC), and Director's Discretionary (DD) programs, respectively (see Operational Performance section). These accomplishments, coupled with the high utilization rates (overall and capability usage), represent the fulfillment of the promise of both leadership-class machines: efficient facilitation of leadership-class computational applications. Table ES.1 presents a summary of the 2019 OLCF metric targets and the associated results. More information can be found in the Operational Performance section for each OLCF resource. The scientific accomplishments of OLCF users are a strong indication of long-term operational success, with publications this year in such notable journals and publications as Nature, Nature Physics, Nature Plants, Physical Review X, Journal of the American Physical Society, Cell, Nano Letters, and Trends in Biotechnology. Crucial domain-specific discoveries facilitated by resources at the OLCF are described in the High Performance Computing Facility Operational Assessment 2019 Oak Ridge Leadership Computing Facility (OAR) Strategic Results section. For example, researchers used Summit to pinpoint and understand the production of proteins from genetic information, including mutations and the functional expression of disease (Section 8.2).

97 MATHEMATICS AND COMPUTING↗

HPC4Mfg with Samsung: Making semiconductor devices cool through HPC ab initio simulations

For decades, the semiconductor technology has followed the Moore’s law, making newer devices more powerful and energy efficient. Recently, however, it has reached a point where the performance and the energy efficiency of the device do not improve with the shrinking device size. One of the fundamental reasons of this deviation from the past trend is the interconnect resistance, which becomes larger with the shrinking size. The devices size is so small that the quantum mechanical effects can no longer be ignored and the traditional continuum simulation tools such as TCAD become inadequate. In this project, Samsung Semiconductor Inc. and Lawrence Berkeley National Laboratory has collaborated to perform first of kind device-scale ab initio simulations to optimize materials and interconnect morphology to minimize interconnect resistance. We have tested the use of LS3DF method and the PEtot_trans approach on top of the folded spectrum method (FSM) Escan code to calculate the scattering state, and to study various effects influence the interconnect conductivity. We found that, the LS3DF can be used to calculate such metallic system. On the other hand, the use of Escan code to solve the linear equation is not practical due to the slow convergence. We have implemented a Chebyshev filter technique to calculate a few hundred eigen states near the scattering state energy E, then use these eigen states as preconditioner to solve the linear equation. We have used this approach to study the different factors which affect the interconnect conductivity, including the shape, the point defect, the temperature, and the grain boundary.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Simulation-based characterization of the variability of earthquake risk to buildings in the near-field

Recent advancements in high performance computing platforms and computational workflow for regional-scale simulations are enabling unprecedented modeling of fault-to-structure earthquake processes. Regional simulations resolving ground motions at frequencies relevant to engineered systems are becoming computationally viable and provide a new capability to improve understanding of the geographical distribution and intensity of risk to buildings and critical infrastructure. As computational capabilities advance, it is essential to move beyond illustrative single rupture realizations for scenario earthquake events towards the development of a full suite of rupture realizations that appropriately characterize the range of risk to building systems. The work described in this article investigates the application of a suite of fault rupture realizations with the objective of assessing near-fault, site-specific seismic demand variability for building structures. A representative high-performance regional-scale computational model is utilized to execute ground motion and building response simulations based on 18 kinematic rupture realizations of an M7 strike-slip scenario earthquake. The fault rupture models for the scenario earthquake are created by systematically perturbing the hypocenter location and stochastically generating rupture parameters (slip, rise time, rake angle) to represent a breadth of ground motion intensities resulting from the spatial and temporal variabilities of an earthquake rupture process. The resulting seismic demand variability for three-story (short period) and forty-story (long period) steel moment-resisting frame buildings is characterized in terms of the median and distribution of peak inter-story drift ratio for a range of near-fault sites. The full suite of 18 fault rupture realizations and approximately 280,000 nonlinear dynamic building simulations indicate that the three-story building undergoes higher median seismic demand and significantly greater variability of demand at a given site than the forty-story building, which has important implications for the level of certainty in predicting building performance during an earthquake. The simulations performed provide deeper insight into the relationship between fault rupture parameterization and building response, which is essential information for developing a representative suite of rupture realizations for specific earthquake scenarios.

58 GEOSCIENCES↗

Highly-scalable GPU-accelerated compressible reacting flow solver for modeling high-speed flows

Emerging supercomputing systems utilize a combination of central processing units (CPUs) and graphics processing units (GPUs) in an effort to reach exascale capabilities while minimizing the energy footprint of operating such systems. Such heterogeneous machines introduce new challenges for fluids solvers because the hardware architecture and operation of a GPU are fundamentally different from conventional CPUs. In this work, a general approach for efficient implementation of finite-volume based reacting flow solvers on such heterogeneous systems is presented. Three main challenges, namely, data access pattern, thread divergence, and thread safety, are addressed. Since compressible reacting flows require special methods to deal with chemical reactions, hyperbolic and nonlinear convection terms, and the presence of turbulence, specific algorithms that ensure GPU-based efficiency are developed. The approach is demonstrated on the widely available OpenFOAM open source software by modifying core algorithms for GPU accessibility. The scalability of the resulting solver, is demonstrated using practical test cases, including flow through a scramjet engine and the dynamics of a rotating detonation engine. Here, the solver provides near-ideal scaleup on a large number of GPUs (>3000), and extremely efficient use of the GPUs, with throughput nearly a constant even when processing a large number of control volumes.

42 ENGINEERING↗

Understanding power and energy utilization in large scale production physics simulation codes

Power is an often-cited reason for the move to advanced architectures on the path to Exascale computing. Here, this is due to practical considerations related to delivering enough power to successfully site and operate these machines, as well as concerns about energy usage while running large simulations. Since obtaining accurate power measurements can be challenging, it may be tempting to use the processor thermal design power (TDP) as a surrogate due to its simplicity and availability. However, TDP is not indicative of typical power usage while running simulations. Using commodity and advanced technology systems at Lawrence Livermore and Sandia National Labs, we performed a series of experiments to measure power and energy usage in running simulation codes. These experiments indicate that large scale Lawrence Livermore simulation codes are significantly more efficient than a simple processor TDP model might suggest.

HPC↗

An Early Investigation of the HHL Quantum Linear Solver for Scientific Applications

In this paper, we explore using the Harrow–Hassidim–Lloyd (HHL) algorithm to address scientific and engineering problems through quantum computing, utilizing the NWQSim simulation package on a high-performance computing platform. Focusing on domains such as power-grid management and climate projection, we demonstrate the correlations of the accuracy of quantum phase estimation, along with various properties of coefficient matrices, on the final solution and quantum resource cost in iterative and non-iterative numerical methods such as the Newton–Raphson method and finite difference method, as well as their impacts on quantum error correction costs using the Microsoft Azure Quantum resource estimator. We summarize the exponential resource cost from quantum phase estimation before and after quantum error correction and illustrate a potential way to reduce the demands on physical qubits. This work lays down a preliminary step for future investigations, urging a closer examination of quantum algorithms’ scalability and efficiency in domain applications.

hybrid software for QC-HPC↗

Brochure on the 2024 ASCR Workshop on Energy-Efficient Computing for Science

Large-scale computing has enabled numerous scientific discoveries, including ground-breaking achievements facilitated by the US Department of Energy (DOE) supercomputers and advances in applied mathematics and computer science. While important advances were made in energy efficiency to enable exascale computing, continued efforts are needed to dramatically improve the energy efficiency of the next generation of high-performance computing (HPC) systems and, more broadly, AI data centers. Without substantial improvements in energy efficiency, the energy consumption associated with computing could become a limiting factor for future scientific discovery, national security, and technological advancement.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

An intra-node HPC network architecture with nanosecond-scale photonic switches

We propose a single-stage network architecture for intra-node connectivity that makes use of nanosecond-scale photonic switches. Although buffering at the switch points is of vital importance for complex multi-stage networks, this is not the case for smaller-scale single-stage networks where the end nodes are located only one hop apart. By limiting the buffering to the end points, the proposed architecture manages to minimize the required electro-optic and opto-electronic conversions, leading in this way to both low end-to-end latency and better energy efficiency. Combining these advantages with nanosecond-scale switching times can allow for high-throughput operation even for frequent switch reconfigurations. The performance of the proposed architecture is evaluated via discrete-event simulations for a wide range of synthetic-traffic cases. The simulation results show that high-throughput operation of ≥90% can be achieved even for small message sizes, i.e., 32 KB for all-to-all communication and 2 KB for uniform random traffic, at a data rate of 400 Gb/s and a switch reconfiguration time of ≤72 ns. Moreover, if 100-ns reconfiguration times are achievable as opposed to 150-ns, then for the all-to-all traffic case a 16% and 28% reduction in completion time can be achieved for message sizes of 8 KB and 1 KB, respectively. In a forthcoming era of optically interfaced processors and accelerators, nanosecond-scale photonic switches appear as a highly promising solution for keeping up with the intra-node bandwidth scaling due to their high-bandwidth, low-latency and fast-switching capabilities.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Development and Experimental Validation of a High-Power DC Distribution Testbed for Advanced Charging Infrastructure and Energy Management

This paper presents the development of a hardware testbed for DC-distributed high-power charging (HPC) stations. As DC distributed solutions emerge as a viable solution to optimize HPC site operations, challenges such as interoperability, protection, and seamless integration of distributed energy resources (DER) persist. These issues underscore the need for a robust testing facility to investigate compliance of available commercial off-the-shelf (COTS) market devices. The developed testbed features a DC-distributed charging hub including a charger, emulated energy storage system (ESS), and site level communication and controller implementation. It facilitates the testing of COTS hardware, charger prototypes, standards validation and site energy management system (SEMS) controllers at rated power. This paper details the development of the charging infrastructure platform, implementation of communication system, validation of different SEMS algorithms, and understanding improvements required for future expansion. Using the developed testbed, interoperability gaps for SEMS implementation with multi-vehicle concurrent charging via a multi-port charger are experimentally observed. Aimed at supporting the transition to large-scale EV charging infrastructure deployment and DER integration, this testbed plays a crucial role in conformity testing of COTS device interoperability.

24 POWER TRANSMISSION AND DISTRIBUTION↗

US Department of Energy, Office of Science High Performance Computing Facility Operational Assessment 2021: Oak Ridge Leadership Computing Facility

Oak Ridge National Laboratory’s (ORNL’s) Leadership Computing Facility (OLCF) continues to surpass its operational target goals of supporting users; delivering fast, reliable computational ecosystems; creating innovative solutions for high-performance computing (HPC) needs; contributing to the community to build the next generation HPC workforce, and managing risks, safety, and security associated with operating some of the most powerful computers in the world. The results can be seen in the cutting-edge science conducted by users and the praise from the research community. Calendar year (CY) 2021 saw continued excellence in research supported by the OLCF’s leadership-class computing resources, including Summit (the nation’s most powerful supercomputer), the global scratch file system Alpine, the Scalable Protected Infrastructure (SPI), the Exploratory Visualization Environment for Research in Science and Technology (EVEREST), and the archival mass-storage resource High-Performance Storage System (HPSS). While maintaining access and exceptional user support for Summit, the OLCF continued to make progress on the installation and deployment of Frontier, which will be the nation’s first exascale system when it comes online at the start of CY 2023. Users have already begun running and optimizing scientific codes on Crusher, the OLCF test and development system equipped with Frontier’s architecture. Throughout the year, the OLCF maintained a strong culture of operational excellence, including risk management, workplace safety, and cybersecurity. The OLCF’s rigorous risk management strategy anticipated and mitigated risks, and at this time there are no high-priority operational risks. Similarly, ORNL and the OLCF were committed to operating under the US Department of Energy’s (DOE’s) safety regulations that ensure a safe workplace. Technical staff tracked and monitored existing threats and vulnerabilities within the OLCF while continually developing tools and practices to enhance operations without increasing the facility’s risk. CY 2021 was filled with outstanding results and accomplishments, including a very high rating from users on overall satisfaction for the eighth consecutive year; a tremendous number of node hours delivered to 1,671 researchers on Summit; and the successful delivery of the allocation split of roughly 60%, 20%, and 20% of core-hours offered for the Innovative and Novel Computational Impact on Theory and Experiment (INCITE), Advanced Scientific Computing Research Leadership Computing Challenge (ALCC), and Director’s Discretionary (DD) programs, respectively (Section 2). COVID-19 research remained a focus in 2021, and the ALCC and DD programs allocated over 1 million Summit hours to the COVID-19 High Performance Computing Consortium. These accomplishments, coupled with the high utilization rates (i.e., overall and capability usage), represent the fulfillment of the promise of leadership class machines: efficient facilitation of leadership-class computational applications.

97 MATHEMATICS AND COMPUTING↗

Massively parallel phase-field simulations targeting exascale

The interface thickness in the phase-field (PF) method limits its simulation scales. Consequently, large-scale PF simulations become prohibitively expensive for resolving the extremely fine microstructures that typically form during rapid solidification processing. This challenge is significant in predicting microstructure evolution in metal additive manufacturing and has been identified by the United States Department of Energy’s Exascale Computing Project. Here, to address this, we develop a multi-GPU and MPI-based massively parallel simulation code, utilizing state-of-the-art algorithms, software, and libraries, for large-scale three-dimensional (3D) PF simulations. We report the first GPU-parallel PF simulations on Frontier (currently the second TOP500 exascale cluster) and Summit machines, taking dendritic growth as an example problem. We evaluate the parallel performance of our implementation using scaling studies with more than 24 000 GPUs (among the largest known computations to date) and the acceleration performance using large-scale simulations of dendritic growth in 3D. Finally, massively parallel GPUs in these supercomputers enabled the first coupled multiscale simulations of laser melting and subsequent dendritic solidification on the scale of a full melt-pool, demonstrating the feasibility of performing PF simulations with a point total over 2 billion grid points within an acceptable time.

Exascale↗

MOSIQS: Persistent Memory Object Storage With Metadata Indexing and Querying for Scientific Computing

Scientific applications often require high-bandwidth shared storage to perform joint simulations and collaborative data analytics. Shared memory pools provide a chance to satisfy such needs. Recently, a high-speed network such as Gen-Z utilizing persistent memory (PM) offers an opportunity to create a shared memory pool connected to compute nodes. However, there are several challenges to use scientific applications on the shared memory pool directly such as scalability, failure-atomicity, and lack of scientific metadata-based search and query. In this paper, we propose MOSIQS, a persistent memory object storage framework with metadata indexing and querying for scientific computing. We design MOSIQS based on the key idea that memory objects on PM pool can live beyond the application lifetime and can become the sharing currency for applications and scientists. MOSIQS provides an aggregate memory pool atop an array of persistent memory devices to store and access memory objects to accelerate scientific computing. MOSIQS uses a lightweight persistent memory key-value store to manage the metadata of memory objects, which enables memory object sharing. To facilitate metadata search and query over millions of memory objects resident on memory pool, we introduce Group Split and Merge (GSM), a novel persistent index data structure designed primarily for scientific datasets. GSM splits and merges dynamically to minimize the query search space and maintains low query processing time while overcoming the index storage overhead. MOSIQS is implemented on top of PMDK. We evaluate the proposed approach on many-core server with an array of real PM devices. Experimental results show that MOSIQS gains a 100% write performance improvement and executes multi-attribute queries efficiently with 2.7× less index storage overhead offering significant potential to speed up scientific computing applications.

97 MATHEMATICS AND COMPUTING↗

Height Above Nearest Drainage (HAND) and Hydraulic Property Table for CONUS

The continental flood inundation mapping (CFIM) framework is a high-performance computing (HPC)-based computational framework for the Height Above Nearest Drainage (HAND)-based inundation mapping methodology. Using the 10m Digital Elevation Model (DEM) data produced by U.S. Geological Survey (USGS) 3DEP (the 3-D Elevation Program) and the NHDPlus hydrography dataset produced by USGS and the U.S. Environmental Protection Agency (EPA), a hydrological terrain raster called HAND is computed for HUC6 units in the conterminous U.S. (CONUS). The value of each raster cell in HAND is an approximation of the relative elevation between the cell and its nearest water stream. Derived from HAND, a hydraulic property table is established to calculate river geometry properties for each of the 2.7 million river reaches covered by NHDPlus (5.5 million kilometers in total length). This table is a lookup table for water depth given an input stream flow value. Such lookup is available between water depth 0m and 25m at 1-foot interval. The flood inundation map is then computed by using HAND and this lookup table based on the near real-time water forecast from the National Water Model (NWM) at the National Oceanic and Atmospheric Administration (NOAA). HAND and the Hydraulic Property Table version 0.2. is created to correspond to data updates in the USGS 1/3 arcsec DEM, the USGS National Hydrography Dataset, including its Water Boundary Dataset, and the NHDPlus medium resolution dataset. This dataset comprises 331 HUC6 units for CONUS (excluding the five great lakes units), each is a downloadable zip file. Version 0.1 was computed at the National Center for Supercomputing Applications (NCSA) at the University of Illinois at Urban-Champaign in 2016 and is currently hosted at the Texas Advanced Computing Center (TACC). Please cite this data DOI and the following publications and data DOIs when you use HAND and the hydraulic property table: Liu, Yan Y., David R. Maidment, David G. Tarboton, Xing Zheng, and Shaowen Wang. 'A CyberGIS integration and computation framework for high-resolution continental-scale flood inundation mapping.' JAWRA Journal of the American Water Resources Association 54, no. 4 (2018): 770-784. DOI: 10.1111/1752-1688.12660 Zheng, Xing, David G. Tarboton, David R. Maidment, Yan Y. Liu, and Paola Passalacqua. 'River channel geometry and rating curve estimation using height above the nearest drainage.' JAWRA Journal of the American Water Resources Association 54, no. 4 (2018): 785-806. DOI: 10.1111/1752-1688.12661 Liu, Y. (2018). Height Above Nearest Drainage (HAND) for CONUS - v0.1, HydroShare, https://doi.org/10.4211/hs.69f7d237675c4c73938481904358c789 LICENSE FOR USE -- MAPS AND DATA DISCLAIMER This resource is shared under the Creative Commons Attribution CC BY, http://creativecommons.org/licenses/by/4.0/ MAPS AND DATA DISCLAIMER The Oak Ridge National Laboratory (ORNL) shall not be held liable for improper or incorrect use of the data described or information contained on this map or associated series of maps. The data and related map graphics are not legal, land survey or engineering documents and are not intended to be used as such. ORNL gives no warranty, express or implied, as to the accuracy, reliability, utility or completeness of this information. The user of these maps and data assumes all responsibility and risk for the use of the maps and data. ORNL disclaims all warranties, representations or endorsements either express or implied, with regard to the information contained in this map product, including, but not limited to, all implied warranties of merchantability, fitness for a particular purpose or non-infringement. This preliminary map product is for research and review purposes only. It is not intended to be used for emergency management operational or life safety decisions at the local or regional governmental level or by the general public. Users requiring information regarding hazardous conditions or meteorological conditions for specific geographic areas should consult directly with their city or county emergency management office.

54 ENVIRONMENTAL SCIENCES↗

Height Above Nearest Drainage (HAND) and Hydraulic Property Table for CONUS - Version 0.21 (20200601)

The continental flood inundation mapping (CFIM) framework is a high-performance computing (HPC)-based computational framework for the Height Above Nearest Drainage (HAND)-based inundation mapping methodology. Using the 10m Digital Elevation Model (DEM) data produced by U.S. Geological Survey (USGS) 3DEP (the 3-D Elevation Program) and the NHDPlus hydrography dataset produced by USGS and the U.S. Environmental Protection Agency (EPA), a hydrological terrain raster called HAND is computed for HUC6 units in the conterminous U.S. (CONUS). The value of each raster cell in HAND is an approximation of the relative elevation between the cell and its nearest water stream. Derived from HAND, a hydraulic property table is established to calculate river geometry properties for each of the 2.7 million river reaches covered by NHDPlus (5.5 million kilometers in total length). This table is a lookup table for water depth given an input stream flow value. Such lookup is available between water depth 0m and 25m at 1-foot interval. The flood inundation map is then computed by using HAND and this lookup table based on the near real-time water forecast from the National Water Model (NWM) at the National Oceanic and Atmospheric Administration (NOAA). HAND and the Hydraulic Property Table version v0.21 is an upgrade to v0.2. Version 0.21 applies NHD HR etching to DEMs for all HUC6 units before running the HAND workflow. This new version also fixed the issue of void areas along HUC boundary. This dataset comprises 331 HUC6 units for CONUS (excluding the five great lakes units), each is a downloadable zip file. Version 0.1 was computed at the National Center for Supercomputing Applications (NCSA) at the University of Illinois at Urban-Champaign in 2016 and is currently hosted at the Texas Advanced Computing Center (TACC). Please cite this data DOI and the following publications and data DOIs when you use HAND and the hydraulic property table: Liu, Yan Y. Height Above Nearest Drainage (HAND) and Hydraulic Property Table for CONUS - Version 0.2. (20200301). Oak Ridge Leadership Computing Facility. (2020). DOI: 10.13139/ORNLNCCS/1608331 Liu, Yan Y., David R. Maidment, David G. Tarboton, Xing Zheng, and Shaowen Wang. 'A CyberGIS integration and computation framework for high-resolution continental-scale flood inundation mapping.' JAWRA Journal of the American Water Resources Association 54, no. 4 (2018): 770-784. DOI: 10.1111/1752-1688.12660 Zheng, Xing, David G. Tarboton, David R. Maidment, Yan Y. Liu, and Paola Passalacqua. 'River channel geometry and rating curve estimation using height above the nearest drainage.' JAWRA Journal of the American Water Resources Association 54, no. 4 (2018): 785-806. DOI: 10.1111/1752-1688.12661 Liu, Y. (2018). Height Above Nearest Drainage (HAND) for CONUS, HydroShare, https://doi.org/10.4211/hs.69f7d237675c4c73938481904358c789 LICENSE FOR USE -- MAPS AND DATA DISCLAIMER This resource is shared under the Creative Commons Attribution CC BY, http://creativecommons.org/licenses/by/4.0/ MAPS AND DATA DISCLAIMER The Oak Ridge National Laboratory (ORNL) shall not be held liable for improper or incorrect use of the data described or information contained on this map or associated series of maps. The data and related map graphics are not legal, land survey or engineering documents and are not intended to be used as such. ORNL gives no warranty, express or implied, as to the accuracy, reliability, utility or completeness of this information. The user of these maps and data assumes all responsibility and risk for the use of the maps and data. ORNL disclaims all warranties, representations or endorsements either express or implied, with regard to the information contained in this map product, including, but not limited to, all implied warranties of merchantability, fitness for a particular purpose or non-infringement. This preliminary map product is for research and review purposes only. It is not intended to be used for emergency management operational or life safety decisions at the local or regional governmental level or by the general public. Users requiring information regarding hazardous conditions or meteorological conditions for specific geographic areas should consult directly with their city or county emergency management office.

54 ENVIRONMENTAL SCIENCES↗

Machine learning-driven predictive resource management in complex science workflows

Here, the collaborative efforts of large communities in science experiments, often comprising thousands of global members, reflect a monumental commitment to exploration and discovery. Recently, advanced and complex data processing has gained increasing importance in science experiments. Data processing workflows typically consist of multiple intricate steps, and the precise specification of resource requirements is crucial for each step to allocate optimal resources for effective processing. Estimating resource requirements in advance is challenging due to a wide range of analysis scenarios, varying skill levels among community members, and the continuously increasing spectrum of computing options. One practical approach to mitigate these challenges involves initially processing a subset of each step to measure precise resource utilization from actual processing profiles before completing the entire step. While this two-staged approach enables processing on optimal resources for most of the workflow, it has drawbacks such as initial inaccuracies leading to potential failures and suboptimal resource usage, along with overhead from waiting for initial processing completion, which is critical for fast-turnaround analyses. In this context, our study introduces a novel pipeline of machine learning models within a comprehensive workflow management system, the Production and Distributed Analysis (PanDA) system. These models employ advanced machine learning techniques to predict key resource requirements, overcoming challenges posed by limited upfront knowledge of characteristics at each step. Accurate forecasts of resource requirements enable informed and proactive decision-making in workflow management, enhancing the efficiency of handling diverse, complex workflows across heterogeneous resources.

97 MATHEMATICS AND COMPUTING↗

Model America - data and models of every U.S. building

The 5-year goal of the 'Model America' concept was to generate a model of every building in the United States. This data repository delivers on that goal. Oak Ridge National Laboratory (ORNL) has developed the Automatic Building Energy Modeling (AutoBEM) software suite to process multiple types of data, extract building-specific descriptors, generate building energy models, and simulate them on High Performance Computing (HPC) resources. For more information, see AutoBEM-related publications (bit.ly/AutoBEM). There were 125,714,640 buildings detected in the United States and this dataset contains 122,930,327 (97.8%) buildings which resulted in a successful simulation. Future, annual updates have been proposed that may include additional buildings, data improvements, or other algorithmic enhancements. This dataset of 122.9 million buildings includes: Models (state_county.zip) - OpenStudio (v3.1.0) and EnergyPlus (v9.4) building energy models. Please note that the download requires the free Globus Connect Personal (https://www.globus.org/globus-connect-personal); Each model has approximately 3,000 building input descriptors that can be extracted. Please see the EnergyPlus(v9.4) 2,784-page Input/Output Reference Guide (https://energyplus.net/sites/all/modules/custom/nrel_custom/pdfs/pdfs_v9.4.0/InputOutputReference.pdf) for everything that can be retrieved or simulated from these models. These models were derived from the following metadata, which is not included in this dataset: 1. ID - unique building ID 2. County - county name 3. State - state name 4. CZ - ASHRAE Climate Zone designation 5. Clim_Zone - text label of climate zone 6. est_year - estimated year of construction 7. est_commercial - estimated building type (0=residential, 1=commercial) 8. Centroid - building center location in latitude/longitude (from Footprint2D) 9. Footprint2D - building polygon of 2D footprint (lat1/lon1_lat2/lon2_...) 10. Height - building height (meters) 11. Area2D - footprint area (ft2) 12. BuildingType - DOE prototype building designation (IECC=residential) as implemented by OpenStudio-standards 13. WWR_surfaces - percent of each facade (pair of points from Footprint2D) covered by fenestration/windows (average 14.5% for residential, 40% for commercial buildings) 14. NumFloors - number of floors (above-grade) 15. Area - estimate of total conditioned floor area (ft2) 16. Standard - building vintage. These models are made free and openly available in hopes of stimulating any simulation-informed use case. Data is provided as-is with no warranties, express or implied, regarding fitness for a particular purpose. We wish to thank our sponsors which include Oak Ridge National Laboratory (ORNL) Laboratory Directed Research and Development (LDRD), U.S. Dept. of Energy's (DOE) Building Technologies Office (BTO), Office of Electricity (OE), Biological and Environmental Research (BER), and National Nuclear Security Administration (NNSA). This research used resources of the Argonne Leadership Computing Facility, which is a DOE Office of Science User Facility supported under Contract DE-AC02-06CH11357. Please cite as: New, Joshua R., Adams, Mark, Bass, Brett, Berres, Anne, and Clinton, Nicholas (2021). 'Model America - data and models of every U.S. building. [Data set].' Constellation, doi.ccs.ornl.gov/ui/doi/339, April 14, 2021

24 POWER TRANSMISSION AND DISTRIBUTION↗