Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “High performance Computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Development of a High Performance Computing Accounts Administration Web Application

INL's High Performance Computing (HPC) accounts administration tool is a web application that that allows an administrator to change user information, create an HPC account, create or modify an LDAP group or project, and modify the members in a group. An administrator can search for a user, obtain the relevant information about the user, and update their account information. To create or modify an LDAP group or project, the administrator can filter a list of groups or projects according to the name, owner, or administrator that created the project or group, choose a name accordingly, and specify the group or project owner.

97 MATHEMATICS AND COMPUTING↗

High Performance Computing Systems Tools, Visualization, and Management

High Performance Computing (HPC) systems are complex setups of servers, storage devices, network switches, and cables that are specifically designed to accommodate hundreds of users running highly computationally intensive applications at a time. These applications require numerous softwares, licenses, and various levels of storage as well. All of these resources must be monitored and managed by HPC administrators, which presents a daunting task. In this project, I created numerous software tools as part of an HPC Visualization and Management system, which is now used by HPC administrators on a daily basis.

97 MATHEMATICS AND COMPUTING↗

aphBO-2GP-3B: a budgeted asynchronous parallel multi-acquisition functions for constrained Bayesian optimization on high-performing computing architecture

High-fidelity complex engineering simulations are often predictive, but also computationally expensive and often require substantial computational efforts. The mitigation of computational burden is usually enabled through parallelism in high-performance cluster (HPC) architecture. Optimization problems associated with these applications is a challenging problem due to the high computational cost of the high-fidelity simulations. In this paper, an asynchronous parallel constrained Bayesian optimization method is proposed to efficiently solve the computationally expensive simulation-based optimization problems on the HPC platform, with a budgeted computational resource, where the maximum number of simulations is a constant. The advantage of this method are three-fold. Firstly, the efficiency of the Bayesian optimization is improved, where multiple input locations are evaluated parallel in an asynchronous manner to accelerate the optimization convergence with respect to physical runtime. This efficiency feature is further improved so that when each of the inputs is finished, another input is queried without waiting for the whole batch to complete. Second, the proposed method can handle both known and unknown constraints. Third, the proposed method samples several acquisition functions based on their rewards using a modified GP-Hedge scheme. The proposed framework is termed aphBO-2GP-3B, which means asynchronous parallel hedge Bayesian optimization with two Gaussian processes and three batches. The numerical performance of the proposed framework aphBO-2GP-3B is comprehensively benchmarked using 16 numerical examples, compared against other 6 parallel Bayesian optimization variants and 1 parallel Monte Carlo as a baseline, and demonstrated using two real-world high-fidelity expensive industrial applications. The first engineering application is based on finite element analysis (FEA) and the second one is based on computational fluid dynamics (CFD) simulations.

97 MATHEMATICS AND COMPUTING↗

Accelerating Computational Materials Discovery with Machine Learning and Cloud High-Performance Computing: from Large-Scale Screening to Experimental Validation

High-throughput computational materials discovery has promised significant acceleration of the design and discovery of new materials for many years. Despite a surge in interest and activity, the constraints imposed by large-scale computational resources present a significant bottleneck. Furthermore, examples of large-scale computational discovery carried through experimental validation remain scarce, especially for materials with product applicability. In this paper, we demonstrate how this vision became reality by first combining state-of-the-art artificial intelligence (AI) models and traditional physics-based models on cloud high performance computing (HPC) resources to quickly navigate through more than 32 million candidates and predict around half a million potentially stable materials. Focusing on solid-state electrolytes for battery applications, our discovery pipeline further identified 18 promising candidates with new compositions and rediscovered a decade’s worth of collective knowledge in the field as a byproduct. By employing around one thousand virtual machines in the cloud, this process took less than 80 hours. We then synthesized and experimentally characterized the structures and conductivities of our top candidates, the Na x Li 3-x YCl 6 (0.5 ≤ x ≤ 2.5) series, demonstrating the potential of these compounds to serve as solid electrolytes. Additional candidate materials are currently under experimental investigation that could offer more examples of the computational discovery of new phases of Li- and Na-conducting solid electrolytes. We believe this unprecedented approach of synergistically integrating AI models and cloud HPC not only accelerates materials discovery but also showcases the potency of AI-guided experimentation in unlocking transformative scientific breakthroughs with real-world applications.

36 MATERIALS SCIENCE↗

US Department of Energy, Office of Science High Performance Computing Facility Operational Assessment 2021: Oak Ridge Leadership Computing Facility

Oak Ridge National Laboratory’s (ORNL’s) Leadership Computing Facility (OLCF) continues to surpass its operational target goals of supporting users; delivering fast, reliable computational ecosystems; creating innovative solutions for high-performance computing (HPC) needs; contributing to the community to build the next generation HPC workforce, and managing risks, safety, and security associated with operating some of the most powerful computers in the world. The results can be seen in the cutting-edge science conducted by users and the praise from the research community. Calendar year (CY) 2021 saw continued excellence in research supported by the OLCF’s leadership-class computing resources, including Summit (the nation’s most powerful supercomputer), the global scratch file system Alpine, the Scalable Protected Infrastructure (SPI), the Exploratory Visualization Environment for Research in Science and Technology (EVEREST), and the archival mass-storage resource High-Performance Storage System (HPSS). While maintaining access and exceptional user support for Summit, the OLCF continued to make progress on the installation and deployment of Frontier, which will be the nation’s first exascale system when it comes online at the start of CY 2023. Users have already begun running and optimizing scientific codes on Crusher, the OLCF test and development system equipped with Frontier’s architecture. Throughout the year, the OLCF maintained a strong culture of operational excellence, including risk management, workplace safety, and cybersecurity. The OLCF’s rigorous risk management strategy anticipated and mitigated risks, and at this time there are no high-priority operational risks. Similarly, ORNL and the OLCF were committed to operating under the US Department of Energy’s (DOE’s) safety regulations that ensure a safe workplace. Technical staff tracked and monitored existing threats and vulnerabilities within the OLCF while continually developing tools and practices to enhance operations without increasing the facility’s risk. CY 2021 was filled with outstanding results and accomplishments, including a very high rating from users on overall satisfaction for the eighth consecutive year; a tremendous number of node hours delivered to 1,671 researchers on Summit; and the successful delivery of the allocation split of roughly 60%, 20%, and 20% of core-hours offered for the Innovative and Novel Computational Impact on Theory and Experiment (INCITE), Advanced Scientific Computing Research Leadership Computing Challenge (ALCC), and Director’s Discretionary (DD) programs, respectively (Section 2). COVID-19 research remained a focus in 2021, and the ALCC and DD programs allocated over 1 million Summit hours to the COVID-19 High Performance Computing Consortium. These accomplishments, coupled with the high utilization rates (i.e., overall and capability usage), represent the fulfillment of the promise of leadership class machines: efficient facilitation of leadership-class computational applications.

97 MATHEMATICS AND COMPUTING↗

The Installation of Direct Water-Cooling Systems to Reduce Cooling Energy Requirements for High-Performance Computing Centers

A large cluster of High-Performance Computing (HPC) equipment at the Lawrence Livermore National Laboratory in California was retrofitted with an Asetek cooling system. The Asetek system is a hybrid scheme with water-cooled cold plates on high-heat-producing components in the information technology (IT) equipment, and with the remainder of the heat being removed by conventional air-cooling systems. In order to determine energy savings of the Asetek system, data were gathered and analyzed two ways: using top-down statistical models, and bottom-up engineering models. The cluster, “Cabernet”, rejected its heat into a facilities cooling water loop which in turn rejected the heat into the same chilled water system serving the computer-room air handlers (CRAHs) that provided the air-based cooling for the room. Because the “before” and “after” cases both reject their heat into the chilled water system, the only savings is due to reduction in CRAH fan power. The top-down analysis showed a 4% overall energy savings for the data center (power usage effectiveness (PUE) —the ratio of total data center energy to IT energy— dropped from 1.60 to 1.53, lower is better); the bottom-up analysis showed a 3% overall energy savings (PUE from 1.70 to 1.66) and an 11% savings for the Cabernet system by itself (partial PUE of 1.51). Greater savings, on the order of 15-20%, would be possible if the chilled water system was not used for rejecting the heat from the Asetek system. About 37% of the heat from the Cab system was rejected to the cooling water, lower than at other installations.

Earni, Shankar↗

Reproduced Computational Results Report for “Ginkgo: A Modern Linear Operator Algebra Framework for High Performance Computing”

The article titled “Ginkgo: A Modern Linear Operator Algebra Framework for High Performance Computing” by Anzt et al. presents a modern, linear operator centric, C++ library for sparse linear algebra. Experimental results in the article demonstrate that Ginkgo is a flexible and user-friendly framework capable of achieving high-performance on state-of-the-art GPU architectures. In this report, the Ginkgo library is installed and a subset of the experimental results are reproduced. Specifically, the experiment that shows the achieved memory bandwidth of the Ginkgo Krylov linear solvers on NVIDIA A100 and AMD MI100 GPUs is redone and the results are compared to what presented in the published article. Upon completion of the comparison, the published results are deemed reproducible.

97 MATHEMATICS AND COMPUTING↗

Nuclear Science User Facilities High Performance Computing: Provide a Science Gateway for HPC Users

Idaho National Laboratory (INL), supported by the Department of Energy Office of Nuclear Energy (DOE-NE) through the Nuclear Science User Facilities (NSUF), provides direct access to the Barracuda Virtual Reactor and 18 Multiphysics Object-Oriented Simulation Environment (MOOSE) applications via a web-based science gateway developed using Open Ondemand on the INL high performance computing (HPC) systems. This gateway features the computational tools of the Nuclear Computational Resource Center (NCRC) and the computing resources of the INL high performance computing systems. These computational tools are a key foundation of collaboration and innovation in nuclear energy systems research. High performance computing resources and INL staff directly support the mission and objectives of DOE-NE. The Barracuda Virtual Reactor was the first science gateway deployed in Jan 2021 to support NSUF users. In July 2021, the science gateway was expanded to support access to NCRC codes for use across all supported INL HPC systems. The HPC science gateway currently supports 20 total applications. The gateway also includes access to training resources specific to NCRC tools.

99 GENERAL AND MISCELLANEOUS↗

ORNL Campus Sustainability and Decarbonization using Waste Heat Recovery from the Oak Ridge Leadership Computing Facility’s High-Performance Computing Data Center

Heat pumps are a clean and efficient technology that can be powered by renewable electricity to transfer heat using a refrigerant from one place to another by different heat sources, making buildings clean and environmentally friendly. With the support of the ORNL Laboratory Modernization Division, this project explored and evaluated an innovative solution that uses water-water cost-effective midtemperature heat pump (MTHP) technology to leverage the low-grade waste heat from ORNL Frontier and the data center to deliver 85°C hot water, which replaces hot steam generated using natural gas combustion boilers for water heating or space heating in the buildings of ORNL campus. Two scenarios were studied. In the first scenario, which considered the 5600-5700-5800 complex only, Carrier’s commercial 1,000 kW MTHP technology achieves more than 6,640 MWh/year energy savings, an emission reduction of 858 TCO2e/year CO2, and a payback time of 4.85 years. In the second scenario, which considered the 5600-5700-5800 complex and Buildings 5100, 5200, and 5300, the CO2 emission reduction is 1,483 TCO2e/year, the operating cost savings are $0.21 million annually, and the payback time is 3.74 years. Additionally, a comprehensive HP ShowCase Tool was developed for evaluating the optimal solution to improve sustainability and decarbonization of the buildings on the ORNL campus. The tool is an Excel-based tool integrated with VBA (Visual Basic for Applications) coding. The tool includes collected ORNL campus building information and an MTHP library, which comprises collected commercial and ORNL-defined MTHPs. The tool was used to evaluate the sustainability and decarbonization of the ORNL campus. The tool can be widely used or referenced for heat pump solutions and building decarbonization renovation strategies to modernize ORNL facilities and energy use–intensive equipment to enable efficient, sustainable, and resilient operations in the future.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

High Performance Computing Innovation Center Open Source Developer Tools

The High Performance Computing Innovation Center (HPCIC) aims to ease the transition for developers to use open source software provided by the lab. HPCIC Developer Tools is a collection of software, containers, cloud configurations, and associated documentation that make it easy to deploy tutorials or small apps to demonstrate lab-developed software. For example, building a tutorial container that includes lab software and interactive interfaces; a command line or web-based tool that accepts user preferences for the tutorial; supporting tools and software development kits (SDKs) for developer interactions or productivity in different languages embraced by the larger developer community such as Go, Rust, and Python; and automation in version control to support continued update of software and associated resources. These tools are best developed in an open source environment such as GitHub, not only to champion the lab's open source software, but for purposes of branding and demonstrating the lab's leadership in open source. Such an effort that brings in more developers to use and contribute to lab software can further improve the quality of the software, and developer experience at the lab.

Beckingsale, DavidA↗

LLNL Response to the DOE ASCR RFI, "Stewardship of Software for Scientific and High-Performance Computing"

For decades, Lawrence Livermore National Laboratory (LLNL) has been engaged in significant research, development, and support for software to enable scientific computing and, particularly, the use of high performance computing (HPC) in the NNSA mission space. In particular, the move in the mid-1990’s to simulation as a leading component of stockpile stewardship through the ASCI and the successor ASC programs, as well as the need for reliable data acquisition and control software for the National Ignition Facility, have been important drivers in building expertise in production-quality software development at LLNL. LLNL has also been a leader in the DOE SciDAC FASTMath Institute and the DOE Exascale Computing Project (ECP), both of which have striven to make scientific computing software – in particular, the enabling technologies underpinning simulation capabilities – more widely adopted and sustainable. As such, we believe that our experience can inform the broader goal of software stewardship for scientific and high-performance computing. LLNL strongly supports the formation of a new DOE ASCR program element in software stewardship and sustainment. Historically, DOE ASCR has funded applied mathematics and computer science research that has led to the development of important new capabilities and algorithms that are expressed as artifacts in research software. Such frameworks, libraries, and tools have seldom been directly funded to address the important issues of code maintenance, documentation, robustness, and community building. Software engineering and support have typically been done on the side in support of the ASCR-driven research products. DOE funding priorities have been slow to recognize that good software engineering, the kind that ensures research investments have more adoption and longevity, requires significant resources. Based upon our experiences, we have prepared this response to highlight the concerns and issues we believe to be important as DOE ASCR considers its role in scientific software stewardship. We believe that role is important and will require a significant investment of new funding to legitimately support the technologies past and future DOE ASCR investments have and will produce to facilitate their uptake and adoption in the broader scientific computing community. Following a summary of our involvement in scientific software development, the remainder our response is organized around the nine topics specifically identified in the RFI.

97 MATHEMATICS AND COMPUTING↗

Microgrid Integration with High Performance Computing Systems for Microreactor Operation

Multiple nuclear microreactor concepts are currently being developed across several sizes and fuel types with high performance computing (HPC) systems anticipated to be end-users of the power. Nuclear microreactors are small in size, portable, produce less than 10 MW electric, operate autonomously, and have a refueling interval of as many as 10 years. However, their load-follow is also generally limited to 10%/minute or worse whereas the power variance in HPC systems easily exceeds this constraint under normal operations. This study explores an approach that requires no load-follow from the microreactor but integrates the HPC system with a microgrid built from commercial-off-the-shelf components. Three typical HPC architectures are explored in the context of microgrid operation in this study. Components of power quality and transient response are empirically measured for five different HPC load-follow response levels using a self-contained mobile datacenter connected to the microgrid capable of integration with a nuclear microreactor.

microgrids↗

Advanced High-Performance Computational Modeling of the Seismic Response of High-Hazard and/or Nuclear Facilities and Critical Infrastructure at the NNSS

New methods for predicting the amplitude and variability of ground shaking from earthquakes (and explosions) are needed for seismic hazard analysis for buildings, nuclear power plants, and critical infrastructure at the NNSS. We are comparing existing 1-D and new 3-D geophysical methods for estimating the shear-wave velocity structure in the upper 30 meters of the ground surface (Vs30), which plays a major role in ground motion amplification and seismic response of buildings. We evaluate the performance of these methodologies at the U1a Complex at the NNSS and develop simple 1-D and high-resolution 3-D Vs30 models. We then emplace these high-resolution models into a background seismic velocity model. We will collaborate with Lawrence Livermore National Laboratory (LLNL) to conduct numerical modeling of the ground shaking at the NNSS using their high-performance computing technology and state-of-the-art ground motion simulation methodology. The primary work that was completed in FY 2019 was to acquire the seismic systems and familiarize staff at the NNSS with their use. We also worked on developing a collection plan with the Device Assembly Facility (DAF) at the NNSS, but due to time constraints and other ongoing projects at the DAF, we had to use U1a as a backup. We were able to coordinate the seismic survey, and we will complete the collection of seismic data in FY 2020. Additionally, during FY 2019, we completed the geologic framework model (GFM) for the U1a Complex and modeled the Yucca fault. LLNL worked with us through FY 2019 to prepare the data files for their modeling software and tested the software for reliability. In FY 2020 we will develop the end-to-end capability so that any facility could easily be modeled and the expected shaking from a local earthquake understood. The work in FY 2020 will include building fault models from the GFM and finalizing the velocity model analysis. The final simulations will be run for multiple rupture models, and final assessments will demonstrate the seismic hazard at the U1a Complex.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Privacy Preservation from High-Performance Computing to Autonomous Science [Industrial and Governmental Activities]

High-Performance Computing (HPC) and Leadership-Class Supercomputing are driving forces behind scientific advancements, enabling researchers to tackle complex challenges in physics, chemistry, biology, and engineering. These systems power vast simulations and data analyses, fueling discoveries in fields ranging from materials science to climate modeling. However, their use often involves processing sensitive data—such as proprietary industry simulations, biomedical records, and national security computations—posing significant privacy concerns. In conclusion, this issue is amplified in collaborative environments like Department of Energy (DOE) user facilities, where HPC resources are shared across institutions to foster innovation.

Kotevska, Olivera [Oak Ridge National Laboratory (↗

Transforming Energy Through Computational Excellence: NREL HPC Resources for High Performance Computing for Energy Innovation (HPC4EI) Program

NREL hosts computing facilities for the U.S. Department of Energy's Office of Energy Efficiency and Renewable Energy (EERE). In 2024, NREL introduced Kestrel, the 3rd generation, EERE-sponsored supercomputer dedicated to renewable energy and energy efficiency research. Kestrel has already been used for hundreds of research projects by NREL, other national laboratories, and university partners. This includes HPC4EI-sponsored industrial partnerships.

high-performance computing↗

Anomalous And Normal High Performance Computing Datacenter Activities

This code provides annotation and an interface to 10+ hours of video activities in a high performance datacenter with 20+ different types of anomalous activities. The purpose is to enable machine learning for video surveillance systems in high performance computing centers. This is the first code of this type addressing the space of high performance computing datacenters.

Anderson, Matthew↗

Probing Particle Impingement in Boilers Using High-Performance Computing with Parallel CPUs and GPUs

The major goals of the project are to calculate and analyze particle impingement within boilers, quantify effects of particulates in boilers, and predict damage rates of boilers under different cycling modes. Collectively, these initiatives develop insight into existing coal plant challenges using advanced modeling tools, particularly those leveraging high-performance computing resources. High-performance CFD computing forms a central theme in this project that will employ a high degree of coordination and communication between these initiatives to realize a final, rigorously sound, and validated computational capability upon completion. These results will create a holistic, comprehensive, systems-level assessment of damage rates under different cycling modes. Together, these objectives will develop critical insight into damage mechanisms in existing coal plant challenges for accurately and efficiently assessing operating performance in fossil energy power plants.

20 FOSSIL-FUELED POWER PLANTS↗

The COVID-19 High-Performance Computing Consortium

In March of 2020, recognizing the potential of High Performance Computing (HPC) to accelerate understanding and the pace of scientific discovery in the fight to stop COVID-19, the HPC community assembled the largest collection of worldwide HPC resources to enable COVID-19 researchers worldwide to advance their critical efforts. Amazingly, the COVID-19 HPC Consortium was formed within one week through the joint effort of the Office of Science and Technology Policy (OSTP), the U.S. Department of Energy (DOE), the National Science Foundation (NSF), and IBM to create a unique public–private partnership between government, industry, and academic leaders. This article is the Consortium's story–how the Consortium was created, its founding members, what it provides, how it works, and its accomplishments. We will reflect on the lessons learned from the creation and operation of the Consortium and describe how the features of the Consortium could be sustained as a National Strategic Computing Reserve to ensure the nation is prepared for future crises.

97 MATHEMATICS AND COMPUTING↗