Profiling and Optimization [Slides]
Intended for Simulating Physics using Efficient and Effective code Development (SPEED) program. This is a LANL internal lecture series for computational science application developers.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Intended for Simulating Physics using Efficient and Effective code Development (SPEED) program. This is a LANL internal lecture series for computational science application developers.
Future heterogeneous Domain-Specific System-on-Chips (DSSoC) will be extraordinarily complex in terms of processors, memory hierarchies, and interconnection networks.To manage this complexity, architects, system software designers, and application developers need programming technologies that are flexible, accurate, efficient, and productive. These technologies will need to be as independent of any one specific architecture as is practical, because the sheer dimensionality and scale of the complexity will not allow porting and optimizing applications foreach given DSSoC. To address these issues, we are developing Cosmic Castle, a performance portable programming toolchain for streaming applications on heterogeneous architectures. The primary focus of Cosmic Castle is on enabling efficient and performant code generation through the smart compiler and intelligent runtime system. This paper presents the preliminary evaluation of our ongoing work toward Cosmic Castle. Specifically, we detail our code porting efforts and evaluate various benchmarks on the Qualcomm Snapdragon SoC using tools developed through Cosmic Castle.
This report examines the software supply chain security posture of mobile applications developed for consumer whole-house battery and energy-management products. While these applications are not currently integrated with critical infrastructure, their growing role in connected energy domain spaces underscores the importance of understanding the external dependencies, permission structures, and runtime behaviors that could introduce systemic risk; particularly, if adoption expands into more critical environments.
Tritium management is critical for the safety, sustainability, and economics of fusion energy systems, and advanced and reliable modeling tools help accelerate the development of tritium technologies. This paper presents the Tritium Migration Analysis Program, Version 8 (TMAP8), an open-source, MOOSE-based application developed to provide state-of-the-art tritium transport and fuel cycle modeling capabilities. TMAP8 aims to expand the capabilities of previous versions (i.e., TMAP4 and TMAP7) by leveraging modern computational techniques, ensuring high software quality assurance standards (key to building trust), and enabling multispecies, multiscale, and multiphysics simulations for integrated tritium transport modeling in complex geometries. This paper outlines TMAP8’s scope and rigorous development practices, emphasizing its transparency, accessibility, modularity, and reliability. We present the current suite of verification and validation cases based on those from TMAP4, demonstrating TMAP8’s accuracy and reliability against analytical solutions and experimental data. Additionally, the paper showcases TMAP8’s integrated fuel cycle modeling capabilities, highlighting its applicability at various scales and levels. The TMAP8 code and documentation are openly available, promoting collaborative development and widespread adoption within the fusion community. Future work will soon expand TMAP8’s verification and validation suite to include those from TMAP7 and other recent experimental studies for validation.
This paper focuses on the progress being made in water-cooled small modular reactor (SMR) advanced integral effects and separate effects experiments for reactor licensing. SMRs, considered modular in design, are mostly factory-built and then shipped to the reactor site. Of the several types of SMR designs available, water-cooled SMRs are likely to receive regulatory approval faster than others, as most of the technologies involved (e.g., the fuel and coolant technologies) are matured. However, the unique safety systems of SMRs, which depend on specific SMR design features, require integral and separate effects experiments to achieve reactor licensing. Thus, of the many SMR designs being proposed, only a few have successfully undergone licensing and reached the final development and demonstration stage. Many nuclear vendors and newcomer companies are investing millions of dollars to develop integral and separate effects testing facilities for preparing final safety analysis reports to include in licensing applications. Development and analysis of these experimental facilities is costly and takes about four to five years. The unique challenges involved can be reduced when stakeholders synergistically apply lessons learned, knowing the critical role played by advancements in experiments that support the licensing of SMR safety systems. Identification of knowledge/research gaps with the phenomena identification and ranking table (PIRT) and designing experimental facilities focusing on the phenomena of interest (POI) and figures of merit (FOMs) are pivotal to select the critical path to successful design demonstration and licensing application.
This study describes the application development of multiple-input multiple-output radios to provide persistent mobile ad hoc network (MANET) for the Department of Homeland Security. By using Man Portable Unit (MPU5) fifth generation radios (manufactured by Persistent Systems) with the Android Team Awareness Kit (ATAK), an Android smartphone geospatial infrastructure and military situational awareness application, the Remote Sensing Laboratory has developed a MANET connectivity to monitor deployed nuclear/radiological search operation assets.
The control system at Fermilab is undergoing an unprecedented modernization effort. Hundreds of legacy applications originally developed with technology from the early nineties will be replaced with a suite of modern web applications. The selection of an user interface framework is a key technology decision that will impact on all the applications developed over the next decade. The Controls department at Fermilab has decided to use Google’s Dart language and Flutter framework for future application development. In this paper we will discuss the decision process that selected Dart/Flutter and the development of early general purpose control system applications with the framework.
Transparent polycrystalline ceramics are of significant importance for a wide range of scientific and industrial applications. Developing a deeper understanding of their thermodynamic behavior is essential for achieving their maximum output performance in technological applications. This study provides a systematic investigation into the thermal excitation-induced fluorescence kinetics and waveguide characteristics of Magnesium Aluminate Spinel (MgAl 2 O 4 ) single crystals, with a focus on the intricate relationship between structure and properties. High-temperature extreme environments through irradiation with swift heavy ions 645.0 MeV Xe and 352.8 MeV Fe ions were created; the atomic deposition energy threshold (E th ) associated with disorder morphologies was assessed between 0.91 and 0.99 eV atom –1 . The track prediction model was developed to support the theoretical prediction of track formation. Electronic energy loss (E ele ) disrupts the balance of the initial structure through the thermal spike effect, leading to the formation of absorption-related F and F + color centers. These defects enhance photoluminescence in the visible spectrum and result in an effective modulation of the intrinsic bandgap. Moreover, as ion beams penetrate into the material, the uneven damage distribution induces the formation of waveguide structures. In conclusion, these findings provide valuable insights into the fabrication of functional devices through irradiation technologies and the structural changes of MgAl 2 O 4 at high temperatures within extreme environments.
As part of the US Department of Energy’s Exascale Computing Project (ECP), Rice University has been refining its HPCToolkit performance tools to better support measurement and analysis of applications executing on exascale supercomputers. To efficiently collect performance measurements of GPU-accelerated applications, HPCToolkit employs novel non-blocking data structures to communicate performance measurements between tool threads and application threads. To attribute performance information in detail to source lines, loop nests, and inlined call chains, HPCToolkit performs parallel analysis of large CPU and GPU binaries involved in the execution of an exascale application to rapidly recover mappings between machine instructions and source code. To analyze terabytes of performance measurements gathered during executions at exascale, HPCToolkit employs distributed-memory parallelism, multithreading, sparse data structures, and out-of-core streaming analysis algorithms. To support interactive exploration of profiles up to terabytes in size, HPCToolkit’s hpcviewer graphical user interface uses out-of-core methods to visualize performance data. The result of these efforts is that HPCToolkit now supports collection, analysis, and presentation of profiles and traces of GPU-accelerated applications at exascale. These improvements have enabled HPCToolkit to efficiently measure, analyze and explore terabytes of performance data for executions using as many as 64K MPI ranks and 64K GPU tiles on ORNL’s Frontier supercomputer. HPCToolkit’s support for measurement and analysis of GPU-accelerated applications has been employed to study a collection of open-science applications developed as part of ECP. This paper reports on these experiences, which provided insight into opportunities for tuning applications, strengths and weaknesses of HPCToolkit itself, as well as unexpected behaviors in executions at exascale.
We describe the accomplishments jointly achieved by Kitware and Sandia over the fiscal years 2016 through 2020 to benefit the Advanced Scientific Computed (ASC) Advanced Technology Development and Mitigation (ATDM) project. As a result of our collaboration, we have improved the Trilinos and ATDM application developer experience by decreasing the time to build, making it easier to identify and resolve build and test defects, and addressing other issues . We have also reduced the turnaround time for continuous integration (CI) results. For example, the combined improvements likely cut the wall clock time to run automated builds of Trilinos posting to CDash by approximately 6x or more in many cases. We primarily achieved these benefits by contributing changes to the Kitware CMake/CTest/CDash suite of open source software development support tools. As a result, ASC developers can now spend more time improving code and less time chasing bugs. And, without this work, one can argue that the stabilization of Trilinos for the ATDM platforms would not have been feasible which would have had a large negative impact on an important internal FY20 L1 milestone.
An accurate model of a power distribution system is the foundation for model-based applications that ensure efficient and reliable grid operation in an advanced distribution management system (ADMS) environment. However, these models are error-prone and comprehensive model validation is challenging due to lack of standards-based systems, data originating from disparate databases and other sources, and the constantly evolving nature of modern power distribution systems. In this paper, a novel framework for comprehensive model validation is described. The proposed application, the Model Validator, ensures that a model is both consistent and feasible by validating the derivative static and operational network model. A modular architecture for the application has been implemented and integrated with an open-source standards-based platform for ADMS application development, GridAPPS-D, allowing new validation capability to be added with minimal time and effort. The Model Validator application is demonstrated on the IEEE 13-bus, 123-bus, and 8500-node test cases over three validation scenarios.
This Technical Collaboration Project focused on packaging large tow (≥ 10 grams/meter) textile carbon fibers, for which currently there is no suitable method. Prior to this project, Oak Ridge National Laboratory (ORNL) used a crude packaging approach to deliver (approximately) 100-meter packages for evaluation by intermediates producers, composites fabricators and other collaborators. That packaging approach was simply a hoop wrap on a standard 75 mm (3 inch) cardboard tube, with paper interleaved between layers and little or no tension control during spooling. While this approach was very effective for initial fiber development and sample fiber and small composite sample evaluation, these packages were inadequate for larger-scale demonstrations and more extensive applications development. The paper-interleaved packages were inconsistent, wasteful, and too small for commercial utilization; they were designed as a quick and simple means to package material for delivery into prototyping operations and were never expected to satisfy industrial requirements for commercial production. This project aimed to develop a commercially relevant approach for packaging large tow textile carbon fibers for robust delivery into downstream commercial intermediates and composites production operations. The development of robust packaging equipment that delivers packages meeting industrial requirements is essential to commercial implementation. In this project, the team identified packaging requirements, created a design, and prototyped a new packaging concept at ORNL’s Carbon Fiber Technology Facility (CFTF) and resulting packages were tested in production of non-crimp fabric (NCF).
Most parallel scientific programs contain compiler directives (pragmas) such as those from OpenMP (Hermanns in Parallel programming in Fortran 95 using openMP, 2002. School of Aeronautical Engineering, Universidad Politécnica de Madrid, España, 2011), explicit calls to runtime library procedures such as those implementing the Message Passing Interface (MPI) (in A message-passing interface standard version 4.0, 2021. https://www.mpi-forum.org/docs/mpi-4.0/mpi40-report.pdf), or compiler-specific language extensions such as those provided by CUDA (Ruetsch and Fatica in CUDA Fortran for scientists and engineers: best practices for efficient CUDA Fortran programming, Elsevier, 2013). By contrast, the recent Fortran standards empower developers to express parallel algorithms without directly referencing lower-level parallel programming models (Numrich in Parallel programming with co-arrays, CRC Press, 2018, and Curcic in Modern Fortran: building efficient parallel applications, Manning Publications, 2020). Fortran’s parallel features place the language within the Partitioned Global Address Space (PGAS) class of programming models. When writing programs that exploit data parallelism, application developers often find it straightforward to develop custom parallel algorithms. Problems involving complex, heterogeneous, staged calculations, however, pose much greater challenges. Such applications require careful coordination of tasks in a manner that respects dependencies prescribed by a directed acyclic graph. When rolling one’s own solution proves difficult, extending a customizable framework becomes attractive. Further, the paper presents the design, implementation, and use of the Framework for Extensible Asynchronous Task Scheduling (FEATS), which we believe to be the first task scheduling tool written in modern Fortran. We describe the benefits and compromises associated with choosing Fortran as the implementation language, and we propose ways in which future Fortran standards can best support the use case in this paper.
As Internet of Things devices and cloud-based platforms become more mature, Energy Management and Information Systems (EMIS) are increasingly gaining momentum in the building industry. In large commercial buildings, Fault-Detection and Diagnostic (FDD) and energy information systems (EIS) are now established technologies with tens of providers and thousands of deployment sites across North America. The new frontier for the EMIS technology is now represented by control systems that use advanced system optimization (ASO) methods to improve the operations of the HVAC system. Given the complexity of the integration of such systems with the existing building automation systems (BAS) and the higher risk involved with direct control of the HVAC, these systems are still emerging in the market. This paper presents the results of a project in which a start-up company partnered with a research institution to develop a cloud-based software EMIS solution and deployed it in a university campus in California. The software system included advanced sensing, data acquisition, storage and advanced control and analytics applications developed on top of the native BAS. The new platform controls ten buildings on the campus and the FDD and the ASO applications deployed on this platform were able to generate energy savings of up to 35% and 25% in certain buildings for each functionality respectively. Where the platform did not save energy, it improved building service (air quality). Lessons learned include the importance of collaborating with and training the building operators and evaluating whether the legacy system can work reliably with the new technology.
Continuous advancements in scientific and engineering understanding of earthquake phenomena, combined with the associated development of representative physics-based models, is providing a foundation for high-performance, fault-to-structure earthquake simulations. However, regional-scale applications of high-performance models have been challenged by the computational requirements at the resolutions required for engineering risk assessments. The EarthQuake SIMulation (EQSIM) framework, a software application development under the US Department of Energy (DOE) Exascale Computing Project, is focused on overcoming the existing computational barriers and enabling routine regional-scale simulations at resolutions relevant to a breadth of engineered systems. This multidisciplinary software development—drawing upon expertise in geophysics, engineering, applied math and computer science—is preparing the advanced computational workflow necessary to fully exploit the DOE’s exaflop computer platforms coming online in the 2023 to 2024 timeframe. Achievement of the computational performance required for high-resolution regional models containing upward of hundreds of billions to trillions of model grid points requires numerical efficiency in every phase of a regional simulation. This includes run time start-up and regional model generation, effective distribution of the computational workload across thousands of computer nodes, efficient coupling of regional geophysics and local engineering models, and application-tailored highly efficient transfer, storage, and interrogation of very large volumes of simulation data. This article summarizes the most recent advancements and refinements incorporated in the workflow design for the EQSIM integrated fault-to-structure framework, which are based on extensive numerical testing across multiple graphics processing unit (GPU)-accelerated platforms, and demonstrates the computational performance achieved on the world’s first exaflop computer platform through representative regional-scale earthquake simulations for the San Francisco Bay Area in California, USA.
The mapping of computational needs onto execution resources is, by and large, a manual task, and users are frequently guided simply by intuition and past experiences. We present a queueing theory based performance model for streaming data applications that takes steps towards a better understanding of resource mapping decisions, thereby assisting application developers to make good mapping choices. The performance model (and associated cost model) are agnostic to the specific properties of the compute resource and application, simply characterizing them by their achievable data throughput. We illustrate the model with a pair of applications, one chosen from the field of computational biology and the second is a classic machine learning problem.
Most parallel scientific programs contain compiler directives (pragmas) such as those from OpenMP, explicit calls to runtime library procedures such as those implementing the Message Passing Interface (MPI), or compiler-specific language extensions such as those provided by CUDA. By contrast, the recent Fortran standards empower developers to express parallel algorithms without directly referencing lower-level parallel programming models. Fortran’s parallel features place the language within the Partitioned Global Address Space (PGAS) class of programming models. When writing programs that exploit data-parallelism, application developers often find it straightforward to develop custom parallel algorithms. Problems involving complex, heterogeneous, staged calculations, however, pose much greater challenges. Such applications require careful coordination of tasks in a manner that respects dependencies prescribed by a directed acyclic graph. When rolling one’s own solution proves difficult, extending a customizable framework becomes attractive. The paper presents the design, implementation, and use of the Framework for Extensible Asynchronous Task Scheduling (FEATS), which we believe to be the first task-scheduling tool written in modern Fortran. We describe the benefits and compromises associated with choosing Fortran as the implementation language, and we propose ways in which future Fortran standards can best support the use case in this paper.