Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Mathematical software performance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Priority research directions for in situ data management: Enabling scientific discovery from diverse data sources

In January 2019, the US Department of Energy, Office of Science program in Advanced Scientific Computing Research, convened a workshop to identify priority research directions (PRDs) for in situ data management (ISDM). A fundamental finding of this workshop is that the methodologies used to manage data among a variety of tasks in situ can be used to facilitate scientific discovery from many different data sources—simulation, experiment, and sensors, for example—and that being able to do so at numerous computing scales will benefit real-time decision-making, design optimization, and data-driven scientific discovery. This article describes six PRDs identified by the workshop, which highlight the components and capabilities needed for ISDM to be successful for a wide variety of applications—making ISDM capabilities more pervasive, controllable, composable, and transparent, with a focus on greater coordination with the software stack and a diversity of fundamentally new data algorithms.

97 MATHEMATICS AND COMPUTING↗

Software Quality Assurance Plan ANSYS LSDYNA Version 2023R1

ANSYS Inc. develops and markets engineering simulation software and services used in the aerospace, automotive, manufacturing, electronics, biomedical, energy, defense, and many other industries. ANSYS is dedicated to engineering simulation and is the world’s leading software provider. ANSYS was founded in 1970 and is headquartered in Canonsburg, Pennsylvania. ANSYS provides an engineering analysis tool combining structural, thermal, computational fluid dynamics, acoustic and electromagnetic simulation capabilities. ANSYS LS-DYNA is the most used explicit simulation program capable of simulating the response of materials to short periods of severe loading. Its many elements, contact formulations, material models, and other controls can be used to simulate complex models with control over all the details of the problem. ANSYS LS-DYNA has a vast array of capabilities to simulate extreme deformation problems using its explicit solver. Engineers can tackle simulations involving material failure and look at how the failure progresses through a part or through a system. Models with large amounts of parts or surfaces interacting with each other are also easily handled, and the interactions and load passing between complex behaviors are modeled accurately. Using computers with higher numbers of CPU cores can drastically reduce solution times. In addition, many consulting firms and hundreds of universities use ANSYS for analysis, research, and educational purposes. ANSYS is recognized worldwide as one of the most widely used and capable programs of its type. ANSYS has successfully passed over 100 customer quality system audits against American Society of Mechanical Engineers (ASME) NQA-1 and 10 CFR Part 50, Appendix B, since the company was founded, over 60 of which have been since 1997. ANSYS has successfully passed over 100 International Organization for Standardization (ISO) 9001 assessments. ANSYS design analysis software is the first created within a quality system with ISO 9001 certification, which is the internationally accepted quality standard. Product development, testing, maintenance, and support processes also meet the US Nuclear Regulatory Commission’s (NRC’s) quality requirements, as they have for nearly four decades. ANSYS staff perform more than 60,000 software verification tests before releasing each new product. ASME NQA-1-2012 (Subpart 2.7 is specific to software) is the industry- and NRC-accepted approach (consensus standard) for meeting 10 CFR Part 50, Appendix B, requirements.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Pre-exascale accelerated application development: The ORNL Summit experience

High-performance computing (HPC) increasingly relies on heterogeneous architectures to achieve higher performance. In the Oak Ridge Leadership Facility (OLCF), Oak Ridge, TN, USA, this trend continues as its latest supercomputer, Summit, entered production in early 2019. The combination of IBM POWER9 CPU and NVIDIA V100 GPU, along with a fast NVLink2 interconnect and other latest technologies, pushes system performance to a new height and breaks the exascale barrier by certain measures. Due to Summit's powerful GPUs and much higher GPU–CPU ratio, offloading to accelerators becomes a requirement for any application, which intends to effectively use the system. To facilitate navigating a complex landscape of competing heterogeneous architectures, a collection of applications from a wide spectrum of scientific domains is selected for early adoption on Summit. In this article, the experience and lessons learned are summarized, in the hope of providing useful guidance to address new programming challenges, such as scalability, performance portability, and software maintainability, for future application development efforts on heterogeneous HPC systems.

97 MATHEMATICS AND COMPUTING↗

PHASM: A Toolkit for Creating AI Surrogate Models within Legacy Codebases

PHASM (“Parallel Hardware viA Surrogate Models”) is a software toolkit for creating AI-based surrogate models of scientific code. AI-based surrogate models are widely used for creating fast and inverse simulations. PHASM anticipates an additional future use case: adapting legacy code to modern hardware. While data centers are investing in heterogeneous hardware such as GPUs and FPGAs, many established scientific codebases remain unable to take advantage of the hardware’s higher parallelism without undergoing a costly rewrite. An alternative is to train a AI-based surrogate model to mimic computationally intensive functions in the code, and run the surrogate instead. PHASM formalizes a development lifecycle for such surrogate models, including discovering functions amenable to replacement with a surrogate model, predicting the resulting performance, identifying the function’s space of inputs and outputs, binding the model to the code, and managing model versions. A suite of software tools for facilitating these steps was written and validated against a set of model problems.

97 MATHEMATICS AND COMPUTING↗

UltraSep Acoustic Separation Platform

UltraSep is an intelligent ultrasonic separation platform that transforms solid–liquid separation through real-time eigenfrequency resonance locking and ultra-low power energy optimization. By dynamically matching ultrasonic output to system resonance while maximizing particulate removal per unit of applied energy, UltraSep replaces centrifugation and fouling-prone filtration with precision-controlled acoustic forces that significantly reduce power consumption, mechanical complexity, and operating cost while improving recovery performance. This integrated platform unites patented resonance-based acoustic control and energy-per-removal optimization with chemistry-enhanced separation and proprietary system software into a scalable, high-impact commercial technology.

42 ENGINEERING↗

Stabilized bases for high-order, interpolation semi-Lagrangian, element-based tracer transport

In a computational fluid model of the atmosphere, the advective transport of trace species, or tracers, can be computationally expensive. For efficiency, models often use semi-Lagrangian advection methods. High-order interpolation semi-Lagrangian (ISL) methods, in particular, can be extremely efficient, if the problem of property preservation specific to them can be addressed. Atmosphere models often use geometrically and logically nonuniform grids for efficiency and, as a result, element-based discretizations. Such grids and discretizations make stability a particular problem for ISL methods. Generally, high-order, element-based ISL methods that use the natural polynomial interpolant associated with a nodal finite-element discretization are unstable. Here, we derive new bases having order of accuracy up to nine, with positive nodal weights, that stabilize the element-based ISL method. We use these bases to construct the linear advection operator in the property-preserving Interpolation Semi-Lagrangian Element-based Transport (Islet) method. Then we discuss key software implementation details. Finally, we show performance results for the Energy Exascale Earth System Model's atmosphere dynamical core, comparing the original and new transport methods. These simulations used up to 27,600 Graphical Processing Units (GPU) on the Oak Ridge Leadership Computing Facility's Summit supercomputer.

97 MATHEMATICS AND COMPUTING↗

TEMPI: An Interposed MPI Library with Canonical Representation of MPI Datatypes [Poster]

TEMPI provides a transparent non-contiguous data-handling layer compatible with various MPIs. MPI Datatypes are a powerful abstraction for allowing an MPI implementation to operate on non-contiguous data. CUDA-aware MPI implementations must also manage transfer of such data between the host system and GPU. The non-unique and recursive nature of MPI datatypes mean that providing fast GPU handling is a challenge. The same noncontiguous pattern may be described in a variety of ways, all of which should be treated equivalently by an implementation. This work introduces a novel technique to do this for strided datatypes. Methods for transferring non-contiguous data between the CPU and GPU depends on the properties of the data layout. This work shows that a simple performance model can accurately select the fastest method. Unfortunately, the combination of MPI software and system hardware available may not provide sufficient performance. The contributions of this work are deployed on OLCF Summit through an interposer library which does not require privileged access to the system to use

97 MATHEMATICS AND COMPUTING↗

Modifications to Sandia's MDT and WNTR tools for ERMA

ERMA is leveraging Sandia’s Microgrid Design Toolkit (MDT) [1] and adding significant new features to it. Development of the MDT was primarily funded by the Department of Energy, Office of Electricity Microgrid Program with some significant support coming from the U.S. Marine Corps. The MDT is a software program that runs on a Microsoft Windows PC. It is an amalgamation of several other software capabilities developed at Sandia and subsequently specialized for the purpose of microgrid design. The software capabilities include the Technology Management Optimization (TMO) application for optimal trade-space exploration, the Microgrid Performance and Reliability Model (PRM) for simulation of microgrid operations, and the Microgrid Sizing Capability (MSC) for preliminary sizing studies of distributed energy resources in a microgrid.

97 MATHEMATICS AND COMPUTING↗

Laboratory Instrument Software Controlled Spread Spectrum Time Domain Reflectometry for Electrical Cable Testing

This research discusses development of a software-controlled laboratory instrument based spread spectrum time domain reflectometry system (SSTDR). This constitutes one task within PNNL’s Light Water Sustainability Program (LWRS) whose mission includes advancing nondestructive examination (NDE) techniques for off-line and on-line in-situ cable condition monitoring. In 2022, PNNL evaluated SSTDR for detection and characterization of a number of cable anomalies (Glass et al. 2022). The review included comparison of SSTDR to Frequency Domain Reflectometry (FDR) techniques which have enjoyed encouraging feedback and are starting to be used in nuclear power plants for periodic cable condition monitoring of cable systems as part of the plant’s overall cable aging management program. The FDR test introduces a broad-band chirp onto the cable at the cable end then listens for any reflection from a change of impedance along the cable caused by a damaged conductor or insulation, splices, contact with moisture, or other cable anomalies. The signal is captured in the frequency domain then transformed back to the time domain using an inverse Fourier transform (IFT). Based on the velocity of propagation, the impedance response signal is plotted against distance along the cable. Peak locations along the X-axis indicate the distance along the cable where a portion of the signal has been reflected back to the instrument as a result of a cable anomaly. The FDR test is considered the gold standard of reflectometry however it does require the cable to be de-energized to perform the test. The LIVEWIRE commercial SSTDR produces a similar plot to the FDR however all processing is in the time domain. A pseudo-random noise code (PN code) is input onto the cable conductor and the instrument listens for any reflected response from cable anomalies. The SSTDR processes the signal as an autocorrelation comparing the input PN code to any reflected signal detected. The autocorrelation analysis for thermal aging, water and water ingress detection, ground fault and phase-to-phase fault detection at various locations along the cable and with the cable attached and detached from a motor load, and on both energized and un-energized conditions were performed. These results were contrasted to Frequency Domain Reflectometry (FDR) measurements of the un-energized cable. Results were encouraging but indicated more work was warranted – particularly with the SSTDR, it seemed that the insulation damage would likely be better evaluated with multiple bandwidth cable tests particularly including larger bandwidths than were possible with the current commercial instrument. The commercial instrument’s bandwidth was set at 6, 12, 24, and 48MHz but note that SSTDR and FDR definitions of bandwidth trend similarly but are not the same. The FDR response could be more broadly adjusted, and the bandwidth of 100 to 500 MHz produced the best responses. FDR responses to anomalies were clearer than SSTDR responses and indications were that a broader bandwidth SSTDR may lead to improved SSTDR detection capability. This project used a laboratory instrument based SSTDR (primarily using an Arbitrary Waveform Generator (AWG) and a digital oscilloscope plus Python in-house software) that allowed software adjustment of the SSTDR bandwidth, window functions applied to the exciting Pseudo-random Noise (PN) code plus and other aspects of the SSTDR signal processing. Hereafter, this will be referred to as the PNNL SSTDR. Evaluating specific performance of the PNNL SSTDR is left to a separate report. This report documents hardware and software development to produce the SSTDR cable test system.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Mesoflow: An Open-Source Reacting Flow Solver for Catalysis at Mesoscale

We present the capabilities and software performance metrics of our open-source continuum solver for catalysis, Mesoflow, developed specifically for modeling transport and chemistry at the mesoscale. Our solver utilizes Cartesian block-structured adaptive mesh refinement to resolve complex catalyst surface morphologies directly obtained from X-ray tomography data. An immersed boundary based formulation enables rapid representation of complex geometries prevalent in most mesoporous catalyst interfaces. The solver is developed on top of open-source performance portable library, AMReX, providing parallel execution capabilities on current and upcoming high-performance-computing (HPC) architectures. Our flexible software framework enables integration of complex chemical mechanisms at heterogenous interfaces and time-split algorithms for circumventing highly disparate reaction and flow time-scales. Our current studies indicate a ten-fold performance gain by using graphics-processing-units (GPUs) compared to a single processor for representative problem sizes (2 million cell mesh). We will also present a brief introduction on how to build and use this software for application problems pertaining to catalytic upgrading and gas transport within porous catalyst particles.

adaptive meshing↗

VTK-m User's Guide (V.1.6)

High-performance computing relies on ever finer threading. Advances in processor technology include ever greater numbers of cores, hyperthreading, accelerators with integrated blocks of cores, and special vectorized instructions, all of which require more software parallelism to achieve peak performance. Traditional visualization solutions cannot support this extreme level of concurrency. Extreme scale systems require a new programming model and a fundamental change in how we design algorithms. To address these issues we created VTK-m: the visualization toolkit for multi-/many-core architectures.

97 MATHEMATICS AND COMPUTING↗

PFLOTRAN Development FY2022

The Spent Fuel & Waste Science and Technology (SFWST) Campaign of the U.S. Department of Energy (DOE) Office of Nuclear Energy (NE), Office of Spent Fuel & Waste Disposition (SFWD) is conducting research and development (R&D) on geologic disposal of spent nuclear fuel (SNF) and high-level nuclear waste (HLW). A high priority for SFWST disposal R&D is to develop a disposal system modeling and analysis capability for evaluating disposal system performance for nuclear waste in geologic media. This report describes fiscal year (FY) 2022 accomplishments by the PFLOTRAN Development group of the SFWST Campaign. The mission of this group is to develop a geologic disposal system modeling capability for nuclear waste that can be used to probabilistically assess the performance of generic disposal concepts. In FY 2022, the PFLOTRAN development team made several advancements to our software infrastructure, code performance, and process modeling capabilities.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Experimental measurements and mathematical modeling of cold plate for aviation thermal management

Herein, this study, which has been motivated by the recent applications of the cold plate device in aviation thermal management, reports on physics-based mathematical models derived from the conservation laws of mass, momentum, and energy, and empiricism-based models. One of the objectives of the present work is to report on an elaborate and successful experimental work carried out on an additively manufactured device for the purpose of rigorously validating the numerical predictions. The excellent agreement between the numerical predictions and measured performance provides much needed confidence in the implementation, in the software package, of the offset-strip fin passage correlations, as well as in the software implementation of user-defined wavy fin correlations for aerospace heat exchangers and cold plates operating with ram air at Reynolds numbers below 8000. The contributions of this work can also be found in the development of a new and accurate thermal-hydraulic analysis procedure, referred to in this paper as plate-fin analogy. Results from this procedure are compared with those from thermal resistance network. The comparative study in this paper of the bulk and discrete enthalpy flux method is also new, as is the relative assessment of four off-set strip fin thermal-hydraulic models.

42 ENGINEERING↗

PIPER: Performance Insight for Programmers and Exascale Runtimes (Final Technical Report)

This project concentrated on the development of novel performance tools and analysis techniques for large scale parallel systems and applications, eventually targeting exascale platforms. Activities conducted included work to novel root cause detection approaches and an auto-tuning framework for parallel systems. The work has led to several software solutions, which are available as open source.

97 MATHEMATICS AND COMPUTING↗

Intelligent Partitioning based Fully Parallel AC Security-Constrained Optimal Power Flow

Today’s power grid is becoming more diverse and integrated with high-level distributed energy resources and smart control technologies that is creating a new set of grid management challenges in terms of large-scale, nonlinear, and non-convex problem modeling, complex and time-consuming computation, as well as difficult uncertainty handling. This project focused on solving a challenging multi-period security-constrained generation scheduling problem, which is of great importance for maximizing the social welfare of real-time dispatch, day-ahead market, as well as weekly planning of power systems. Our developed software explored parallel optimization algorithms for complex and realistic power system models, and develop fast, efficient, and robust grid optimization solutions on the high-performance computing platform that will enable increased grid economics, flexibility, resilience, as well as energy security in the United States.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Performance on HPC Platforms Is Possible Without C++

Computing at large scales has become extremely challenging due to increasing heterogeneity in both hardware and software. More and more scientific workflows must tackle a range of scales and use machine learning and AI intertwined with more traditional numerical modeling methods, placing more demands on computational platforms. These constraints indicate a need to fundamentally rethink the way computational science is done and the tools that are needed to enable these complex workflows. The current set of C++-based solutions may not suffice, and relying exclusively upon C++ may not be the best option, especially because several newer languages and boutique solutions offer more robust design features to tackle the challenges of heterogeneity. In June 2023, we held a mini symposium that explored the use of newer languages and heterogeneity solutions that are not tied to C++ and that offer options beyond template metaprogramming and Parallel. For for performance and portability. In conclusion, we describe some of the presentations and discussion from the mini symposium in this article.

97 MATHEMATICS AND COMPUTING↗

MAPPRAISER: A massively parallel map-making framework for multi-kilo pixel CMB experiments

Forthcoming cosmic microwave background (CMB) polarized anisotropy experiments have the potential to revolutionize our understanding of the Universe and fundamental physics. The sought-after, tale-telling signatures will be however distributed over voluminous data sets which these experiments will collect. These data sets will need to be efficiently processed and unwanted contributions due to astrophysical, environmental, and instrumental effects characterized and efficiently mitigated in order to uncover the signatures. This poses a significant challenge to data analysis methods, techniques, and software tools which will not only have to be able to cope with huge volumes of data but to do so with unprecedented precision driven by the demanding science goals posed for the new experiments. A keystone of efficient CMB data analysis is solvers of very large linear systems of equations. Such systems appear in very diverse contexts throughout CMB data analysis pipelines, however they typically display similar algebraic structures and can therefore be solved using similar numerical techniques. Linear systems arising in the so-called map-making problem are one of the most prominent and common ones. In this work we present a massively parallel, flexible and extensible framework, comprised of a numerical library, MIDAPACK, and a high level code, MAPPRAISER, which provide tools for solving efficiently such systems. Here, the framework implements iterative solvers based on conjugate gradient techniques: enlarged and preconditioned using different preconditioners. We demonstrate the framework on simulated examples reflecting basic characteristics of the forthcoming data sets issued by ground-based and satellite-borne instruments, executing it on as many as 16,384 compute cores. The software is developed as an open source project freely available to the community at: https://github.com/B3Dcmb/midapack.

79 ASTRONOMY AND ASTROPHYSICS↗

PtychoShelves , a versatile high-level framework for high-performance analysis of ptychographic data

Over the past decade, ptychography has been proven to be a robust tool for non-destructive high-resolution quantitative electron, X-ray and optical microscopy. It allows for quantitative reconstruction of the specimen's transmissivity, as well as recovery of the illuminating wavefront. Additionally, various algorithms have been developed to account for systematic errors and improved convergence. With fast ptychographic microscopes and more advanced algorithms, both the complexity of the reconstruction task and the data volume increase significantly. PtychoShelves is a software package which combines high-level modularity for easy and fast changes to the data-processing pipeline, and high-performance computing on CPUs and GPUs.

97 MATHEMATICS AND COMPUTING↗