Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Mathematical software performance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Toward performance-portable PETSc for GPU-based exascale systems

The Portable Extensible Toolkit for Scientific computation (PETSc) library delivers scalable solvers for nonlinear time-dependent differential and algebraic equations and for numerical optimization. The PETSc design for performance portability addresses fundamental GPU accelerator challenges and stresses flexibility and extensibility by separating the programming model used by the application from that used by the library, and it enables application developers to use their preferred programming model, such as Kokkos, RAJA, SYCL, HIP, CUDA, or OpenCL, on upcoming exascale systems. Furthermore, a blueprint for using GPUs from PETSc-based codes is provided, and case studies emphasize the flexibility and high performance achieved on current GPU-based systems.

97 MATHEMATICS AND COMPUTING↗

From NWChem to NWChemEx: Evolving with the Computational Chemistry Landscape

Since the advent of the first computers, chemists have been at the forefront of using computers to understand and solve complex chemical problems. As the hardware and software have evolved, so have the theoretical and computational chemistry methods and algorithms. Parallel computers clearly changed the common computing paradigm in the late 1970s and 80s, and the field has again seen a paradigm shift with the advent of graphical processing units. This review explores the challenges and some of the solutions in transforming software from the terascale to the petascale and now to the upcoming exascale computers. While discussing the field in general, NWChem and its redesign, NWChemEx, will be highlighted as one of the early codesign projects to take advantage of massively parallel computers and emerging software standards to enable large scientific challenges to be tackled.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

2019 Computing Sciences Strategic Plan

Computing has transformed nearly every aspect of scientific inquiry — across disciplines and across scales — from the behavior of subatomic particles to the formation of structures in the early universe, from the assembly of the human genome to the evolution of earth systems. Over the past two decades, computing has become an integral part of how Berkeley Lab is “Bringing Science Solutions to the World.” Advances in computing and mathematics have been key, with new mathematical models of complex physical phenomena, new methods for analyzing complex data, new algorithms for accuracy and scaling and sophisticated software systems that encapsulate these techniques into open, reusable tools. The performance of NERSC computers and the ESnet network have grown by several orders of magnitude, along with our understanding of how to map scientific computations and workflows onto these systems. From research to facility operations, the passion, talent and dedication of the Computing Sciences Area staff has been the cornerstone of our success. The plan outlined in this document describes the next step in a journey to expand the influence and impact of our efforts, building an increasingly connected global enterprise for science that places more powerful instruments in the hands of scientists, along with more powerful methods and tools for modeling, analysis and prediction.

97 MATHEMATICS AND COMPUTING↗

Performance efficient macromolecular mechanics via sub-nanometer shape based coarse graining

Dimensionality reduction via coarse grain modeling is a valuable tool in biomolecular research. For large assemblies, ultra coarse models are often knowledge-based, relying on a priori information to parameterize models thus hindering general predictive capability. Here, we present substantial advances to the shape based coarse graining (SBCG) method, which we refer to as SBCG2. SBCG2 utilizes a revitalized formulation of the topology representing network which makes high-granularity modeling possible, preserving atomistic details that maintain assembly characteristics. Further, we present a method of granularity selection based on charge density Fourier Shell Correlation and have additionally developed a refinement method to optimize, adjust and validate high-granularity models. We demonstrate our approach with the conical HIV-1 capsid and heteromultimeric cofilin-2 bound actin filaments. Our approach is available in the Visual Molecular Dynamics (VMD) software suite, and employs a CHARMM-compatible Hamiltonian that enables high-performance simulation in the GPU-resident NAMD3 molecular dynamics engine.

59 BASIC BIOLOGICAL SCIENCES↗

Streaming Generalized Canonical Polyadic Tensor Decompositions

In this paper, we develop a method which we call OnlineGCP for computing the Generalized Canonical Polyadic (GCP) tensor decomposition of streaming data. GCP differs from traditional canonical polyadic (CP) tensor decompositions as it allows for arbitrary objective functions which the CP model attempts to minimize. This approach can provide better fits and more interpretable models when the observed tensor data is strongly non-Gaussian. In the streaming case, tensor data is gradually observed over time and the algorithm must incrementally update a GCP factorization with limited access to prior data. In this work, we extend the GCP formalism to the streaming context by deriving a GCP optimization problem to be solved as new tensor data is observed, formulate a tunable history term to balance reconstruction of recently observed data with data observed in the past, develop a scalable solution strategy based on segregated solves using stochastic gradient descent methods, describe a software implementation that provides performance and portability to contemporary CPU and GPU architectures and integrates with Matlab for enhanced usability, and demonstrate the utility and performance of the approach and software on several synthetic and real tensor data sets.

97 MATHEMATICS AND COMPUTING↗

Using Likwid and Byfl to Benchmark Hardware Performance [Poster]

Benchmark Study conducted focusing on CPU and program performance analysis. Performance data gathered using 2 different programs and comparisons made based on performance. After comparisons are made, conclusions can be drawn and improvements are made upon hardware and software.

97 MATHEMATICS AND COMPUTING↗

A computational modeling framework for pre-clinical evaluation of cardiac mapping systems

There are a variety of difficulties in evaluating clinical cardiac mapping systems, most notably the inability to record the transmembrane potential throughout the entire heart during patient procedures which prevents the comparison to a relevant “gold standard”. Cardiac mapping systems are comprised of hardware and software elements including sophisticated mathematical algorithms, both of which continue to undergo rapid innovation. The purpose of this study is to develop a computational modeling framework to evaluate the performance of cardiac mapping systems. The framework enables rigorous evaluation of a mapping system’s ability to localize and characterize (i.e., focal or reentrant) arrhythmogenic sources in the heart. The main component of our tool is a library of computer simulations of various dynamic patterns throughout the entire heart in which the type and location of the arrhythmogenic sources are known. Our framework allows for performance evaluation for various electrode configurations, heart geometries, arrhythmias, and electrogram noise levels and involves blind comparison of mapping systems against a “silver standard” comprised of computer simulations in which the precise transmembrane potential patterns throughout the heart are known. A feasibility study was performed using simulations of patterns in the human left atria and three hypothetical virtual catheter electrode arrays. Activation times (AcT) and patterns (AcP) were computed for three virtual electrode arrays: two basket arrays with good and poor contact and one high-resolution grid with uniform spacing. The average root mean squared difference of AcTs of electrograms and those of the nearest endocardial action potential was less than 1 ms and therefore appears to be a poor performance metric. In an effort to standardize performance evaluation of mapping systems a novel performance metric is introduced based on the number of AcPs identified correctly and those considered spurious as well as misclassifications of arrhythmia type; spatial and temporal localization accuracy of correctly identified patterns was also quantified. This approach provides a rigorous quantitative analysis of cardiac mapping system performance. Proof of concept of this computational evaluation framework suggests that it could help safeguard that mapping systems perform as expected as well as provide estimates of system accuracy.

59 BASIC BIOLOGICAL SCIENCES↗

GradientGraph

Under this SBIR Phase II, Reservoir Labs has developed G2 Analytics, a new technology that allows network operators to analyze bottleneck and flow performance with high precision. G2 delivers a new analytical approach and framework to resolve a variety of key problems found in modern communication networks, including: traffic engineering, routing, flow scheduling, network design, capacity planning, resiliency analysis, network slicing, or service level agreement (SLA) management, among others. G2 leverages the bottleneck structure of congestion-controlled communication networks, a recent mathematical discovery by the Reservoir team [RL19b, RL20a, RL20b, RL21a]. Bottleneck structures reveal how perturbations on flows and links propagate through the network, providing an analytical framework to measure (qualitatively and quantitatively) the ripple effects induced as they traverse the network. Leveraging the mathematics of bottleneck structures, Reservoir Labs is developing the G2 technology to provide network operators with a framework to design, optimize and troubleshoot network performance. This delivery includes the G2 software stack.

Yellamraju, Sruthi↗

Performance Improvements for the Griffin Transport Solvers

Griffin is a Multiphysics Object-Oriented Simulation Environment based reactor multiphysics analysis application jointly developed by Idaho National Laboratory and Argonne National Laboratory. Griffin includes a variety of deterministic radiation transport solvers for fixed source, k-eigenvalue, adjoint, and subcritical multiplication, as well as transient solvers for point-kinetics, improved quasi-static, and spatial dynamics. A code assessment performed in FY-20 identified two significant issues with the transport solvers in Griffin: first, the primary heterogeneous SN (discrete ordinates) transport solver based on continuous finite element methods required significant mesh refinement and higher memory usage compared to solvers based on the method of characteristic for equivalent accuracy. Second, the homogeneous PN (spherical harmonics expansion) transport solver did not adequately support polynomial refinement, which is a feature usually required for problems with spatial homogenization and pronounced streaming, typical in fast or gas-cooled reactor systems. To address the first issue, the development effort focused on the more promising discontinuous finite element method (DFEM)-based SN transport solver in Griffin. The addition of an asynchronous parallel transport sweeper and coarse mesh finite difference (CMFD) acceleration have rendered a superior heterogeneous SN transport capability for multiphysics problems that requires far less computing resources in terms of both CPU time and memory usage. This is demonstrated with typical thermal- and fast-spectrum reactor benchmark problems, including 2D Transient Reactor Test, 3D Advanced Burner Test Reactor (ABTR), and 2D and 3D Empire microreactor. For the second issue, the development effort focused on a new transport solver based on the hybrid finite element PN method (HFEM-PN), equivalent to the variational nodal method, as well as a new diffusion solver based on HFEM-Diffusion. This solver is intended for homogenized domains with multiphysics coupling (i.e., supports mesh displacement, seamless temperature feedback, etc.). Initial calculations with the HFEM-Diffusion implementation show very good parallel efficiency for the residual evaluations with the 2D ABTR benchmark. A future development effort will be centered on further improvements to the CMFD, HFEM-PN, and DFEM diffusion solvers to ensure Griffin meets performance and software quality assurance requirements for advanced reactor design and analysis.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

A general approach to seismic inversion with automatic differentiation

Imaging Earth structure or seismic sources from seismic data involves minimizing a target misfit function, and is commonly solved through gradient-based optimization. The adjoint-state method has been developed to compute the gradient efficiently; however, its implementation can be time-consuming and difficult. We develop a general seismic inversion framework to calculate gradients using reverse-mode automatic differentiation. The central idea is that adjoint-state methods and reverse-mode automatic differentiation are mathematically equivalent. Here, the mapping between numerical PDE simulation and deep learning allows us to build a seismic inverse modeling library, ADSeismic, based on deep learning frameworks, which supports high performance reverse-mode automatic differentiation on CPUs and GPUs. We demonstrate the performance of ADSeismic on inverse problems related to velocity model estimation, rupture imaging, earthquake location, and source time function retrieval. ADSeismic has the potential to solve a wide variety of inverse modeling applications within a unified framework.

58 GEOSCIENCES↗

ExTreeM: Scalable Augmented Merge Tree Computation via Extremum Graphs

Over the last decade merge trees have been proven to support a plethora of visualization and analysis tasks since they effectively abstract complex datasets. Here, this paper describes the ExTreeM-Algorithm: A scalable algorithm for the computation of merge trees via extremum graphs. The core idea of ExTreeM is to first derive the extremum graph G of an input scalar field f defined on a cell complex K, and subsequently compute the unaugmented merge tree of f on G instead of K; which are equivalent. Any merge tree algorithm can be carried out significantly faster on G, since K in general contains substantially more cells than G. To further speed up computation, ExTreeM includes a tailored procedure to derive merge trees of extremum graphs. The computation of the fully augmented merge tree, i.e., a merge tree domain segmentation of K, can then be performed in an optional post-processing step. All steps of ExTreeM consist of procedures with high parallel efficiency, and we provide a formal proof of its correctness. Our experiments, performed on publicly available datasets, report a speedup of up to one order of magnitude over the state-of-the-art algorithms included in the TTK and VTK-m software libraries, while also requiring significantly less memory and exhibiting excellent scaling behavior.

97 MATHEMATICS AND COMPUTING↗

On the Efficient Evaluation of the Exchange Correlation Potential on Graphics Processing Unit Clusters

The predominance of Kohn–Sham density functional theory (KS-DFT) for the theoretical treatment of large experimentally relevant systems in molecular chemistry and materials science relies primarily on the existence of efficient software implementations which are capable of leveraging the latest advances in modern high-performance computing (HPC). With recent trends in HPC leading toward increasing reliance on heterogeneous accelerator-based architectures such as graphics processing units (GPU), existing code bases must embrace these architectural advances to maintain the high levels of performance that have come to be expected for these methods. In this work, we purpose a three-level parallelism scheme for the distributed numerical integration of the exchange-correlation (XC) potential in the Gaussian basis set discretization of the Kohn–Sham equations on large computing clusters consisting of multiple GPUs per compute node. In addition, we purpose and demonstrate the efficacy of the use of batched kernels, including batched level-3 BLAS operations, in achieving high levels of performance on the GPU. We demonstrate the performance and scalability of the implementation of the purposed method in the NWChemEx software package by comparing to the existing scalable CPU XC integration in NWChem.

97 MATHEMATICS AND COMPUTING↗

An extension to V ORO ++ for multithreaded computation of Voronoi cells

V ORO ++ is a software library written in C++ for computing the Voronoi tessellation, a technique in computational geometry that is widely used for analyzing systems of particles. V ORO ++ was released in 2009 and is based on computing the Voronoi cell for each particle individually. Here, we take advantage of modern computer hardware, and extend the original serial version to allow for multithreaded computation of Voronoi cells via the OpenMP application programming interface. We test the performance of the code, and demonstrate that it can achieve parallel efficiencies greater than 95% in many cases. Further, the multithreaded extension follows standard OpenMP programming paradigms, allowing it to be incorporated into other programs. We provide an example of this using the VoroTop software library, performing a multithreaded Voronoi cell topology analysis of up to 102.4 million particles.

97 MATHEMATICS AND COMPUTING↗

Turbo FRMAC Implemetation of IAEA Radiological Assessment Methodologies for Nuclear and Radiological Emergencies.

This report documents the findings of an assessment of the Turbo FRMAC software's ability to implement International Atomic Energy Agency (IAEA) guidance for calculating Operational Intervention Levels (OIL) 1 & 2 for nuclear and radiological emergencies. The IAEA OIL and U.S. Federal Radiological Monitoring and Assessment Center (FRMAC) Derived Response Level methodology and implementation in respective tools were compared, as demonstrated through benchmarking activities for a nuclear power plant source term and potential radionuclides of concern for radiological dispersal devices. This comparison revealed some shortcomings in Turbo FRMACs ability to perform IAEA OIL calculations and resulted in recommended software modifications to be considered for future development.

61 RADIATION PROTECTION AND DOSIMETRY↗

PyGDH: Python Grid Discretization Helper

Mathematical models expressed in the form of discretized equations play an important role inmany scientific disciplines. In our experience, few domain scientists have sufficient backgroundin numerical computing (or the time required to acquire such a background) to use manyflexible and powerful but complex open source packages, such as FEniCS (Alnæs et al., 2015)and OpenFOAM (The OpenFOAM Foundation Ltd, n.d.). Many user-friendly open sourcepackages, such as FiPy (J. E. Guyer & Warren, 2009), and many commercial packages, suchas COMSOL Multiphysics (COMSOL AB, n.d.) and Simcenter STAR-CCM+ (Siemens DigitalIndustries Software, n.d.), provide limited flexibility in the equations that users can express.Additionally, the use of commercial packages, which by nature do not perform calculationstransparently, can hinder reproducibility, which is vital to the scientific process. PyGDH(“pigged”) is a Python 2 / Python 3 (Python Software Foundation, 1991–2020) package thatis meant to be accessible to scientists who might not be specialists in scientific computing,while approaching the level of flexibility associated with writing dedicated programs tailoredto solving specific problems. The PyGDH User’s Guide provides detailed instructions forcreating numerical models, including a brief introduction to necessary command line andPython (Python Software Foundation, 1991–2020) skills, and discussions of discretization andvalidation. Note that PyGDH emphasizes flexibility and simplicity over performance, and wasnot designed for high-performance applications or models describing complex spatial domains.

97 MATHEMATICS AND COMPUTING↗