Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Structures and Algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Cabana: A Performance Portable Library for Particle-Based Simulations

Particle-based simulations are ubiquitous throughout many fields of computational science and engineering, spanning the atomistic level with molecular dynamics (MD), to mesoscale particle-in-cell (PIC) simulations for solid mechanics, device-scale modeling with PIC methods for plasma physics, and massive N-body cosmology simulations of galaxy structures, with many other methods in between (Hockney & Eastwood, 1989). While these methods use particles to represent significantly different entities with completely different physical models, many low-level details are shared including performant algorithms for short- and/or long-range particle interactions, multi-node particle communication patterns, and other data management tasks such as particle sorting and neighbor list construction. Cabana is a performance portable library for particle-based simulations, developed as part of the Co-Design Center for Particle Applications (CoPA) within the Exascale Computing Project (ECP) (Alexander et al., 2020). The CoPA project and its full development scope, including ECP partner applications, algorithm development, and similar software libraries for quantum MD, is described in (Mniszewski et al., 2021). Cabana uses the Kokkos library for on-node parallelism (Edwards et al., 2014; Trott et al., 2022), enabling simulation on multi-core CPU and GPU architectures, and MPI for GPU-aware, multi-node communication. Cabana provides particle simulation capabilities on almost all current Kokkos backends, including serial execution, OpenMP (including OpenMP-Target for GPUs), CUDA (NVIDIA GPUs), HIP (AMD GPUs), and SYCL (Intel GPUs), providing a clear path for the coming generation of accelerator-based exascale hardware. Cabana builds on Kokkos by providing new particle data structures and particle algorithms resulting in a similar execution policy-based, node-level programming model that is intended to be used in addition to the core Kokkos library within an application. Cabana is designed as an application and physics agnostic, but particle-specific toolkit which can either be used to generate a new application, or to be used as needed in existing applications at various levels of invasiveness including through interfaces that wrap user memory in existing data structures.

97 MATHEMATICS AND COMPUTING↗

Towards Superior Software Portability with SHAD and HPX C++ Libraries

As hardware architectures and software stacks complexity grows, development productivity, performance and software portability, quickly evolve from desirable features to actual needs. SHAD, the Scalable High-performance Algorithms and Data-structures C++ library is designed to mitigate these issues: it provides general purpose building blocks as well as high-level custom utilities, and offers a shared-memory programming abstraction which facilitates the programming of complex systems, scaling up to High Performance Computing clusters. SHAD’s portability is achieved through an abstract runtime interface, which decouples the upper layers of the library and hides the low level details of the underlying architecture. This layer enables SHAD to interface with different runtime/threading systems, e.g. Intel TBB and Global Memory and Threading (GMT). However, current backends targeting distributed systems, rely on a centralized controller which may possibly limit scalability up to hundreds of nodes and creates a network hot spot due to all to one communication for synchronization, and possibly resulting in degraded performance at high process counts. In this research, we explore HPX, the C++ standard library for parallelism and concurrency, as an additional backend in support of the SHAD library, and present the methodologies in support of local and remote task executions in SHAD with respect to HPX. Finally, we evaluate the proposed system by comparing against existing backends of SHAD and analyzing their performance on C++ Standard Template Library algorithms.

Wu, Nanmiao↗

Implementing a neural network interatomic model with performance portability for emerging exascale architectures

The two main thrusts of computational science are increasingly accurate predictions and faster calculations; to this end, the zeitgeist in molecular dynamics (MD) simulations is pursuing machine learned and data driven interatomic models, e.g. neural network potentials, and novel hardware architectures, e.g. GPUs. Current implementations of neural network potentials are orders of magnitude slower than traditional interatomic models and while looming exascale computing offers the ability to run large, accurate simulations with these models, achieving portable performance for MD with new and varied exascale hardware requires rethinking traditional algorithms, using novel data structures, and library solutions. We re-implement a neural network interatomic model in CabanaMD, an MD proxy application, built on libraries developed for performance portability. Our implementation shows significantly improved thread scaling in this complex kernel as compared to a current LAMMPS implementation, across both strong and weak scaling. Our single-source solution enables simulations up to 20 million atoms on a single CPU node and 4 million atoms with improved performance on a single GPU. Furthermore, we also explore parallelism and data layout choices (using flexible data structures called AoSoAs) and their effect on performance, seeing up to ~50% and ~5% improvements in performance on a GPU by choosing the right level of parallelism and data layout respectively.

97 MATHEMATICS AND COMPUTING↗

MemFriend: Understanding Memory Performance with Spatial-Temporal Affinity

In HPC applications, memory access behavior is one of the main factors affecting performance. Improving an application’s memory access behavior involves optimizing data layout and/or restructuring code, and requires studying spatial-temporal data locality. Existing data locality analyses focus on single-location metrics and are restricted to evaluating temporal locality. We introduce spatial-temporal affinity metrics that quantify temporal access proximity, forward access correlation, and nearby access correlation between pairs of memory locations. We describe methods for distinguishing between potential vs. realized affinity and for reasoning about affinity at multiple resolutions (3D, 2D, 1D). Finally, we construct spatial-temporal affinity signatures that classify memory behavior and that be used to reason about changes in software (data relayout, code refactoring) or hardware (caching, prefetching). We describe methods for signature visualization, interpretation, and quantitative comparison of signatures. We evaluate our methodology using applications with variants that contrast data structures, data layouts and algorithms. We show that spatial-temporal affinity analysis provides novel insights and enables predictive reasoning about application performance when contrasted with reuse distance analysis.

Suriyakumar, Yasodhadevi↗

Correlated Anion Disorder in Heteroanionic Cubic TiOF 2

Resolving anion configurations in heteroanionic materials is crucial for understanding and controlling their properties. For anion-disordered oxyfluorides, conventional Bragg diffraction cannot fully resolve the anionic structure, necessitating alternative structure determination methods. We have investigated the anionic structure of anion-disordered cubic (ReO 3 -type) TiOF 2 using X-ray pair distribution function (PDF), 19 F MAS NMR analysis, density functional theory (DFT), cluster expansion modeling, and genetic-algorithm structure prediction. Our computational data predict short-range anion ordering in TiOF 2 , characterized by predominant cis-[O 2 F 4 ] titanium coordination, resulting in correlated anion disorder at longer ranges. To validate our predictions, we generated partially disordered supercells using genetic-algorithm structure prediction and computed simulated X-ray PDF data and 19 F MAS NMR spectra, which we compared directly to experimental data. To construct our simulated 19 F NMR spectra, we derived new transformation functions for mapping calculated magnetic shieldings to predicted magnetic chemical shifts in titanium (oxy)fluorides, obtained by fitting DFT-calculated magnetic shieldings to previously published experimental chemical shift data for TiF 4 . We find good agreement between our simulated and experimental data, which supports our computationally predicted structural model and demonstrates the effectiveness of complementary experimental and computational techniques in resolving anionic structure in anion-disordered oxyfluorides. From additional DFT calculations, we predict that increasing anion disorder makes lithium intercalation more favorable by, on average, up to 2 eV, highlighting the significant effect of variations in short-range order on the intercalation properties of anion-disordered materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Algorithms and file structures to extend and enhance liquid chromatography and ion mobility mass spectrometry workflows (CRADA Final Report)

The purpose of this project was to continue supporting customizations of algorithms and raw data file structures to enhance software workflows for liquid chromatography (LC), mass spectrometry (MS) and ion mobility mass spectrometry (IM-MS)-based protein and metabolite characterization. PNNL worked with Agilent to design, implement, evaluate, and demonstrate new algorithms and integrated them as functionalities into the PNNL-PreProcessor software. The project augmented PNNL’s capabilities to analyze complex proteomics and metabolomics samples. These capabilities are directly beneficial to DOE and PNNL efforts to characterize and analyze these compounds in microbial and plant communities. The project assisted Agilent in further developing improved instrument-software solutions combining liquid chromatography and ion mobility with mass spectrometry for widespread applications in life sciences and other fields.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Algorithms and file structures to enhance software workflows for ion mobility mass spectrometry (IM-MS)

Support customizations of algorithms and raw data file structures to enhance software workflows for liquid chromatography (LC), mass spectrometry (MS) and ion mobility mass spectrometry (IM-MS)-based metabolite characterization. Evaluate and improve the integration of ion mobility to existing MS analysis methods of the Mass Profiler Professional workflow (Mass Profiler, ID Browser and Mass Profiler Professional).

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Algorithms and file structures to enhance software workflows for ion mobility mass spectrometry (IM-MS): CRADA 410 (Final Report)

This document is the final report for CRADA 410 (Project No. 72496). The purpose of this project was supporting customizations of algorithms and raw data file structures to enhance software workflows for liquid chromatography (LC), mass spectrometry (MS) and ion mobility mass spectrometry (IM-MS)-based metabolite characterization. PNNL worked with Agilent to evaluate and improve the integration of ion mobility into existing MS data analysis methods of Agilent software tools. The project augmented PNNL’s capabilities to analyze complex omics samples. These capabilities are directly beneficial to DOE and PNNL efforts to characterize and analyze compounds in microbial and plant communities. The project assisted Agilent in further developing improved instrument-software solutions combining ion mobility with mass spectrometry for widespread applications in life sciences and other fields.

97 MATHEMATICS AND COMPUTING↗

SPACE: 3D parallel solvers for Vlasov-Maxwell and Vlasov-Poisson equations for relativistic plasmas with atomic transformations

A parallel, relativistic, three-dimensional particle-in-cell code SPACE has been developed for the simulation of electromagnetic fields, relativistic particle beams, and plasmas. In addition to the standard second-order Particle-in-Cell (PIC) algorithm, SPACE includes efficient novel algorithms to resolve atomic physics processes such as multi-level ionization of plasma atoms, recombination, and electron attachment to dopants in dense neutral gases. SPACE also contains a highly adaptive particle-based method, called Adaptive Particle-in-Cloud (AP-Cloud), for solving the Vlasov-Poisson problems. It eliminates the traditional Cartesian mesh of PIC and replaces it with an adaptive octree data structure. The code's algorithms, structure, capabilities, parallelization strategy, and performance have been discussed. Additionally, typical examples of SPACE applications to accelerator science and engineering problems are described.

43 PARTICLE ACCELERATORS↗

A machine learning photon detection algorithm for coherent x-ray ultrafast fluctuation analysis

X-ray free electron laser experiments have brought unique capabilities and opened new directions in research, such as creating new states of matter or directly measuring atomic motion. One such area is the ability to use finely spaced sets of coherent x-ray pulses to be compared after scattering from a dynamic system at different times. This enables the study of fluctuations in many-body quantum systems at the level of the ultrafast pulse durations, but this method has been limited to a select number of examples and required complex and advanced analytical tools. By applying a new methodology to this problem, we have made qualitative advances in three separate areas that will likely also find application to new fields. As compared to the “droplet-type” models, which typically are used to estimate the photon distributions on pixelated detectors to obtain the coherent x-ray speckle patterns, our algorithm achieves an order of magnitude speedup on CPU hardware and two orders of magnitude improvement on GPU hardware. We also find that it retains accuracy in low-contrast conditions, which is the typical regime for many experiments in structural dynamics. Finally, it can predict photon distributions in high average-intensity applications, a regime which up until now has not been accessible. Our artificial intelligence-assisted algorithm will enable a wider adoption of x-ray coherence spectroscopies, by both automating previously challenging analyses and enabling new experiments that were not otherwise feasible without the developments described in this work.

47 OTHER INSTRUMENTATION↗

SciPy 1.0: fundamental algorithms for scientific computing in Python

SciPy is an open-source scientific computing library for the Python programming language. Since its initial release in 2001, SciPy has become a de facto standard for leveraging scientific algorithms in Python, with over 600 unique code contributors, thousands of dependent packages, over 100,000 dependent repositories and millions of downloads per year. In this work, we provide an overview of the capabilities and development practices of SciPy 1.0 and highlight some recent technical developments.

97 MATHEMATICS AND COMPUTING↗

RANGE: A robust adaptive nature-inspired global explorer of potential energy surfaces

With the growing demand for realistic representations of chemical structures and the advent of exascale computing, the intelligent sampling of potential energy surfaces and efficient identification of global minima have become more essential but also more feasible. Building on prior studies demonstrating the efficiency of the Artificial Bee Colony (ABC) swarm intelligence algorithm, we report a hybrid metaheuristic framework that integrates the adaptive exploration capabilities of ABC coupled with the exploitation strengths of genetic algorithms (GA) in a scalable, Python-based implementation. The resulting tool, RANGE (Robust Adaptive Nature-inspired Global Explorer), provides seamless interfaces to multiple potential energy evaluators, either directly or via widely used Python libraries, and is designed for high-performance computing environments. We describe the implementation details of RANGE and evaluate its performance, relative to ABC- or GA-alone based algorithms, on a variety of chemical systems, including molecular clusters and heterogeneous surfaces. In conclusion, our results demonstrate RANGE’s efficiency, robustness, and broad applicability in addressing challenging global optimization problems in computational chemistry and materials science.

Algorithms and data structure↗

Computing water flow through complex landscapes – Part 3: Fill–Spill–Merge: flow routing in depression hierarchies

Abstract. Depressions – inwardly draining regions – are common to many landscapes. When there is sufficient moisture, depressions take the form of lakes and wetlands; otherwise, they may be dry. Hydrological flow models used in geomorphology, hydrology, planetary science, soil and water conservation, and other fields often eliminate depressions through filling or breaching; however, this can produce unrealistic results. Models that retain depressions, on the other hand, are often undesirably expensive to run. In previous work we began to address this by developing a depression hierarchy data structure to capture the full topographic complexity of depressions in a region. Here, we extend this work by presenting the Fill–Spill–Merge algorithm that utilizes our depression hierarchy data structure to rapidly process and distribute runoff. Runoff fills depressions, which then overflow and spill into their neighbors. If both a depression and its neighbor fill, they merge. We provide a detailed explanation of the algorithm and results from two sample study areas. In these case studies, the algorithm runs 90–2600 times faster (with a reduction in compute time of 2000–63 000 times) than the commonly used Jacobi iteration and produces a more accurate output. Complete, well-commented, open-source code with 97 % test coverage is available on GitHub and Zenodo.

58 GEOSCIENCES↗

Understanding, discovery, and synthesis of 2D materials enabled by machine learning

Machine learning (ML) is becoming an effective tool for studying 2D materials. Taking as input computed or experimental materials data, ML algorithms predict the structural, electronic, mechanical, and chemical properties of 2D materials that have yet to be discovered. Such predictions expand investigations on how to synthesize 2D materials and use them in various applications, as well as greatly reduce the time and cost to discover and understand 2D materials. This tutorial review focuses on the understanding, discovery, and synthesis of 2D materials enabled by or benefiting from various ML techniques. Here, we introduce the most recent efforts to adopt ML in various fields of study regarding 2D materials and provide an outlook for future research opportunities. The adoption of ML is anticipated to accelerate and transform the study of 2D materials and their heterostructures.

2D Materials↗

Convergence acceleration of Monte Carlo many-body perturbation methods by direct sampling

In the Monte Carlo many-body perturbation (MC-MP) method, the conventional correlation-correction formula, which is a long sum of products of low-dimensional integrals, is first recast into a short sum of high-dimensional integrals over electron-pair and imaginary-time coordinates. These high-dimensional integrals are then evaluated by the Monte Carlo method with random coordinates generated by the Metropolis–Hasting algorithm according to a suitable distribution. The latter algorithm, while advantageous in its ability to sample nearly any distribution, introduces autocorrelation in sampled coordinates, which in turn increases the statistical uncertainty of the integrals and thus the computational cost. It also involves wasteful rejected moves and an initial “burn-in” step as well as displays hysteresis. Here, an algorithm is proposed that directly produces a random sequence of electron-pair coordinates for the same distribution used in the MC-MP method, which is free from autocorrelation, rejected moves, a burn-in step, or hysteresis. Furthermore, this direct-sampling algorithm is shown to accelerate second- (MC-MP2) and third-order Monte Carlo many-body perturbation (MC-MP3) calculations by up to 222% and 38%, respectively.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Impact of Numerical Hydrodynamics in Turbulent Mixing Transition Simulations

Underresolved simulations are unavoidable in high Reynolds (Re) and Mach (Ma) number turbulent flow applications at scale. Implicit large-Eddy simulation (ILES) often becomes the effective strategy to capture the dominating effects of convectively driven flow instabilities. We evaluate the impact of three distinct numerical strategies in simulations of transition and turbulence decay with ILES: the Harten–Lax–van Leer (HLL) Riemann solver applying Strang splitting and a Lagrange-plus-Remap formalism to solve the directional sweep—denoted split; the Harten–Lax–Van Leer-Contact (HLLC) Riemann solver using a directionally unsplit strategy and parabolic reconstruction—denoted unsplit; and the HLLC Riemann solver using unsplit and a low-Ma correction (LMC)—denoted unsplit*. Three case studies are considered: (1) a shock tube problem prototyping shock-driven turbulent mixing, (2) the Taylor–Green Vortex (TGV) prototyping transition to turbulence, and, (3) an homogeneous isotropic turbulence (HIT) case, focusing on the impact of discretization on transition and decay from fixed well-characterized initial conditions. Significantly more accurate predictions are provided by the unsplit schemes, in particular, when augmented with the LMC. For given resolution, only the unsplit schemes predict the turbulent mixing transition after reshock observed in the shock tube experiments. Relevant comparisons of ILES based on Euler and Navier–Stokes equations addressing potential occurrence of low-Re regimes in the applications are presented. Unsplit schemes are instrumental in allowing to capture the spatial development of the TGV flow and its validation at prescribed Re with significantly less resolution. HIT analysis confirms higher simulated turbulence Re and increased small-scale content associated with the unsplit discretizations.

74 ATOMIC AND MOLECULAR PHYSICS↗

Imaging pyrometry for most color cameras using a triple pass filter

A simple combination of the Planck blackbody emission law, optical filters, and digital image processing is demonstrated to enable most commercial color cameras (still and video) to be used as an imaging pyrometer for flames and explosions. The hardware and data processing described take advantage of the color filter array (CFA) that is deposited on the surface of the light sensor array present in most digital color cameras. In this work, a triple-pass optical filter incorporated into the camera lens allows light in three 10-nm wide bandpass regions to reach the CFA/light sensor array. These bandpass regions are centered over the maxima in the blue, green, and red transmission regions of the CFA, minimizing the spectral overlap of these regions normally present. A computer algorithm is used to retrieve the blue, green, and red image matrices from camera memory and correct for remaining spectral overlap. A second algorithm calibrates the corrected intensities to a gray body emitter of known temperature, producing a color intensity correction factor for the camera/filter system. The Wien approximation to the Planck blackbody emission law is used to construct temperature images from the three color (blue, green, red) matrices. A short pass filter set eliminates light of wavelengths longer than 750 nm, providing reasonable accuracy (±10%) for temperatures between 1200 and 6000 K. The effectiveness of this system is demonstrated by measuring the temperature of several systems for which the temperature is known.

47 OTHER INSTRUMENTATION↗