Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Dynamic clustering algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Scheduling Operations for Massive Heterogeneous Clusters

High-performance computing (HPC) programming has become increasingly difficult with the advent of hybrid supercomputers consisting of multicore CPUs and accelerator boards such as the GPU. Manual tuning of software to achieve high performance on this type of machine has been performed by programmers. This is needlessly difficult and prone to being invalidated by new hardware, new software, or changes in the underlying code. A system was developed for task-based representation of programs, which when coupled with a scheduler and runtime system, allows for many benefits, including higher performance and utilization of computational resources, easier programming and porting, and adaptations of code during runtime. The system consists of a method of representing computer algorithms as a series of data-dependent tasks. The series forms a graph, which can be scheduled for execution on many nodes of a supercomputer efficiently by a computer algorithm. The schedule is executed by a dispatch component, which is tailored to understand all of the hardware types that may be available within the system. The scheduler is informed by a cluster mapping tool, which generates a topology of available resources and their strengths and communication costs. Software is decoupled from its hardware, which aids in porting to future architectures. A computer algorithm schedules all operations, which for systems of high complexity (i.e., most NASA codes), cannot be performed optimally by a human. The system aids in reducing repetitive code, such as communication code, and aids in the reduction of redundant code across projects. It adds new features to code automatically, such as recovering from a lost node or the ability to modify the code while running. In this project, the innovators at the time of this reporting intend to develop two distinct technologies that build upon each other and both of which serve as building blocks for more efficient HPC usage. First is the scheduling and dynamic execution framework, and the second is scalable linear algebra libraries that are built directly on the former.

Humphrey, John↗

Short- and medium-range orders in Al90Tb10 glass and their relation to the structures of competing crystalline phases

Molecular dynamics simulations using an interatomic potential developed by artificial neural network deep machine learning are performed to study the local structural order in Al90Tb10 metallic glass. We show that more than 80% of the Tb-centered clusters in Al90Tb10 glass have short-range order (SRO) with their 17 first coordination shell atoms stacked in a ‘3661’ or ‘15551’ sequence. Medium-range order (MRO) in Bergman-type packing extended out to the second and third coordination shells is also clearly observed. Analysis of the network formed by the ‘3661’ and ‘15551’ clusters show that ~82% of such SRO units share their faces or vertexes, while only ~6% of neighboring SRO pairs are interpenetrating. Such a network topology is consistent with the Bergman-type MRO around the Tb-centers. Moreover, crystal structure searches using genetic algorithm and the neural network interatomic potential reveal several low-energy metastable crystalline structures in the composition range close to Al 90 Tb 10 . Some of these crystalline structures have the ‘3661’ SRO while others have the ‘15551’ SRO. While the crystalline structures with the ‘3661’ SRO also exhibit the MRO very similar to that observed in the glass, the ones with the ‘15551’ SRO have very different atomic packing in the second and third shells around the Tb centers from that of the Bergman-type MRO observed in the glassy phase.

36 MATERIALS SCIENCE↗

Continuous-variable quantum computation of the O(3) model in 1+1 dimensions

We formulate the $O(3)$ non-linear sigma model in $1+1$ dimensions as a limit of a three-component scalar field theory restricted to the unit sphere in the large squeezing limit. This allows us to describe the model in terms of the continuous variable (CV) approach to quantum computing. Here we construct the ground state and excited states using the coupled cluster ansatz and find excellent agreement with the exact diagonalization results for a small number of lattice sites. We then present the simulation protocol for the time evolution of the model using CV gates, estimate the discretization error, and present numerical results obtained from a photonic quantum simulator. We expect that the methods developed in this work will be useful for exploring interesting dynamics for a wide class of sigma models and gauge theories, as well as for simulating scattering events on quantum hardware in the coming decade.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Fast and Scalable FFT-Based GPU-Accelerated Algorithms for Block-Triangular Toeplitz Matrices with Application to Linear Inverse Problems Governed by Autonomous Dynamical Systems

In this work, we present an efficient and scalable algorithm for performing matrix-vector multiplications (matvecs) for block Toeplitz matrices. Such matrices, which are shift-invariant with respect to their blocks, arise in the context of solving inverse problems governed by autonomous systems, and time-invariant systems in particular. In this article, we consider inverse problems that infer unknown parameters from observational data of a linear time-invariant dynamical system given in the form of partial differential equations (PDEs). Matrix-free Newton-conjugate-gradient methods are often the gold standard for solving these inverse problems, but they require numerous actions of the Hessian on a vector. Matrix-free adjoint-based Hessian matvecs require solution of a pair of linearized forward/adjoint PDE solves per Hessian action, which may be prohibitive for large-scale inverse problems. Time invariance of the forward PDE problem leads to a block Toeplitz structure of the discretized parameter-to-observable (p2o) map defining the mapping from inputs (parameters) to outputs (observables) of the PDEs. This block Toeplitz structure enables us to exploit two key properties: (1) compact storage of the p2o map and its adjoint, and (2) efficient fast Fourier transform–based Hessian matvecs. The proposed algorithm is mapped onto large multi-GPU clusters and achieves more than 80% of peak bandwidth on NVIDIA A100 GPUs. Excellent weak scaling is shown for up to 48 A100 GPUs. For the targeted problems, the implementation executes Hessian matvecs within fractions of a second, which is orders of magnitude faster than can be achieved by conventional matrix-free Hessian matvecs via forward/adjoint PDE solves.

97 MATHEMATICS AND COMPUTING↗

Conserved unique peptide patterns (CUPP) online platform 2.0: implementation of +1000 JGI fungal genomes

Carbohydrate-processing enzymes, CAZymes, are classified into families based on sequence and three-dimensional fold. Because many CAZyme families contain members of diverse molecular function (different EC-numbers), sophisticated tools are required to further delineate these enzymes. Such delineation is provided by the peptide-based clustering method CUPP, Conserved Unique Peptide Patterns. CUPP operates synergistically with the CAZy family/subfamily categorizations to allow systematic exploration of CAZymes by defining small protein groups with shared sequence motifs. The updated CUPP library contains 21,930 of such motif groups including 3,842,628 proteins. The new implementation of the CUPP-webserver, https://cupp.info/, now includes all published fungal and algal genomes from the Joint Genome Institute (JGI), genome resources MycoCosm and PhycoCosm, dynamically subdivided into motif groups of CAZymes. This allows users to browse the JGI portals for specific predicted functions or specific protein families from genome sequences. Thus, a genome can be searched for proteins having specific characteristics. All JGI proteins have a hyperlink to a summary page which links to the predicted gene splicing including which regions have RNA support. The new CUPP implementation also includes an update of the annotation algorithm that uses only a fourth of the RAM while enabling multi-threading, providing an annotation speed below 1 ms/protein.

59 BASIC BIOLOGICAL SCIENCES↗

Model Based Validation of Intelligent Powertrain Strategies for Connected and Automated Vehicles

Systems incorporating Vehicle to Everything (V2X) and conventional cellular based communication in vehicles can significantly help improve energy consumption via a combination of intelligent powertrain control strategies, smarter routing algorithms and driving in such a way as to minimize fuel economy and the emission of carbon dioxide, known as "eco-driving." In projects led by the Southwest Research Institute (SwRI), large-scale traffic simulations are created to model real-world scenarios with dynamic behavior that is reactive to imposed changes. Coupled with high fidelity powertrain models, the closed loop framework enables research and development of such Connected and Automated Vehicle (CAV) enabled technologies at scale. This paper will discuss a traffic system simulation environment that was built based on the High Street urban corridor in Columbus, Ohio. Eco-driving strategies were tested at scale on a variety of powertrain platforms – internal combustion engines, hybrid electric and fully electric vehicles. Furthermore, the paper will focus on hybrid electric powertrain modeling along with details on how the powertrain model was leveraged to develop a sophisticated clustering scheme to help down-select speed traces from large scale simulation studies for validation on vehicle dynamometer. Nominal energy consumption improvement around 12% was observed with good match between simulation studies and vehicle testing.

33 ADVANCED PROPULSION SYSTEMS↗

A Hardware-in-the-Loop Testbed for Spacecraft Formation Flying Applications

The Formation Flying Test Bed (FFTB) at NASA Goddard Space Flight Center (GSFC) is being developed as a modular, hybrid dynamic simulation facility employed for end-to-end guidance, navigation, and control (GN&C) analysis and design for formation flying clusters and constellations of satellites. The FFTB will support critical hardware and software technology development to enable current and future missions for NASA, other government agencies, and external customers for a wide range of missions, particularly those involving distributed spacecraft operations. The initial capabilities of the FFTB are based upon an integration of high fidelity hardware and software simulation, emulation, and test platforms developed at GSFC in recent years; including a high-fidelity GPS simulator which has been a fundamental component of the Guidance, Navigation, and Control Center's GPS Test Facility. The FFTB will be continuously evolving over the next several years from a too[ with initial capabilities in GPS navigation hardware/software- in-the- loop analysis and closed loop GPS-based orbit control algorithm assessment to one with cross-link communications and relative navigation analysis and simulation capability. Eventually the FFT13 will provide full capability to support all aspects of multi-sensor, absolute and relative position determination and control, in all (attitude and orbit) degrees of freedom, as well as information management for satellite clusters and constellations. In this paper we focus on the architecture for the FFT13 as a general GN&C analysis environment for the spacecraft formation flying community inside and outside of NASA GSFC and we briefly reference some current and future activities which will drive the requirements and development.

Leitner, Jesse↗

Gradients in giant branch morphology in the core of 47 Tucanae

I describe an algorithm which uses the high spatial resolution of the Hubble Space Telescope to complement the high spatial-to-noise, approximately symmetric point response function, relatively large spatial coverage, and standard filters available from ground based images of crowded fields. Applying this technique to the central regions of the globular cluster 47 Tucanae, I find that the morphology of the giant branch in the core is significantly different from that in more distant regions (r approximately equals 5 to 10 core radii) of the cluster. In particular, there appear to be fewer bright giants in the core, along with an enhanced `asymptotic giant branch' (AGB) sequence. Depletion of giants has been observed in the cores of other dense clusters, and may be due to `stripping' of large stars by stellar encounters and/or mass transfer in binary systems. Central concentrations of true asymptotic giant branch stars are not expected to result from dynamical processes; possibly some of these stars may be evolved blue stragglers.

Bailyn, Charles D.↗

A Generative Model for Realistic Galaxy Cluster X-Ray Morphologies

Abstract The X-ray morphologies of clusters of galaxies display significant variations, reflecting their dynamical histories and the nonlinear dependence of X-ray emissivity on the density of the intracluster gas. Qualitative and quantitative assessments of X-ray morphology have long been considered a proxy for determining whether clusters are dynamically active or “relaxed.” Conversely, the use of circularly or elliptically symmetric models for cluster emission can be complicated by the variety of complex features realized in nature, spanning scales from megaparsecs down to the resolution limit of current X-ray observatories. In this work, we use mock X-ray images from simulated clusters from The Three Hundred project to define a basis set of cluster image features. We take advantage of the clusters’ approximate self-similarity to minimize the differences between images before encoding the remaining diversity through a distribution of high-order polynomial coefficients. Principal component analysis then provides an orthogonal basis for this distribution, corresponding to natural perturbations from an average model. This representation allows novel, realistically complex X-ray cluster images to be easily generated, and we provide code to do so. The approach provides a simple way to generate training data for cluster image analysis algorithms and could be straightforwardly adapted to generate clusters displaying specific types of features or selected by physical characteristics available in the original simulations.

79 ASTRONOMY AND ASTROPHYSICS↗

Computational strategies for three-dimensional flow simulations on distributed computer systems

This research effort is directed towards an examination of issues involved in porting large computational fluid dynamics codes in use within the industry to a distributed computing environment. This effort addresses strategies for implementing the distributed computing in a device independent fashion and load balancing. A flow solver called TEAM presently in use at Lockheed Aeronautical Systems Company was acquired to start this effort. The following tasks were completed: (1) The TEAM code was ported to a number of distributed computing platforms including a cluster of HP workstations located in the School of Aerospace Engineering at Georgia Tech; a cluster of DEC Alpha Workstations in the Graphics visualization lab located at Georgia Tech; a cluster of SGI workstations located at NASA Ames Research Center; and an IBM SP-2 system located at NASA ARC. (2) A number of communication strategies were implemented. Specifically, the manager-worker strategy and the worker-worker strategy were tested. (3) A variety of load balancing strategies were investigated. Specifically, the static load balancing, task queue balancing and the Crutchfield algorithm were coded and evaluated. (4) The classical explicit Runge-Kutta scheme in the TEAM solver was replaced with an LU implicit scheme. And (5) the implicit TEAM-PVM solver was extensively validated through studies of unsteady transonic flow over an F-5 wing, undergoing combined bending and torsional motion. These investigations are documented in extensive detail in the dissertation, 'Computational Strategies for Three-Dimensional Flow Simulations on Distributed Computing Systems', enclosed as an appendix.

Sankar, Lakshmi N.↗

HipBone: A performance-portable graphics processing unit-accelerated C++ version of the NekBone benchmark

We present hipBone, an open-source performance-portable proxy application for the Nek5000 (and NekRS) computational fluid dynamics applications. HipBone is a fully GPU-accelerated C++ implementation of the original NekBone CPU proxy application with several novel algorithmic and implementation improvements which optimize its performance on modern fine-grain parallel GPU accelerators. Our optimizations include a conversion to store the degrees of freedom of the problem in assembled form in order to reduce the amount of data moved during the main iteration and a portable implementation of the main Poisson operator kernel. We demonstrate near-roofline performance of the operator kernel on three different modern GPU accelerators from two different vendors. We present a novel algorithm for splitting the application of the Poisson operator on GPUs which aggressively hides MPI communication required for both halo exchange and assembly. Our implementation of nearest-neighbor MPI communication then leverages several different routing algorithms and GPU-Direct RDMA capabilities, when available, which improves scalability of the benchmark. We demonstrate the performance of hipBone on three different clusters housed at Oak Ridge National Laboratory, namely, the Summit supercomputer and the Frontier early-access clusters, Spock and Crusher. Our tests demonstrate both portability across different clusters and very good scaling efficiency, especially on large problems.

Computer Science↗

Simulation of Guidance, Navigation, and Control Systems for Formation Flying Missions

Concepts for missions of distributed spacecraft flying in formation abound. From high resolution interferometry to spatially distributed in-situ measurements, these mission concepts levy a myriad of guidance, navigation, and control (GNC) requirements on the spacecraft/formation as a single system. A critical step toward assessing and meeting these challenges lies in realistically simulating distributed spacecraft systems. The Formation Flying TestBed (FFTB) at NASA Goddard Space Flight Center's (GSFC) Guidance, Navigation, and Control Center is a hardware-in-the-loop simulation and development facility focused on GNC issues relevant to formation flying systems. The FFTB provides a realistic simulation of the vehicle dynamics and control for formation flying missions in order to: (1) conduct feasibility analyses of mission requirements, (2) conduct and answer mission and spacecraft design trades, and (3) serve as a host for GNC software and hardware development and testing. The initial capabilities of the FFTB are based upon an integration of high fidelity hardware and software simulation, emulation, and test platforms developed or employed at GSFC in recent years, including a high-fidelity Global Positioning System (GPS) simulator which has been a fundamental component of the GNC Center's GPS Test Facility. The FFTB will be continuously evolving over the next several years from a tool with capabilities in GPS navigation hardware/software-in-the-loop analysis and closed loop GPS-based orbit control algorithm assessment. Eventually, it will include full capability to support all aspects of multi-sensor, absolute and relative state determination and control, in all (attitude and orbit) degrees of freedom, as well as information management for satellite clusters and constellations. A detailed description of the FFTB architecture is presented in the paper.

Burns, Rich↗

Loci-STREAM Version 0.9

Loci-STREAM is an evolving computational fluid dynamics (CFD) software tool for simulating possibly chemically reacting, possibly unsteady flows in diverse settings, including rocket engines, turbomachines, oil refineries, etc. Loci-STREAM implements a pressure- based flow-solving algorithm that utilizes unstructured grids. (The benefit of low memory usage by pressure-based algorithms is well recognized by experts in the field.) The algorithm is robust for flows at all speeds from zero to hypersonic. The flexibility of arbitrary polyhedral grids enables accurate, efficient simulation of flows in complex geometries, including those of plume-impingement problems. The present version - Loci-STREAM version 0.9 - includes an interface with the Portable, Extensible Toolkit for Scientific Computation (PETSc) library for access to enhanced linear-equation-solving programs therein that accelerate convergence toward a solution. The name "Loci" reflects the creation of this software within the Loci computational framework, which was developed at Mississippi State University for the primary purpose of simplifying the writing of complex multidisciplinary application programs to run in distributed-memory computing environments including clusters of personal computers. Loci has been designed to relieve application programmers of the details of programming for distributed-memory computers.

Wright, Jeffrey↗

Hijacking a rapid and scalable metagenomic method reveals subgenome dynamics and evolution in polyploid plants

Premise: The genomes of polyploid plants archive the evolutionary events leading to their present forms. However, plant polyploid genomes present numerous hurdles to the genome comparison algorithms for classification of polyploid types and exploring genome dynamics. Methods: Here, the problem of intra- and inter-genome comparison for examining polyploid genomes is reframed as a metagenomic problem, enabling the use of the rapid and scalable MinHashing approach. To determine how types of polyploidy are described by this metagenomic approach, plant genomes were examined from across the polyploid spectrum for both k-mer composition and frequency with a range of k-mer sizes. In this approach, no subgenome-specific k-mers are identified; rather, whole-chromosome k-mer subspaces were utilized. Results: Given chromosome-scale genome assemblies with sufficient subgenome-specific repetitive element content, literature-verified subgenomic and genomic evolutionary relationships were revealed, including distinguishing auto- from allopolyploidy and putative progenitor genome assignment. The sequences responsible were the rapidly evolving landscape of transposable elements. An investigation into the MinHashing parameters revealed that the downsampled k-mer space (genomic signatures) produced excellent approximations of sequence similarity. Furthermore, the clustering approach used for comparison of the genomic signatures is scrutinized to ensure applicability of the metagenomics-based method. Discussion: The easily implementable and highly computationally efficient MinHashing-based sequence comparison strategy enables comparative subgenomics and genomics for large and complex polyploid plant genomes. Such comparisons provide evidence for polyploidy-type subgenomic assignments. In cases where subgenome-specific repeat signal may not be adequate given a chromosomes' global k-mer profile, alternative methods that are more specific but more computationally complex outperform this approach.

59 BASIC BIOLOGICAL SCIENCES↗

A Decentralized Approach for Modeling Organized Convection Based on Thermal Populations on Microgrids

Abstract In this study, a spectral model for convective transport is coupled to a thermal population model on a two‐dimensional horizontal “microgrid,” covering the typical gridbox size of general circulation models. The goal is to explore new ways of representing impacts of spatial organization in cumulus cloud fields. The thermals are considered the smallest building block of convection, with thermal life cycle and movement represented through binomial functions. Thermals interact through two simple rules, reflecting pulsating growth and environmental deformation. Long‐lived thermal clusters thus form on the microgrid, exhibiting scale growth and spacing that represent simple forms of spatial organization and memory. Size distributions of cluster number are diagnosed from the microgrid through an online clustering algorithm, and provided as input to a spectral multiplume eddy‐diffusivity mass flux scheme. This yields a decentralized transport system, in that the thermal clusters acting as independent but interacting nodes that carry information about spatial structure. The main objectives of this study are (a) to seek proof of concept of this approach, and (b) to gain insight into impacts of spatial organization on convective transport. Single‐column model experiments demonstrate satisfactory skill in reproducing two observed cases of continental shallow convection. Metrics expressing self‐organization and spatial organization match well with large‐eddy simulation results. We find that in this coupled system, spatial organization impacts convective transport primarily through the scale break in the size distribution of cluster number. The rooting of saturated plumes in the subcloud mixed layer plays a key role in this process.

54 ENVIRONMENTAL SCIENCES↗

Gene Expression Dynamics Inspector (GEDI): for integrative analysis of expression profiles

Genome-wide expression profiles contain global patterns that evade visual detection in current gene clustering analysis. Here, a Gene Expression Dynamics Inspector (GEDI) is described that uses self-organizing maps to translate high-dimensional expression profiles of time courses or sample classes into animated, coherent and robust mosaics images. GEDI facilitates identification of interesting patterns of molecular activity simultaneously across gene, time and sample space without prior assumption of any structure in the data, and then permits the user to retrieve genes of interest. Important changes in genome-wide activities may be quickly identified based on 'Gestalt' recognition and hence, GEDI may be especially useful for non-specialist end users, such as physicians. AVAILABILITY: GEDI v1.0 is written in Matlab, and binary Matlab.dll files which require Matlab to run can be downloaded for free by academic institutions at http://www.chip.org/~ge/gedihome.html Supplementary information: http://www.chip.org/~ge/gedihome.html.

Database Management Systems↗

Code Validation Study for Base Flows

New and old rocket launch concepts recommend the clustering of motors for improved lift capability. The flowfield of the base region of the rocket is very complex and can contain high temperature plume gases. These hot gases can cause catastrophic problems if not adequately designed for. To assess the base region characteristics, advanced computational fluid dynamics (CFD) is being used. As a precursor to these calculations the CFD code requires validation on base flows. The primary objective of this code validation study was to establish a high level of confidence in predicting base flows with the USA CFD code. USA has been extensively validated for fundamental flows and other applications. However, base heating flows have a number of unique characteristics so it was necessary to extend the existing validation for this class of problems. In preparation for the planned NLS 1.5 Stage base heating analysis, six case sets were studied to extend the USA code validation data base. This presentation gives a cursive review of three of these cases. The cases presented include a 2D axi-symmetric study, a 3D real nozzle study, and a 3D multi-species study. The results of all the studies show good general agreement with data with no adjustments to the base numerical algorithms or physical models in the code. The study proved the capability of the USA code for modeling base flows within the accuracy of available data.

Ascoli, Edward P.↗

Comparing Synoptic Pattern Evolution for Flash‐Flood‐Producing and Non‐Flash‐Flood‐Producing Mesoscale Convective Systems in the United States

Understanding how the short-term evolution of synoptic weather patterns influence Mesoscale Convective Systems (MCSs) is essential, as these systems are responsible for over half of central U.S. flash floods, leading to substantial socioeconomic and water resource management impacts. This study analyzes long-term MCS data, flash flood reports, and atmospheric reanalyses from 2007 to 2017 using a machine learning clustering algorithm to examine how the synoptic weather patterns evolve prior to MCS initiation. While the clusters reflect seasonal and regional differences in MCS occurrence, they do not consistently distinguish between MCSs that do and do not produce flash floods. Systems in the southern Great Plains are more flood-prone when a synoptic-scale forcing, located near the system, drives strong water vapor transport from the nearby moisture source. More generally under different synoptic weather patterns, a broader precipitating area is the most dominant factor governing MCS flash flood potential.

atmospheric dynamics↗