Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “local clustering algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Metals and Quantum Materials with Spin-orbit Interactions by Quantum Monte Carlo methods

The key goals of this project were as follows: 1) Analysis and benchmarks of electron correlation effects recovered in the fixed-node approximation that is inherent to quantum Monte Carlo (QMC) method as applied to metallic states; 2) development of new algorithms for electron spin-degrees of freedom to be treated as explicit quantum variables; 3) designing electronic structure QMC algorithm for efficient evaluation of spin-orbit effects in systems with heavy atoms; 4) adapting the algorithm to complex wave functions and developing corresponding fixed-phase approximation; 5) design and testing of algorithm for valence-only non-local spin-orbit operators; 6) analysis of fixed-node vs fixed-phase errors and their comparisons. The key accomplishments: i) We carried out a systematic study of Li systems by the fixed-node diffusion Monte Carlo method. This involved Li atom, molecule, cluster and solid calculated by the full range of QMC methods including fixed-node QMC.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Data Analysis of the 2020 Central Idaho Mainshock-Aftershock Sequence

In an effort to inform the Senior Seismic Hazard Analysis Committee for the Idaho National Laboratory, we provide an improved aftershock catalog related to the March 31, 2020, Mw6.5 Stanley, Idaho earthquake from picks related to a temporary network of two real-time and 15 non-telemetered seismometers within the epicentral area. From the permanent and temporary (XP) real-time network, the USGS cataloged 1,946 aftershocks between April 1, 2020 and October 31, 2020. To improve aftershock location and magnitudes, we manually picked arrival times of P and S waves from off-line stations in the XP temporary network, generated a new crustal velocity model, and independently relocated each event using the HypoDD double-difference earthquake algorithm. We created our new velocity model from existing broadband and active source seismic campaign data that were acquired near the epicentral region prior to the 2020 earthquake. We compare arrival time differences, epicentral locations and depths between aftershocks recorded with the two catalogs. We find the addition of local stations provides tighter aftershock clustering that suggests an improved aftershock locations. To detect lower magnitude events, we employed deep learning. Our method solves common problems associated with detecting many events that have a low signal-to-noise ratio. From the machine learning database, we detected more than 74,000 aftershocks. Based on the number of identified earthquakes and Gutenberg-Richter relationships derived from the USGS catalog, we estimate that we have reduced the completion magnitude for the Stanley earthquake sequence to below M1 using this machine learning approach. We located each aftershock with our new velocity model. Our new velocity model and picks suggests aftershocks occurred mostly at shallower depths than assessed in the USGS catalog. These aftershocks align along two linear trends that suggest the activation of two unnamed primary faults.

58 GEOSCIENCES↗

Initialization and Restart in Stochastic Local Search: Computing a Most Probable Explanation in Bayesian Networks

For hard computational problems, stochastic local search has proven to be a competitive approach to finding optimal or approximately optimal problem solutions. Two key research questions for stochastic local search algorithms are: Which algorithms are effective for initialization? When should the search process be restarted? In the present work we investigate these research questions in the context of approximate computation of most probable explanations (MPEs) in Bayesian networks (BNs). We introduce a novel approach, based on the Viterbi algorithm, to explanation initialization in BNs. While the Viterbi algorithm works on sequences and trees, our approach works on BNs with arbitrary topologies. We also give a novel formalization of stochastic local search, with focus on initialization and restart, using probability theory and mixture models. Experimentally, we apply our methods to the problem of MPE computation, using a stochastic local search algorithm known as Stochastic Greedy Search. By carefully optimizing both initialization and restart, we reduce the MPE search time for application BNs by several orders of magnitude compared to using uniform at random initialization without restart. On several BNs from applications, the performance of Stochastic Greedy Search is competitive with clique tree clustering, a state-of-the-art exact algorithm used for MPE computation in BNs.

Mengshoel, Ole J.↗

Personalized Tucker Decomposition: Modeling Commonality and Peculiarity on Tensor Data

In this paper, we propose a personalized Tucker decomposition (perTucker) to address the limitations of traditional tensor decomposition methods in capturing heterogeneity across different datasets. perTucker decomposes tensor data into shared global components and personalized local components. We introduce an order orthogonality assumption and develop a proximal gradient regularized block coordinate descent algorithm guaranteed to converge to a stationary point. The unique and common representations learned by perTucker reveal intrinsic statistical patterns in data and provide valuable information for a wide range of downstream analytics, including anomaly detection, source classification, and clustering. We demonstrate perTucker’s effectiveness through a simulation study and two case studies on solar flare detection and tonnage signal classification.

14 SOLAR ENERGY↗

Classification of Ascension Island and Natal Ozonesondes Using Self-Organizing Maps

Ozone profiles from balloon-borne ozonesondes are used for development of satellite algorithms and in chemistry-climate model initialization, assimilation and evaluation. An important issue in the application of these profiles is how best to treat variations where varying photochemical and dynamical influences can cause the ozone mixing ratio in the tropospheric segments of the profile to change by of a factor of 2-3 within a day. Clustering techniques are an ideal way to approach the statistical classification of profile data and we apply self-organizing maps to tropical tropospheric SHADOZ data, hypothesizing that the data will sort according to various influences on ozone, namely anthropogenic sources like biomass burning, meteorological conditions, and stratospheric or extra-tropical intrusions. Self-organizing maps, that use a learning algorithm to reveal the most prominent features of a data set according to a specified number of clusters, have been determined for the 1998-2009 SHADOZ profiles over Ascension Island (512 profiles, 7.98 deg. S, 14.42 deg. W) and Natal, Brazil (425 profiles, 5.42degS, 35.38degW). The 2 × 2 self-organizing map, which creates 4 clusters, reveals that deviations from the average ozone in the free troposphere include both increased ozone resulting from seasonal biomass burning in Africa and locally reduced ozone brought about by convective lifting of unpolluted boundary-layer air. Expanding to a 4 × 4 self-organizing map shows how biomass burning influences the yearly cycle of tropospheric ozone at Ascension Island and captures the seasonality of ozone at both Ascension Island and Natal. Comparing Ascension Island and Natal using a 4 × 4 self-organizing map at each site reveals similarities in mid-tropospheric ozone, but shows differences in lower-tropospheric ozone due to Ascension Island being closer to African biomass burning and more affected by descent from the mean Walker circulation, with less convective activity, than Natal.

algorithms↗

Adaptive grid methods for RLV environment assessment and nozzle analysis

Rapid access to highly accurate data about complex configurations is needed for multi-disciplinary optimization and design. In order to efficiently meet these requirements a closer coupling between the analysis algorithms and the discretization process is needed. In some cases, such as free surface, temporally varying geometries, and fluid structure interaction, the need is unavoidable. In other cases the need is to rapidly generate and modify high quality grids. Techniques such as unstructured and/or solution-adaptive methods can be used to speed the grid generation process and to automatically cluster mesh points in regions of interest. Global features of the flow can be significantly affected by isolated regions of inadequately resolved flow. These regions may not exhibit high gradients and can be difficult to detect. Thus excessive resolution in certain regions does not necessarily increase the accuracy of the overall solution. Several approaches have been employed for both structured and unstructured grid adaption. The most widely used involve grid point redistribution, local grid point enrichment/derefinement or local modification of the actual flow solver. However, the success of any one of these methods ultimately depends on the feature detection algorithm used to determine solution domain regions which require a fine mesh for their accurate representation. Typically, weight functions are constructed to mimic the local truncation error and may require substantial user input. Most problems of engineering interest involve multi-block grids and widely disparate length scales. Hence, it is desirable that the adaptive grid feature detection algorithm be developed to recognize flow structures of different type as well as differing intensity, and adequately address scaling and normalization across blocks. These weight functions can then be used to construct blending functions for algebraic redistribution, interpolation functions for unstructured grid generation, forcing functions to attract/repel points in an elliptic system, or to trigger local refinement, based upon application of an equidistribution principle. The popularity of solution-adaptive techniques is growing in tandem with unstructured methods. The difficultly of precisely controlling mesh densities and orientations with current unstructured grid generation systems has driven the use of solution-adaptive meshing. Use of derivatives of density or pressure are widely used for construction of such weight functions, and have been proven very successful for inviscid flows with shocks. However, less success has been realized for flowfields with viscous layers, vortices or shocks of disparate strength. It is difficult to maintain the appropriate mesh point spacing in the various regions which require a fine spacing for adequate resolution. Mesh points often migrate from important regions due to refinement of dominant features. An example of this is the well know tendency of adaptive methods to increase the resolution of shocks in the flowfield around airfoils, but in the incorrect location due to inadequate resolution of the stagnation region. This problem has been the motivation for this research.

Thornburg, Hugh J.↗

An Extension of the Athena++ Code Framework for Radiation-magnetohydrodynamics in General Relativity Using a Finite-solid-angle Discretization

We extend the general-relativistic magnetohydrodynamics (GRMHD) capabilities of Athena++ to incorporate radiation. The intensity field in each finite-volume cell is discretized in angle, with explicit transport in both space and angle properly accounting for the effects of gravity on null geodesics, and with matter and radiation coupled in a locally implicit fashion. Here we describe the numerical procedure in detail, verifying its correctness with a suite of tests. Motivated in particular by black hole accretion in the high-accretion-rate, thin-disk regime, we demonstrate the application of the method to this problem. With excellent scaling on flagship computing clusters, the port of the algorithm to the GPU-enabled AthenaK code now allows the simulation of many previously intractable radiation-GRMHD systems.

79 ASTRONOMY AND ASTROPHYSICS↗

VoroClust: Scalable Clustering for Remote Sensing

Although supervised machine learning provides a powerful framework for image classification and segmentation, it requires comprehensive consistent datasets, which are not available for many remote-sensing applications. Remote-sensing datasets are expensive to collect, and each is acquired under different environmental conditions or with significant variations in system operating parameters. Unsupervised clustering algorithms analyze the structure of each dataset independently, rather than drawing on similarities with existing “training” examples, and are thus well suited for practical remote-sensing applications. We introduce VoroClust, a fast density-based unsupervised clustering algorithm applicable to high-resolution and high-dimensional data. VoroClust runs as fast as distance-based clustering methods, while capturing complex regional geometries at least as well as current-density-based methods. It uses a data-centered sphere cover to reduce computational demands, while still capturing data topology. It then propagates clusters outward from local peaks in density. We show that VoroClust provides fast state-of-the-art clustering for both high-resolution polarimetric synthetic aperture radar and high-dimensional hyperspectral imaging datasets.

42 ENGINEERING↗

Transonic Navier-Stokes computations of strake-generated vortex interactions for a fighter-like configuration

Transonic Euler/Navier-Stokes computations are accomplished for wing-body flow fields using a computer program called Transonic Navier-Stokes (TNS). The wing-body grids are generated using a program called ZONER, which subdivides a coarse grid about a fighter-like aircraft configuration into smaller zones, which are tailored to local grid requirements. These zones can be either finely clustered for capture of viscous effects, or coarsely clustered for inviscid portions of the flow field. Different equation sets may be solved in the different zone types. This modular approach also affords the opportunity to modify a local region of the grid without recomputing the global grid. This capability speeds up the design optimization process when quick modifications to the geometry definition are desired. The solution algorithm embodied in TNS is implicit, and is capable of capturing pressure gradients associated with shocks. The algebraic turbulence model employed has proven adequate for viscous interactions with moderate separation. Results confirm that the TNS program can successfully be used to simulate transonic viscous flows about complicated 3-D geometries.

Reznick, Steve↗

NASA’s Prototype Spectral Water Inversion Processor and Emulator (SWIPE): Towards Global Coastal and Inland Water Quality and Algal Biodiversity Monitoring

Degradation of Earth’s inland water resources due to anthropogenic perturbations and climate anomalies at both local and global scales continues to place human health at substantial risk. There is now a growing necessity to develop pragmatic approaches that allow timely and effective extrapolation of local processes, to spatially resolved global products, and to promote operational and sustainable resource policy management. This presentation will provide updates on NASA’s prototype open-source aquatic modeling platform, Spectral Water Inversion Processor and Emulator (SWIPE), which is a comprehensive, multi-faceted modeling platform for both forward and inverse modeling of diverse aquatic ecosystems from the benthos to top-of-atmosphere (TOA). SWIPE provides a cohesive application which leverages recent advancements in particle modeling, Big Data analytics, and machine learning to develop a high-fidelity synthetic training ground for sensitivity studies and algorithm development for multispectral or upcoming hyperspectral missions. Some of the prominent features of SWIPE to be discussed include: 1. Advanced hyperspectral modeling of globally diverse algal and non-algal particles using a novel two-layer coated sphere scattering model and radiative transfer modeling, 2. Massive, highly detailed synthetic spectral libraries of Analysis-Ready-Data (ARD) which include spectral libraries of particle microphysics, water biogeophysical and optical properties, as well as surface and TOA reflectances at 1 nm resolution, 3. An ensemble of pre-built analytic, machine learning, and deep learning inversion algorithms for various water quality and biodiversity related retrieval parameters and uncertainty quantification, 4. Sensor-agnostic water quality inversion at wide ranging spatial and spectral resolutions including a codebase for seamless application in the Google Earth Engine and NASA Earth Exchange (NEX) for planetary scale analysis. SWIPE will be a fully open-source platform based in python with comprehensive documentation, tutorials, and options for distributed computing on high performance computing clusters or on single, local machines. Further, we will discuss how we envision SWIPE contributing towards a global analysis of coastal and inland water quality dynamics.

top-of-atmosphere (TOA)↗

Scalable, In-situ Data Clustering Data Analysis for Extreme Scale Scientific Computing (Final Report)

The objective of this project is to address challenges in the design and development of scalable in-situ data clustering and analytics algorithms and software. Our goal is to develop parallel software consisting of a set of spatio-temporal data clustering and anomaly detection functions, both of which are very important for large-scale analysis and have wide applicability for in-situ runs as well as post-processing analysis. Our design principles for in-situ analysis consider the following: (1) identify parts of the computation can be done close to the data within the nodes, while it is still in memory; (2) extract analysis components can (and should) be performed in remote staging and analysis nodes; (3) develop error-bound approximation methods for applications tolerable for small errors; (4) identify the type of derived distributions and statistics, for spatio-temporal data, that can be kept locally in order to both accelerate computations and meet energy constraints in subsequent iterations and phases; (5) use a self-describing data format so that data can be consistent and understood among local storage (memory and SSDs) and at staging and analysis nodes, thereby providing portability and flexibility; (6) develop service-oriented functions that can schedule in-situ and post-hoc analysis tasks based on the dynamic requirements of applications. Our development focus is to produce the parallel data analysis software/library that will be scalable, reusable, extensible, and generic for applications in different disciplines. The software will be able to run in-situ with the simulations as well as post-hoc analysis. This approach will satisfy many synergistic requirements for data intensive applications executed on data coming from instruments and experiments. In particular, the proposed multilevel approach is directly applicable to perform design tradeoffs for running part of the algorithms near the instruments and the rest on remote (analysis) systems.

97 MATHEMATICS AND COMPUTING↗

Extracting Independent Local Oscillatory Geophysical Signals by Geodetic Tropospheric Delay

Zenith Tropospheric Delay (ZTD) due to water vapor derived from space geodetic techniques and numerical weather prediction simulated-reanalysis data exhibits non-linear and non-stationary properties akin to those in the crucial geophysical signals of interest to the research community. These time series, once decomposed into additive (and stochastic) components, have information about the long term global change (the trend) and other interpretable (quasi-) periodic components such as seasonal cycles and noise. Such stochastic component(s) could be a function that exhibits at most one extremum within a data span or a monotonic function within a certain temporal span. In this contribution, we examine the use of the combined Ensemble Empirical Mode Decomposition (EEMD) and Independent Component Analysis (ICA): the EEMD-ICA algorithm to extract the independent local oscillatory stochastic components in the tropospheric delay derived from the European Centre for Medium-Range Weather Forecasts (ECMWF) over six geodetic sites (HartRAO, Hobart26, Wettzell, Gilcreek, Westford, and Tsukub32). The proposed methodology allows independent geophysical processes to be extracted and assessed. Analysis of the quality index of the Independent Components (ICs) derived for each cluster of local oscillatory components (also called the Intrinsic Mode Functions (IMFs)) for all the geodetic stations considered in the study demonstrate that they are strongly site dependent. Such strong dependency seems to suggest that the localized geophysical signals embedded in the ZTD over the geodetic sites are not correlated. Further, from the viewpoint of non-linear dynamical systems, four geophysical signals the Quasi-Biennial Oscillation (QBO) index derived from the NCEP/NCAR reanalysis, the Southern Oscillation Index (SOI) anomaly from NCEP, the SIDC monthly Sun Spot Number (SSN), and the Length of Day (LoD) are linked to the extracted signal components from ZTD. Results from the synchronization analysis show that ZTD and the geophysical signals exhibit (albeit subtle) site dependent phase synchronization index.

Botai, O. J.↗

A new self-adaptive reconstruction method to identify defects through Wigner–Seitz approach

A new self-adaptive reconstruction method based on local atomic structure at any given molecular dynamics (MD) step has been developed in this article. The method can be used in Wigner–Seitz defect analysis approach to correctly and efficiently explore the information of both point defects and complex defect clusters (e.g. dislocation loops and voids) formed after a displacement cascade where the cascade interacts with grain boundaries and/or dislocations. The algorithm and validation are provided in detail. Results for identification of radiation defects during and after cascades interacting with a dislocation network show that the new method can well recognize all simple and complex defects and defect clusters. Thus, this new method provides a totally new way to explore the density and size of radiation defects at atomic scale after complex MD evolution processes, providing correct information to understand and predict radiation damage in materials through atomic simulations.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

An Algorithm for Characterization of Fiber Aggregation in Composite Microstructures

Composite structures are susceptible to localized flaws, or variability, that drive global failure. To capture this variance, a multiscale model needs to be introduced that not only accurately represents the statistical nature of the composite microstructure but is also efficient. At the microscale, fiber aggregation creates local stress concentrations, where failure is likely to occur sooner than expected. A method of rapid microstructure generation, in conjunction with cluster characterization, is examined to develop accurate reproductions of 2D composite cross-sections. Parameterization of shape and size of fiber clusters is used to characterize representative volume elements. The discrete element method is used to generate pseudo-microstructures. The clustering parameters from the pseudo- and actual microstructure arrangements (obtained from micrographs) can be compared to determine the validity of the representative volume element generated with the discrete element method. Results show a promising approach to evaluating randomness of fiber distributions and how to accurately recreate microstructures for strength analysis.

Carbon fiber↗

Distributed Tomographic Reconstruction with Quantization

Conventional tomographic reconstruction typically depends on centralized servers for both data storage and computation, leading to concerns about memory limitations and data privacy. Distributed reconstruction algorithms mitigate these issues by partitioning data across multiple nodes, reducing server load and enhancing privacy. However, these algorithms often encounter challenges related to memory constraints and communication overhead between nodes. In this paper, we introduce a decentralized Alternating Directions Method of Multipliers (ADMM) with configurable quantization. By distributing local objectives across nodes, our approach is highly scalable and can efficiently reconstruct images while adapting to available resources. To overcome communication bottlenecks, we propose two quantization techniques based on K-means clustering and JPEG compression. Numerical experiments with benchmark images illustrate the tradeoffs between communication efficiency, memory use, and reconstruction accuracy.

Miao, Runxuan↗

NuLattice: Ab initio computations of atomic nuclei on lattices

Here, we introduce NuLattice, a Python software package for ab initio computations of atomic nuclei on lattices. The computational tools consist of Hartree Fock, the coupled cluster method, the in-medium similarity renormalization group, and full configuration interaction. At present, the employed interactions are from pion-less effective field theory at leading order and consist of two-body and three-body contacts. We present results for light nuclei 2 H, 3,4 He, 8 Be, 12 C, and 16 O. NuLattice algorithms exploit the sparsity and locality of lattice interactions, and as a result computations can be run on laptops.

Rothman, Maxwell [Univ. of Tennessee, Knoxville, T↗

GPU acceleration of Swendsen–Wang dynamics

When simulating a lattice system near its critical temperature, local algorithms for modeling the system’s evolution can introduce very large autocorrelation times into sampled data. Here, this critical slowing down places restrictions on the analysis that can be completed in a timely manner of the behavior of systems around the critical point. Because it is often desirable to study such systems around this point, a new algorithm must be introduced. Therefore, we turn to cluster algorithms, such as the Swendsen–Wang algorithm and the Wolff clustering algorithm. They incorporate global updates which generate new lattice configurations with little correlation to previous states, even near the critical point. We look to accelerate the rate at which these algorithm are capable of running by implementing and benchmarking a parallel implementation of each algorithm designed to run on GPUs under NVIDIA’s CUDA framework. A 17 and 90 fold increase in the computational rate was, respectively, experienced when measured against the equivalent algorithm implemented in serial code.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A Spatiotemporal Indexing Approach for Efficient Processing of Big Array-Based Climate Data with MapReduce

Climate observations and model simulations are producing vast amounts of array-based spatiotemporal data. Efficient processing of these data is essential for assessing global challenges such as climate change, natural disasters, and diseases. This is challenging not only because of the large data volume, but also because of the intrinsic high-dimensional nature of geoscience data. To tackle this challenge, we propose a spatiotemporal indexing approach to efficiently manage and process big climate data with MapReduce in a highly scalable environment. Using this approach, big climate data are directly stored in a Hadoop Distributed File System in its original, native file format. A spatiotemporal index is built to bridge the logical array-based data model and the physical data layout, which enables fast data retrieval when performing spatiotemporal queries. Based on the index, a data-partitioning algorithm is applied to enable MapReduce to achieve high data locality, as well as balancing the workload. The proposed indexing approach is evaluated using the National Aeronautics and Space Administration (NASA) Modern-Era Retrospective Analysis for Research and Applications (MERRA) climate reanalysis dataset. The experimental results show that the index can significantly accelerate querying and processing (10 speedup compared to the baseline test using the same computing cluster), while keeping the index-to-data ratio small (0.0328). The applicability of the indexing approach is demonstrated by a climate anomaly detection deployed on a NASA Hadoop cluster. This approach is also able to support efficient processing of general array-based spatiotemporal data in various geoscience domains without special configuration on a Hadoop cluster.

big data↗