Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Structures and Algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Bayesian Blocks: A New Method to Analyze Photon Counting Data

A Bayesian analysis of photon-counting data leads to a new time-domain algorithm for detecting localized structures (bursts), revealing pulse shapes, and generally characterizing intensity variations. The raw counting data -- time-tag events (TTE), time-to-spill (TTS) data, or binned counts -- is converted to a maximum likelihood segmentation of the observation into time intervals during which the photon arrival rate is perceptibly constant -- i.e. has a fixed intensity without statistically significant variations. The resulting structures, Bayesian Blocks, can be thought of as bins with arbitrary spacing determined by the data. The method itself sets no lower limit to the time scale on which variability can be detected. We have applied the method to RXTE data on Cyg X-1, yielding information on this source's short-time-scale variability.

Scargle, Jeffrey D.↗

Imaging Complex Subsurface Structures for Geothermal Exploration at Pirouette Mountain and Eleven-Mile Canyon in Nevada

Accurate imaging of subsurface complex structures with faults is crucial for geothermal exploration because faults are generally the primary conduit of hydrothermal flow. It is very challenging to image geothermal exploration areas because of complex geologic structures with various faults and noisy surface seismic data with strong and coherent ground-roll noise. In addition, fracture zones and most geologic formations behave as anisotropic media for seismic-wave propagation. Properly suppressing ground-roll noise and accounting for subsurface anisotropic properties are essential for high-resolution imaging of subsurface structures and faults for geothermal exploration. We develop a novel wavenumber-adaptive bandpass filter to suppress the ground-roll noise without affecting useful seismic signals. This filter adaptively exploits both characteristics of the lower frequency and the smaller velocity of the ground-roll noise than those of the signals. Consequently, this filter can effectively differentiate the ground-roll noise from the signal. We use our novel filter to attenuate the ground-roll noise in seismic data along five survey lines acquired by the U.S. Navy Geothermal Program Office at Pirouette Mountain and Eleven-Mile Canyon in Nevada, United States. We then apply our novel anisotropic least-squares reverse-time migration algorithm to the resulting data for imaging subsurface structures at the Pirouette Mountain and Eleven-Mile Canyon geothermal exploration areas. The migration method employs an efficient implicit wavefield-separation scheme to reduce image artifacts and improve the image quality. Our results demonstrate that our wavenumber-adaptive bandpass filtering method successfully suppresses the strong and coherent ground-roll noise in the land seismic data, and our anisotropic least-squares reverse-time migration produces high-resolution subsurface images of Pirouette Mountain and Eleven-Mile Canyon, facilitating accurate fault interpretation for geothermal exploration.

15 GEOTHERMAL ENERGY↗

Development of Steady-State and Dynamic Mass and Energy Constrained Neural Networks for Distributed Chemical Systems Using Noisy Transient Data

The paper presents the development of algorithms for mass and energy constrained neural network models that can exactly conserve the overall mass and energy of distributed chemical process systems, even though the noisy transient data used for optimal model training violate the same. In contrast to approximately satisfying mass and energy balance constraints of a system by soft penalization of objective function, algorithms have been developed for solving equality-constrained nonlinear optimization problems, thus providing the guarantee of exactly satisfying the system mass and energy conservation laws. For developing dynamic mass-energy constrained network models for distributed systems, hybrid series and parallel dynamic-static neural networks have been leveraged. The developed algorithms for solving both the training and forward problems are validated using both steady-state and dynamic data in the presence of various noise characteristics. The developed data-driven algorithms are flexible to exactly satisfy mass and energy balance constraints for dynamic chemical processes if the system holdup information is available. The proposed network structures and algorithms are applied to the development of data-driven lumped and distributed models of an adiabatic superheater/reheater system, a nonisothermal continuous stirred tank reactor, as well as an electrically heated plug-flow reactor system where one form of energy gets transformed to another. It has been observed that the mass-energy constrained neural networks yield a root mean squared error of <1% with respect to the system truth for the case studies evaluated in this work.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Performance issues for iterative solvers in device simulation

Due to memory limitations, iterative methods have become the method of choice for large scale semiconductor device simulation. However, it is well known that these methods still suffer from reliability problems. The linear systems which appear in numerical simulation of semiconductor devices are notoriously ill-conditioned. In order to produce robust algorithms for practical problems, careful attention must be given to many implementation issues. This paper concentrates on strategies for developing robust preconditioners. In addition, effective data structures and convergence check issues are also discussed. These algorithms are compared with a standard direct sparse matrix solver on a variety of problems.

Fan, Qing↗

VIIRS Deep Blue Aerosol Products Over Land: Extending the EOS Long‐Term Aerosol Data Records

A primary goal of the Deep Blue (DB) project is to create consistent long‐term aerosol data records, suitable for climate studies, using multiple satellite instruments. In order to continue Earth Observing System (EOS)‐era aerosol products into the Joint Polar Satellite System era, we have successfully ported the DB algorithm to process data from the Visible Infrared Imaging Radiometer Suite (VIIRS). Although the basic structure of the VIIRS algorithm is similar to that for the Moderate Resolution Imaging Spectroradiometer (MODIS), many enhancements have been made compared to the MODIS collection 6 (C6) version. Most have also been implemented in the latest MODIS Collection 6.1 (C6.1). For example, a new smoke mask was developed based on the spectral curvature of measured reflectance to distinguish biomass burning smoke from weakly absorbing urban/industrial aerosols. Consequently, a new aerosol‐type flag was added into the VIIRS DB data set. In addition, new dust models have been developed to account for the nonsphericity of mineral dust. As a result, a discontinuity in the retrieved aerosol optical depth (AOD) of Saharan dust plumes seen in MODIS C6 products near the boundary between North Africa and the Atlantic has been much reduced. We have also evaluated the VIIRS and MODIS Terra/Aqua C6.1 AOD against Aerosol Robotic Network data. VIIRS and MODIS retrievals show similar performance; around 80% of matchups agree with Aerosol Robotic Network within the expected error of ±(0.05 + 20)%, indicating that DB can provide consistent AOD through the historical EOS and present Joint Polar Satellite System eras.

aerosols↗

Get Non-Real: Randomized Sketching for High-Dimensional Non-Real Valued Data (Final Report)

In our final report for DE-C0022186, we describe the work we did on this grant towards the goals we proposed. Our first goal was characterizing fundamental limits for sketching of discrete high-dimensional matrices with low-dimensional structures. Our second main goal was designing algorithms for data reconstruction from sketches. We focus on approaches that are either specifically designed for non-real-valued data (binary, finite field) or that will translate more readily to that setting.

97 MATHEMATICS AND COMPUTING↗

Data Summarization and Inference at Scale

This is the final report for the DOE ASCR grant SC-0022260, Data Summarization and Inference at Scale, PI: Alex Pothen, Purdue University. The goal of the project was to solve data-intensive and compute-intensive problems in the physical sciences, engineering, information science, data science, etc. by designing and implementing new algorithms that could work with a subset of the data. The four subgoals were: (a) The solution of problems where the data is too large to be stored in the memory of a computer. In this streaming model of computation, the data arrives as a stream of elements to the computer, each element is processed as it arrives, and a decision is made to discard the data or to store it; only a small subset of the data proportional to the size of the output solution is stored, and when all the data has been streamed, a solution to the problem is computed from the stored subset. (b) The use of machine learning methods to compute solutions to data-intensive problems. The use of GPUs is critical to obtain high performance on machine learning tasks, but their memory sizes are smaller relative to that of CPUs. For large-scale problems, the data is sampled many times, and small samples are used with repetition, for robustness, to compute solutions to inference tasks. This sampling reduces the memory required to solve the problem, but attention is needed to avoid slow convergence to the solutions, and reduced accuracy of inference. We propose submodular optimization, Large Language Models, and physics-informed neural networks to enable GPU computations here. (c) Modeling and visualization of high-dimensional data using interpretable features. Clinical proteomic data sets from immunology for the detection of cancer and other diseases are temporal and high-dimensional, and algorithms for visualizing these data sets using clinically interpretable features are lacking. We propose methods that compute distances based on the optimal transportation problem and graph edit distances to address this problem. We also propose the use of optimal transport-based distances, spatial statistics, and network structure to classify image data sets, We apply these algorithms to electron micrographs of the peripheral nervous system in the digestive tract. (d) The design of data-intensive algorithms on emerging architectures, specifically, noisy, intermediate-scale quantum (NISQ) devices. Quantum computers offer the possibility of exploring large solution spaces due to the principle of superposition, but current quantum computers are limited by few qubits, short coherence times due to noise, poor interconections among the qubits, etc. We propose the use of the divide and conquer paradigm to solve large-scale problems, wherein collections of small subproblems are solved on the quantum devices, and the solutions to the subproblems are integrated into a solution for the original problem on a classical computer.

97 MATHEMATICS AND COMPUTING↗

Robust Group Subspace Recovery: A New Approach for Multi-Modality Data Fusion

Robust Subspace Recovery (RoSuRe) algorithm was recently introduced as a principled and numerically efficient algorithm that unfolds underlying Unions of Subspaces (UoS) structure, present in the data. The union of Subspaces (UoS) is capable of identifying more complex trends in data sets than simple linear models. In this work, we build on and extend RoSuRe to prospect the structure of different data modalities individually. We propose a novel multi-modal data fusion approach based on group sparsity which we refer to as Robust Group Subspace Recovery (RoGSuRe). Relying on a bi-sparsity pursuit paradigm and non-smooth optimization techniques, the introduced framework learns a new joint representation of the time series from different data modalities, respecting an underlying UoS model. We subsequently integrate the obtained structures to form a unified subspace structure. The proposed approach exploits the structural dependencies between the different modalities data to cluster the associated target objects. The resulting fusion of the unlabeled sensors’ data from experiments on audio and magnetic data has shown that our method is competitive with other state of the art subspace clustering methods. The resulting UoS structure is employed to classify newly observed data points, highlighting the abstraction capacity of the proposed method.

47 OTHER INSTRUMENTATION↗

Optimal sensor placement for reconstructing wind pressure field around buildings using compressed sensing

Deciding how to optimally deploy sensors in a large, complex, and spatially extended structure is critical to ensure that the surface pressure field is accurately captured for subsequent analysis and design. In some cases, reconstruction of missing data is required in downstream tasks such as the development of digital twins. Here, this paper presents a data-driven sparse sensor selection algorithm, aiming to provide the most information contents for reconstructing aerodynamic characteristics of wind pressures over tall building structures parsimoniously. The algorithm first fits a set of basis functions to the training data, then applies a computationally efficient QR algorithm that ranks existing pressure sensors in order of importance based on the state reconstruction to this tailored basis. The findings of this study show that the proposed algorithm successfully re- constructs the aerodynamic characteristics of tall buildings from sparse measurement locations, generating stable and optimal solutions across a range of conditions. As a result, this study serves as a promising first step toward leveraging the success of data-driven and machine learning algorithms to supplement traditional genetic algorithms currently used in wind engineering.

42 ENGINEERING↗

The MolSSI QCArchive project: An open-source platform to compute, organize, and share quantum chemistry data

The Molecular Sciences Software Institute's (MolSSI) Quantum Chemistry Archive (QCArchive) project is an umbrella name that covers both a central server hosted by MolSSI for community data and the Python-based software infrastructure that powers automated computation and storage of quantum chemistry (QC) results. The MolSSI-hosted central server provides the computational molecular sciences community a location to freely access tens of millions of QC computations for machine learning, methodology assessment, force-field fitting, and more through a Python interface. Facile, user-friendly mining of the centrally archived quantum chemical data also can be achieved through web applications found at the website. The software infrastructure can be used as a standalone platform to compute, structure, and distribute hundreds of millions of QC computations for individuals or groups of researchers at any scale. The QCArchiveInfrastructure is open-source (BSD-3C), code repositories can be found at github, and releases can be downloaded via PyPI and Conda. This article is categorized under: Electronic Structure Theory > Ab Initio Electronic Structure Methods Software > Quantum Chemistry Data Science > Computer Algorithms and Programming

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Phoenix: A Scalable Streaming Hypergraph Analysis Framework

We present Phoenix, a scalable hypergraph analytics framework for data analytics and knowledge discovery that was implemented on the leadership class computing platforms at Oak Ridge National Laboratory (ORNL). Our software framework comprises a distributed implementation of a streaming server architecture which acts as a gateway for various hypergraph generators/external sources to connect. Phoenix has the capability to utilize diverse hypergraph generators, including HyGen, a very large-scale hypergraph generator developed by ORNL. Phoenix incorporates specific algorithms for efficient data representation by exploiting hidden structures of the hypergraphs. Our experimental results demonstrate Phoenix’s scalable and stable performance on massively parallel computing platforms. Phoenix’s superior performance is due to the merging of high-performance computing with data analytic.

Kurte, Kuldeep↗

The CSM testbed matrix processors internal logic and dataflow descriptions

This report constitutes the final report for subtask 1 of Task 5 of NASA Contract NAS1-18444, Computational Structural Mechanics (CSM) Research. This report contains a detailed description of the coded workings of selected CSM Testbed matrix processors (i.e., TOPO, K, INV, SSOL) and of the arithmetic utility processor AUS. These processors and the current sparse matrix data structures are studied and documented. Items examined include: details of the data structures, interdependence of data structures, data-blocking logic in the data structures, processor data flow and architecture, and processor algorithmic logic flow.

Regelbrugge, Marc E.↗

The design and implementation of a parallel unstructured Euler solver using software primitives

This paper is concerned with the implementation of a 3D unstructured-grid Euler-solver on massively parallel distributed-memory computer architectures. The goal is to minimize solution time by achieving high computational rates with a numerically efficient algorithm. An unstructured multigrid algorithm with an edge-based data-structure has been adopted, and a number of optimizations have been devised and implemented in order to accelerate the parallel computational rates. The implementation is carried out by creating a set of software tools, which ease the implementation of computational problems on parallel architecture machines by relieving the user of the low-level machine specific issues. The quantitative effect of the various optimizations are demonstrated, and we show that the combined effect of these optimizations leads to roughly a factor of three performance improvement. The overall solution efficiency is compared with that obtained on the CRAY-YMP vector supercomputer.

Das, R.↗

A deep learning-enhanced framework for multiphysics joint inversion

Joint inversion has drawn considerable attention due to the availability of multiple geophysical data sets, ever-increasing computational resources, the development of advanced algorithms, and its ability to reduce inversion uncertainty. A key issue of joint inversion is to develop effective strategies to link different geophysical data in a unified mathematical framework, in which the information obtained from different models can complement each other. We have developed a deep learning-enhanced joint inversion framework to simultaneously reconstruct different physical models by fusing different types of geophysical data. Traditionally, structure similarity constraints are pursued by joint inversion algorithms using manually crafted formulations (e.g., cross gradient). The constraint is constructed by a deep neural network (DNN) during the learning process. The framework is designed to combine the DNN and a traditional independent inversion workflow and improve the joint inversion result iteratively. The network can be easily extended to incorporate multiphysics without structural changes. Numerical experiments on the joint inversion of 2D DC resistivity data and seismic traveltime are used to validate our method. In addition, this learning-based framework demonstrates excellent generalization abilities when tested on data sets using different geologic structures. It also can handle different sensing configurations and nonconforming discretization.

Geochemistry & Geophysics↗

Biolink Model: A universal schema for knowledge graphs in clinical, biomedical, and translational science

Abstract Within clinical, biomedical, and translational science, an increasing number of projects are adopting graphs for knowledge representation. Graph‐based data models elucidate the interconnectedness among core biomedical concepts, enable data structures to be easily updated, and support intuitive queries, visualizations, and inference algorithms. However, knowledge discovery across these “knowledge graphs” (KGs) has remained difficult. Data set heterogeneity and complexity; the proliferation of ad hoc data formats; poor compliance with guidelines on findability, accessibility, interoperability, and reusability; and, in particular, the lack of a universally accepted, open‐access model for standardization across biomedical KGs has left the task of reconciling data sources to downstream consumers. Biolink Model is an open‐source data model that can be used to formalize the relationships between data structures in translational science. It incorporates object‐oriented classification and graph‐oriented features. The core of the model is a set of hierarchical, interconnected classes (or categories) and relationships between them (or predicates) representing biomedical entities such as gene, disease, chemical, anatomic structure, and phenotype. The model provides class and edge attributes and associations that guide how entities should relate to one another. Here, we highlight the need for a standardized data model for KGs, describe Biolink Model, and compare it with other models. We demonstrate the utility of Biolink Model in various initiatives, including the Biomedical Data Translator Consortium and the Monarch Initiative, and show how it has supported easier integration and interoperability of biomedical KGs, bringing together knowledge from multiple sources and helping to realize the goals of translational science.

60 APPLIED LIFE SCIENCES↗

G-Mapper: Learning a Cover in the Mapper Construction

The Mapper algorithm is a visualization technique in topological data analysis (TDA) that outputs a graph reflecting the structure of a given dataset. However, the Mapper algorithm requires tuning several parameters in order to generate a “nice” Mapper graph. This paper focuses on selecting the cover parameter. We present an algorithm that optimizes the cover of a Mapper graph by splitting a cover repeatedly according to a statistical test for normality. Our algorithm is based on G-means clustering, which searches for the optimal number of clusters in 𝑘-means by iteratively applying the Anderson–Darling test. Our splitting procedure employs a Gaussian mixture model to carefully choose the cover according to the distribution of the given data. In conclusion, experiments for synthetic and real-world datasets demonstrate that our algorithm generates covers so that the Mapper graphs retain the essence of the datasets, while also running significantly faster than a previous iterative method.

G-means clustering↗

Data reduction using cubic rational B-splines

A geometric method is proposed for fitting rational cubic B-spline curves to data that represent smooth curves including intersection or silhouette lines. The algorithm is based on the convex hull and the variation diminishing properties of Bezier/B-spline curves. The algorithm has the following structure: it tries to fit one Bezier segment to the entire data set and if it is impossible it subdivides the data set and reconsiders the subset. After accepting the subset the algorithm tries to find the longest run of points within a tolerance and then approximates this set with a Bezier cubic segment. The algorithm uses this procedure repeatedly to the rest of the data points until all points are fitted. It is concluded that the algorithm delivers fitting curves which approximate the data with high accuracy even in cases with large tolerances.

Chou, Jin J.↗