Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “algorithms and data structure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

On the suitability of the connection machine for direct particle simulation

The algorithmic structure was examined of the vectorizable Stanford particle simulation (SPS) method and the structure is reformulated in data parallel form. Some of the SPS algorithms can be directly translated to data parallel, but several of the vectorizable algorithms have no direct data parallel equivalent. This requires the development of new, strictly data parallel algorithms. In particular, a new sorting algorithm is developed to identify collision candidates in the simulation and a master/slave algorithm is developed to minimize communication cost in large table look up. Validation of the method is undertaken through test calculations for thermal relaxation of a gas, shock wave profiles, and shock reflection from a stationary wall. A qualitative measure is provided of the performance of the Connection Machine for direct particle simulation. The massively parallel architecture of the Connection Machine is found quite suitable for this type of calculation. However, there are difficulties in taking full advantage of this architecture because of lack of a broad based tradition of data parallel programming. An important outcome of this work has been new data parallel algorithms specifically of use for direct particle simulation but which also expand the data parallel diction.

Dagum, Leonard↗

Recent advances and progress towards an integrated interdisciplinary thermal-structural finite element technology

An integrated finite element approach is presented for interdisciplinary thermal-structural problems. Of the various numerical approaches, finite element methods with direct time integration procedures are most widely used for these nonlinear problems. Traditionally, combined thermal-structural analysis is performed sequentially by transferring data between thermal and structural analysis. This approach is generally effective and routinely used. However, to solve the combined thermal-structural problems, this approach results in cumbersome data transfer, incompatible algorithmic representations, and different discretized element formulations. The integrated approach discussed in this paper effectively combines thermal and structural fields, thus overcoming the above major shortcomings. The approach follows Lax-Wendroff type finite element formulations with flux and stress based representations. As a consequence, this integrated approach uses common algorithmic representations and element formulations. Illustrative test examples show that the approach is effective for integrated thermal-structural problems.

Namburu, Raju R.↗

Sensitivity of Latent Heating Profiles to Environmental Conditions: Implications for TRMM and Climate Research

The Tropical Rainfall Measuring Mission (TRMM) as a part of NASA's Earth System Enterprise is the first mission dedicated to measuring tropical rainfall through microwave and visible sensors, and includes the first spaceborne rain radar. Tropical rainfall comprises two-thirds of global rainfall. It is also the primary distributor of heat through the atmosphere's circulation. It is this circulation that defines Earth's weather and climate. Understanding rainfall and its variability is crucial to understanding and predicting global climate change. Weather and climate models need an accurate assessment of the latent heating released as tropical rainfall occurs. Currently, cloud model-based algorithms are used to derive latent heating based on rainfall structure. Ultimately, these algorithms can be applied to actual data from TRMM. This study investigates key underlying assumptions used in developing the latent heating algorithms. For example, the standard algorithm is highly dependent on a system's rainfall amount and structure. It also depends on an a priori database of model-derived latent heating profiles based on the aforementioned rainfall characteristics. Unanswered questions remain concerning the sensitivity of latent heating profiles to environmental conditions (both thermodynamic and kinematic), regionality, and seasonality. This study investigates and quantifies such sensitivities and seeks to determine the optimal latent heating profile database based on the results. Ultimately, the study seeks to produce an optimized latent heating algorithm based not only on rainfall structure but also hydrometeor profiles.

Shepherd, J. Marshall↗

A Frequency-Domain Substructure System Identification Algorithm

A new frequency-domain system identification algorithm is presented for system identification of substructures, such as payloads to be flown aboard the Space Shuttle. In the vibration test, all interface degrees of freedom where the substructure is connected to the carrier structure are either subjected to active excitation or are supported by a test stand with the reaction forces measured. The measured frequency-response data is used to obtain a linear, viscous-damped model with all interface-degree of freedom entries included. This model can then be used to validate analytical substructure models. This procedure makes it possible to obtain not only the fixed-interface modal data associated with a Craig-Bampton substructure model, but also the data associated with constraint modes. With this proposed algorithm, multiple-boundary-condition tests are not required, and test-stand dynamics is accounted for without requiring a separate modal test or finite element modeling of the test stand. Numerical simulations are used in examining the algorithm's ability to estimate valid reduced-order structural models. The algorithm's performance when frequency-response data covering narrow and broad frequency bandwidths is used as input is explored. Its performance when noise is added to the frequency-response data and the use of different least squares solution techniques are also examined. The identified reduced-order models are also compared for accuracy with other test-analysis models and a formulation for a Craig-Bampton test-analysis model is also presented.

Blades, Eric L.↗

Trajectory design via unsupervised probabilistic learning on optimal manifolds

Abstract This article illustrates the use of unsupervised probabilistic learning techniques for the analysis of planetary reentry trajectories. A three-degree-of-freedom model was employed to generate optimal trajectories that comprise the training datasets. The algorithm first extracts the intrinsic structure in the data via a diffusion map approach. We find that data resides on manifolds of much lower dimensionality compared to the high-dimensional state space that describes each trajectory. Using the diffusion coordinates on the graph of training samples, the probabilistic framework subsequently augments the original data with samples that are statistically consistent with the original set. The augmented samples are then used to construct conditional statistics that are ultimately assembled in a path planning algorithm. In this framework, the controls are determined stage by stage during the flight to adapt to changing mission objectives in real-time.

42 ENGINEERING↗

Model Structures and Algorithms for Identification of Aerodynamic Models for Flight Dynamics Applications

This paper describes model structures and parameter estimation algorithms suitable for the identification of unsteady aerodynamic models from input-output data. The model structures presented are state space models and include linear time-invariant (LTI) models and linear parameter-varying (LPV) models. They cover a wide range of local and parameter dependent identification problems arising in unsteady aerodynamics and nonlinear flight dynamics. We present a residue algorithm for estimating model parameters from data. The algorithm can incorporate apriori information and is described in detail. The algorithms are evaluated on the F-16XL wind-tunnel test data from NAS Langley Research Center. Results of numerical evaluation are presented. The paper concludes with a discussion major issues and directions for future work.

Prasanth, Ravi K.↗

Photogrammetric Metrology for the James Webb Space Telescope Integrated Science Instrument Module

The James Webb Space Telescope (JWST) is a 6.6m diameter, segmented, deployable telescope for cryogenic IR space astronomy (approximately 40K). The JWST Observatory architecture includes the Optical Telescope Element and the Integrated Science Instrument Module (ISIM) element that contains four science instruments (SI) including a Guider. The ISM optical metering structure is a roughly 2.2x1.7x2.2m, asymmetric frame that is composed of carbon fiber and resin tubes bonded to invar end fittings and composite gussets and clips. The structure supports the SIs, isolates the SIs from the OTE, and supports thermal and electrical subsystems. The structure is attached to the OTE structure via strut-like kinematic mounts. The ISIM structure must meet its requirements at the approximately 40K cryogenic operating temperature. The SIs are aligned to the structure's coordinate system under ambient, clean room conditions using laser tracker and theodolite metrology. The ISIM structure is thermally cycled for stress relief and in order to measure temperature-induced mechanical, structural changes. These ambient-to-cryogenic changes in the alignment of SI and OTE-related interfaces are an important component in the JWST Observatory alignment plan and must be verified. We report on the planning for and preliminary testing of a cryogenic metrology system for ISIM based on photogrammetry. Photogrammetry is the measurement of the location of custom targets via triangulation using images obtained at a suite of digital camera locations and orientations. We describe metrology system requirements, plans, and ambient photogrammetric measurements of a mock-up of the ISIM structure to design targeting and obtain resolution estimates. We compare these measurements with those taken from a well known ambient metrology system, namely, the Leica laser tracker system. We also describe the data reduction algorithm planned to interpret cryogenic data from the Flight structure. Photogrammetry was selected from an informal trade study of cryogenic metrology systems because its resolution meets sub-allocations to ISIM alignment requirements and it is a non-contact method that can in principle measure six degrees of freedom changes in target location. In addition, photogrammetry targets can be readily related to targets used for ambient surveys of the structure. By thermally isolating the photogrammetry camera during testing, metrology can be performed in situ during thermal cycling. Photogrammetry also has a small but significant cryogenic heritage in astronomical instrumentation metrology. It was used to validate the displacement/deformation predictions of the reflectors and the feed horns during thermal/vacuum testing (90K) for the Microwave Anisotropy Probe (MAP). It also was used during thermal vacuum testing (100K) to verify shape and component alignment at operational temperature of the High Gain Antenna for New Horizons. With tighter alignment requirements and lower operating temperatures than the aforementioned observatories, ISIM presents new challenges in the development of this metrology system.

Nowak, Maria↗

Refining HPCToolkit for application performance analysis at exascale

As part of the US Department of Energy’s Exascale Computing Project (ECP), Rice University has been refining its HPCToolkit performance tools to better support measurement and analysis of applications executing on exascale supercomputers. To efficiently collect performance measurements of GPU-accelerated applications, HPCToolkit employs novel non-blocking data structures to communicate performance measurements between tool threads and application threads. To attribute performance information in detail to source lines, loop nests, and inlined call chains, HPCToolkit performs parallel analysis of large CPU and GPU binaries involved in the execution of an exascale application to rapidly recover mappings between machine instructions and source code. To analyze terabytes of performance measurements gathered during executions at exascale, HPCToolkit employs distributed-memory parallelism, multithreading, sparse data structures, and out-of-core streaming analysis algorithms. To support interactive exploration of profiles up to terabytes in size, HPCToolkit’s hpcviewer graphical user interface uses out-of-core methods to visualize performance data. The result of these efforts is that HPCToolkit now supports collection, analysis, and presentation of profiles and traces of GPU-accelerated applications at exascale. These improvements have enabled HPCToolkit to efficiently measure, analyze and explore terabytes of performance data for executions using as many as 64K MPI ranks and 64K GPU tiles on ORNL’s Frontier supercomputer. HPCToolkit’s support for measurement and analysis of GPU-accelerated applications has been employed to study a collection of open-science applications developed as part of ECP. This paper reports on these experiences, which provided insight into opportunities for tuning applications, strengths and weaknesses of HPCToolkit itself, as well as unexpected behaviors in executions at exascale.

Adhianto, Laksono↗

Autonomous Polycrystalline Material Decomposition For Hyperspectral Neutron Tomography

Hyperspectral neutron tomography is an effective method for analyzing crystalline material samples with complex compositions in a non-destructive manner. Since the counts in the hyperspectral neutron radiographs directly depend on the neutron cross-sections, materials may exhibit contrasting neutron responses across wavelengths. Therefore, it is possible to extract the unique signatures associated with each material and use them to separate the crystalline phases simultaneously.We introduce an autonomous material decomposition (AMD) algorithm to automatically characterize and localize polycrystalline structures using Bragg edges with contrasting neutron responses from hyperspectral data. The algorithm estimates the linear attenuation coefficient spectra from the measured radiographs and then uses these spectra to perform polycrystalline material decomposition and reconstructs 3D material volumes to localize materials in the spatial domain. Our results demonstrate that the method can accurately estimate both the linear attenuation coefficient spectra and associated reconstructions on both simulated and experimental neutron data.

Samin nur chowdhury, Mohammad↗

Development of Steady-State and Dynamic Mass and Energy Constrained Neural Networks for Distributed Chemical Systems Using Noisy Transient Data

The paper presents the development of algorithms for mass and energy constrained neural network models that can exactly conserve the overall mass and energy of distributed chemical process systems, even though the noisy transient data used for optimal model training violate the same. In contrast to approximately satisfying mass and energy balance constraints of a system by soft penalization of objective function, algorithms have been developed for solving equality-constrained nonlinear optimization problems, thus providing the guarantee of exactly satisfying the system mass and energy conservation laws. For developing dynamic mass-energy constrained network models for distributed systems, hybrid series and parallel dynamic-static neural networks have been leveraged. The developed algorithms for solving both the training and forward problems are validated using both steady-state and dynamic data in the presence of various noise characteristics. The developed data-driven algorithms are flexible to exactly satisfy mass and energy balance constraints for dynamic chemical processes if the system holdup information is available. The proposed network structures and algorithms are applied to the development of data-driven lumped and distributed models of an adiabatic superheater/reheater system, a nonisothermal continuous stirred tank reactor, as well as an electrically heated plug-flow reactor system where one form of energy gets transformed to another. It has been observed that the mass-energy constrained neural networks yield a root mean squared error of <1% with respect to the system truth for the case studies evaluated in this work.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Performance issues for iterative solvers in device simulation

Due to memory limitations, iterative methods have become the method of choice for large scale semiconductor device simulation. However, it is well known that these methods still suffer from reliability problems. The linear systems which appear in numerical simulation of semiconductor devices are notoriously ill-conditioned. In order to produce robust algorithms for practical problems, careful attention must be given to many implementation issues. This paper concentrates on strategies for developing robust preconditioners. In addition, effective data structures and convergence check issues are also discussed. These algorithms are compared with a standard direct sparse matrix solver on a variety of problems.

Fan, Qing↗

VIIRS Deep Blue Aerosol Products Over Land: Extending the EOS Long‐Term Aerosol Data Records

A primary goal of the Deep Blue (DB) project is to create consistent long‐term aerosol data records, suitable for climate studies, using multiple satellite instruments. In order to continue Earth Observing System (EOS)‐era aerosol products into the Joint Polar Satellite System era, we have successfully ported the DB algorithm to process data from the Visible Infrared Imaging Radiometer Suite (VIIRS). Although the basic structure of the VIIRS algorithm is similar to that for the Moderate Resolution Imaging Spectroradiometer (MODIS), many enhancements have been made compared to the MODIS collection 6 (C6) version. Most have also been implemented in the latest MODIS Collection 6.1 (C6.1). For example, a new smoke mask was developed based on the spectral curvature of measured reflectance to distinguish biomass burning smoke from weakly absorbing urban/industrial aerosols. Consequently, a new aerosol‐type flag was added into the VIIRS DB data set. In addition, new dust models have been developed to account for the nonsphericity of mineral dust. As a result, a discontinuity in the retrieved aerosol optical depth (AOD) of Saharan dust plumes seen in MODIS C6 products near the boundary between North Africa and the Atlantic has been much reduced. We have also evaluated the VIIRS and MODIS Terra/Aqua C6.1 AOD against Aerosol Robotic Network data. VIIRS and MODIS retrievals show similar performance; around 80% of matchups agree with Aerosol Robotic Network within the expected error of ±(0.05 + 20)%, indicating that DB can provide consistent AOD through the historical EOS and present Joint Polar Satellite System eras.

aerosols↗

A statistical evaluation and comparison of VISSR Atmospheric Sounder (VAS) data

In order to account for the temporal and spatial discrepancies between the VAS and rawinsonde soundings, the rawinsonde data were adjusted to a common hour of release where the new observation time corresponded to the satellite scan time. Both the satellite and rawinsonde observations of the basic atmospheric parameters (T Td, and Z) were objectively analyzed to a uniform grid maintaining the same mesoscale structure in each data set. The performance of each retrieval algorithm in producing accurate and representative soundings was evaluated using statistical parameters such as the mean, standard deviation, and root mean square of the difference fields for each parameter and grid level. Horizontal structure was also qualitatively evaluated by examining atmospheric features on constant pressure surfaces. An analysis of the vertical structure of the atmosphere were also performed by looking at colocated and grid mean vertical profiles of both the satellite and rawinsonde data sets. Highlights of these results are presented.

Jedlovec, G. J.↗

Get Non-Real: Randomized Sketching for High-Dimensional Non-Real Valued Data (Final Report)

In our final report for DE-C0022186, we describe the work we did on this grant towards the goals we proposed. Our first goal was characterizing fundamental limits for sketching of discrete high-dimensional matrices with low-dimensional structures. Our second main goal was designing algorithms for data reconstruction from sketches. We focus on approaches that are either specifically designed for non-real-valued data (binary, finite field) or that will translate more readily to that setting.

97 MATHEMATICS AND COMPUTING↗

Data Summarization and Inference at Scale

This is the final report for the DOE ASCR grant SC-0022260, Data Summarization and Inference at Scale, PI: Alex Pothen, Purdue University. The goal of the project was to solve data-intensive and compute-intensive problems in the physical sciences, engineering, information science, data science, etc. by designing and implementing new algorithms that could work with a subset of the data. The four subgoals were: (a) The solution of problems where the data is too large to be stored in the memory of a computer. In this streaming model of computation, the data arrives as a stream of elements to the computer, each element is processed as it arrives, and a decision is made to discard the data or to store it; only a small subset of the data proportional to the size of the output solution is stored, and when all the data has been streamed, a solution to the problem is computed from the stored subset. (b) The use of machine learning methods to compute solutions to data-intensive problems. The use of GPUs is critical to obtain high performance on machine learning tasks, but their memory sizes are smaller relative to that of CPUs. For large-scale problems, the data is sampled many times, and small samples are used with repetition, for robustness, to compute solutions to inference tasks. This sampling reduces the memory required to solve the problem, but attention is needed to avoid slow convergence to the solutions, and reduced accuracy of inference. We propose submodular optimization, Large Language Models, and physics-informed neural networks to enable GPU computations here. (c) Modeling and visualization of high-dimensional data using interpretable features. Clinical proteomic data sets from immunology for the detection of cancer and other diseases are temporal and high-dimensional, and algorithms for visualizing these data sets using clinically interpretable features are lacking. We propose methods that compute distances based on the optimal transportation problem and graph edit distances to address this problem. We also propose the use of optimal transport-based distances, spatial statistics, and network structure to classify image data sets, We apply these algorithms to electron micrographs of the peripheral nervous system in the digestive tract. (d) The design of data-intensive algorithms on emerging architectures, specifically, noisy, intermediate-scale quantum (NISQ) devices. Quantum computers offer the possibility of exploring large solution spaces due to the principle of superposition, but current quantum computers are limited by few qubits, short coherence times due to noise, poor interconections among the qubits, etc. We propose the use of the divide and conquer paradigm to solve large-scale problems, wherein collections of small subproblems are solved on the quantum devices, and the solutions to the subproblems are integrated into a solution for the original problem on a classical computer.

97 MATHEMATICS AND COMPUTING↗

Optimal sensor placement for reconstructing wind pressure field around buildings using compressed sensing

Deciding how to optimally deploy sensors in a large, complex, and spatially extended structure is critical to ensure that the surface pressure field is accurately captured for subsequent analysis and design. In some cases, reconstruction of missing data is required in downstream tasks such as the development of digital twins. Here, this paper presents a data-driven sparse sensor selection algorithm, aiming to provide the most information contents for reconstructing aerodynamic characteristics of wind pressures over tall building structures parsimoniously. The algorithm first fits a set of basis functions to the training data, then applies a computationally efficient QR algorithm that ranks existing pressure sensors in order of importance based on the state reconstruction to this tailored basis. The findings of this study show that the proposed algorithm successfully re- constructs the aerodynamic characteristics of tall buildings from sparse measurement locations, generating stable and optimal solutions across a range of conditions. As a result, this study serves as a promising first step toward leveraging the success of data-driven and machine learning algorithms to supplement traditional genetic algorithms currently used in wind engineering.

42 ENGINEERING↗

The CSM testbed matrix processors internal logic and dataflow descriptions

This report constitutes the final report for subtask 1 of Task 5 of NASA Contract NAS1-18444, Computational Structural Mechanics (CSM) Research. This report contains a detailed description of the coded workings of selected CSM Testbed matrix processors (i.e., TOPO, K, INV, SSOL) and of the arithmetic utility processor AUS. These processors and the current sparse matrix data structures are studied and documented. Items examined include: details of the data structures, interdependence of data structures, data-blocking logic in the data structures, processor data flow and architecture, and processor algorithmic logic flow.

Regelbrugge, Marc E.↗

The design and implementation of a parallel unstructured Euler solver using software primitives

This paper is concerned with the implementation of a 3D unstructured-grid Euler-solver on massively parallel distributed-memory computer architectures. The goal is to minimize solution time by achieving high computational rates with a numerically efficient algorithm. An unstructured multigrid algorithm with an edge-based data-structure has been adopted, and a number of optimizations have been devised and implemented in order to accelerate the parallel computational rates. The implementation is carried out by creating a set of software tools, which ease the implementation of computational problems on parallel architecture machines by relieving the user of the low-level machine specific issues. The quantitative effect of the various optimizations are demonstrated, and we show that the combined effect of these optimizations leads to roughly a factor of three performance improvement. The overall solution efficiency is compared with that obtained on the CRAY-YMP vector supercomputer.

Das, R.↗