Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Scarce data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Leveraging explainable AI to characterize floating-point exceptions in linear solvers

Linear solver packages are central to many scientific, engineering, and machine learning applications. When floating-point exceptions occur in these solvers, e.g., division by zero or overflow, numerical results are compromised and become unreliable. Existing static and dynamic analysis tools can detect such exceptions, but they do not explain why the exceptions occur in terms of the solver inputs. Here, we present a study to characterize the inputs that cause numerical exceptions in linear solver packages. Our approach uses explainable AI (XAI) to find the most relevant characteristics of input matrices that explain the occurrence of exceptions in the solvers. Since training data in this domain is scarce, we perform extensive data gathering and data augmentation to obtain exception-inducing inputs. Our approach uses a repair strategy on the features blamed by XAI to validate that such features indeed explain the exceptions. We compare the LIME and SHAP XAI techniques using a dozen matrix features with three classifiers. We evaluate the approach on three widely used linear solver packages and find that some input characteristics can explain the occurrence of exceptions 100% of the time, in specific solvers and preconditioners.

Explainable AI↗

A Fast Implementation of the ISOCLUS Algorithm

Unsupervised clustering is a fundamental tool in numerous image processing and remote sensing applications. For example, unsupervised clustering is often used to obtain vegetation maps of an area of interest. This approach is useful when reliable training data are either scarce or expensive, and when relatively little a priori information about the data is available. Unsupervised clustering methods play a significant role in the pursuit of unsupervised classification. One of the most popular and widely used clustering schemes for remote sensing applications is the ISOCLUS algorithm, which is based on the ISODATA method. The algorithm is given a set of n data points (or samples) in d-dimensional space, an integer k indicating the initial number of clusters, and a number of additional parameters. The general goal is to compute a set of cluster centers in d-space. Although there is no specific optimization criterion, the algorithm is similar in spirit to the well known k-means clustering method in which the objective is to minimize the average squared distance of each point to its nearest center, called the average distortion. One significant feature of ISOCLUS over k-means is that clusters may be merged or split, and so the final number of clusters may be different from the number k supplied as part of the input. This algorithm will be described in later in this paper. The ISOCLUS algorithm can run very slowly, particularly on large data sets. Given its wide use in remote sensing, its efficient computation is an important goal. We have developed a fast implementation of the ISOCLUS algorithm. Our improvement is based on a recent acceleration to the k-means algorithm, the filtering algorithm, by Kanungo et al.. They showed that, by storing the data in a kd-tree, it was possible to significantly reduce the running time of k-means. We have adapted this method for the ISOCLUS algorithm. For technical reasons, which are explained later, it is necessary to make a minor modification to the ISOCLUS specification. We provide empirical evidence, on both synthetic and Landsat image data sets, that our algorithm's performance is essentially the same as that of ISOCLUS, but with significantly lower running times. We show that our algorithm runs from 3 to 30 times faster than a straightforward implementation of ISOCLUS. Our adaptation of the filtering algorithm involves the efficient computation of a number of cluster statistics that are needed for ISOCLUS, but not for k-means.

Memarsadeghi, Nargess↗

Image processing workflow yielding high contrast synchrotron nanoscale computed tomography data from Ni-YSZ electrodes

The operating lifetime of Ni-YSZ fuel electrodes used in solid oxide electrolysis cells and fuel cells (SOECs and SOFCs) is limited by Ni redistribution, one of the primary degradation mechanisms that must be overcome to extend the longevity and maximize the performance of SOECs and SOFCs. To achieve this, 3D microstructural data is needed to relate both initial performance and performance loss over time to microstructural properties and their evolution throughout operation under various conditions. However, 3D microstructure data remains relatively scarce within the literature due to multiple challenges in acquiring and analyzing such data reliably. This work presents a workflow for acquiring and processing synchrotron X-ray nanoscale computed tomography (nano-CT) data from Ni-YSZ electrodes. Parameters for each step in the nano-CT workflow are described up to the final result (a 3D reconstruction), with particular emphasis on image alignment using freely available software. Following the results of a parametric sweep of the image alignment step, high contrast, low signal-to-noise 3D nano-CT data is obtained with relatively short compute times. While the exact methods best suited to samples with different microstructural qualities, or similar Ni-YSZ nano-CT data obtained from other sources may deviate from the solution found herein, this work also generalizes the decision points and evaluation of each step to provide a starting point to adapt this workflow to other datasets.

08 HYDROGEN↗

Applications of Data Assimilation to Analysis of the Ocean on Large Scales

It is commonplace to begin talks on this topic by noting that oceanographic data are too scarce and sparse to provide complete initial and boundary conditions for large-scale ocean models. Even considering the availability of remotely-sensed data such as radar altimetry from the TOPEX and ERS-1 satellites, a glance at a map of available subsurface data should convince most observers that this is still the case. Data are still too sparse for comprehensive treatment of interannual to interdecadal climate change through the use of models, since the new data sets have not been around for very long. In view of the dearth of data, we must note that the overall picture is changing rapidly. Recently, there have been a number of large scale ocean analysis and prediction efforts, some of which now run on an operational or at least quasi-operational basis, most notably the model based analyses of the tropical oceans. These programs are modeled on numerical weather prediction. Aside from the success of the global tide models, assimilation of data in the tropics, in support of prediction and analysis of seasonal to interannual climate change, is probably the area of large scale ocean modeling and data assimilation in which the most progress has been made. Climate change is a problem which is particularly suited to advanced data assimilation methods. Linear models are useful, and the linear theory can be exploited. For the most part, the data are sufficiently sparse that implementation of advanced methods is worthwhile. As an example of a large scale data assimilation experiment with a recent extensive data set, we present results of a tropical ocean experiment in which the Kalman filter was used to assimilate three years of altimetric data from Geosat into a coarsely resolved linearized long wave shallow water model. Since nonlinear processes dominate the local dynamic signal outside the tropics, subsurface dynamical quantities cannot be reliably inferred from surface height anomalies. Because of its potential for large scale synoptic coverage of the deep ocean, acoustic travel time data should be a natural complement to satellite altimetry. Satellite data give us vertical integrals associated with thermodynamic and dynamic processes.

Miller, Robert N.↗

Critical Heat Flux of Liquid Hydrogen, Liquid Methane, and Liquid Oxygen: A Review of Available Data and Predictive Tools

Available experimental data dealing with critical heat flux (CHF) of liquid hydrogen (LH 2 ), liquid methane (LCH 4 ), and liquid oxygen (LO 2 ) in pool and flow boiling are compiled. The compiled data are compared with widely used correlations. Experimental pool boiling CHF data for the aforementioned cryogens are scarce. Based on only 25 data points found in five independent sources, the correlation of Sun and Lienhard (1970) is recommended for predicting the pool CHF of LH 2 . Only two experiments with useful CHF data for the pool boiling of LCH 4 could be found. Four different correlations including the correlation of Lurie and Noyes (1964) can predict the pool boiling CHF of LCH 4 within a factor of two for more than 70% of the data. Furthermore, based on the 19 data points taken from only two available sources, the correlation of Sun and Lienhard (1970) is recommended for the prediction of pool CHF of LO 2 . Flow boiling CHF data for LH 2 could be found in seven experimental studies, five of them from the same source. Based on the 91 data points, it is suggested that the correlation of Katto and Ohno (1984) be used to predict the flow CHF of LH 2 . No useful data could be found for flow boiling CHF of LCH 4 or LO 2 . The available databases for flow boiling of LCH 4 and LO 2 are generally deficient in all boiling regimes. This deficiency is particularly serious with respect to flow boiling.

Multi-Phase Flow↗

Critical Heat Flux of Liquid Hydrogen, Liquid Methane, and Liquid Oxygen: A Review of Available Data and Predictive Tools

Available experimental data dealing with critical heat flux (CHF) of liquid hydrogen (LH2), liquid methane (LCH4), and liquid oxygen (LO2) in pool and flow boiling are compiled. The compiled data are compared with widely used correlations. Experimental pool boiling CHF data for the aforementioned cryogens are scarce. Based on only 25 data points found in five independent sources, the correlation of Sun and Lienhard (1970) is recommended for predicting the pool CHF of LH2. Only two experiments with useful CHF data for the pool boiling of LCH4 could be found. Four different correlations including the correlation of Lurie and Noyes (1964) can predict the pool boiling CHF of LCH4 within a factor of two for more than 70% of the data. Furthermore, based on the 19 data points taken from only two available sources, the correlation of Sun and Lienhard (1970) is recommended for the prediction of pool CHF of LO2. Flow boiling CHF data for LH2 could be found in seven experimental studies, five of them from the same source. Based on the 91 data points, it is suggested that the correlation of Katto and Ohno (1984) be used to predict the flow CHF of LH2. No useful data could be found for flow boiling CHF of LCH4 or LO2. The available databases for flow boiling of LCH4 and LO2 are generally deficient in all boiling regimes. This deficiency is particularly serious with respect to flow boiling.

Multi-Phase Flow↗

New Directions in Tropical Phenology

Earth’s most speciose biomes are in the tropics, yet tropical plant phenology remains poorly understood. Tropical phenological data are comparatively scarce and viewed through the lens of a ‘temperate phenological paradigm’ expecting phenological traits to respond to strong, predictably annual shifts in climate (e.g., between subfreezing and frost-free periods). Digitized herbarium data greatly expand existing phenological data for tropical plants; and circular data, statistics, and models are more appropriate for analyzing tropical (and temperate) phenological datasets. Phylogenetic information, which remains seldom applied in phenological investigations, provides new insights into phenological responses of large groups of related species to climate. Consistent combined use of herbarium data, circular statistical distributions, and robust phylogenies will rapidly advance our understanding of tropical – and temperate – phenology.

tropical phenology↗

Error Localization Examples: Looking for a Needle in a Hay-stack

Finite element models (FEM) are routinely developed and used during fabrication of high dollar-value hardware. NASA as part of the pre-flight certification of launch vehicles routinely conducts vibration and static tests to calibrate models used for flight-risk assessments. As part of the calibration process, certain areas in the model are modified, using engineering judgment and sensitivity analysis, to match the test results. Unfortunately, tools to identify problem areas in the FEM using test data directly are scarce. Over the years, Error Localization Algorithms (ELA) have been proposed with very limited success. Recently, the Analytical Dynamics Model Improvement (ADMI) algorithm, which computes closed-form mass and stiffness corrections to match the test data exactly, have been shown to be effective for error localization. The paper will present several FEM example problems where ELA is used with simulated test data to determine FEM problem areas. For each example, the correct answer is shown along with ELA results. It is shown that the ELA process is able to identify general problem areas in the FEM, which are consistent with known model perturbations. However, in most cases the ELA identified area of improvement is larger than the true answer. Nonetheless, with proper optimization tools, calibration results using the ELA identified areas provide excellent results.

error localization↗

Error Localization Examples: Looking for a Needle in a Haystack

Finite element models (FEM) are routinely developed and used during fabrication of high dollar-value hardware. NASA, as part of the pre-flight certification of launch vehicles, routinely conducts vibration and static tests to calibrate models used for flight-risk assessments. During model calibration, certain areas of the model are modified, using engineering judgment and sensitivity analysis, to match the test results. Unfortunately, tools to identify problem areas in the FEM using test data directly are scarce and infrequently applied. Over the years, error localization algorithms have been proposed with very limited success. Recently, the Analytical Dynamics Model Improvement (ADMI) algorithm, which computes closed-form mass and stiffness corrections to match the test data exactly, have been shown to be an effective Error Localization Algorithm (ELA). The paper discusses three examples where ELA is used with simulated test data to locate problem areas. To gain confidence in the approach, the exact answer is shown along with ELA results. Results show that ELA is able to identify general problem areas consistent with known problem areas. In all examples, the ELA identified area is larger than the exact problem area. Nonetheless, with proper optimization tools, calibration results using the ELA identified areas provide excellent results.

model calibration↗

Incorporating Biological Knowledge into Evaluation of Casual Regulatory Hypothesis

Biological data can be scarce and costly to obtain. The small number of samples available typically limits statistical power and makes reliable inference of causal relations extremely difficult. However, we argue that statistical power can be increased substantially by incorporating prior knowledge and data from diverse sources. We present a Bayesian framework that combines information from different sources and we show empirically that this lets one make correct causal inferences with small sample sizes that otherwise would be impossible.

Chrisman, Lonnie↗

Pluminate: Quantifying aerosol injection behavior from simulation, experimentation and observations

Marine aerosol injections are a key component in further understanding of both the potentials of deliberate injection for marine cloud brightening (MCB), a potential climate intervention (CI) strategy, and key aerosol-cloud interaction behaviors that currently form the largest uncertainty in global climate model (GCM) predictions of our climate. Since the rate of spread of aerosols in a marine environment directly translates to the effectiveness and ability of aerosol injections in impacting cloud radiative forcing, it is crucial to understand the spatial and temporal extent of injected-aerosol effects following direct injection into marine environments. The ubiquity of ship-injected aerosol tracks from satellite imagery renders observational validation of new parameterizations possible in 2D, however, 3D compatible data is more scarce, and necessary for the development of subgrid scale parameterizations of aerosol-cloud interactions in GCMs. This report introduces two novel parameterizations of atmospheric aerosol injection behavior suitable for both 3D (GCM-compatible) and 2D (observation-related) modeling. Their applicability is highlighted using a wealth of different observational data: small and larger scale salt-aerosol injection experiments conducted at SNL, 3D large eddy simulations of ship-injected aerosol tracks and 2D satellite images of ship tracks. The power of experimental data in enhancing knowledge of aerosol-cloud interactions is in particular emphasized by studying key aerosol microphysical and optical properties as observed through their mixing in cloud-like environments.

54 ENVIRONMENTAL SCIENCES↗

A new activity index for comets

An activity index, AI, is derived from observational data to measure the increase of activity in magnitudes for comets when brightest near perihelion as compared to their inactive reflective brightness at great solar distances. Because the observational data are still instrumentally limited in the latter case and because many comets carry particulate clouds about them at great solar distances, the application of the activity index is still limited. A tentative application is made for the comets observed by Max Beyer over a period of nearly 40 years, providing a uniform magnitude system for the near-perihelion observations. In all, 32 determinations are made for long-period (L-P) comets and 15 for short-period (S-P). Although the correlations are scarcely definitive, the data suggest that the faintest comets are just as active as the brightest and that the S-P comets are almost as active as those with periods (P) exceeding 10(exp 4) years or those with orbital inclinations of i less than 120 deg. Comets in the range 10(exp 2) less than P less than 10(exp 4) yr. or with i greater than 120 deg appear to be somewhat more active than the others. There is no evidence to suggest aging among the L-P comets or to suggest other than a common nature for comets generally.

Whipple, Fred L.↗

Audacity of huge: overcoming challenges of data scarcity and data quality for machine learning in computational materials discovery

Machine learning (ML)-accelerated discovery requires large amounts of high-fidelity data to reveal predictive structure–property relationships. For many properties of interest in materials discovery, the challenging nature and high cost of data generation has resulted in a data landscape that is both scarcely populated and of dubious quality. Data-driven techniques starting to overcome these limitations include the use of consensus across functionals in density functional theory, the development of new functionals or accelerated electronic structure theories, and the detection of where computationally demanding methods are most necessary. When properties cannot be reliably simulated, large experimental data sets can be used to train ML models. In the absence of manual curation, increasingly sophisticated natural language processing and automated image analysis are making it possible to learn structure–property relationships from the literature. Finally, models trained on these data sets will improve as they incorporate community feedback.

36 MATERIALS SCIENCE↗

Comprehensive assessment of metrology techniques for heliostat efficiency and performance evaluation

Concentrating solar power plants, specifically central receiver type systems and their heliostat field, are struggling with negative reputation in the USA, due to perceived underperformance and reliability issues. This is in part due to a lack of standards for performance assessment as well as overly simplified techno-economical models. A better understanding of influences and losses along the solar radiation path from the sun, across the solar collector to the receiver, increases the fidelity of heliostat efficiency assessment as well as solar field performance predictions. Such data are currently scarce and require a complete set of metrology capabilities to evaluate direct solar irradiance, sun shape, atmospheric attenuation, reflectance, collector shape, slope errors and total beam dispersion. In preparation for establishing a 3rd party metrology platform in collaboration with Sandia National Labs, NLR conducted a scoping study on available metrology. We present an extensive overview of techniques and commercial systems for each category. Our work includes an analysis to increase understanding of strengths and limitations of the many techniques used for surface shape and slope measurement. This applies to a controlled, indoor or outdoor laboratory environment assessing a single heliostat.

14 SOLAR ENERGY↗

Multi-Scale Modeling of the Evolution of Structure and Properties in Materials for Nuclear Energy Applications [Slides]

Nuclear energy is an important component of an overall strategy to address climate change. Idaho National Laboratory (INL) is the U.S. Department of Energy’s primary facility for research and development in nuclear science and technology for energy generation, supporting the improvement and life extension of the existing reactor fleet and the development and licensing of new reactor designs. Computational modeling is an important component of these activities, particularly in the area of materials for nuclear applications, where experimental data can be very challenging and expensive to acquire, and where data is especially scarce for new reactor designs. INL has used multi-scale modeling – linking atomistic, mesoscale, and engineering scales – to improve the ability to predict the performance of materials for nuclear energy applications. In this talk, I will give an overview of the approach and tools used, and several examples of application, including performance of nuclear fuels, understanding radiation-driven formation of nanoscale void and gas bubble superlattices, and powder densification through electric field assisted sintering.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Safety Risk Knowledge Elicitation in Support of Aeronautical R and D Portfolio Management: A Case Study

Aviation is a problem domain characterized by a high level of system complexity and uncertainty. Safety risk analysis in such a domain is especially challenging given the multitude of operations and diverse stakeholders. The Federal Aviation Administration (FAA) projects that by 2025 air traffic will increase by more than 50 percent with 1.1 billion passengers a year and more than 85,000 flights every 24 hours contributing to further delays and congestion in the sky (Circelli, 2011). This increased system complexity necessitates the application of structured safety risk analysis methods to understand and eliminate where possible, reduce, and/or mitigate risk factors. The use of expert judgments for probabilistic safety analysis in such a complex domain is necessary especially when evaluating the projected impact of future technologies, capabilities, and procedures for which current operational data may be scarce. Management of an R&D product portfolio in such a dynamic domain needs a systematic process to elicit these expert judgments, process modeling results, perform sensitivity analyses, and efficiently communicate the modeling results to decision makers. In this paper a case study focusing on the application of an R&D portfolio of aeronautical products intended to mitigate aircraft Loss of Control (LOC) accidents is presented. In particular, the knowledge elicitation process with three subject matter experts who contributed to the safety risk model is emphasized. The application and refinement of a verbal-numerical scale for conditional probability elicitation in a Bayesian Belief Network (BBN) is discussed. The preliminary findings from this initial step of a three-part elicitation are important to project management practitioners as they illustrate the vital contribution of systematic knowledge elicitation in complex domains.

Shih, Ann T.↗

Northeast US Ecological Forecasting: Modeling Invasive Plant Habitat Suitability to Support Management Efforts in the American Northeast

Invasive plant species threaten environmental and economic interests when they spread into new areas, outcompete native species, and disrupt ecosystem services. If the spread is not controlled early, species can become well-established and increasingly difficult to manage. The National Park Service (NPS) Invasive Plant Management Teams (IPMTs) strive for an “early detection, rapid response” approach to reducing invasive species spread. Management teams can better prioritize their work with the help of species distribution models (SDMs), which map habitat suitability by combining species occurrences with environmental predictor variables. Scarce invaded range data for newly arrived invasive species presents a particular challenge for producing accurate models. To improve future modeling efforts, this project compared SDM methods using different spatial scales to model two plant species invasive to the Northeast US: the well-established Japanese stiltgrass (Microstegium vimineum) and newer invasive species wavyleaf basketgrass (Oplismenus undulatifolius). The team used NASA Earth observations and climate datasets to model occurrence data and predictor layers at a US-specific extent (90m2 spatial resolution) and global extent (1 km2 spatial resolution). Landsat 5 Thematic Mapper (TM), Landsat 7 Enhanced Thematic Mapper Plus (ETM+), and Landsat 8 Operational Land Imager (OLI) provided data for US Normalized Difference Moisture Indices (NDMI), while global NDMI and topographic predictor layers were derived from Shuttle Radar Topography Mission (SRTM) and Terra Moderate Resolution Imaging Spectroradiometer (MODIS). The resulting models indicated important predictor variables for each species and explored the benefits and tradeoffs of using global data to model habitat suitability for new-arrival invasive species.

Rebecca Ohman↗