Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “random sampling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Estimation of 3D Woven Design Sensitivities Using a Rapid Multiscale Analysis Technique

Highly-refined finite element models of three-dimension (3D) woven composite systems currently require excessive computational demands that limit their use in sensitivity analysis, uncertainty quantification, and optimization. An alternative analysis methodology was developed using the NASA Multiscale Analysis Tool (NASMAT) where multiscale models of a 3D woven composite (including inter-tow matrix voids and constituent failure) can be completed on a single central processing unit(CPU)on the order of ~30 s. To develop inputs and validation data for the NASMAT model, coupon and acid-digesting testing and x-ray computed tomography were performed. The NASMAT inputs were parameterized using a set of 25 input variables and distributions. These inputs were randomly sampled to generate a total of 100,000 NASMAT analyses that could be used to understand the influence of different material and geometric properties on the warp and weft-direction stiffness and strength. These analyses (including pre/post-processing) were performed in less than eight hours on a 120 CPU cluster. The computational efficiency of the NASMAT model enabled a sensitivity analysis to be performed, and dominant input variables were able to be identified. Key results were consistent with theoretical and experimental observations for the specific 3D woven system studied in this work.

NASMAT↗

Estimation of 3D Woven Design Sensitivities Using a Rapid Multiscale Analysis Technique

Highly-refined finite element models of three-dimension (3D) woven composite systems currently require excessive computational demands that limit their use in sensitivity analysis, uncertainty quantification, and optimization. An alternative analysis methodology was developed using the NASA Multiscale Analysis Tool (NASMAT) where multiscale models of a 3D woven composite (including inter-tow matrix voids and constituent failure) can be completed on a single central processing unit (CPU) on the order of ~30s. To develop inputs and validation data for the NASMAT model, coupon and acid-digesting testing and x-ray computed tomography were performed. The NASMAT inputs were parameterized using a set of 25 input variables and distributions. These inputs were randomly sampled to generate a total of 100,000 NASMAT analyses that could be used to understand the influence of different material and geometric properties on the warp and weft-direction stiffness and strength. These analyses (including pre/post-processing) were performed in less than eight hours on a 120 CPU cluster. The computational efficiency of the NASMAT model enabled a sensitivity analysis to be performed, and dominant input variables were able to be identified. Key results were consistent with theoretical and experimental observations for the specific 3D woven system studied in this work.

NASMAT↗

Self-Consistent Implementation of a Zero-Equation Transport Model Into a Predictive Model for a Hall Effect Thruster

The performance of an axisymmetric multi-fluid Hall thruster code that incorporates a self-consistent, data-driven closure model for the anomalous electron transport is investigated. Five different operating conditions of the H9 magnetically shielded Hall thruster are simulated with the Jet Propulsion Laboratory’s Hall2De. In order to capture the inherent uncertainty associated with the closure model, O(100) simulations are run for each condition, each of using a coefficient set sampled randomly from a probability distribution. The results of these simulations provide probabilistic predictions of thruster performance quantities including thrust, and discharge current, as well as several component efficiencies and centerline plasma properties. The model is found to yield converged solutions at all conditions, with large 10 kHzrange oscillations and performance trends with voltage and flow rate similar to experiment. The model under-predicts the thrust by 15-25% and over-predicts the discharge current by 20% on average compared to experiments at the same discharge voltage and mass flow rate. This performance discrepancy is due to lower beam utilization, mass utilization, and divergence efficiency than experiment, resulting from high Hall parameters in the acceleration region, which lead to a protracted ion acceleration region. The physical processes underlying this result are discussed in the context of future data-driven modeling efforts.

Jorns, Benjamin A.↗

Transcriptomics-based Machine Learning (ML) Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% was shown on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% was shown on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗

Assessing Risk Due to Small Sample Size in Probability of Detection Analysis Using Tolerance Intervals

Small sample size (e.g.6-30) poses risk in results of probability of detection (POD) analysis using tolerance intervals. This method is also called as the limited sample or LS POD. The analysis is performed either during NDE procedure qualification or for assessment of reliability of an NDE procedure. The risk is primarily due to sampling error. Smaller samples are not likely to be random to the population or representative of the population. The small samples are likely to be biased. Biased samples have smaller standard deviation compared to the population. POD analysis with small biased sample can lead to overestimation of POD. Many sampling schemes are available in statistics to mitigate sampling risk. Primary objective of POD analysis is to determine a decision threshold from signal response measurements of a sample such that it is less than or equal to population decision threshold for 90% POD. Sampling error implies that this NDE reliability condition is violated. One of sampling types is called a representative sample. Representative samples reduce variance in POD estimates but also reduce magnitude of the error. Sampling sensitivity analysis for some sampling types is performed here using repetitive random sampling or Monte Carlo method. Six sampling types are considered for comparison. Some of the sampling types are similar to drawing a representative sample. LS POD model assumes random sampling. Therefore, random sampling is used as a basis for comparison with each sampling type. The sampling types used in the analysis are, A. Nominal and worst-case sampling, B. Worst-case sampling, C. Nominal case sampling, D. Random sampling, E. Random target, and sub-target sampling. F. Nominal target and sub-target sampling. Results of Monte Carlo simulation indicate that type F sampling can mitigate sampling risk and is also more practical to implement. Type A sampling may also mitigate the sampling risk, but it may be less practical to implement.

Ajay M Koshti↗

Oscillating-flow regenerator test rig: Woven screen and metal felt results

We present correlating expressions, in terms of Reynolds or Peclet numbers, for friction factors, Nusselt numbers, enhanced axial conduction ratios, and overall heat flux ratios in four porous regenerator samples representative of stirling cycle regenerators: two woven screen samples and two random wire samples. Error estimates and comparison of data with others suggest our correlations are reliable, but we need to test more samples over a range of porosities before our results will become generally useful.

Gedeon, D.↗

Benchmarking Bayesian Optimization Frameworks and Acquisition Strategies for Materials Discovery and Autonomous Laboratories

Bayesian optimization (BO) can accelerate materials discovery by guiding expensive experiments toward the most promising processing conditions. We systematically compare five BO surrogate and framework combinations (Gaussian processes in Ax, Gaussian processes and Monte-Carlo neural networks in BayBE, random forests in Lolopy, and tree-structured Parzen (TPE) estimators in Hyperopt) on three benchmarks that mimic common materials design tasks (a discrete solid-electrolyte composition space, a hybrid discrete/continuous laminate-composite design problem solved with micromechanics modeling, and the continuous Ishigami analytic function which is a standard optimization benchmark). Each BO surrogate is paired with posterior mean, probability of improvement, and expected improvement acquisition functions and run for 100 trials from randomized initial samples with uniform random search providing a control. Across five random seeds per setting, BayBE’s Gaussian-process surrogate with expected improvement consistently reached ≥95 % of the known optimum in the fewest evaluations, while Lolopy’s random forest matched or exceeded GP performance on purely categorical or mixed spaces at a higher computational cost. Posterior mean alone often stagnated at local optima, underscoring the need for exploration, whereas probability and expected improvement balanced exploration and exploitation leading to better optimization in fewer trials. Execution times ranged from milliseconds for TPE to minutes for neural-network and random-forest surrogates. These results establish baseline expectations for BO in automated materials laboratories and highlight expected improvement with Gaussian processes as a reliable first choice, with random forests offering a strong alternative when categorical variables dominate. The benchmark suite and code are released to facilitate future surrogate, acquisition, and constraint-handling research in data-driven materials optimization.

Bayesian optimization↗

The Terrestrial Organism and Biogeochemistry Spatial Sampling Design for the National Ecological Observatory Network

The National Ecological Observatory Network (NEON) seeks to facilitate ecological prediction at a continental scale by measuring processes that drive change and responses at sites across the United States for thirty years. The spatial distribution of observations of terrestrial organisms and soil within NEON sites is determined according to a “design‐based” sample design that relies on the randomization of sampling locations. Development of the sample design was guided by high‐level NEON objectives and the multitude of data products that will be subjected to numerous analytical approaches to address the causes and consequences of ecological change. A requirement framework permeates the NEON design, ensuring traceability from each facet of the design to the high‐level requirements that make the NEON mission statement actionable. Requirements were developed for the terrestrial sample design to guide the key components of the design: Randomizing the sample locations ensures the unbiased collection of data, is appropriate for organisms and soil, and provides data suitable for a variety of analyses. Stratification increases efficiency and allows sampling to focus on those parts of the landscape measured by other NEON observation platforms. Attention to the sample size and spatial plot allocation ensures that data products will be sufficient to inform questions asked of the data and the NEON objectives. Establishing a framework with the capacity for re‐evaluate and design iteration allows for adaption to unexpected challenges and optimization of the sample design based on early data returns. The utility of the NEON sampling design is highlighted by its application across terrestrial systems. The data generated from this unique design will be used to quantify patterns in: the abundance and diversity of small mammals, breeding birds, insects, and soil microbes; vegetation structure, biomass, productivity, and diversity; and soil biogeochemistry.

National Ecological Observatory Network↗

Lunar glass compositions - Apollo 16 core sections 60002 and 60004

Approximately 500 glasses between 1 mm and 125 microns in size have been analyzed from fourteen samples from the Apollo 16 core sections 60002 and 60004. The majority of glasses have compositions comparable to those found in previous studies of lunar surface soils; however, two new and distinct glass compositions that are probably derived in part from mare material occur in the core samples. The major glass composition in all samples is that of Highland Basalt glass, but it also appears that high-K Fra Mauro Basalt (KREEP) glass is more common at the Apollo 16 site than was previously thought. The relative abundance of glasses within the core samples is random in distribution: each sample is characterized by a particular assemblage and distribution of the constituent glass compositions.

Meyer, H. O. A.↗

On the stability of robotic systems with random communication rates

Control problems of sampled data systems which are subject to random sample rate variations and delays are studied. Due to the rapid growth of the use of computers more and more systems are controlled digitally. Complex systems such as space telerobotic systems require the integration of a number of subsystems at different hierarchical levels. While many subsystems may run on a single processor, some subsystems require their own processor or processors. The subsystems are integrated into functioning systems through communications. Communications between processes sharing a single processor are also subject to random delays due to memory management and interrupt latency. Communications between processors involve random delays due to network access and to data collisions. Furthermore, all control processes involve delays due to casual factors in measuring devices and to signal processing. Traditionally, sampling rates are chosen to meet the worst case communication delay. Such a strategy is wasteful as the processors are then idle a great proportion of the time; sample rates are not as high as possible resulting in poor performance or in the over specification of control processors; there is the possibility of missing data no matter how low the sample rate is picked. Asymptotical stability with probability one for randomly sampled multi-dimensional linear systems is studied. A sufficient condition for the stability is obtained. This condition is so simple that it can be applied to practical systems. A design procedure is also shown.

Kobayashi, H.↗

Lessons from 18 Years of Hyperspectral Infrared Sounder Data

By the end of 2013 NASA and EUMETSAT will have accumulated more than 11 years of AIRS, 6 years of IASI and one year of CrIS data. All three instruments were nominally specified to support the NWC for short term weather forecasting with a five year lifetime, but continue to exceed the accuracy requirement needed for weather forecasting alone. This allows use of their data for a much broader range of applications, including the calibration of broad-band instruments in space and climate research. We illustrate calibration aspects with examples from AIRS, IASI and CrIS using spatially uniform clear conditions, simultaneous nadir overpasses and random nadir samples. The differences between AIRS, IASI and CrIS for the purpose of weather forecasting are small and we expect that the excellent forecast impact demonstrated by the combination of AIRS and IASI will be continued by the combination of CrIS and IASI. Clear data are useful for calibration, but contain no climate signal. The analysis of random nadir samples from AIRS and CrIS identifies larger biases for observation of extreme conditions, represented by 1% and 99%tile data than for non-extreme observations. This is relevant for climate analysis. Resolution of these differences require further work, since they can complicate the continuation of trends established by AIRS with CrIS data, at least for extrema. The unequaled stability of the AIRS data allows us to evaluate trends using random nadir sampled data. We see an increasing frequency in severe storms over land, a decreasing frequency over ocean. The 11 years of AIRS data are too short to tell if these trends are significant from a climate change viewpoint, or if they are parts of multi-decadal oscillations.

CRIS↗

Model-based quantification of image quality

In 1982, Park and Schowengerdt published an end-to-end analysis of a digital imaging system quantifying three principal degradation components: (1) image blur - blurring caused by the acquisition system, (2) aliasing - caused by insufficient sampling, and (3) reconstruction blur - blurring caused by the imperfect interpolative reconstruction. This analysis, which measures degradation as the square of the radiometric error, includes the sample-scene phase as an explicit random parameter and characterizes the image degradation caused by imperfect acquisition and reconstruction together with the effects of undersampling and random sample-scene phases. In a recent paper Mitchell and Netravelli displayed the visual effects of the above mentioned degradations and presented subjective analysis about their relative importance in determining image quality. The primary aim of the research is to use the analysis of Park and Schowengerdt to correlate their mathematical criteria for measuring image degradations with subjective visual criteria. Insight gained from this research can be exploited in the end-to-end design of optical systems, so that system parameters (transfer functions of the acquisition and display systems) can be designed relative to each other, to obtain the best possible results using quantitative measurements.

Hazra, Rajeeb↗

X-Ray Diffraction Reference Intensity Ratios of Amorphous and Poorly Crystalline Phases: Implications for CheMin on the Mars Science Laboratory

The CheMin instrument on the Mars Science Laboratory (MSL) rover Curiosity is an X-ray diffraction (XRD) and X-ray fluorescence (XRF) instrument capable of providing the mineralogical and chemical compositions of rocks and soils on the surface of Mars. CheMin uses a microfocus X-ray tube with a Co target, transmission geometry, and an energy-discriminating X-ray sensitive CCD to produce simultaneous 2-D XRD patterns and energy-dispersive X-ray histograms from powdered samples. Piezoelectric vibration of the cell is used to randomize the sample to reduce preferred orientation effects. Instrument details are provided in [1, 2, 3]. Analyses of rock and soil samples by the Mars Exploration Rovers (MER) show nanophase ferric oxide (npOx) is a significant component of the Martian global soil [4] and is thought to be one of the major contributing phases that the Curiosity rover will encounter if a soil sample is analyzed in Gale Crater. Because of the nature of this material, npOx will likely contribute to an X-ray amorphous or short-order component of a XRD pattern measured by the CheMin instrument.

Morris, R. V.↗

Statistical methods for efficient design of community surveys of response to noise: Random coefficients regression models

Research studies of residents' responses to noise consist of interviews with samples of individuals who are drawn from a number of different compact study areas. The statistical techniques developed provide a basis for those sample design decisions. These techniques are suitable for a wide range of sample survey applications. A sample may consist of a random sample of residents selected from a sample of compact study areas, or in a more complex design, of a sample of residents selected from a sample of larger areas (e.g., cities). The techniques may be applied to estimates of the effects on annoyance of noise level, numbers of noise events, the time-of-day of the events, ambient noise levels, or other factors. Methods are provided for determining, in advance, how accurately these effects can be estimated for different sample sizes and study designs. Using a simple cost function, they also provide for optimum allocation of the sample across the stages of the design for estimating these effects. These techniques are developed via a regression model in which the regression coefficients are assumed to be random, with components of variance associated with the various stages of a multi-stage sample design.

Tomberlin, T. J.↗

Investigation of electrolyte measurement in diluted whole blood using spectroscopic and chemometric methods

The feasibility of using near-infrared (NIR) spectroscopy in combination with partial least-squares (PLS) regression was explored to measure electrolyte concentration in whole blood samples. Spectra were collected from diluted blood samples containing randomized, clinically relevant concentrations of Na+, K+, and Ca2+. Sodium was also studied in lysed blood. Reference measurements were made from the same samples using a standard clinical chemistry instrument. Partial least squares (PLS) was used to develop calibration models for each ion with acceptable results (Na+, R2 = 0.86, CVSEP = 9.5 mmol/L; K+, R2 = 0.54, CVSEP = 1.4 mmol/L; Ca2+, R2 = 0.56, CVSEP = 0.18 mmol/L). Slightly improved results were obtained using a narrower wavelength region (470-925 nm) where hemoglobin, but not water, absorbed indicating that ionic interaction with hemoglobin is as effective as water in causing measurable spectral variation. Good models were also achieved for sodium in lysed blood, illustrating that cell swelling, which is correlated with sodium concentration, is not required for calibration model development.

Non-NASA Center↗