Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “random sampling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Machine Learning for Advanced Building Construction: Preprint

High-efficiency retrofits can play a key role in reducing carbon emissions associated with buildings if processes can be scaled-up to reduce cost, time, and disruption. Here we demonstrate an artificial intelligence/computer vision (AI/CV)- enabled framework for converting exterior build scans and dimensional data directly into manufacturing and installation specifications for overclad panels. In our workflow point clouds associated with LiDAR-scanned buildings are segmented into a facade feature space, vectorized features are extracted using an iterative random-sampling consensus algorithm, and from this representation an optimal panel design plan satisfying manufacturing constraints is generated. This system and the corresponding construction process is demonstrated on a test facade structure constructed at the National Renewable Energy Laboratory (NREL). We also include a brief summary of a techno-economic study designed to estimate the potential energy and cost impact of this new system.

building retrofits↗

Machine Learning for Advanced Building Construction

High-efficiency retrofits can play a key role in reducing carbon emissions associated with buildings if processes can be scaled-up to reduce cost, time, and disruption. Here we demonstrate an artificial intelligence/computer vision (AI/CV)-enabled framework for converting exterior build scans and dimensional data directly into manufacturing and installation specifications for overclad panels. In our workflow point clouds associated with LiDAR-scanned buildings are segmented into a facade feature space, vectorized features are extracted using an iterative random-sampling consensus algorithm, and from this representation an optimal panel design plan satisfying manufacturing constraints is generated. This system and the corresponding construction process is demonstrated on a test facade structure constructed at the National Renewable Energy Laboratory (NREL). We also include a brief summary of a techno-economic study designed to estimate the potential energy and cost impact of this new system.

build scans↗

An influence of manufacturing tolerances on Pin-Cell k-infinity of MOX fuel using data from the FUBILA experiment program

An influence of manufacturing tolerances (MTs) was evaluated on the pin-cell k-infinity of MOX fuel through the random sampling of CASMO5 calculations. The data from the FUBILA experiment program was used as the manufacturing parameters. For the uncertainties of element/isotope mass fractions, their covariance matrices were calculated by the generalized least square method. The total k-infinity uncertainty was 120-250 pcm (percent mille). From the breakdown of k-infinity uncertainty per materials, the MTs of fuel pellet had a dominant influence. The individual influences were also evaluated for element/isotope mass fractions, an inner/outer diameter, and a density. Those of element/isotope mass fractions were less than a few dozen pcm. Since the perturbations of fuel pellet diameter and density caused large variations to the total amount of heavy metals and effect of spatial self-shielding, they had a negative correlation with the k-infinity perturbation. On the AG3 (Al-Mg alloy) over-cladding, the inner/outer diameter perturbations had both negative and positive correlations. The negative one was due to decrease in amount of light water and the positive one was due to increase in amount of AG3. Since the effect of neutron slowing-down by light water had a dominant influence, the influence of the former negative correlation was larger than that of the latter positive correlation. It is, therefore, concluded that the consideration of the MTs that had a large influence on the effect of neutron slowing-down is important for precise quantification of the pin-cell k-infinity uncertainty of MOX fuel. (author)

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Robust Multi-fidelity Bayesian Optimization with Deep Kernel and Partition

Multi-fidelity Bayesian optimization (MFBO) is a powerful approach that utilizes lowfidelity, cost-effective sources to expedite the exploration and exploitation of a high-fidelity objective function. Existing MFBO methods with theoretical foundations either lack justification for performance improvements over single-fidelity optimization or rely on strong assumptions about the relationships between fidelity sources to construct surrogate models and direct queries to low-fidelity sources. To mitigate the dependency on cross-fidelity assumptions while maintaining the advantages of low-fidelity queries, we introduce a random sampling and partition-based MFBO framework with deep kernel learning. This framework is robust to cross-fidelity model misspecification and explicitly illustrates the benefits of low-fidelity queries. Our results demonstrate that the proposed algorithm effectively manages complex cross-fidelity relationships and efficiently optimizes the target fidelity function.

Zhang, Fengxue [University of Chicago, Illinois, U↗

Transcriptomics-based Machine Learning (ML) Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% was shown on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% was shown on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗

Assessing Risk Due to Small Sample Size in Probability of Detection Analysis Using Tolerance Intervals

Small sample size (e.g.6-30) poses risk in results of probability of detection (POD) analysis using tolerance intervals. This method is also called as the limited sample or LS POD. The analysis is performed either during NDE procedure qualification or for assessment of reliability of an NDE procedure. The risk is primarily due to sampling error. Smaller samples are not likely to be random to the population or representative of the population. The small samples are likely to be biased. Biased samples have smaller standard deviation compared to the population. POD analysis with small biased sample can lead to overestimation of POD. Many sampling schemes are available in statistics to mitigate sampling risk. Primary objective of POD analysis is to determine a decision threshold from signal response measurements of a sample such that it is less than or equal to population decision threshold for 90% POD. Sampling error implies that this NDE reliability condition is violated. One of sampling types is called a representative sample. Representative samples reduce variance in POD estimates but also reduce magnitude of the error. Sampling sensitivity analysis for some sampling types is performed here using repetitive random sampling or Monte Carlo method. Six sampling types are considered for comparison. Some of the sampling types are similar to drawing a representative sample. LS POD model assumes random sampling. Therefore, random sampling is used as a basis for comparison with each sampling type. The sampling types used in the analysis are, A. Nominal and worst-case sampling, B. Worst-case sampling, C. Nominal case sampling, D. Random sampling, E. Random target, and sub-target sampling. F. Nominal target and sub-target sampling. Results of Monte Carlo simulation indicate that type F sampling can mitigate sampling risk and is also more practical to implement. Type A sampling may also mitigate the sampling risk, but it may be less practical to implement.

Ajay M Koshti↗

A new efficient grain growth model using a random Gaussian-sampled mode filter

This paper presents the use of a Gaussian neighborhood mode filter for predicting grain growth in a manner similar to the solutions obtained by a Monte Carlo Potts model. This flexible grain growth model can quickly utilize modern, computationally optimized data science strategies on graphics processing units to simulate grain growth up to 100 times faster than the state-of-the-art, publicly available Monte Carlo Potts model. We show that, given the correct neighborhood, the mode filter can replicate normal grain growth in two or three dimensions. In addition, the paper briefly demonstrates the ability to model limited anisotropic in grain boundary energy and mobility. Anisotropic grain boundary energy is modeled by defining a weighted mode filter operation. Anisotropic grain boundary mobility is modeled by scaling and orienting the Gaussian neighborhood in a particular direction.

Anisotropy↗

Eco-Friendly High-Performance Carbon Building Material Development from Coal

The development of eco-friendly high-performance building materials based on coal-derived materials is greatly beneficial for promoting environmental sustainability and the coal industry. This project focused on advancing technology to create innovative and scalable construction materials from two domestic sources: (1) coal-derived pyrolyzed char (PC), and (2) solvent-extracted coal deposit, extracts, and residue (CDER). It aimed to simultaneously develop and produce two types of coal-derived building products: (1) char-based concrete bricks (CCB) for wall construction in buildings, and (2) carbon-based structural units (CSU) intended as alternatives to traditional wood, concrete, or steel framing structures. The project team has successfully developed methods for manufacturing CCB samples. Extensive experiments were conducted to assess various manufacturing techniques for these samples. This included measuring their physical properties such as thermal conductivity (ranging from 0.26 to 0.38 W/mK) and bulk density (ranging from 0.85 to 1.15 g/cm 3 ), along with their mechanical properties (e.g., compressive strength ranging from 14 to 16 MPa at 28 days) for samples with a PC composition of 70%. Quality assurance and quality control tests, focusing on density and compressive strength, were also performed on samples selected randomly from pilot-scale production, and cured for a long term (over three months). The average long-term compressive strength of two randomly selected CCB samples was measured as 15.1 MPa, which was higher than the required target strength of 14 MPa. This indicates that the manufactured CCB samples from the preliminary pilot scale production run have good quality control and long-term strength performance.

01 COAL, LIGNITE, AND PEAT↗

Oscillating-flow regenerator test rig: Woven screen and metal felt results

We present correlating expressions, in terms of Reynolds or Peclet numbers, for friction factors, Nusselt numbers, enhanced axial conduction ratios, and overall heat flux ratios in four porous regenerator samples representative of stirling cycle regenerators: two woven screen samples and two random wire samples. Error estimates and comparison of data with others suggest our correlations are reliable, but we need to test more samples over a range of porosities before our results will become generally useful.

Gedeon, D.↗

Benchmarking Bayesian Optimization Frameworks and Acquisition Strategies for Materials Discovery and Autonomous Laboratories

Bayesian optimization (BO) can accelerate materials discovery by guiding expensive experiments toward the most promising processing conditions. We systematically compare five BO surrogate and framework combinations (Gaussian processes in Ax, Gaussian processes and Monte-Carlo neural networks in BayBE, random forests in Lolopy, and tree-structured Parzen (TPE) estimators in Hyperopt) on three benchmarks that mimic common materials design tasks (a discrete solid-electrolyte composition space, a hybrid discrete/continuous laminate-composite design problem solved with micromechanics modeling, and the continuous Ishigami analytic function which is a standard optimization benchmark). Each BO surrogate is paired with posterior mean, probability of improvement, and expected improvement acquisition functions and run for 100 trials from randomized initial samples with uniform random search providing a control. Across five random seeds per setting, BayBE’s Gaussian-process surrogate with expected improvement consistently reached ≥95 % of the known optimum in the fewest evaluations, while Lolopy’s random forest matched or exceeded GP performance on purely categorical or mixed spaces at a higher computational cost. Posterior mean alone often stagnated at local optima, underscoring the need for exploration, whereas probability and expected improvement balanced exploration and exploitation leading to better optimization in fewer trials. Execution times ranged from milliseconds for TPE to minutes for neural-network and random-forest surrogates. These results establish baseline expectations for BO in automated materials laboratories and highlight expected improvement with Gaussian processes as a reliable first choice, with random forests offering a strong alternative when categorical variables dominate. The benchmark suite and code are released to facilitate future surrogate, acquisition, and constraint-handling research in data-driven materials optimization.

Bayesian optimization↗

The Terrestrial Organism and Biogeochemistry Spatial Sampling Design for the National Ecological Observatory Network

The National Ecological Observatory Network (NEON) seeks to facilitate ecological prediction at a continental scale by measuring processes that drive change and responses at sites across the United States for thirty years. The spatial distribution of observations of terrestrial organisms and soil within NEON sites is determined according to a “design‐based” sample design that relies on the randomization of sampling locations. Development of the sample design was guided by high‐level NEON objectives and the multitude of data products that will be subjected to numerous analytical approaches to address the causes and consequences of ecological change. A requirement framework permeates the NEON design, ensuring traceability from each facet of the design to the high‐level requirements that make the NEON mission statement actionable. Requirements were developed for the terrestrial sample design to guide the key components of the design: Randomizing the sample locations ensures the unbiased collection of data, is appropriate for organisms and soil, and provides data suitable for a variety of analyses. Stratification increases efficiency and allows sampling to focus on those parts of the landscape measured by other NEON observation platforms. Attention to the sample size and spatial plot allocation ensures that data products will be sufficient to inform questions asked of the data and the NEON objectives. Establishing a framework with the capacity for re‐evaluate and design iteration allows for adaption to unexpected challenges and optimization of the sample design based on early data returns. The utility of the NEON sampling design is highlighted by its application across terrestrial systems. The data generated from this unique design will be used to quantify patterns in: the abundance and diversity of small mammals, breeding birds, insects, and soil microbes; vegetation structure, biomass, productivity, and diversity; and soil biogeochemistry.

National Ecological Observatory Network↗

Spoofing Cross-Entropy Measure in Boson Sampling

Cross-entropy (XE) measure is a widely used benchmark to demonstrate quantum computational advantage from sampling problems, such as random circuit sampling using superconducting qubits and boson sampling (BS). We present a heuristic classical algorithm that attains a better XE than the current BS experiments in a verifiable regime and is likely to attain a better XE score than the near-future BS experiments in a reasonable running time. The key idea behind the algorithm is that there exist distributions that correlate with the ideal BS probability distribution and that can be efficiently computed. The correlation and the computability of the distribution enable us to postselect heavy outcomes of the ideal probability distribution without computing the ideal probability, which essentially leads to a large XE. Our method scores a better XE than the recent Gaussian BS experiments when implemented at intermediate, verifiable system sizes. Much like current state-of-the-art experiments, we cannot verify that our spoofer works for quantum-advantage-size systems. However, we demonstrate that our approach works for much larger system sizes in fermion sampling, where we can efficiently compute output probabilities. Finally, we provide analytic evidence that the classical algorithm is likely to spoof noisy BS efficiently.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Lunar glass compositions - Apollo 16 core sections 60002 and 60004

Approximately 500 glasses between 1 mm and 125 microns in size have been analyzed from fourteen samples from the Apollo 16 core sections 60002 and 60004. The majority of glasses have compositions comparable to those found in previous studies of lunar surface soils; however, two new and distinct glass compositions that are probably derived in part from mare material occur in the core samples. The major glass composition in all samples is that of Highland Basalt glass, but it also appears that high-K Fra Mauro Basalt (KREEP) glass is more common at the Apollo 16 site than was previously thought. The relative abundance of glasses within the core samples is random in distribution: each sample is characterized by a particular assemblage and distribution of the constituent glass compositions.

Meyer, H. O. A.↗

On the stability of robotic systems with random communication rates

Control problems of sampled data systems which are subject to random sample rate variations and delays are studied. Due to the rapid growth of the use of computers more and more systems are controlled digitally. Complex systems such as space telerobotic systems require the integration of a number of subsystems at different hierarchical levels. While many subsystems may run on a single processor, some subsystems require their own processor or processors. The subsystems are integrated into functioning systems through communications. Communications between processes sharing a single processor are also subject to random delays due to memory management and interrupt latency. Communications between processors involve random delays due to network access and to data collisions. Furthermore, all control processes involve delays due to casual factors in measuring devices and to signal processing. Traditionally, sampling rates are chosen to meet the worst case communication delay. Such a strategy is wasteful as the processors are then idle a great proportion of the time; sample rates are not as high as possible resulting in poor performance or in the over specification of control processors; there is the possibility of missing data no matter how low the sample rate is picked. Asymptotical stability with probability one for randomly sampled multi-dimensional linear systems is studied. A sufficient condition for the stability is obtained. This condition is so simple that it can be applied to practical systems. A design procedure is also shown.

Kobayashi, H.↗