Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Bootstrap aggregating”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Privacy-Preserving Knowledge Transfer with Bootstrap Aggregation of Teacher Ensembles

There is a need to transfer knowledge among institutions and organizations to save effort in annotation and labeling or in enhancing task performance. However, knowledge transfer is difficult because of restrictions that are in place to ensure data security and privacy. Institutions are not allowed to exchange data or perform any activity that may expose personal information. With the leverage of a differential privacy algorithm in a high-performance computing environment, we propose a new training protocol, Bootstrap Aggregation of Teacher Ensembles (BATE), which is applicable to various types of machine learning models. The BATE algorithm is based on and provides enhancements to the PATE algorithm, maintaining competitive task performance scores on complex datasets with underrepresented class labels.We conducted a proof-of-the-concept study of the information extraction from cancer pathology report data from four cancer registries and performed comparisons between four scenarios: no collaboration, no privacy-preserving collaboration, the PATE algorithm, and the proposed BATE algorithm. The results showed that the BATE algorithm maintained competitive macro-averaged F1 scores, demonstrating that the suggested algorithm is an effective yet privacy-preserving method for machine learning and deep learning solutions.

Yoon, Hong-Jun↗

Allan Variance is Bootstrap Aggregation for Spectral Estimation

Characterization of clocks and inertial sensors, such as accelerometers and gyroscopes, typically includes Allan variance analysis. Allan variance is ubiquitous in timing and navigation communities which may appear niche compared with generalized spectral analysis. This note provides some motivation for Allan Variance for audiences more familiar with spectral analysis.

Walker, Michael Ray [Sandia National Laboratories ↗

Unraveling the depth-dependent causal dynamics of methanogenesis and methanotrophy in a high-latitude fen peatland

The dynamics of methane (CH 4 ) cycling in high-latitude peatlands through different pathways of methanogenesis and methanotrophy are still poorly understood due to the spatiotemporal complexity of microbial activities and biogeochemical processes. Additionally, long-term in situ measurements within soil columns are limited and associated with large uncertainties in microbial substrates (e.g. dissolved organic carbon, acetate, hydrogen). To better understand CH 4 cycling dynamics, we first applied an advanced biogeochemical model, ecosys , to explicitly simulate methanogenesis, methanotrophy, and CH 4 transport in a high-latitude fen (within the Stordalen Mire, northern Sweden). Next, to explore the vertical heterogeneity in CH 4 cycling, we applied the PCMCI/PCMCI+ causal detection framework with a bootstrap aggregation method to the modeling results, characterizing causal relationships among regulating factors (e.g. temperature, microbial biomass, soil substrate concentrations) through acetoclastic methanogenesis, hydrogenotrophic methanogenesis, and methanotrophy, across three depth intervals (0–10 cm, 10–20 cm, 20–30 cm). Our results indicate that temperature, microbial biomass, and methanogenesis and methanotrophy substrates exhibit significant vertical variations within the soil column. Soil temperature demonstrates strong causal relationships with both biomass and substrate concentrations at the shallower depth (0–10 cm), while these causal relationships decrease significantly at the deeper depth within the two methanogenesis pathways. In contrast, soil substrate concentrations show significantly greater causal relationships with depth, suggesting the substantial influence of substrates on CH 4 cycling. CH 4 production is found to peak in August, while CH 4 oxidation peaks predominantly in October, showing a lag response between production and oxidation. Overall, this research provides important insights into the causal mechanisms modulating CH 4 cycling across different depths, which will improve carbon cycling predictions, and guide the future field measurement strategies.

54 ENVIRONMENTAL SCIENCES↗

Real-time deep-learning inversion of seismic full waveform data for CO 2 saturation and uncertainty in geological carbon storage monitoring

Deep-learning inversion has recently drawn attention in geological carbon storage research due to its potential of imaging and monitoring carbon storage in real time, significantly improving efficiency and safety of carbon storage operations. We present a deep-learning full waveform inversion method that after the neural network has been trained can image CO 2 saturation and its uncertainty in real time. Our deep-learning inversion method is based on the U-Net architecture with the neural network trained on pairs of synthetic seismic data and CO 2 saturation models. Accordingly, our training establishes a mapping relationship between seismic data and CO 2 saturation models and once fully trained directly estimates CO 2 saturation as a function of subsurface location. We further quantify uncertainties of CO 2 saturation estimates using the Monte Carlo dropout method and a bootstrap aggregating method. For this proof-of-concept study, the CO 2 training models and data are derived from the Kimberlina 1.2 model, a hypothetical 3D geological carbon storage model that is constructed based on various geological and hydrological data from the Southern San Joaquin Basin, California. We perform deep-learning inversion experiments using noise-free and noisy training and test data sets and compare the results. Our modelling experiments show that (1) the deep-learning inversion can estimate 2D distributions of CO 2 fairly well even in the presence of Gaussian random noise and (2) both CO 2 saturation imaging and uncertainty quantification can be done in real time. Our results suggest that the deep-learning inversion method can serve as a robust real-time monitoring tool for geological carbon storage and/or other time-varying reservoir/aquifer properties that result from injection, extraction, and/or other subsurface transport phenomena.

58 GEOSCIENCES↗

Deep learning multiphysics network for imaging CO 2 saturation and estimating uncertainty in geological carbon storage

Multiphysics inversion exploits different types of geophysical data that often complement each other and aims to improve overall imaging resolution and reduce uncertainties in geophysical interpretation. Despite the advantages, traditional multiphysics inversion is challenging because it requires a large amount of computational time and intensive human interactions for preprocessing data and finding trade-off parameters. These issues make it nearly impossible for traditional multiphysics inversion to be applied as a real-time monitoring tool for geological carbon storage. In this paper, we present a deep learning (DL) multiphysics network for imaging CO 2 saturation in real time. The multiphysics network consists of three encoders for analysing seismic, electromagnetic and gravity data and shares one decoder for combining imaging capabilities of the different geophysical data for better predicting CO 2 saturation. The network is trained on pairs of CO 2 label models and multiphysics data so that it can directly image CO 2 saturation. Here we use the bootstrap aggregating method to enhance the imaging accuracy and estimate uncertainties associated with CO 2 saturation images. Using realistic CO 2 label models and multiphysics data derived from the Kimberlina CO 2 storage model, we evaluate the performance of the deep learning multiphysics network and compare its imaging results to those from the deep learning single-physics networks. Our modelling experiments show that the deep learning multiphysics network for seismic, electromagnetic, and gravity data not only improves the imaging accuracy but also reduces uncertainties associated with CO 2 saturation images. Our results also suggest that the deep learning multiphysics network for the non-seismic data (i.e., electromagnetic and gravity) can be used as an effective low-cost monitoring tool in between regular seismic monitoring.

58 GEOSCIENCES↗

Predicting variations of the least principal stress with depth: Application to unconventional oil and gas reservoirs using a log-based viscoelastic stress relaxation model

Knowledge of layer-to-layer variations of the least principal stress, S hmin , with depth is essential for optimization of multi-stage hydraulic fracturing in unconventional reservoirs. Utilizing a geomechanical model based on viscoelastic stress relaxation in relatively clay rich rocks, we present a new method for predicting continuous S hmin variations with depth. The method utilizes geophysical log data and S hmin measurements from routine diagnostic fracture injection tests (DFITs) at several depths for calibration. We consider a case study in the Wolfcamp formation in the Midland Basin, where both geophysical logs and values of S hmin from DFITs are available. We compute a continuous stress profile as a function of the well logs that fits all of the DFITs well. We utilized several machine learning technologies, such as bootstrap aggregation (or bagging), to improve the generalization of the model and demonstrate that the excellent fit between predicted and observed stress values is not the result of over-fitting the calibration points. The model is then validated by accurately predicting hold-out stress measurements from four wells within the study area and, without recalibration, accurately predicting stress as a function of depth in an offset pad about 6 miles away.

58 GEOSCIENCES↗

Stream Temperature Predictions for River Basin Management in the Pacific Northwest and Mid-Atlantic Regions Using Machine Learning

Stream temperature (Ts) is an important water quality parameter that affects ecosystem health and human water use for beneficial purposes. Accurate Ts predictions at different spatial and temporal scales can inform water management decisions that account for the effects of changing climate and extreme events. In particular, widespread predictions of Ts in unmonitored stream reaches can enable decision makers to be responsive to changes caused by unforeseen disturbances. In this study, we demonstrate the use of classical machine learning (ML) models, support vector regression and gradient boosted trees (XGBoost), for monthly Ts predictions in 78 pristine and human-impacted catchments of the Mid-Atlantic and Pacific Northwest hydrologic regions spanning different geologies, climate, and land use. The ML models were trained using long-term monitoring data from 1980–2020 for three scenarios: (1) temporal predictions at a single site, (2) temporal predictions for multiple sites within a region, and (3) spatiotemporal predictions in unmonitored basins (PUB). In the first two scenarios, the ML models predicted Ts with median root mean squared errors (RMSE) of 0.69–0.84 °C and 0.92–1.02 °C across different model types for the temporal predictions at single and multiple sites respectively. For the PUB scenario, we used a bootstrap aggregation approach using models trained with different subsets of data, for which an ensemble XGBoost implementation outperformed all other modeling configurations (median RMSE 0.62 °C).The ML models improved median monthly Ts estimates compared to baseline statistical multi-linear regression models by 15–48% depending on the site and scenario. Air temperature was found to be the primary driver of monthly Ts for all sites, with secondary influence of month of the year (seasonality) and solar radiation, while discharge was a significant predictor at only 10 sites. The predictive performance of the ML models was robust to configuration changes in model setup and inputs, but was influenced by the distance to the nearest dam with RMSE <1 °C at sites situated greater than 16 and 44 km from a dam for the temporal single site and regional scenarios, and over 1.4 km from a dam for the PUB scenario. Our results show that classical ML models with solely meteorological inputs can be used for spatial and temporal predictions of monthly Ts in pristine and managed basins with reasonable (<1 °C) accuracy for most locations.

54 ENVIRONMENTAL SCIENCES↗

Modeling Weather Impact on Airport Arrival Miles-in-Trail Restrictions

When the demand for either a region of airspace or an airport approaches or exceeds the available capacity, miles-in-trail (MIT) restrictions are the most frequently issued traffic management initiatives (TMIs) that are used to mitigate these imbalances. Miles-intrail operations require aircraft in a traffic stream to meet a specific inter-aircraft separation in exchange for maintaining a safe and orderly flow within the stream. This stream of aircraft can be departing an airport, over a common fix, through a sector, on a specific route or arriving at an airport. This study begins by providing a high-level overview of the distribution and causes of arrival MIT restrictions for the top ten airports in the United States. This is followed by an in-depth analysis of the frequency, duration and cause of MIT restrictions impacting the Hartsfield-Jackson Atlanta International Airport (ATL) from 2009 through 2011. Then, machine-learning methods for predicting (1) situations in which MIT restrictions for ATL arrivals are implemented under low demand scenarios, and (2) days in which a large number of MIT restrictions are required to properly manage and control ATL arrivals are presented. More specifically, these predictions were accomplished by using an ensemble of decision trees with Bootstrap aggregation (BDT) and supervised machine learning was used to train the BDT binary classification models. The models were subsequently validated using data cross validation methods. When predicting the occurrence of arrival MIT restrictions under low demand situations, the model was able to achieve over all accuracy rates ranging from 84% to 90%, with false alarm ratios ranging from 10% to 15%. In the second set of studies designed to predict days on which a high number of MIT restrictions were required, overall accuracy rates of 80% were achieved with false alarm ratios of 20%. Overall, the predictions proposed by the model give better MIT usage information than what has been currently provided under current day operations. Traffic flow managers can use these predictions to identify potential MIT restrictions to eliminate (e.g., those occurring during low arrival demand periods), and to determine the days in which a significant number of restrictions may be required

Operation↗

A Step-Down Test Procedure for Wavelet Shrinkage Using Bootstrapping

Wavelet thresholding (or shrinkage) attempts to remove the noises existing in the signals while preserving inherent pattern characteristics in the reconstruction of true signals. For data-denoising purpose, we present a new wavelet thresholding procedure which employs the step-down testing idea of identifying active contrasts in unreplicated fractional factorial experiments. The proposed method employs bootstrapping methods to a step-down test for thresholding wavelet coefficients. By introducing the concept of a false discovery error rate in testing wavelet coefficients, we shrink the wavelet coefficients with p -values higher than the error rate. The error rate controls the expected proportion of wrongly accepted coefficients among chosen wavelet coefficients. Bootstrap samples are used to approximate the p -value for computational efficiency. We also present some guidelines for selecting the values of hyper-parameters which affect the performance in the step-down thresholding procedure. Based on some common testing signals and an air-conditioner sounds example, the comparison of our proposed procedure with other thresholding methods in the literature is performed. The analytical results show that the proposed procedure has a potential in data-denoising and data-reduction in a variety of signal reconstruction applications.

42 ENGINEERING↗

A bootstrapping approach to social media quantification

Abstract This work considers the use of classifiers in a downstream aggregation task estimating class proportions, such as estimating the percentage of reviews for a movie with positive sentiment. We derive the bias and variance of the class proportion estimator when taking classification error into account to determine how to best trade off different error types when tuning a classifier for these tasks. Additionally, we propose a method for constructing confidence intervals that correctly adjusts for classification error when estimating these statistics. We conduct experiments on four document classification tasks comparing our methods to prior approaches across classifier thresholds, sample sizes, and label distributions. Prior approaches have focused on providing the most accurate point estimate while this work focuses on the creation of correct confidence intervals that appropriately account for classifier error. Compared to the prior approaches, our methods provide lower error and more accurate confidence intervals.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Global Variability in Sonic Boom Exposure due to Macroscopic Effects

Supersonic flight over land has been prohibited since 1973 due to the loudness of sonic booms. NASA is building the X-59 aircraft as part of its Quesst mission to demonstrate low-loudness shaped sonic booms, or “sonic thumps.” The Quesst mission will gather human perception data via a series of community noise surveys across the USA. The noise dose and perceptual response data will be provided to the International Civil Aviation Organization (ICAO) and the Federal Aviation Administration for use in determining potential future supersonic aircraft noise certification standards, effectively changing the prohibition from a speed limit to a noise limit. These noise regulations must be globally effective, as long travel distances see the largest benefit to supersonic flight. The state of the atmosphere through which a sonic boom travels affects the size of the region exposed to sound, the “carpet width” (CW), as well as the loudness. The focus of this dissertation is to understand and quantify the expected loudness and CW of sonic booms due to the macroscopic atmospheric effects around the world. A pair of large-scale propagation simulation studies were conducted using the NASA PCBoom code to compare predicted sonic boom loudness and CW statistics first across the USA and then across the world. For the USA study, near-field data of the X-59 in steady cruise was propagated at 4 cardinal headings at 138 locations through 5 years of Climate Forecast System Version 2 (CFSv2) atmospheric profiles. Results of a bootstrap forest predictor screening model indicated the importance of climate zone, latitude, ground elevation, season, and heading. It also noted the unimportance of time of day for predicting loudness and CW. The data is visualized in aggregate, and then broken out geographically, by season and heading, and by climate zone. Multiple linear regression models were fit to the data from the 138 locations so that estimates of the loudness and CW can be produced anywhere in the US. The results can aid in planning when and where to fly the X-59. For the global study, near-field data from three aircraft, the X-59 in a quiet and loud configuration, B-58, and Concorde, were propagated at four cardinal headings through data from three atmospheric models, the CFSv2, the Global Forecast System (GFS), and the ECMWF Reanalysis Version 5 (ERA5), at 100 global locations over 1 year. Results of a bootstrap forest predictor screening model indicated the importance of climate zone, ground elevation, season, and heading. Similar to the US study, the model indicated time of day was not an important predictor. The model also indicated that choice of weather model was not important, so the atmospheric model data are effectively interchangeable. The ERA5 model was chosen for use in an extension of the study to include 18 additional locations to ensure sampling of every climate zone. Loudness and CW results are shown in aggregate, and split geographically and by heading, season, and climate. Multiple linear regression models were fit to the data from the 118 locations so that estimates of loudness and CW can be produced around the world. N-waves and shaped booms did not have the same global variability. Koppen-Geiger climate zones were used as the climate zone definition for the global study. These are available as present-day and future climate projections. Making use of the multiple linear regression models, the future climate zones were input to estimate the effect of the changing climate on sonic boom loudness and CW. Results indicate that a changing climate would have little impact on the effectiveness of noise regulations.

X-59↗

Conservation management decreases surface runoff and soil erosion

Conservation management practices – including agroforestry, cover cropping, no-till, reduced tillage, and residue return – have been applied for decades to control surface runoff and soil erosion, yet results have not been integrated and evaluated across cropping systems. In this study we collected data comparing agricultural production with and without conservation management strategies. We used a bootstrap resampling analysis to explore interactions between practice type, soil texture, surface runoff, and soil erosion. We then used a correlation analysis to relate changes in surface runoff and soil erosion to 13 other soil health and agronomic indicators, including soil organic carbon, soil aggregation, infiltration, porosity, subsurface leaching, and cash crop yield. Across all conservation management practices, surface runoff and erosion had respective mean decreases of 67% and 80% compared with controls. Use of cover cropping provided the largest decreases in erosion and surface runoff, thus emphasizing the importance of maintaining continuous vegetative cover on soils. Coarse- and medium-textured soils had greater decreases in both erosion and runoff than fine-textured soils. Changes in surface runoff and soil erosion under conservation management were highly correlated with soil organic carbon, aggregation, porosity, infiltration, leaching, and yield, showing that conservation practices help drive important interactions between these different facets of soil health. This study offers the first large-scale comparison of how different conservation agriculture practices reduce surface runoff and soil erosion, and at the same time provides new insight into how these interactions influence the improvement or loss of soil health.

58 GEOSCIENCES↗

A Framework for Inverse Prediction Using Functional Response Data

Inverse prediction models have commonly been developed to handle scalar data from physical experiments. However, it is not uncommon for data to be collected in functional form. When data are collected in functional form, it must be aggregated to fit the form of traditional methods, which often results in a loss of information. For expensive experiments, this loss of information can be costly. In this study, we introduce the functional inverse prediction (FIP) framework, a general approach which uses the full information in functional response data to provide inverse predictions with probabilistic prediction uncertainties obtained with the bootstrap. The FIP framework is a general methodology that can be modified by practitioners to accommodate many different applications and types of data. We demonstrate the framework, highlighting points of flexibility, with a simulation example and applications to weather data and to nuclear forensics. Results show how functional models can improve the accuracy and precision of predictions.

42 ENGINEERING↗

Generalized Bootstrap AMG and AIR-AMG for coupled PDE systems with a focus on spacetime discretizatoins

The Pennsylvania State University (“Subcontractor”) has worked on the design of new algebraic, parallel, multilevel methods that obtain the full space and time solution of systems of PDEs. In particular, the PI and his collaborators explored semi-intrusive approaches based on algebraic multigrid (AMG). The focus of the research has been on the development of these techniques for the Euler equations in 1d and 2d. The overall research focused on the development of adaptive AIR (approximate ideal restriction) AMG solvers for these problems. The PI also explored the use of smoothed aggregation and root-node energy-based AMG solvers for these problems.

97 MATHEMATICS AND COMPUTING↗

Comparing traditional and Bayesian approaches to ecological meta‐analysis

Abstract Despite the wide application of meta‐analysis in ecology, some of the traditional methods used for meta‐analysis may not perform well given the type of data characteristic of ecological meta‐analyses. We reviewed published meta‐analyses on the ecological impacts of global climate change, evaluating the number of replicates used in the primary studies ( n i ) and the number of studies or records ( k ) that were aggregated to calculate a mean effect size. We used the results of the review in a simulation experiment to assess the performance of conventional frequentist and Bayesian meta‐analysis methods for estimating a mean effect size and its uncertainty interval. Our literature review showed that n i and k were highly variable, distributions were right‐skewed and were generally small (median n i = 5, median k = 44). Our simulations show that the choice of method for calculating uncertainty intervals was critical for obtaining appropriate coverage (close to the nominal value of 0.95). When k was low (<40), 95% coverage was achieved by a confidence interval (CI) based on the t distribution that uses an adjusted standard error (the Hartung–Knapp–Sidik–Jonkman, HKSJ), or by a Bayesian credible interval, whereas bootstrap or z distribution CIs had lower coverage. Despite the importance of the method to calculate the uncertainty interval, 39% of the meta‐analyses reviewed did not report the method used, and of the 61% that did, 94% used a potentially problematic method, which may be a consequence of software defaults. In general, for a simple random‐effects meta‐analysis, the performance of the best frequentist and Bayesian methods was similar for the same combinations of factors ( k and mean replication), though the Bayesian approach had higher than nominal (>95%) coverage for the mean effect when k was very low ( k < 15). Our literature review suggests that many meta‐analyses that used z distribution or bootstrapping CIs may have overestimated the statistical significance of their results when the number of studies was low; more appropriate methods need to be adopted in ecological meta‐analyses.

Pappalardo, Paula↗

Data from: Comparing traditional and Bayesian approaches to ecological meta-analysis

Despite the wide application of meta-analysis in ecology, some of the traditional methods used for meta-analysis may not perform well given the type of data characteristic of ecological meta-analyses. We reviewed published meta-analyses on the ecological impacts of global climate change, evaluating the number of replicates used in the primary studies (ni) and the number of studies or records (k) that were aggregated to calculate a mean effect size. We used the results of the review in a simulation experiment to assess the performance of conventional frequentist and Bayesian meta-analysis methods for estimating a mean effect size and its uncertainty interval. Our literature review showed that ni and k were highly variable, distributions were right-skewed, and were generally small (median ni =5, median k=44). Our simulations show that the choice of method for calculating uncertainty intervals was critical for obtaining appropriate coverage (close to the nominal value of 0.95). When k was low (<40), 95% coverage was achieved by a confidence interval based on the t-distribution that uses an adjusted standard error (the Hartung-Knapp-Sidik-Jonkman, HKSJ), or by a Bayesian credible interval, whereas bootstrap or z-distribution confidence intervals had lower coverage. Despite the importance of the method to calculate the uncertainty interval, 39% of the meta-analyses reviewed did not report the method used, and of the 61% that did, 94% used a potentially problematic method, which may be a consequence of software defaults. In general, for a simple random-effects meta-analysis, the performance of the best frequentist and Bayesian methods were similar for the same combinations of factors (k and mean replication), though the Bayesian approaches had higher than nominal (>95%) coverage for the mean effect when k was very low (k<15). Our literature review suggests that many meta-analyses that used z-distribution or bootstrapping confidence intervals may have over-estimated the statistical significance of their results when the number of studies was low; more appropriate methods need to be adopted in ecological meta-analyses.

54 ENVIRONMENTAL SCIENCES↗