Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Multivariate regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Upscaling Soil Organic Carbon Measurements at the Continental Scale Using Multivariate Clustering Analysis and Machine Learning

Abstract Estimates of soil organic carbon (SOC) stocks are essential for many environmental applications. However, significant inconsistencies exist in SOC stock estimates for the U.S. across current SOC maps. We propose a framework that combines unsupervised multivariate geographic clustering (MGC) and supervised Random Forests regression, improving SOC maps by capturing heterogeneous relationships with SOC drivers. We first used MGC to divide the U.S. into 20 SOC regions based on the similarity of covariates (soil biogeochemical, bioclimatic, biological, and physiographic variables). Subsequently, separate Random Forests models were trained for each SOC region, utilizing environmental covariates and SOC observations. Our estimated SOC stocks for the U.S. (52.6 ± 3.2 Pg for 0–30 cm and 108.3 ± 8.2 Pg for 0–100 cm depth) were within the range estimated by existing products like Harmonized World Soil Database, HWSD (46.7 Pg for 0–30 cm and 90.7 Pg for 0–100 cm depth) and SoilGrids 2.0 (45.7 Pg for 0–30 cm and 133.0 Pg for 0–100 cm depth). However, independent validation with soil profile data from the National Ecological Observatory Network showed that our approach ( R 2 = 0.51) outperformed the estimates obtained from Harmonized World Soil Database ( R 2 = 0.23) and SoilGrids 2.0 ( R 2 = 0.39) for the topsoil (0–30 cm). Uncertainty analysis (e.g., low representativeness and high coefficients of variation) identified regions requiring more measurements, such as Alaska and the deserts of the U.S. Southwest. Our approach effectively captures the heterogeneous relationships between widely available predictors and the current SOC baseline across regions, offering reliable SOC estimates at 1 km resolution for benchmarking Earth system models.

58 GEOSCIENCES↗

Data augmentation for disruption prediction via robust surrogate models

The goal of this work is to generate large statistically representative datasets to train machine learning models for disruption prediction provided by data from few existing discharges. Such a comprehensive training database is important to achieve satisfying and reliable prediction results in artificial neural network classifiers. Here, we aim for a robust augmentation of the training database for multivariate time series data using Student-t process regression. We apply Student-t process regression in a state space formulation via Bayesian filtering to tackle challenges imposed by outliers and noise in the training data set and to reduce the computational complexity. Thus, the method can also be used if the time resolution is high. We use an uncorrelated model for each dimension and impose correlations afterwards via coloring transformations. We demonstrate the efficacy of our approach on plasma diagnostics data of three different disruption classes from the DIII-D tokamak. To evaluate if the distribution of the generated data is similar to the training data, we additionally perform statistical analyses using methods from time series analysis, descriptive statistics, and classic machine learning clustering algorithms.

97 MATHEMATICS AND COMPUTING↗

mvBayesR

SAND2025-11559O The mvBayesR tool performs multivariate Bayesian analysis on generic data. It includes tools for regression modeling, diagnosis, basis decomposition, sensitivity analysis, and visualization. The tool compiles state-of-the-art methodology into one easy-to-use package. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Tucker, James [Sandia National Lab. (SNL-CA), Live↗

mvBayesPy

SAND2025-11476O The mvBayesPy tool is a Python package that performs multivariate Bayesian analysis on generic data. It includes tools for regression modeling, diagnosis, basis decomposition, sensitivity analysis and visualization. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Tucker, James [Sandia National Lab. (SNL-CA), Live↗

MSW Variability Mapping and Conversion to Biofuel

MSW (Municipal Solid Waste) is a form of biomass which consists of categorized components of waste/trash. The general categories are paper, yard trash, construction & debris, appliances, tires, glass, metals, aluminum & steel cans, plastics, organics, inorganics, and HHW (Household Hazardous Waste). This project focuses on the factors within a region or population that contribute to variability in the composition of MSW and in turn MSW’s convertibility to biofuel. A list of contributors was determined (Social Vulnerability Index, Access to Public Transportation, Racial Distribution, GDP, Personal Income) and then JMP was used to perform a Multivariate analysis to determine correlations and a Partial Least-Squares regression to determine Variable Importance Plots for each MSW category. In addition to data analysis, the convertibility of MSW to biofuel was studied via microwave pyrolysis system in order to separate and characterize the various gaseous and bio-oil products.

09 BIOMASS FUELS↗

Bayesian projection pursuit regression

In projection pursuit regression (PPR), a univariate response variable is approximated by the sum of $M$ “ridge functions,” which are flexible functions of one-dimensional projections of a multivariate input variable. Traditionally, optimization routines are used to choose the projection directions and ridge functions via a sequential algorithm, and $M$ is typically chosen via cross-validation. Here, we introduce a novel Bayesian version of PPR, which has the benefit of accurate uncertainty quantification. To infer appropriate projection directions and ridge functions, we apply novel adaptations of methods used for the single ridge function case ($M$=1), called the Bayesian Single Index Model; and use a Reversible Jump Markov chain Monte Carlo algorithm to infer the number of ridge functions $M$. We evaluate the predictive ability of our model in 20 simulated scenarios and for 23 real datasets, in a bake-off against an array of state-of-the-art regression methods. Finally, we generalize this methodology and demonstrate the ability to accurately model multivariate response variables. Its effective performance indicates that Bayesian Projection Pursuit Regression is a valuable addition to the existing regression toolbox.

97 MATHEMATICS AND COMPUTING↗

Analyzing Wildland Fire Smoke Emissions Data Using Compositional Data Techniques

By conservation of mass, the mass of wildland fuel that is pyrolyzed and combusted must equal the mass of smoke emissions, residual char and ash. For a given set of conditions, these amounts are fixed. This places a constraint on smoke emissions data which violates statistical assumptions for many of the methods currently used to analyze these data such as linear regression, analysis of variance, and t-tests. These data are inherently multivariate and non-negative parts of a whole. This paper introduces the field of compositional data analysis to the emissions community and provides examples of appropriate statistical treatment of emissions data. It is shown that modified combustion efficiency should not be used as a predictor variable for other smoke emissions because it is not an independent variable. An alternative method based on compositional linear trends to estimate trace gas composition using CO and CO2 is presented. The data used in this paper resulted from projects the DOD/DOE/EPA Strategic 586 Environmental Research and Development Program projects RC-1648 and 1649. The senior 587 author appreciates the guidance and R scripts provided by Prof. Girty at San Diego State 588 University to estimate linear trends by perturbation. J. P.-A. was supported by the Spanish 589 Ministry of Science, Innovation and Universities under the project CODAMET (RTI2018-590 095518-B-C21, 2019-2021). The data used in this study have been previously published and are 591 available in the original publications. DRW conceived the initial manuscript (70 percent) and 592 performed the bulk of the data analysis. JPA provided statistical guidance and compositional data 593 expertise and contributed 20 percent of the manuscript. TJJ and HJ were extensively involved in 594 the study that provided the data. TJJ provide smoke emissions expertise and HJ provided 595 combustion expertise. The authors declare that they have no conflict of interest. The use of trade 596 or firm names in this publication is for reader information and does not imply endorsement by 597 the U.S. Department of Agriculture of any product or service.

simplex, compositional data analysis, balance, log↗

Monitoring Noble Gases (Xe and Kr) and Aerosols (Cs and Rb) in a Molten Salt Reactor Surrogate Off-Gas Stream Using Laser-Induced Breakdown Spectroscopy (LIBS)

In this study with surrogate materials we show that laser-induced breakdown spectroscopy (LIBS) is a robust tool with promising capability toward monitoring gaseous (Xe and Kr) and aerosol (Cs and Rb) species in an off-gas stream from a molten salt reactor (MSR). MSRs will continually evolve fission products into the cover gas flowing across the reactor headspace. The cover gas entrains Xe and Kr gases, along with aerosol particles, before passing into an off-gas treatment system. Univariate models of Xe and Kr peaks showed a strong correlation to concentration indicated by their coefficients of determination of 0.983 and 0.997, respectively. Multivariate models were built for all four analytes using partial least squares regression coupled with preprocessing steps including normalization, trimming, and/or genetic algorithm derived filters. The models were evaluated by predicting the concentrations of the analytes in four validation samples, in which all calibration models were successfully validated at a confidence interval of 99.9%. Finally, pressure controllers were used to regulate the mass flow rate of Kr flowing into the measurement cell in sinusoidal and stepwise waveforms to test the real-time monitoring capabilities of the regression models. Both univariate and partial least squares Kr models were able to successfully quantify the gas concentration in the real-time evaluation. The root mean squared error of prediction (RMSEP) values for these real-time tests were calculated to be 0.051, 0.060, and 0.121 mol% demonstrating the measurement systems’ capability to perform online monitoring with acceptable accuracy.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Microbial community diversity changes during voltage reversal repair in a 12-unit microbial fuel cell

Microbial fuel cell stacks (MFC-Stack) are often confronted with voltage reversals, likely due to an interplay between microbial community dynamics and insufficient electric circuit balancing. Herein, we provide new insight into voltage reversals by examining the microbiomes of twelve MFC units of a 12-liter Pilot-MFC-Stack during repair. Different biofilm repair methods (self-healing, electrostimulation, and re-acclimatization upon cross-inoculation) were used to evaluate the microbial community response. In addition, MFC-Stack simulation was performed based on Kirchhoff’s Second Law to predict values for source potentials and post-evaluate internal resistances. Analysis of the 16S rRNA amplicon sequencing data suggests that the biofilm repair methods could slowly heal damaged biofilms. Notably, severely voltage reversed MFC units had low electrogen relative abundances (18%) and positive anode potentials, while strong bioanodes and contained more than 50% electrogens and had negative anode potentials. Between-community analyses (beta diversity ordination and multinomial regression) of the voltage reversed MFC units revealed differences among biofilms in contrast to healthy/strong MFC units. Permutational multivariate analysis of variance (PERMANOVA) confirmed that reversed biofilms were, indeed, significantly (p < 0.05) different from stronger ones. Overall, these analyses demonstrated the utility of combining electrotechnical and microbial community analyses, especially beta diversity ordination and multinomial regression, to understand problematic MFC units and the potential success of a biofilm repair method. Finally, thicker biofilms were usually healthier and stronger, although thickness was no guarantee for proper structure and power function as all factors were interdependent. There was an evolutionary trend that strong anodes became stronger/healthier and others weaker. This spontaneous trend has to be considered to avoid irreversible voltage reversals and to repair electrogenic biofilms in an MFC-Stack.

59 BASIC BIOLOGICAL SCIENCES↗

Machine Learning Approach for Spatiotemporal Multivariate Optimization of Environmental Monitoring Sensor Locations

Abstract Long-term environmental monitoring is critical for managing the soil and groundwater at contaminated sites. Recent improvements in state-of-the-art sensor technology, communication networks, and artificial intelligence have created opportunities to modernize this monitoring activity for automated, fast, robust, and predictive monitoring. In such modernization, it is required that sensor locations be optimized to capture the spatiotemporal dynamics of all monitoring variables as well as to make it cost-effective. The legacy monitoring datasets of the target area are important to perform this optimization. In this study, we have developed a machine-learning approach to optimize sensor locations for soil and groundwater monitoring based on ensemble supervised learning and majority voting. For spatial optimization, Gaussian process regression (GPR) is used for spatial interpolation, while the majority voting is applied to accommodate the multivariate temporal dimension. Results show that the algorithms significantly outperform the random selection of the sensor locations for predictive spatiotemporal interpolation. While the method has been applied to a four-dimensional dataset (with two-dimensional space, time, and multiple contaminants), we anticipate that it can be generalizable to higher-dimensional datasets for environmental monitoring sensor location optimization.

Siddiquee, Masudur R.↗

Raman Spectroscopy Coupled with Chemometric Analysis for Speciation and Quantitative Analysis of Aqueous Phosphoric Acid Systems

Complex chemical systems that exhibit varied and matrix-dependent speciation are notoriously difficult to monitor and characterize on-line and in real-time. Optical spectroscopy is an ideal tool for in situ characterization of chemical species that can enable quantification as well as species identification. Chemometric modeling, a multivariate method, has been successfully paired with optical spectroscopy to enable measurement of analyte concentrations even in complex solutions where univariate methods such as Beer’s law analysis fail. Here, Raman spectroscopy is used to quantify the concentration of phosphoric acid and its three deprotonated forms during a titration. In this system, univariate approaches would be difficult to apply due to multiple species being present simultaneously within the solution as pH is varied. Locally Weighted Regression (LWR) modeling was used to determine phosphate concentration from spectral signature. LWR results, in tandem with Multivariate Curve Resolution modeling, provide direct measurement of the concentration of each phosphate species using only the Raman signal. Furthermore, results are presented within the context of fundamental solution chemistry, including Pitzer equations to compensate for activity coefficients and non-idealities associated with high ionic strength systems.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Uncertain characterization of reservoir fluids due to brittleness of equation of state regression

Equations of state (EoS) play a central role in modeling the phase equilibrium of fluid mixtures. Their parameterization involves fitting a model to experimental data, i.e., solving a nonlinear, non-convex, multivariate optimization problem. The latter requires one to select design variables, domains of definition for each variable, and weights assigned to individual measurements. We demonstrate that subjective choices of an optimization algorithm and an initial guess also impact the regression process. Consequently, EoS predictions are fundamentally uncertain even after the EoS tuning to a limited set of experimental data points. We demonstrate this observation for two hydrocarbon reservoir fluids, in which five properties of the heaviest carbon fraction are treated as design variables. While all the optimization algorithms and initial guesses match experimental data for the gas and liquid properties, the resulting EoS parameterizations lead to dramatically different predictions of the fluid’s thermophysical behavior in the unsampled pressure and temperature regions. In conclusion, we propose the probabilistic treatment of design variables to quantify the predictive uncertainty of the resulting fluid models.

15 GEOTHERMAL ENERGY↗

Dietary B group vitamin intake and the bladder cancer risk: a pooled analysis of prospective cohort studies

Abstract Purpose Diet may play an essential role in the aetiology of bladder cancer (BC). The B group complex vitamins involve diverse biological functions that could be influential in cancer prevention. The aim of the present study was to investigate the association between various components of the B group vitamin complex and BC risk. Methods Dietary data were pooled from four cohort studies. Food item intake was converted to daily intakes of B group vitamins and pooled multivariate hazard ratios (HRs), with corresponding 95% confidence intervals (CIs), were obtained using Cox-regression models. Dose–response relationships were examined using a nonparametric test for trend. Results In total, 2915 BC cases and 530,012 non-cases were included in the analyses. The present study showed an increased BC risk for moderate intake of vitamin B1 (HR B1 : 1.13, 95% CI: 1.00–1.20). In men, moderate intake of the vitamins B1, B2, energy-related vitamins and high intake of vitamin B1 were associated with an increased BC risk (HR (95% CI): 1.13 (1.02–1.26), 1.14 (1.02–1.26), 1.13 (1.02–1.26; 1.13 (1.02–1.26), respectively). In women, high intake of all vitamins and vitamin combinations, except for the entire complex, showed an inverse association (HR (95% CI): 0.80 (0.67–0.97), 0.83 (0.70–1.00); 0.77 (0.63–0.93), 0.73 (0.61–0.88), 0.82 (0.68–0.99), 0.79 (0.66–0.95), 0.80 (0.66–0.96), 0.74 (0.62–0.89), 0.76 (0.63–0.92), respectively). Dose–response analyses showed an increased BC risk for higher intake of vitamin B1 and B12. Conclusion Our findings highlight the importance of future research on the food sources of B group vitamins in the context of the overall and sex-stratified diet.

60 APPLIED LIFE SCIENCES↗

Gaining Perspective on Unconventional Well Design Choices through Play-level Application of Machine Learning Modeling

The recent development of unconventional oil and gas (O&G) reservoirs has led to an abundant hydrocarbon supply, both domestically and globally. However, there is a continued push to develop new and innovative approaches to improve exploration and extraction efficiencies and overall well productivity moving forward. Substantial improvements in unconventional O&G development are expected through optimized well completion and stimulation strategies aimed at maximizing well productivity. Optimizing well designs will require tailoring to the distinctive geologic conditions present for any newly placed well. To better evaluate the impact of well design attributes and their associated interactions on productivity in a major unconventional play, multivariate machine learning-based models that use empirical datasets were developed. A gradient boosted regression tree (GBRT) algorithm was applied. GBRT has been narrowly investigated for O&G applications but enables straightforward parametric importance and influence evaluation, as well as assessment of parameter interaction effects. Models were trained on well design and locational parameters that serve as a proxy for variable geologic conditions to estimate two types of productivity indicator response variables strongly correlated to estimated ultimate recovery (EUR). The dataset utilized consists of over 7,000 well observations that cover the majority of the productive region of the Marcellus Shale. Model performance was evaluated and algorithm parameters tuned by analyzing the goodness-of-fit for simulated results against observed data in a cross-validation approach. Models were found capable of 73–79 percent prediction accuracy on held out testing data of gas equivalent production and can be used to inform future well design and placement decisions for increasing EUR per well and improving overall field-level recovery. Study results indicate that Marcellus well performance improves most with upscaling perforated interval lengths and water and proppant volumes per foot; but relative productivity improvements are spatially dependent across the play. Finally, optimal combinations of water and proppant on well performance were found to vary depending on well location, emphasizing the utility of data-driven models capable of broad application across a play of interest for informing tailored well design approaches prior to their field deployment.

04 OIL SHALES AND TAR SANDS↗

Crowd-based spatial risk assessment of urban flooding: Results from a municipal flood hotline in Detroit, MI

Climate change is increasing the frequency and intensity of extreme precipitation events, raising the risk of urban flood disasters. This study uses a crowd-sourced municipal call database to characterize the spatial distribution of flood risk in Detroit, MI. Call data including dates and addresses were obtained from the City of Detroit Department of Public Works for 2021. Calls were mapped and aggregated to census tract counts and merged with neighborhood-level data. Associations of predictors with flood calls were tested using spatial regression models. Flooding calls were located throughout the city but were concentrated in specific areas. Multivariate models of census tract level call counts indicated that increased poverty and Black, immigrant, and older residents were positively associated with flood calls, while increased elevation was associated with protective effects. Longer distances from waste water interceptors were associated with higher risk for calls. Crowd-sourced flood hotline call data can be used for effective spatial flood risk assessment. Though flooding occurs throughout the city of Detroit, infrastructural, neighborhood, and household factors influence flooding extent. Limitations included the self-reported nature of calls. Future modeling efforts might include input from local stakeholders to improve spatial risk assessment.

54 ENVIRONMENTAL SCIENCES↗

Vegetable intake and the risk of bladder cancer in the BLadder Cancer Epidemiology and Nutritional Determinants (BLEND) international study

Although a potential inverse association between vegetable intake and bladder cancer risk has been reported, epidemiological evidence is inconsistent. This research aimed to elucidate the association between vegetable intake and bladder cancer risk by conducting a pooled analysis of data from prospective cohort studies. Vegetable intake in relation to bladder cancer risk was examined by pooling individual-level data from 13 cohort studies, comprising 3203 cases among a total of 555,685 participants. Pooled multivariate hazard ratios (HRs), with corresponding 95% confidence intervals (CIs), were estimated using Cox proportional hazards regression models stratified by cohort for intakes of total vegetable, vegetable subtypes (i.e. non-starchy, starchy, green leafy and cruciferous vegetables) and individual vegetable types. In addition, a diet diversity score was used to assess the association of the varied types of vegetable intake on bladder cancer risk. The association between vegetable intake and bladder cancer risk differed by sex ( P -interaction = 0.011) and smoking status ( P -interaction = 0.038); therefore, analyses were stratified by sex and smoking status. With adjustment of age, sex, smoking, energy intake, ethnicity and other potential dietary factors, we found that higher intake of total and non-starchy vegetables were inversely associated with the risk of bladder cancer among women (comparing the highest with lowest intake tertile: HR = 0.79, 95% CI = 0.64–0.98, P = 0.037 for trend, HR per 1 SD increment = 0.89, 95% CI = 0.81–0.99; HR = 0.78, 95% CI = 0.63–0.97, P = 0.034 for trend, HR per 1 SD increment = 0.88, 95% CI = 0.79–0.98, respectively). However, no evidence of association was observed among men, and the intake of vegetable was not found to be associated with bladder cancer when stratified by smoking status. Moreover, we found no evidence of association for diet diversity with bladder cancer risk. Higher intakes of total and non-starchy vegetable are associated with reduced risk of bladder cancer for women. Further studies are needed to clarify whether these results reflect causal processes and potential underlying mechanisms.

59 BASIC BIOLOGICAL SCIENCES↗

Patterns of Element Incorporation in Calcium Carbonate Biominerals Recapitulate Phylogeny for a Diverse Range of Marine Calcifiers

Elemental ratios in biogenic marine calcium carbonates are widely used in geobiology, environmental science, and paleoenvironmental reconstructions. It is generally accepted that the elemental abundance of biogenic marine carbonates reflects a combination of the abundance of that ion in seawater, the physical properties of seawater, the mineralogy of the biomineral, and the pathways and mechanisms of biomineralization. Here we report measurements of a suite of nine elemental ratios (Li/Ca, B/Ca, Na/Ca, Mg/Ca, Zn/Ca, Sr/Ca, Cd/Ca, Ba/Ca, and U/Ca) in 18 species of benthic marine invertebrates spanning a range of biogenic carbonate polymorph mineralogies (low-Mg calcite, high-Mg calcite, aragonite, mixed mineralogy) and of phyla (including Mollusca, Echinodermata, Arthropoda, Annelida, Cnidaria, Chlorophyta, and Rhodophyta) cultured at a single temperature (25°C) and a range of p CO 2 treatments (ca. 409, 606, 903, and 2856 ppm). This dataset was used to explore various controls over elemental partitioning in biogenic marine carbonates, including species-level and biomineralization-pathway-level controls, the influence of internal pH regulation compared to external pH changes, and biocalcification responses to changes in seawater carbonate chemistry. The dataset also enables exploration of broad scale phylogenetic patterns of elemental partitioning across calcifying species, exhibiting high phylogenetic signals estimated from both uni- and multivariate analyses of the elemental ratio data (univariate: λ = 0–0.889; multivariate: λ = 0.895–0.99). Comparing partial R 2 values returned from non-phylogenetic and phylogenetic regression analyses echo the importance of and show that phylogeny explains the elemental ratio data 1.4–59 times better than mineralogy in five out of nine of the elements analyzed. Therefore, the strong associations between biomineral elemental chemistry and species relatedness suggests mechanistic controls over element incorporation rooted in the evolution of biomineralization mechanisms.

58 GEOSCIENCES↗

Within and Among Fish Species Differences in Simulated Turbine Blade Strike Mortality: Limits on the Use of Surrogacy for Untested Species

Use of surrogacy remains a useful method for prioritizing research on representatives of at-risk groups of fishes, yet quantifiable evidence in support of its use is generally not available. Blade strike impact represents one of the most traumatic stressors experienced by fish during non-volitional movements through hydropower turbines. Here, we use data generated from laboratory trials on blade strike impact experiments to directly test use of surrogacy for salmonid and clupeid fishes. Results of logistic regression indicated that a -taxonomic (genus) variable was not a significant predictor of mortality among large rainbow trout and brook trout. Similar results were found for young-of-the-year shad species, but genus-level taxonomy was a significant predictor of mortality while species was not. Multivariate analysis of morphometric data showed that shad clustered together based on similarities in fish shape which was also closely associated with genus. Logistic regression including size as a major covariate suggested total fish length was not a significant predictor of mortality, yet dose–response data suggest differential susceptibility to lower strike velocities. We suggest that use of surrogacy among species is justifiable but should be avoided within a species since the effects of size remain unclear.

59 BASIC BIOLOGICAL SCIENCES↗