Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Random forests”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Feature-Based PMU Event Classification under Variable PMU Participation and Overlapping Events

Danovo Energy Solution's presented its paper named: Feature-Based PMU Event Classification under Variable PMU Participation and Overlapping Events at the 2026 Georgia Tech Fault & Disturbance Analysis Conference. The full paper can be found at OSTI ID# 3169150 Paper Abstract—Phasor Measurement Units (PMUs) stream time synchronized, high-resolution measurements from the grid, enabling data-driven techniques for event detection and classification. Accurate event classification improves grid reliability and stability. Events can be detected by varying numbers of PMUs and exhibit different durations depending on the event type. This variability challenges standard classifiers that require uniform input sizes. Moreover, multiple events may coincide, which increases classification complexity. Standard classifiers assign each instance to the class with the highest predicted probability, whereas overlapping events may exhibit comparable probabilities across multiple classes. In this study, to handle data size variability, we extract a wide range of time–frequency domain features from all available PMUs for each event into a fixed-length vector, facilitating the application of standard machine learning classifiers, including Random Forest, XGBoost, LightGBM, Support Vector Machine, and Multilayer Perceptron. To account for overlapping events, a probabilistic post-processing step is applied. For a given data instance, if multiple predicted class probabilities exceed 30% and the differences between them are less than 10%, the event is assigned to multiple classes. Experiments using real-world PMU data demonstrate that the Random Forest and XGBoost models achieve the highest accuracy, while the proposed post-processing method yields perfect classification performance on external unseen test sets.

Nematirad, Reza [Danova Energy Solutions]↗

Improving Medication Regimen Recommendation for Parkinson’s Disease Using Sensor Technology

Parkinson’s disease medication treatment planning is generally based on subjective data obtained through clinical, physician-patient interactions. The Personal KinetiGraph™ (PKG) and similar wearable sensors have shown promise in enabling objective, continuous remote health monitoring for Parkinson’s patients. In this proof-of-concept study, we propose to use objective sensor data from the PKG and apply machine learning to cluster patients based on levodopa regimens and response. The resulting clusters are then used to enhance treatment planning by providing improved initial treatment estimates to supplement a physician’s initial assessment. We apply k-means clustering to a dataset of within-subject Parkinson’s medication changes—clinically assessed by the MDS-Unified Parkinson’s Disease Rating Scale-III (MDS-UPDRS-III) and the PKG sensor for movement staging. A random forest classification model was then used to predict patients’ cluster allocation based on their respective demographic information, MDS-UPDRS-III scores, and PKG time-series data. Clinically relevant clusters were partitioned by levodopa dose, medication administration frequency, and total levodopa equivalent daily dose—with the PKG providing similar symptomatic assessments to physician MDS-UPDRS-III scores. A random forest classifier trained on demographic information, MDS-UPDRS-III scores, and PKG time-series data was able to accurately classify subjects of the two most demographically similar clusters with an accuracy of 86.9%, an F1 score of 90.7%, and an AUC of 0.871. A model that relied solely on demographic information and PKG time-series data provided the next best performance with an accuracy of 83.8%, an F1 score of 88.5%, and an AUC of 0.831, hence further enabling fully remote assessments. These computational methods demonstrate the feasibility of using sensor-based data to cluster patients based on their medication responses with further potential to assist with medication recommendations.

59 BASIC BIOLOGICAL SCIENCES↗

Feature-Based PMU Event Classification under Variable PMU Participation and Overlapping Events

This paper is the basis for a presentation help at the 2026 Georgia Tech Fault & Disturbance Analysis Conference, which can be found at OSTI # 3168287 Paper Abstract—Phasor Measurement Units (PMUs) stream time synchronized, high-resolution measurements from the grid, enabling data-driven techniques for event detection and classification. Accurate event classification improves grid reliability and stability. Events can be detected by varying numbers of PMUs and exhibit different durations depending on the event type. This variability challenges standard classifiers that require uniform input sizes. Moreover, multiple events may coincide, which increases classification complexity. Standard classifiers assign each instance to the class with the highest predicted probability, whereas overlapping events may exhibit comparable probabilities across multiple classes. In this study, to handle data size variability, we extract a wide range of time–frequency domain features from all available PMUs for each event into a fixed-length vector, facilitating the application of standard machine learning classifiers, including Random Forest, XGBoost, LightGBM, Support Vector Machine, and Multilayer Perceptron. To account for overlapping events, a probabilistic post-processing step is applied. For a given data instance, if multiple predicted class probabilities exceed 30% and the differences between them are less than 10%, the event is assigned to multiple classes. Experiments using real-world PMU data demonstrate that the Random Forest and XGBoost models achieve the highest accuracy, while the proposed post-processing method yields perfect classification performance on external unseen test sets.

Nematirad, Reza [Danovo Energy Solutions]↗

Predicting the Operational Acceptance of Airborne Flight Reroute Requests Using Data Mining

For tools that generate more efficient flight routes or reroute advisories, it is important to ensure compatibility of automation and autonomy decisions with human objectives so as to ensure acceptability by the human operators. In this paper, the authors developed a proof of concept predictor of operational acceptability for route changes during a flight. Such a capability could have applications in automation tools that identify more efficient routes around airspace impacted by weather or congestion and that better meet airline preferences. The predictor is based on applying data mining techniques, including logistic regression, a decision tree, a support vector machine, a random forest and Adaptive Boost, to historical flight plan amendment data reported during operations and field experiments. Cross validation was used for model development, while nested cross validation was used to validate the models. The model found to have the best performance in predicting air traffic controller acceptance or rejection of a route change, using the available data from Fort Worth Air Traffic Control Center and its adjacent Centers, was the random forest, with an F-score of 0.77. This result indicates that the operational acceptance of reroute requests does indeed have some level of predictability, and that, with suitable data, models can be trained to predict the operational acceptability of reroute requests. Such models may ultimately be used to inform route selection by decision support tools, contributing to the development of increasingly autonomous systems that are capable of routing aircraft with less human input than is currently the case.

Operational Acceptability↗

MArVD2: a machine learning enhanced tool to discriminate between archaeal and bacterial viruses in viral datasets

Abstract Our knowledge of viral sequence space has exploded with advancing sequencing technologies and large-scale sampling and analytical efforts. Though archaea are important and abundant prokaryotes in many systems, our knowledge of archaeal viruses outside of extreme environments is limited. This largely stems from the lack of a robust, high-throughput, and systematic way to distinguish between bacterial and archaeal viruses in datasets of curated viruses. Here we upgrade our prior text-based tool (MArVD) via training and testing a random forest machine learning algorithm against a newly curated dataset of archaeal viruses. After optimization, MArVD2 presented a significant improvement over its predecessor in terms of scalability, usability, and flexibility, and will allow user-defined custom training datasets as archaeal virus discovery progresses. Benchmarking showed that a model trained with viral sequences from the hypersaline, marine, and hot spring environments correctly classified 85% of the archaeal viruses with a false detection rate below 2% using a random forest prediction threshold of 80% in a separate benchmarking dataset from the same habitats.

Vik, Dean (ORCID:000000027546899X)↗

Assessing the Influence of Climate on the Spatial Pattern of West Nile Virus Incidence in the United States

West Nile virus (WNV) is the leading cause of mosquito-borne disease in humans in the United States. Since the introduction of the disease in 1999, incidence levels have stabilized in many regions, allowing for analysis of climate conditions that shape the spatial structure of disease incidence. Our goal was to identify the seasonal climate variables that influence the spatial extent and magnitude of WNV incidence in humans. We developed a predictive model of contemporary mean annual WNV incidence using U.S. county-level case reports from 2005 to 2019 and seasonally averaged climate variables. We used a random forest model that had an out-of-sample model performance of R 2 =0.61. Our model accurately captured the V-shaped area of higher WNV incidence that extends from states on the Canadian border south through the middle of the Great Plains. It also captured a region of moderate WNV incidence in the southern Mississippi Valley. The highest levels of WNV incidence were in regions with dry and cold winters and wet and mild summers. The random forest model classified counties with average winter precipitation levels <23.3 mm/month as having incidence levels over 11 times greater than those of counties that are wetter. Among the climate predictors, winter precipitation, fall precipitation, and winter temperature were the three most important predictive variables. We consider which aspects of the WNV transmission cycle climate conditions may benefit the most and argued that dry and cold winters are climate conditions optimal for the mosquito species key to amplifying WNV transmission. Our statistical model may be useful in projecting shifts in WNV risk in response to climate change.

60 APPLIED LIFE SCIENCES↗

Comparison of Machine Learning Algorithms for Natural Gas Identification with Mixed Potential Electrochemical Sensor Arrays

Mixed-potential electrochemical sensor arrays consisting of indium tin oxide (ITO), La 0.87 Sr 0.13 CrO 3 , Au, and Pt electrodes can detect the leaks from natural gas infrastructure. Algorithms are needed to correctly identify natural gas sources from background natural and anthropogenic sources such as wetlands or agriculture. We report for the first time a comparison of several machine learning methods for mixture identification in the context of natural gas emissions monitoring by mixed potential sensor arrays. Random Forest, Artificial Neural Network, and Nearest Neighbor methods successfully classified air mixtures containing only CH 4 , two types of natural gas simulants, and CH 4 +NH 3 with >98% identification accuracy. The model complexity of these methods were optimized and the degree of robustness against overfitting was determined. Finally, these methods are benchmarked on both desktop PC and single-board computer hardware to simulate their application in a portable internet-of-things sensor package. The combined results show that the random forest method is the preferred method for mixture identification with its high accuracy (>98%), robustness against overfitting with increasing model complexity, and had less than 10 ms training time and less than 0.1 ms inference time on single-board computer hardware.

03 NATURAL GAS↗

Dataset for 'Ombadi, M. & Varadharajan, C. (2022). Urbanization and aridity mediate distinct salinity response to floods in rivers and streams across the Contiguous United States, Water Research'

This package contains data sets and code used to obtain the results in Ombadi, M., & Varadharajan, C. (2022). Urbanization and aridity mediate distinct salinity response to floods in rivers and streams across the Contiguous United States. Water Research, 118664. The folder "data" contains 259 .csv files, each of which has daily time series of concurrent streamflow (Q) and specific conductance (SC) for each of the sites used in this study originally downloaded from the USGS National Water Information System (NWIS; USGS, 2016). The number of data points in each of the files is at least 3650 (i.e. 10 years of daily measurements). The folder "RF_single_data" contains 259 .csv files, each of which include data used to train and test the Random Forest models at individual sites for predicting SC during days of floods. The folder "RF_regional_data" contains 3 .csv files, each of which include scaled data compiled from all sites within each climate zone (arid, temperate and wet). "metadata.csv" contains the physical properties of the 259 catchments corresponding to the sites used in this study; this data was extracted from GAGES-II dataset (Falcone et al., 2010). "RF_implementation.ipynb" is a Jupyter notebook with the code needed to implement the analysis using Random Forest models either for individual sites or for the regional models (for each climate zone). The code utilizes the data in the two folders: "RF_single_data" and "RF_regional_data" and the metadata.csv file.

54 ENVIRONMENTAL SCIENCES↗

Sub-pilot-scale Production of High-Value Products from U.S. Coals

Investigators from the University of Utah, University of Wyoming and Marshall University pursued a program to study the conversion of raw coal to high-value products of carbon fiber and silicon carbide. Team members also developed an initial framework for a data portal that can incorporate laboratory data on coal processing and product quality, and also work with tools for machine learning for data analysis, data visualization and economic assessment. Experimental R&D efforts focused on the conversion of raw coal to coal tar and other byproducts, and the resulting tar intermediates were upgraded to form anisotropic and isotropic pitch materials. These pitch materials were produced from coal using both thermal (pyrolysis) and chemical (mild solvolysis liquefaction) decomposition of raw coal. Four different coals were studied: Utah bituminous coal (Sufco), Wyoming PRB coal (Black Thunder), Illinois bituminous coal (Illinois #6), and West Virginia bituminous coal (Flying Eagle). Both metallurgical-grade coking coals and lower-grade steam coals were investigated, and controlled secondary gas-phase reactions were used during a two-stage pyrolysis process to induce cracking and condensation reactions among the pyrolytic tar species. This approach successfully improved the performance of the lower grade coals for yielding pitch materials, with properties more consistent with a commercial-grade pitch that had previously demonstrated success for quality carbon fiber production. The use of waste plastic materials was also studied, to help improve physical and chemical characteristics of the intermediate tars and final pitch product; in particular, for lowering the pitch softening point to an acceptable level for melt spinning carbon fiber. Mild solvolysis liquefaction was also used as a method for producing pitch for carbon fiber production. As expected, significantly higher pitch yields were obtained using this approach, and waste plastic materials were also successfully used to reduce pitch softening point to an acceptable level. The plastic materials were also utilized to create a solvent for the mild solvolysis process, and this plastic-derived solvent was shown to provide results consistent with more expensive commercial chemical solvents, and could thus avoid the need for costly recovery and recycle of a liquefaction solvent. Additional experimental R&D focused on the production of silicon carbide (β-SiC) from the residual char byproduct from pitch production, and also on the production of carbon fiber from the anisotropic pitch. SiC was successfully synthesized using a mixture of residual char and sandstone at a ratio of 1:1. Reaction temperature and residence time were optimized and yielded a product purity of 81%. For carbon fiber production, the most successful pitch samples were obtained from the mild solvolysis liquefaction approach, combined with the use of a plastic (HDPE)-derived solvent. Fiber properties improved over time as laboratory fiber production methodologies improved, and final yields of carbon fiber were obtained with a diameter of 12.14 ± 1.10 um, Modulus of 173.73 ± 15.25 GPa, and Tensile Strength of 1.04 ± 0.10 GPa. A proof-of-concept Modern Community Research Data Portal (MCRDP) was developed and deployed for coal and coal-derived pitch characterization, with the full support of (i) remote web-based access, (ii) distributed analysis, (iii) interactive visualization and exploration, (iv) shared and long-term data access, (v) advanced query capabilities and (vi) real-time collaboration. The Coal to Products Data Portal “coaltoproducts.org” provides researchers with space to store and share data within a project, tools for analyzing and understanding data for scientific investigation, and the ability to publish data to the broader community for reproducibility. The portal leverages the Material Commons 2.0 (MC) platform developed by the Center for PRedictive Integrated Structural Materials Science (PRISMS) of the University of Michigan, to achieve long-term longevity of data collections and, more importantly, collaborative science. A number of data visualization tools were also assessed and implemented for interrogating the experimental and modeling data. The machine learning portion of this project analyzed datasets from two different coal conversion processes performed on a diverse set of coal samples from both the coal pyrolysis experiments and the solvent liquefaction experiments. The work was initiated by exploring standard regression models on the pyrolysis data, aiming to understand the impact of sample characteristics and processing conditions on key product metrics. Over the course of the project, the focus expanded to include a variety of machine learning tools, delving into both supervised and unsupervised learning methods. Models tested on the pyrolysis data included linear, ridge, lasso, elastic-net, Gaussian process, random forest regression, and AutoSklearn, and the approach was continually refined to enhance predictive accuracy and model interpretability. Similar techniques were applied to the liquefaction data with an additional focus on feature engineering. Along with mesophase content, additional outputs of interest were the pitch yield, softening point, and QI content. Insights derived from these analyses are crucial in determining the factors influencing the quality and yield of coal-derived products. As the work progressed, the research evolved from foundational model comparisons to analyses of random forests, decision paths, and feature importance scores. A thorough market analysis was performed to examine the prospects of coal-based carbon fibers. The best opportunities for coal come from its lower and more stable price relative to petroleum, particularly for subbituminous coals, which is the primary advantage that a coal refinery may have over a petroleum refinery. Before a commercial CTP production facility can be modeled, however, several things need to be understood regarding the nature of the would-be coal refinery. These include the technology to be deployed, the size of facility, the volume(s) of co-product(s), and the waste and emissions profile of the plant. The volume of co-products and waste may be substantial and will require separate market analysis to ensure viability. In the near-term, the importance of coal tar pitch, in the form of carbon pitch, to the aluminum and steel industries is likely to overshadow the alternative use of this material as an input for carbon fiber. The importance of steel and aluminum in building materials, and the need for carbon materials in their manufacturing, will ensure that demand for these products remains for the long run. In addition, carbon fiber may also be the best substitute for steel and aluminum well into the future. While society will eventually be able to shift production of much of its electricity needs to renewables, it will not be able to shift away from fossil fuels for production of high-strength construction and vehicular materials. Demand for carbon fiber is expected to increase quickly, but the volume of carbon fiber and the amount of coal that would be needed to produce even a sizeable share of this market may still be relatively small compared to current coal production. Thus, other coal-based products like graphene, graphite, carbon foams, resins, and carbon-based building products will play important roles in sustaining coal production as coal-fired power generation continues to decline.

01 COAL, LIGNITE, AND PEAT↗

Photometric redshift estimation of BASS DR3 quasars by machine learning

ABSTRACT Correlating Beijing–Arizona Sky Survey (BASS) data release 3 (DR3) catalogue with the ALLWISE data base, the data from optical and infrared information are obtained. The quasars from Sloan Digital Sky Survey are taken as training and test samples while those from LAMOST are considered as external test sample. We propose two schemes to construct the redshift estimation models with XGBoost, CatBoost, and Random Forest. One scheme (namely one-step model) is to predict photometric redshifts directly based on the optimal models created by these three algorithms; the other scheme (namely two-step model) is to first classify the data into low- and high-redshift data sets, and then predict photometric redshifts of these two data sets separately. For one-step model, the performance of these three algorithms on photometric redshift estimation is compared with different training samples, and CatBoost is superior to XGBoost and Random Forest. For two-step model, the performances of these three algorithms on the classification of low and high redshift subsamples are compared, and CatBoost still shows the best performance. Therefore, CatBoost is regarded as the core algorithm of classification and regression in two-step model. In contrast to one-step model, two-step model is optimal when predicting photometric redshift of quasars, especially for high-redshift quasars. Finally, the two models are applied to predict photometric redshifts of all quasar candidates of BASS DR3. The number of high-redshift quasar candidates is 3938 (redshift ≥3.5) and 121 (redshift ≥4.5) by two-step model. The predicted result will be helpful for quasar research and follow-up observation of high-redshift quasars.

79 ASTRONOMY AND ASTROPHYSICS↗

Importance of Depth and Artificial Structure as Predictors of Female Red Snapper Reproductive Parameters

Abstract The Red Snapper Lutjanus campechanus is a structure‐associated species occurring across a wide depth range in the northern Gulf of Mexico. We used the random forest machine learning algorithm to understand which habitat and individual fish characteristics could predict reproductive parameters of female Red Snapper. We evaluated fish captured from 2016 to 2018 on three artificial structure types with various structure heights at depths of 100 m or less. Overall, we found that depth and month were important predictors for most reproductive parameters, but the type of structure (artificial reefs, oil platforms, and rigs‐to‐reefs structures) was not important. Maturity was correctly classified in 88.9% of the cases when using the random forest ensemble model, with important predictors including FL, depth, structure height, and month of collection. Spawning seasonality (measured as gonadosomatic index [GSI]) was correctly classified in 59.5% of the cases when using histology reproductive phase, FL, month, and depth variables. Reproductively active or inactive females were correctly classified in 89.3% of the cases using GSI, month, FL, and depth, while females in the developing versus spawning capable phases were correctly classified in 82.2% of the cases using GSI, FL, month, and depth. Histological indicators that show potential spawning within a 36‐h period were correctly classified 61.5% of the time, with the best predictors being depth, FL, GSI, and month. Stepwise regression indicated that month was the only factor that significantly predicted contrasts in relative batch fecundity, with significantly greater values in August compared to all other months. Our findings suggest that female Red Snapper reproductive effort is not consistently or well predicted by artificial structure type or height but that a combination of fish FL, month, and depth can predict reproductive characteristics of female Red Snapper.

Brown‐Peterson, Nancy J.↗

How accurate is a machine learning-based wind speed extrapolation under a round-robin approach?

As the size of commercial wind turbines keeps increasing, having accurate ways to vertically extrapolate wind speed is essential to obtain a precise characterization of the wind resource for wind energy production. Recently, machine learning has been proposed and applied to extrapolate wind speed to hub heights. However, previous studies trained and tested the machine learning methods at the same site, giving them an unfair advantage over the conventional extrapolation techniques, which are instead more universal. Here, we use data from four sites in Oklahoma to test a round-robin validation approach for machine learning, under which we train a random forest at a site, and test it at a different site, where the model has no prior knowledge of the wind resource. We quantify how the accuracy of this technique varies with distance from the training site, and we find that it outperforms conventional techniques for wind extrapolation at all the considered spatial separations. We then assess how the accuracy of the machine-learning based approach varies when it is used to predict wind speed in a wind farm far wake. Finally, we explore as case study the performance of the random forest in extrapolating winds during a low-level jet event.

17 WIND ENERGY↗

Lithium-Ion Battery Diagnostics Using Electrochemical Impedance via Machine-Learning

Diagnosing battery states such as health, state-of-charge, or temperature is crucial for ensuring the safety and reliability of electrochemical energy storage systems. While some states, such as temperature, may be measured using cheap sensors, accurate diagnosis of battery health metrics usually requires time-consuming performance measurements, making them infeasible for use in real-world operation. These health metrics can be measured during lab-testing and then estimated on-line using predictive life models or via state observer algorithms such as Kalman filters, but these predictive methods should be supplemented by actual measurement of battery health whenever possible to ensure reliability. Rapid measurement of battery health may be done by various types of fast diagnostic techniques such as electrochemical impedance spectroscopy (EIS), which can be performed in only a few minutes and require only a fraction of the energy and power needed for a full charge and discharge measurement. But there is a substantial challenge for estimating battery health using EIS data, as EIS is sensitive to cell temperature, state-of-charge, current, and resting time in addition to health. Thus, utilizing EIS data to predict battery capacity requires correcting for all these additional variables, a task that is extremely difficult to handle analytically. This talk utilizes machine-learning methods to estimate the effectiveness of battery capacity prediction from EIS data, leveraging a data set of hundreds of EIS measurements recorded at varying temperature and state-of-charge throughout a 500-day aging study of 32 commercial, large-format NMC-Graphite lithium-ion batteries. Using EIS as input to machine-learning models is complicated by the nonlinear response of impedance to battery health, temperature, and state-of-charge, as well as the collinearity between the impedance response at neighboring frequencies, which can easily lead to overfit models. To train robust models, features from EIS data need to be extracted from the data or some subset of critical frequencies selected. Many approaches for extracting and selecting features from EIS data from electrochemical analysis and machine-learning fields were identified for analysis: using the entire raw spectra; selection of one, two, or many frequencies from the entire spectra; selecting interesting points from the EIS measurement using domain knowledge; fitting EIS with an equivalent-circuit model; calculating statistics on the raw impedance values; and reducing the dimensionality of the data using unsupervised linear (principal component analysis) and non-linear (uniform manifold approximation and projection) methods. These approaches were rigorously compared using a machine-learning pipeline approach, training linear, Gaussian process, and random forest regression models and quantifying performance using cross-validation as well as a held-out test set. An artificial neural network model trained on the raw spectra was also tested. Promising pipelines were fine-tuned via Bayesian hyperparameter optimization using cross-validation loss and training with class-specific weights to counter data set imbalance. The most reliable method for utilizing impedance in this work was the selection of two optimal frequencies through an exhaustive search, resulting in about 2% mean absolute error on test data for both Gaussian process and random forest model architectures. Interrogation of a variety of models reveals critical frequencies of 100 Hz and 103 Hz for this data set, though the optimal set of frequencies is not necessarily intuitive, i.e., the best performing models are not simply those that use impedance at frequencies that have the highest correlation to the relative discharge capacity. The best performing model is an ensemble model, which is able to predict battery capacity with 1.9% mean absolute error for unseen cells using impedance recorded at a variety of temperatures and states-of-charge.

battery↗

Snow Distribution Patterns Revisited: A Physics-Based and Machine Learning Hybrid Approach to Snow Distribution Mapping in the Sub-Arctic

Snowpack distribution in Arctic and alpine landscapes often occurs in repeating, year-to-year patterns due to local topographic, weather, and vegetation characteristics. Previous studies have suggested that with years of observational data, these snow distribution patterns can be statistically integrated into a snow process modeling workflow. Recent advances in snow hydrology and machine learning (ML) have increased our ability to predict snowpack distribution using in-situ observations, remote sensing data sets, and simple landscape characteristics that can be easily obtained for most environments. Here, we propose a hybrid approach to couple a ML snow distribution pattern (MLSDP) map with a physics-based, snow process model. We trained a random forest ML algorithm on tens of thousands of snow survey observations from a subarctic study area on the Seward Peninsula, Alaska, collected during peak snow water equivalent (SWE). We validated hybrid model outputs using in-situ snow depth and SWE observations, as well as a light detection and ranging data set and a distributed temperature profiling sensor data set. When the hybrid results were compared with the physics-based method, the hybrid method more accurately depicted the spatial patterns of the snowpack, areas of drifting snow, and years when no in-situ observations were used in the random forest ML training data set. The hybrid method also showed improvements in root mean squared error at 61% of locations where time-series estimations of snow depth were observed. These results can be applied to any physics-based model to improve the snow distribution patterning to reflect observed conditions in high latitude and high elevation cold region environments.

54 ENVIRONMENTAL SCIENCES↗

Modeling freight mode choice using machine learning classifiers: a comparative study using Commodity Flow Survey (CFS) data

This study explores the usefulness of machine learning classifiers for modeling freight mode choice. We investigate eight commonly used machine learning classifiers, namely Naïve Bayes, Support Vector Machine, Artificial Neural Network, K-Nearest Neighbors, Classification and Regression Tree, Random Forest, Boosting and Bagging, along with the classical Multinomial Logit model. US 2012 Commodity Flow Survey data are used as the primary data source; we augment it with spatial attributes from secondary data sources. The performance of the classifiers is compared based on prediction accuracy results. The current research also examines the role of sample size and training-testing data split ratios on the predictive ability of the various approaches. In addition, the importance of variables is estimated to determine how the variables influence freight mode choice. The results show that the tree-based ensemble classifiers perform the best. Specifically, Random Forest produces the most accurate predictions, closely followed by Boosting and Bagging. With regard to variable importance, shipment characteristics, such as shipment distance, industry classification of the shipper and shipment size, are the most significant factors for freight mode choice decisions.

42 ENGINEERING↗

A Predictive Model for Survival of Escherichia coli O157:H7 and Generic E. coli in Soil Amended with Untreated Animal Manure

Abstract This study aimed at developing a predictive model that captures the influences of a variety of agricultural and environmental variables and is able to predict the concentrations of enteric bacteria in soil amended with untreated Biological Soil Amendments of Animal Origin (BSAAO) under dynamic conditions. We developed and validated a Random Forest model using data from a longitudinal field study conducted in mid‐Atlantic United States investigating the survival of Escherichia coli O157:H7 and generic E. coli in soils amended with untreated dairy manure, horse manure, or poultry litter. Amendment type, days of rain since the previous sampling day, and soil moisture content were identified as the most influential agricultural and environmental variables impacting concentrations of viable E. coli O157:H7 and generic E. coli recovered from amended soils. Our model results also indicated that E. coli O157:H7 and generic E. coli declined at similar rates in amended soils under dynamic field conditions.The Random Forest model accurately predicted changes in viable E. coli concentrations over time under different agricultural and environmental conditions. Our model also accurately characterized the variability of E. coli concentration in amended soil over time by providing upper and lower prediction bound estimates. Cross‐validation results indicated that our model can be potentially generalized to other geographic regions and incorporated into a risk assessment for evaluating the risks associated with application of untreated BSAAO. Our model can be validated for other regions and predictive performance also can be enhanced when data sets from additional geographic regions become available.

Pang, Hao↗

Line Faults Classification Using Machine Learning on Three Phase Voltages Extracted from Large Dataset of PMU Measurements

An end-to-end supervised learning method is developed to classify transmission line faults in a twoyear field-recorded dataset that includes synchronized measurements of three-phase voltages recorded by 38 Phasor Measurement Units (PMU) sparsely located in in the US Western Grid interconnection. Statistical analysis is performed to extract features from this large dataset to train Support Vector Machine (SVM), Random Forest (RF), and eXtreme Gradient Boosting (XGBoost) classifiers initially. The training further leverages a simulated dataset from a synthetic grid with 12 PMUs to increase the number of faults of types infrequently seen in the field-recorded dataset. Training the classification models with the combined dataset resulted in a classification accuracy of 97.7%. This is a significant improvement over 89.7% to 92.5% accuracy obtained by relying on the field-recorded dataset alone.

47 OTHER INSTRUMENTATION↗

Interpreting Write Performance of Supercomputer I/O Systems with Regression Models

This work seeks to advance the state of the art in HPC I/O performance analysis and interpretation. In particular, we demonstrate effective techniques to: (1) model output performance in the presence of I/O interference from production loads; (2) build features from write patterns and key parameters of the system architecture and configurations; (3) employ suitable machine learning algorithms to improve model accuracy. We train models with five popular regression algorithms and conduct experiments on two distinct production HPC platforms. We find that the lasso and random forest models predict output performance with high accuracy on both of the target systems. We also explore use of the models to guide adaptation in I/O middleware systems, and show potential for improvements of at least 15% from model-guided adaptation on 70% of samples, and improvements up to 10× on some samples for both of the target systems.

Xie, Bing↗