Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Random forests”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

GraphAlign: Graph-Enabled Machine Learning for Seismic Event Filtering

This report summarizes results from a 2 year effort to improve the current automated seismic event processing system by leveraging machine learning models that can operated over the inherent graph data structure of a seismic sensor network. Specifically, the GraphAlign project seeks to utilize prior information on which stations are more likely to detect signals originating from particular geographic regions to inform event filtering. To date, the GraphAlign team has developed a Graphical Neural Network (GNN) model to filter out false events generated by the Global Associator (GA) algorithm. The algorithm operates directly on waveform data that has been associated to an event by building a variable sized graph of station waveforms nodes with edge relations to an event location node. This builds off of previous work where random forest models were used to do the same task using hand crafted features. The GNN model performance was analyzed using an 8 week IMS/IDC dataset, and it was demonstrated that the GNN outperforms the random forest baseline. We provide additional error analysis of which events the GNN model performs well and poorly against concluded by future directions for improvements.

58 GEOSCIENCES↗

Power System Feature-Based Event Classification by Means of Multiple PMU Data

Abstract—Phasor Measurement Units (PMUs) provide time synchronized measurements across the power grid, enabling data driven event detection and classification for enhanced system monitoring and situational awareness. However, variations in event duration, spatial extent, and severity, along with coincident events, pose challenges for conventional classification models that require fixed-size inputs. This paper presents a feature-based framework that aggregates diverse attributes from all available PMUs for each event into a fixed-length vector, facilitating the application of standard machine learning classifiers, including Random Forest, XGBoost, and Multilayer Perceptron. A probabilistic post-processing scheme is further introduced to enable multi-label classification in the presence of overlapping events. Experiments using real-world PMU data demonstrate that the Random Forest model achieves 95% accuracy, while the proposed post-processing method yields an additional 3% improvement.

Nematirad, Reza↗

Spatial patterns of snow distribution in the sub-Arctic

Abstract. The spatial distribution of snow plays a vital role in sub-Arctic and Arctic climate, hydrology, and ecology due to its fundamental influence on the water balance, thermal regimes, vegetation, and carbon flux. However, the spatial distribution of snow is not well understood, and therefore, it is not well modeled, which can lead to substantial uncertainties in snow cover representations. To capture key hydro-ecological controls on snow spatial distribution, we carried out intensive field studies over multiple years for two small (2017–2019; ∼ 2.5 km2) sub-Arctic study sites located on the Seward Peninsula of Alaska. Using an intensive suite of field observations (> 22 000 data points), we developed simple models of the spatial distribution of snow water equivalent (SWE) using factors such as topographic characteristics, vegetation characteristics based on greenness (normalized different vegetation index, NDVI), and a simple metric for approximating winds. The most successful model was random forest, using both study sites and all years, which was able to accurately capture the complexity and variability of snow characteristics across the sites. Approximately 86 % of the SWE distribution could be accounted for, on average, by the random forest model at the study sites. Factors that impacted year-to-year snow distribution included NDVI, elevation, and a metric to represent coarse microtopography (topographic position index, TPI), while slope, wind, and fine microtopography factors were less important. The characterization of the SWE spatial distribution patterns will be used to validate and improve snow distribution modeling in the Department of Energy's Earth system model and for improved understanding of hydrology, topography, and vegetation dynamics in the sub-Arctic and Arctic regions of the globe.

54 ENVIRONMENTAL SCIENCES↗

Upscaling Soil Organic Carbon Measurements at the Continental Scale Using Multivariate Clustering Analysis and Machine Learning

Abstract Estimates of soil organic carbon (SOC) stocks are essential for many environmental applications. However, significant inconsistencies exist in SOC stock estimates for the U.S. across current SOC maps. We propose a framework that combines unsupervised multivariate geographic clustering (MGC) and supervised Random Forests regression, improving SOC maps by capturing heterogeneous relationships with SOC drivers. We first used MGC to divide the U.S. into 20 SOC regions based on the similarity of covariates (soil biogeochemical, bioclimatic, biological, and physiographic variables). Subsequently, separate Random Forests models were trained for each SOC region, utilizing environmental covariates and SOC observations. Our estimated SOC stocks for the U.S. (52.6 ± 3.2 Pg for 0–30 cm and 108.3 ± 8.2 Pg for 0–100 cm depth) were within the range estimated by existing products like Harmonized World Soil Database, HWSD (46.7 Pg for 0–30 cm and 90.7 Pg for 0–100 cm depth) and SoilGrids 2.0 (45.7 Pg for 0–30 cm and 133.0 Pg for 0–100 cm depth). However, independent validation with soil profile data from the National Ecological Observatory Network showed that our approach ( R 2 = 0.51) outperformed the estimates obtained from Harmonized World Soil Database ( R 2 = 0.23) and SoilGrids 2.0 ( R 2 = 0.39) for the topsoil (0–30 cm). Uncertainty analysis (e.g., low representativeness and high coefficients of variation) identified regions requiring more measurements, such as Alaska and the deserts of the U.S. Southwest. Our approach effectively captures the heterogeneous relationships between widely available predictors and the current SOC baseline across regions, offering reliable SOC estimates at 1 km resolution for benchmarking Earth system models.

58 GEOSCIENCES↗

Geographical Insights into Suicide Mortality Through Spatial Machine Learning

Suicide mortality is a leading cause of death in the United States, with an upward trend that emphasizes its significance as a public health issue. Previous research has employed global models like ordinary least squares (OLS) regression and local models such as geographically weighted regression (GWR). While local models are useful for analyzing spatial variations in suicide mortality, they share limitations with traditional global models, particularly about their inability to handle multi-collinearity and non-linear relationships. Machine learning approaches, like random forests (RF), can address some of these limitations but often fail to account for spatial variability. This gap highlights the need for spatial ML models specifically designed to tackle suicide mortality. This research seeks to fill this void by using a geographically weighted random forest model (GWRF) to examine the associations between county-level suicide mortality in the U.S. from 2010 to 2020 and various social and environmental determinants of health. A key aspect of our methodology is disciplined feature selection, which reduces the pool of explanatory variables by about 90%. This refinement enhances the explanatory power of both global (R2 improved from 0.59 to 0.67) and local (R2 improved from 0.64 to 0.67) RF models while reducing their run times. An analysis of the importance scores for these selected features reveals that the drivers of suicide mortality vary by context. Thus, to effectively address regional disparities and inform targeted public health interventions, a holistic approach that incorporates multiple county-level characteristics is essential.

Lebakula, Viswadeep [ORNL] (ORCID:0000000152935914↗

Evaluating county-level lung cancer incidence from environmental radiation exposure, PM 2.5 , and other exposures with regression and machine learning models

Characterizing the interplay between exposures shaping the human exposome is vital for uncovering the etiology of complex diseases. For example, cancer risk is modified by a range of multifactorial external environmental exposures. Environmental, socioeconomic, and lifestyle factors all shape lung cancer risk. However, epidemiological studies of radon aimed at identifying populations at high risk for lung cancer often fail to consider multiple exposures simultaneously. For example, moderating factors, such as PM 2.5 , may affect the transport of radon progeny to lung tissue. This ecological analysis leveraged a population-level dataset from the National Cancer Institute’s Surveillance, Epidemiology, and End-Results data (2013–17) to simultaneously investigate the effect of multiple sources of low-dose radiation (gross γ activity and indoor radon) and PM 2.5 on lung cancer incidence rates in the USA. County-level factors (environmental, sociodemographic, lifestyle) were controlled for, and Poisson regression and random forest models were used to assess the association between radon exposure and lung and bronchus cancer incidence rates. Tree-based machine learning (ML) method perform better than traditional regression: Poisson regression: 6.29/7.13 (mean absolute percentage error, MAPE), 12.70/12.77 (root mean square error, RMSE); Poisson random forest regression: 1.22/1.16 (MAPE), 8.01/8.15 (RMSE). The effect of PM 2.5 increased with the concentration of environmental radon, thereby confirming findings from previous studies that investigated the possible synergistic effect of radon and PM 2.5 on health outcomes. In summary, the results demonstrated (1) a need to consider multiple environmental exposures when assessing radon exposure’s association with lung cancer risk, thereby highlighting (1) the importance of an exposomics framework and (2) that employing ML models may capture the complex interplay between environmental exposures and health, as in the case of indoor radon exposure and lung cancer incidence.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

The Use of Machine Learning Models for Predicting the Dielectric Strength of Gases

Technological advancements in high voltage systems have pushed sulfur hexafluoride (SF6) to its operational limits. Furthermore, this gas has other drawbacks including a high liquefaction temperature and a high global warming potential. Therefore, there has been an urgent need to find alternative gases with high dielectric strength (DS). In this work, density functional theory (DFT) is used to calculate molecular descriptors that are fed into an artificial neural network (ANN) and a random forest (RF). These machine learning (ML) models are then used to predict the DS for hundreds of molecules. A finite element model (FEM) is also used to calculate the electric field profile of multiple simple electrode geometries as the applied voltage to the system is increased. Results indicate that the random forest model has better generalization to unseen data than the neural network. The highest DS value predicted by the RF was 2.16 relative to the experimental DS of SF6. The results also demonstrate how choosing a gas with a higher DS and a geometry with minimal edges and corners can significantly increase the operating voltage of an electrical system. Due to its superior generalization, the RF represents the most promising path toward an accurate DS predictor once sufficient experimental data are available.

Mileski, Matthew [AFIT]↗

The importance of round-robin validation when assessing machine-learning-based vertical extrapolation of wind speeds

The extrapolation of wind speeds measured at a meteorological mast to wind turbine rotor heights is a key component in a bankable wind farm energy assessment and a significant source of uncertainty. Industry-standard methods for extrapolation include the power-law and logarithmic profiles. The emergence of machine-learning applications in wind energy has led to several studies demonstrating substantial improvements in vertical extrapolation accuracy in machine-learning methods over these conventional power-law and logarithmic profile methods. In all cases, these studies assess relative model performance at a measurement site where, critically, the machine-learning algorithm requires knowledge of the rotor-height wind speeds in order to train the model. This prior knowledge provides fundamental advantages to the site-specific machine-learning model over the power-law and log profiles, which, by contrast, are not highly tuned to rotor-height measurements but rather can generalize to any site. Furthermore, there is no practical benefit in applying a machine-learning model at a site where winds at the heights relevant for wind energy production are known; rather, its performance at nearby locations (i.e., across a wind farm site) without rotor-height measurements is of most practical interest. To more fairly and practically compare machine-learning-based extrapolation to standard approaches, we implemented a round-robin extrapolation model comparison, in which a random-forest machine-learning model is trained and evaluated at different sites and then compared against the power-law and logarithmic profiles. We consider 20 months of lidar and sonic anemometer data collected at four sites between 50 and 100 km apart in the central United States. We find that the random forest outperforms the standard extrapolation approaches, especially when incorporating surface measurements as inputs to include the influence of atmospheric stability. When compared at a single site (the traditional comparison approach), the machine-learning improvement in mean absolute error was 28 % and 23 % over the power-law and logarithmic profiles, respectively. Using the round-robin approach proposed here, this improvement drops to 20 % and 14 %, respectively. These latter values better represent practical model performance, and we conclude that round-robin validation should be the standard for machine-learning-based wind speed extrapolation methods.

17 WIND ENERGY↗

Application of machine learning to discover new intermetallic catalysts for the hydrogen evolution and the oxygen reduction reactions

The adsorption energies for hydrogen, oxygen, and hydroxyl were calculated by means of density functional theory on the lowest energy surface of 24 pure metals and 332 binary intermetallic compounds with stoichiometries AB, A 2 B, and A 3 B taking into account the effect of biaxial elastic strains. This information was used to train two random forest regression models, one for the hydrogen adsorption and another for the oxygen and hydroxyl adsorption, based on 9 descriptors that characterized the geometrical and chemical features of the adsorption site as well as the applied strain. All the descriptors for each compound in the models could be obtained from physico-chemical databases. The random forest models were used to predict the adsorption energy for hydrogen, oxygen, and hydroxyl of ≈2700 binary intermetallic compounds with stoichiometries AB, A 2 B, and A 3 B made of metallic elements, excluding those that were environmentally hazardous, radioactive, or toxic. This information was used to search for potential good catalysts for the HER and ORR from the criteria that their adsorption energy for H and O/OH, respectively, should be close to that of Pt. Further, this investigation shows that the suitably trained machine learning models can predict adsorption energies with an accuracy not far away from density functional theory calculations with minimum computational cost from descriptors that are readily available in physico-chemical databases for any compound. Moreover, the strategy presented in this paper can be easily extended to other compounds and catalytic reactions, and is expected to foster the use of ML methods in catalysis.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Forest Soil Carbon Efflux Evaluation Across China: A New Estimate With Machine Learning

Forest soil respiration (Rs) plays an important role in the carbon balance of terrestrial ecosystems. China's forest occupies a large part of the world's forest. However, due to the lack of integrated observation data and appropriate upscaling methodologies, substantial uncertainties exist in the Rs evaluation, which limits our understanding of the carbon balance. Here, we re-evaluated the total soil carbon effluxes across China by combining field observations from 634 published annual Rs with a machine learning technique (i.e., Random Forest (RF)). Our results revealed that the combination of systematic measurements with the RF model allowed a definite estimate. The average annual Rs was 773.6 g C m -2 yr -1 , ranging from 404.6 to 1,772.3 g C m -2 yr -1 . Total forest soil carbon effluxes amounted to 1.17 Pg C yr -1 in China. Geographically, annual Rs showed a clear spatially increasing trend from northeast to southwest. Forest type is an important factor in determining the soil respiration rate. Bamboo and Evergreen broadleaf forests were higher than other types of forests. In conclusion, these results provide a unique insight into the magnitudes and mechanisms of soil CO 2 emissions in China's forest ecosystems.

58 GEOSCIENCES↗

Feature-Based PMU Event Classification under Variable PMU Participation and Overlapping Events

Danovo Energy Solution's presented its paper named: Feature-Based PMU Event Classification under Variable PMU Participation and Overlapping Events at the 2026 Georgia Tech Fault & Disturbance Analysis Conference. The full paper can be found at OSTI ID# 3169150 Paper Abstract—Phasor Measurement Units (PMUs) stream time synchronized, high-resolution measurements from the grid, enabling data-driven techniques for event detection and classification. Accurate event classification improves grid reliability and stability. Events can be detected by varying numbers of PMUs and exhibit different durations depending on the event type. This variability challenges standard classifiers that require uniform input sizes. Moreover, multiple events may coincide, which increases classification complexity. Standard classifiers assign each instance to the class with the highest predicted probability, whereas overlapping events may exhibit comparable probabilities across multiple classes. In this study, to handle data size variability, we extract a wide range of time–frequency domain features from all available PMUs for each event into a fixed-length vector, facilitating the application of standard machine learning classifiers, including Random Forest, XGBoost, LightGBM, Support Vector Machine, and Multilayer Perceptron. To account for overlapping events, a probabilistic post-processing step is applied. For a given data instance, if multiple predicted class probabilities exceed 30% and the differences between them are less than 10%, the event is assigned to multiple classes. Experiments using real-world PMU data demonstrate that the Random Forest and XGBoost models achieve the highest accuracy, while the proposed post-processing method yields perfect classification performance on external unseen test sets.

Nematirad, Reza [Danova Energy Solutions]↗

Improving Medication Regimen Recommendation for Parkinson’s Disease Using Sensor Technology

Parkinson’s disease medication treatment planning is generally based on subjective data obtained through clinical, physician-patient interactions. The Personal KinetiGraph™ (PKG) and similar wearable sensors have shown promise in enabling objective, continuous remote health monitoring for Parkinson’s patients. In this proof-of-concept study, we propose to use objective sensor data from the PKG and apply machine learning to cluster patients based on levodopa regimens and response. The resulting clusters are then used to enhance treatment planning by providing improved initial treatment estimates to supplement a physician’s initial assessment. We apply k-means clustering to a dataset of within-subject Parkinson’s medication changes—clinically assessed by the MDS-Unified Parkinson’s Disease Rating Scale-III (MDS-UPDRS-III) and the PKG sensor for movement staging. A random forest classification model was then used to predict patients’ cluster allocation based on their respective demographic information, MDS-UPDRS-III scores, and PKG time-series data. Clinically relevant clusters were partitioned by levodopa dose, medication administration frequency, and total levodopa equivalent daily dose—with the PKG providing similar symptomatic assessments to physician MDS-UPDRS-III scores. A random forest classifier trained on demographic information, MDS-UPDRS-III scores, and PKG time-series data was able to accurately classify subjects of the two most demographically similar clusters with an accuracy of 86.9%, an F1 score of 90.7%, and an AUC of 0.871. A model that relied solely on demographic information and PKG time-series data provided the next best performance with an accuracy of 83.8%, an F1 score of 88.5%, and an AUC of 0.831, hence further enabling fully remote assessments. These computational methods demonstrate the feasibility of using sensor-based data to cluster patients based on their medication responses with further potential to assist with medication recommendations.

59 BASIC BIOLOGICAL SCIENCES↗

Feature-Based PMU Event Classification under Variable PMU Participation and Overlapping Events

This paper is the basis for a presentation help at the 2026 Georgia Tech Fault & Disturbance Analysis Conference, which can be found at OSTI # 3168287 Paper Abstract—Phasor Measurement Units (PMUs) stream time synchronized, high-resolution measurements from the grid, enabling data-driven techniques for event detection and classification. Accurate event classification improves grid reliability and stability. Events can be detected by varying numbers of PMUs and exhibit different durations depending on the event type. This variability challenges standard classifiers that require uniform input sizes. Moreover, multiple events may coincide, which increases classification complexity. Standard classifiers assign each instance to the class with the highest predicted probability, whereas overlapping events may exhibit comparable probabilities across multiple classes. In this study, to handle data size variability, we extract a wide range of time–frequency domain features from all available PMUs for each event into a fixed-length vector, facilitating the application of standard machine learning classifiers, including Random Forest, XGBoost, LightGBM, Support Vector Machine, and Multilayer Perceptron. To account for overlapping events, a probabilistic post-processing step is applied. For a given data instance, if multiple predicted class probabilities exceed 30% and the differences between them are less than 10%, the event is assigned to multiple classes. Experiments using real-world PMU data demonstrate that the Random Forest and XGBoost models achieve the highest accuracy, while the proposed post-processing method yields perfect classification performance on external unseen test sets.

Nematirad, Reza [Danovo Energy Solutions]↗

MArVD2: a machine learning enhanced tool to discriminate between archaeal and bacterial viruses in viral datasets

Abstract Our knowledge of viral sequence space has exploded with advancing sequencing technologies and large-scale sampling and analytical efforts. Though archaea are important and abundant prokaryotes in many systems, our knowledge of archaeal viruses outside of extreme environments is limited. This largely stems from the lack of a robust, high-throughput, and systematic way to distinguish between bacterial and archaeal viruses in datasets of curated viruses. Here we upgrade our prior text-based tool (MArVD) via training and testing a random forest machine learning algorithm against a newly curated dataset of archaeal viruses. After optimization, MArVD2 presented a significant improvement over its predecessor in terms of scalability, usability, and flexibility, and will allow user-defined custom training datasets as archaeal virus discovery progresses. Benchmarking showed that a model trained with viral sequences from the hypersaline, marine, and hot spring environments correctly classified 85% of the archaeal viruses with a false detection rate below 2% using a random forest prediction threshold of 80% in a separate benchmarking dataset from the same habitats.

Vik, Dean (ORCID:000000027546899X)↗

Assessing the Influence of Climate on the Spatial Pattern of West Nile Virus Incidence in the United States

West Nile virus (WNV) is the leading cause of mosquito-borne disease in humans in the United States. Since the introduction of the disease in 1999, incidence levels have stabilized in many regions, allowing for analysis of climate conditions that shape the spatial structure of disease incidence. Our goal was to identify the seasonal climate variables that influence the spatial extent and magnitude of WNV incidence in humans. We developed a predictive model of contemporary mean annual WNV incidence using U.S. county-level case reports from 2005 to 2019 and seasonally averaged climate variables. We used a random forest model that had an out-of-sample model performance of R 2 =0.61. Our model accurately captured the V-shaped area of higher WNV incidence that extends from states on the Canadian border south through the middle of the Great Plains. It also captured a region of moderate WNV incidence in the southern Mississippi Valley. The highest levels of WNV incidence were in regions with dry and cold winters and wet and mild summers. The random forest model classified counties with average winter precipitation levels <23.3 mm/month as having incidence levels over 11 times greater than those of counties that are wetter. Among the climate predictors, winter precipitation, fall precipitation, and winter temperature were the three most important predictive variables. We consider which aspects of the WNV transmission cycle climate conditions may benefit the most and argued that dry and cold winters are climate conditions optimal for the mosquito species key to amplifying WNV transmission. Our statistical model may be useful in projecting shifts in WNV risk in response to climate change.

60 APPLIED LIFE SCIENCES↗

Comparison of Machine Learning Algorithms for Natural Gas Identification with Mixed Potential Electrochemical Sensor Arrays

Mixed-potential electrochemical sensor arrays consisting of indium tin oxide (ITO), La 0.87 Sr 0.13 CrO 3 , Au, and Pt electrodes can detect the leaks from natural gas infrastructure. Algorithms are needed to correctly identify natural gas sources from background natural and anthropogenic sources such as wetlands or agriculture. We report for the first time a comparison of several machine learning methods for mixture identification in the context of natural gas emissions monitoring by mixed potential sensor arrays. Random Forest, Artificial Neural Network, and Nearest Neighbor methods successfully classified air mixtures containing only CH 4 , two types of natural gas simulants, and CH 4 +NH 3 with >98% identification accuracy. The model complexity of these methods were optimized and the degree of robustness against overfitting was determined. Finally, these methods are benchmarked on both desktop PC and single-board computer hardware to simulate their application in a portable internet-of-things sensor package. The combined results show that the random forest method is the preferred method for mixture identification with its high accuracy (>98%), robustness against overfitting with increasing model complexity, and had less than 10 ms training time and less than 0.1 ms inference time on single-board computer hardware.

03 NATURAL GAS↗

Dataset for 'Ombadi, M. & Varadharajan, C. (2022). Urbanization and aridity mediate distinct salinity response to floods in rivers and streams across the Contiguous United States, Water Research'

This package contains data sets and code used to obtain the results in Ombadi, M., & Varadharajan, C. (2022). Urbanization and aridity mediate distinct salinity response to floods in rivers and streams across the Contiguous United States. Water Research, 118664. The folder "data" contains 259 .csv files, each of which has daily time series of concurrent streamflow (Q) and specific conductance (SC) for each of the sites used in this study originally downloaded from the USGS National Water Information System (NWIS; USGS, 2016). The number of data points in each of the files is at least 3650 (i.e. 10 years of daily measurements). The folder "RF_single_data" contains 259 .csv files, each of which include data used to train and test the Random Forest models at individual sites for predicting SC during days of floods. The folder "RF_regional_data" contains 3 .csv files, each of which include scaled data compiled from all sites within each climate zone (arid, temperate and wet). "metadata.csv" contains the physical properties of the 259 catchments corresponding to the sites used in this study; this data was extracted from GAGES-II dataset (Falcone et al., 2010). "RF_implementation.ipynb" is a Jupyter notebook with the code needed to implement the analysis using Random Forest models either for individual sites or for the regional models (for each climate zone). The code utilizes the data in the two folders: "RF_single_data" and "RF_regional_data" and the metadata.csv file.

54 ENVIRONMENTAL SCIENCES↗

Sub-pilot-scale Production of High-Value Products from U.S. Coals

Investigators from the University of Utah, University of Wyoming and Marshall University pursued a program to study the conversion of raw coal to high-value products of carbon fiber and silicon carbide. Team members also developed an initial framework for a data portal that can incorporate laboratory data on coal processing and product quality, and also work with tools for machine learning for data analysis, data visualization and economic assessment. Experimental R&D efforts focused on the conversion of raw coal to coal tar and other byproducts, and the resulting tar intermediates were upgraded to form anisotropic and isotropic pitch materials. These pitch materials were produced from coal using both thermal (pyrolysis) and chemical (mild solvolysis liquefaction) decomposition of raw coal. Four different coals were studied: Utah bituminous coal (Sufco), Wyoming PRB coal (Black Thunder), Illinois bituminous coal (Illinois #6), and West Virginia bituminous coal (Flying Eagle). Both metallurgical-grade coking coals and lower-grade steam coals were investigated, and controlled secondary gas-phase reactions were used during a two-stage pyrolysis process to induce cracking and condensation reactions among the pyrolytic tar species. This approach successfully improved the performance of the lower grade coals for yielding pitch materials, with properties more consistent with a commercial-grade pitch that had previously demonstrated success for quality carbon fiber production. The use of waste plastic materials was also studied, to help improve physical and chemical characteristics of the intermediate tars and final pitch product; in particular, for lowering the pitch softening point to an acceptable level for melt spinning carbon fiber. Mild solvolysis liquefaction was also used as a method for producing pitch for carbon fiber production. As expected, significantly higher pitch yields were obtained using this approach, and waste plastic materials were also successfully used to reduce pitch softening point to an acceptable level. The plastic materials were also utilized to create a solvent for the mild solvolysis process, and this plastic-derived solvent was shown to provide results consistent with more expensive commercial chemical solvents, and could thus avoid the need for costly recovery and recycle of a liquefaction solvent. Additional experimental R&D focused on the production of silicon carbide (β-SiC) from the residual char byproduct from pitch production, and also on the production of carbon fiber from the anisotropic pitch. SiC was successfully synthesized using a mixture of residual char and sandstone at a ratio of 1:1. Reaction temperature and residence time were optimized and yielded a product purity of 81%. For carbon fiber production, the most successful pitch samples were obtained from the mild solvolysis liquefaction approach, combined with the use of a plastic (HDPE)-derived solvent. Fiber properties improved over time as laboratory fiber production methodologies improved, and final yields of carbon fiber were obtained with a diameter of 12.14 ± 1.10 um, Modulus of 173.73 ± 15.25 GPa, and Tensile Strength of 1.04 ± 0.10 GPa. A proof-of-concept Modern Community Research Data Portal (MCRDP) was developed and deployed for coal and coal-derived pitch characterization, with the full support of (i) remote web-based access, (ii) distributed analysis, (iii) interactive visualization and exploration, (iv) shared and long-term data access, (v) advanced query capabilities and (vi) real-time collaboration. The Coal to Products Data Portal “coaltoproducts.org” provides researchers with space to store and share data within a project, tools for analyzing and understanding data for scientific investigation, and the ability to publish data to the broader community for reproducibility. The portal leverages the Material Commons 2.0 (MC) platform developed by the Center for PRedictive Integrated Structural Materials Science (PRISMS) of the University of Michigan, to achieve long-term longevity of data collections and, more importantly, collaborative science. A number of data visualization tools were also assessed and implemented for interrogating the experimental and modeling data. The machine learning portion of this project analyzed datasets from two different coal conversion processes performed on a diverse set of coal samples from both the coal pyrolysis experiments and the solvent liquefaction experiments. The work was initiated by exploring standard regression models on the pyrolysis data, aiming to understand the impact of sample characteristics and processing conditions on key product metrics. Over the course of the project, the focus expanded to include a variety of machine learning tools, delving into both supervised and unsupervised learning methods. Models tested on the pyrolysis data included linear, ridge, lasso, elastic-net, Gaussian process, random forest regression, and AutoSklearn, and the approach was continually refined to enhance predictive accuracy and model interpretability. Similar techniques were applied to the liquefaction data with an additional focus on feature engineering. Along with mesophase content, additional outputs of interest were the pitch yield, softening point, and QI content. Insights derived from these analyses are crucial in determining the factors influencing the quality and yield of coal-derived products. As the work progressed, the research evolved from foundational model comparisons to analyses of random forests, decision paths, and feature importance scores. A thorough market analysis was performed to examine the prospects of coal-based carbon fibers. The best opportunities for coal come from its lower and more stable price relative to petroleum, particularly for subbituminous coals, which is the primary advantage that a coal refinery may have over a petroleum refinery. Before a commercial CTP production facility can be modeled, however, several things need to be understood regarding the nature of the would-be coal refinery. These include the technology to be deployed, the size of facility, the volume(s) of co-product(s), and the waste and emissions profile of the plant. The volume of co-products and waste may be substantial and will require separate market analysis to ensure viability. In the near-term, the importance of coal tar pitch, in the form of carbon pitch, to the aluminum and steel industries is likely to overshadow the alternative use of this material as an input for carbon fiber. The importance of steel and aluminum in building materials, and the need for carbon materials in their manufacturing, will ensure that demand for these products remains for the long run. In addition, carbon fiber may also be the best substitute for steel and aluminum well into the future. While society will eventually be able to shift production of much of its electricity needs to renewables, it will not be able to shift away from fossil fuels for production of high-strength construction and vehicular materials. Demand for carbon fiber is expected to increase quickly, but the volume of carbon fiber and the amount of coal that would be needed to produce even a sizeable share of this market may still be relatively small compared to current coal production. Thus, other coal-based products like graphene, graphite, carbon foams, resins, and carbon-based building products will play important roles in sustaining coal production as coal-fired power generation continues to decline.

01 COAL, LIGNITE, AND PEAT↗