Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “preprocessed”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Data Readiness for Scientific AI at Scale

This paper examines how Data Readiness for AI (DRAI) principles apply to leadership-scale scientific datasets used to train foundation models. We analyze archetypal workflows across four representative domains—climate, nuclear fusion, bio/health, and materials—to identify common preprocessing patterns and domain-specific constraints. We introduce a two-dimensional readiness framework that combines canonical preprocessing patterns with a five-level operational readiness scale, both tailored to high-performance computing (HPC) environments. This framework helps outline key challenges in transforming large-scale scientific data into formats suitable for scalable AI training. Together, these dimensions form a conceptual maturity matrix that characterizes scientific data readiness and guides infrastructure development toward standardized, cross-domain support for scalable and reproducible AI for science.

Brewer, Wes [ORNL] (ORCID:0000000236393956)↗

Distribution Substation Planning Toolkit (dsp-toolkit) v1.0

The Distribution Substation Planning Toolkit (DSP Toolkit) is a software suite designed to streamline the planning and optimization of distribution substations. This toolkit offers a comprehensive set of tools and APIs for data curation, short-term electric load forecasting, and weather-sensitive load adjustment, making it an essential resource for utility companies, engineers, and researchers. Features • Data Preprocessing and Curation: Efficiently manage and preprocess large datasets to ensure high-quality input for analysis. • Short-Term Load Forecasting: Utilize data-driven models to predict short-term electric loads accurately. • Weather-Sensitive Modeling: Automatically adjust load forecasts based on weather data to predict future peak demands more precisely. Uses The DSP Toolkit is ideal for planning and optimizing distribution substations, providing a user-friendly interface and comprehensive documentation. It is suitable for both novice and experienced users, facilitating efficient and accurate planning processes. Advantages • Efficiency: Automates complex planning tasks, reducing manual effort and minimizing errors. • Scalability: Handles large datasets and complex models, making it suitable for large-scale projects. • Community and Support: Open-source with active community contributions, ensuring continuous improvement and support. • Extensibility: Easily extendable with custom modules and plugins, allowing users to tailor the toolkit to their specific needs. The DSP Toolkit stands out by offering a robust, flexible, and user-friendly solution for distribution substation planning. Public Abstract

Li, Han [Lawrence Berkeley National Laboratory (LB↗

TEMPEST

This repository solves the problem of driver identification through vehicular and biometric data. Through an embedding-based approach and a novel loss function, we're able to distinguish between different drivers' behaviors. This also provides preprocessing for reproducibility of results.The code preprocesses vehicular data, trains neural networks, and outputs predictions.This code introduces a novel embedding-based neural network with a 91% rank-1 accuracy, as well as all code to reproduce training and results.

Musgrove, Kyle↗

Human limits in machine learning: prediction of potato yield and disease using soil microbiome data

Abstract Background The preservation of soil health is a critical challenge in the 21st century due to its significant impact on agriculture, human health, and biodiversity. We provide one of the first comprehensive investigations into the predictive potential of machine learning models for understanding the connections between soil and biological phenotypes. We investigate an integrative framework performing accurate machine learning-based prediction of plant performance from biological, chemical, and physical properties of the soil via two models: random forest and Bayesian neural network. Results Prediction improves when we add environmental features, such as soil properties and microbial density, along with microbiome data. Different preprocessing strategies show that human decisions significantly impact predictive performance. We show that the naive total sum scaling normalization that is commonly used in microbiome research is one of the optimal strategies to maximize predictive power. Also, we find that accurately defined labels are more important than normalization, taxonomic level, or model characteristics. ML performance is limited when humans can’t classify samples accurately. Lastly, we provide domain scientists via a full model selection decision tree to identify the human choices that optimize model prediction power. Conclusions Our study highlights the importance of incorporating diverse environmental features and careful data preprocessing in enhancing the predictive power of machine learning models for soil and biological phenotype connections. This approach can significantly contribute to advancing agricultural practices and soil health management.

Aghdam, Rosa↗

ORBIT-2 Dataset for Scaling Exascale Vision Foundation Models for Weather and Climate Downscaling

This dataset release corresponds to the work conducted in ORBIT-2: Scaling Exascale Vision Foundation Models for Weather and Climate Downscaling, where large-scale AI methods were applied to improve climate and weather resolution. The collection integrates four widely used, publicly available datasets: ERA5, PRISM, DAYMET, and IMERG. To prepare the data for ORBIT-2 model training and evaluation, we applied a preprocessing pipeline that generates paired low-resolution and high-resolution samples, enabling supervised downscaling experiments. The transformation from coarse to fine scales was performed using bilinear regridding, consistent with the procedures described in WeatherBench2, a community benchmark for weather and climate AI models. This dataset supports the development and evaluation of foundation models designed for weather and climate downscaling at exascale. Additional details on methodology and applications can be found in Wang et al., ORBIT-2 (arXiv:2505.04802, 2025).

54 ENVIRONMENTAL SCIENCES↗

Models, data, and scripts associated with “Prediction of Distributed River Sediment Respiration Rates using Community-Generated Data and Machine Learning”

This data package is associated with the publication “Prediction of Distributed River Sediment Respiration Rates using Community-Generated Data and Machine Learning’’ submitted to the Journal of Geophysical Research: Machine Learning and Computation (Scheibe et al. 2024). River sediment respiration observations are expensive and labor intensive to obtain and there is no physical model for predicting this quantity. The Worldwide Hydrobiogeochemisty Observation Network for Dynamic River Systems (WHONDRS) observational data set (Goldman et al.; 2020) is used to train machine learning (ML) models to predict respiration rates at unsampled sites. This repository archives training data, ML models, predictions, and model evaluation results for the purposes of reproducibility of the results in the associated manuscript and community reuse of the ML models trained in this project. One of the key challenges in this work was to find an optimum configuration for machine learning models to work with this feature-rich (i.e. 100+ possible input variables) data set. Here, we used a two-tiered approach to managing the analysis of this complex data set: 1) a stacked ensemble of ML models that can automatically optimize hyperparameters to accelerate the process of model selection and tuning and 2) feature permutation importance to iteratively select the most important features (i.e. inputs) to the ML models. The major elements of this ML workflow are modular, portable, open, and cloud-based, thus making this implementation a potential template for other applications. This data package is associated with the GitHub repository found at Please see the file level metadata (flmd; “sl-archive-whondrs_flmd.csv”) for a list of all files contained in this data package and descriptions for each. Please see the data dictionary (dd; “sl-archive-whondrs_dd.csv”) for a list of all column headers contained within comma separated value (csv) files in this data package and descriptions for each. The GitHub repository is organized into five top-level directories: (1) “input_data” holds the training data for the ML models; (2) “ml_models” holds machine learning models trained on the data in “input_data”; (3) “scripts” contains data preprocessing and postprocessing scripts and intermediate results specific to this data set that bookend the ML workflow; (4) “examples” contains the visualization of the results in this repository including plotting scripts for the manuscript (e.g., model evaluation, FPI results) and scripts for running predictions with the ML models (i.e., reusing the trained ML models); (5) “output_data” holds the overall results of the ML model on that branch. Each trained ML model resides on its own branch in the repository; this means that inputs and outputs can be different branch-to-branch. Furthermore, depending on the number of features used to train the ML models, the preprocessing and postprocessing scripts, and their intermediate results, can also be different branch-to-branch. The “main-*” branches are meant to be starting points (i.e. trunks) for each model branch (i.e. sprouts). Please see the Branch Navigation section in the top-level README.md in the GitHub repository for more details. There is also one hidden directory “.github/workflows”. This hidden directory contains information for how to run the ML workflow as an end-to-end automated GitHub Action but it is not needed for reusing the ML models archived here. Please the top-level README.md in the GitHub repository for more details on the automation.

13C↗

Dataset: "Widespread Drought-driven Declines in Streamflows and Water quality in the Upper Colorado River Basin (1998-2022)"

This data package contains the associated data and scripts for Nagamoto, E., Ombadi, M., Ciulla, F. et al. Widespread drought-driven declines in streamflows and water quality in the Upper Colorado River Basin during 1998-2022. Commun Earth Environ 7, 734 (2026). https://doi.org/10.1038/s43247-026-03890-5. This purpose of this study was to investigate the impact of the 21st century drought on water quantity and quality at catchments throughout the Upper Colorado River Basin (UCRB). We used stream flow, water temperature, specific conductance, air temperature, precipitation, and catchment attribute data for over 200 sites in the UCRB, collected from the National Water Information System using Basin3D (Varadharajan, 2023), GAGESII (Falcone, 2010), and the Google Earth Engine. We identified years of severe drought between 1998 and 2022 using the Standardized Precipitation Evaporation Index (SPEI), then calculated the relative change percentage of the stream flow, water temperature, and specific conductance from drought versus non-drought years. We used the attribute information from GAGESII to investigate what physical traits of catchments are associated streamflow vulnerability (greater relative change) or resilience to drought. We used land cover data from the National Land Cover Database (USGS, 2024) to assess any changes to physical attributes that may not be represented in the static attributes information in GAGESII. To increase data availability, we modeled stream temperature using methods from Willard, 2023. While the study period is water years 1998 to 2022, the raw water quantity and quality data extends to 1950 and the meteorological data extends to 1980. The data and code can be downloaded via the UCRB_drought.zip. Within the zip, the files are organized as follows: - INPUTS: Contains all input data used in UCRB_Drought_Workflow.ipynb - OUTPUTS: Contains all intermediate data created from UCRB_Drought_Workflow.ipynb as well as final products including the calculated Standardized Evapotranspiration Index (SPEI) - climatic_variables: The code used to collect meteorologic data from Google Earth Engine - feature_importance: The code used for the catchment attributes analysis - preprocessing: Code used in UCRB_Drought_Workflow_Preprocessing.ipynb - pyeto: Code used in UCRB_Drought_Workflow_Preprocessing.ipynb - calculations: Code used in UCRB_Drought_Workflow_Impacts.ipynb - plotting: Code used in UCRB_Drought_Workflow_Impacts.ipynb - README.md - UCRB_Drought_Workflow_Preprocessing.ipynb: The code used to prep raw data for the analysis - UCRB_Drought_Workflow_Impact.ipynb: The code which uses the prepped raw data for analysis, and plots all figures - requirements_ucrb-drought_v2.yml: The requirements file to create a virtual environment and Jupyter Lab kernel to run the code The INPUTS folder is organized into the following major directories and sub-directories. The "RDC_WT_SC_RAW" folder contains raw data for streamflow, water temperature, and specific conductance in a ".h5" file. The "NLCD_RAW" folder contains ".csv" files with annual land cover percentages for counties within the UCRB. The "MET_RAW" folder contains a ".csv" file with monthly meteorological data (air temperature and precipitation) for the sites in the UCRB which was obtained from code in the climatic_variables folder. The "GAGESII" folder contains ".csv" files with physical catchment attribute variables for catchments across the country. The "WT_LSTM_data" folder contains ".csv" files with calculated WT (Willard, 2023) and the associated RMSEs. The "Upper_Colorado_River_Basin_Boundary" folder contains geographic data including a shapefile for plotting in the UCRB_Drought_Workflow.ipynb. The "RESERVOIRS_RAW" folder contains ".csv" files for each reservoir in the UCRB with daily reservoir storage. There are also two files in the INPUTS folder that have combined reservoir storage data and reservoir metadata. The OUTPUTS folder is organized into the following major directories and sub-directories. The "RDC_WT_SC_data" folder contains a folder "Water_year" with the associated cleaned data, metadata, and data availability information in ".csv" files, a folder "Median_Relchange" with the relative change comparing drought to non-drought years in ".csv" files, and a folder "Peak95_Min5_Relchange" that has ".csv" files for the relative change in peak (95th %) and minimum (5th %) variables. The "NLCD_data" folder contains the difference in land cover from the beginning to end of the study period and the percentage of the county that is within UCRB bounds can be found in Nagamoto et al (2025)). The "MET_data" folder contains separated monthly air temperature and precipitation data and the calculated PET in ".csv" files. The "SPEI_data" folder contains ".csv" files with calculated SPEI values (one restricted to the study period and the other with information from the entire MET data period). The "Paper_Tables" folder contains two ".csv" files containing site information and data availability and information about the GAGESII trait aggregated categories. The base directory includes the file “flmd.csv” for a list and description of all files and the file “dd.csv” for data dictionaries. Scripts for preprocessing, analysis, and figure generation are located in the associated GitHub repository found at [https://github.com/iNAIADS/drought-impacts/tree/develop/UCRB-drought]. UPDATE 1: Title and code file updated to match submitted manuscript 10-15-2025. UPDATE 2: Code and data files updated to match revised manuscript 3-4-2026. UPDATE 3: Code and data files updated to match revised manuscript 6-7-2026. ** NOTE: DD and FLMD have not been updated yet. UPDATE 4: Added associated Manuscript information and DD and FLMD have been updated. To cite this code, please use the following BibTeX: @misc{nagamoto2025drought, author = {Emily Nagamoto and Fabio Ciulla and Mohammad Ombadi and Jared Willard and Rosemary Carroll and Charuleka Varadharajan}, title = {Dataset: "Widespread Drought-driven Declines in Streamflows and Water quality in the Upper Colorado River Basin (1998-2022)"}, year = {2025}, doi = {10.15485/2551894}, publisher = {ESS-DIVE Repository}, url = {https://data.ess-dive.lbl.gov/datasets/doi:10.15485/2551894} }

54 ENVIRONMENTAL SCIENCES↗

Phasor-Measurement-Unit-Based Data Analytics Using Digital Twin and PhasorAnalytics Software

A major objective of this project was to apply GE’s commercial machine learning and data analytics toolsets to large-scale, real-world, anonymized Phasor Measurement Unit (PMU) datasets in order to extract signatures, correlated and/or causal factors, and precursor patterns associated with significant power system phenomena. The project had a particular emphasis on extraction of insights relevant to asset health monitoring, real-time load modeling and cybersecurity monitoring. Additionally, the team was directed to undertake a comprehensive data quality analysis for the provided datasets and encouraged to estimate the ‘machine-learning readiness’ of the datasets by documenting any major obstacles to the application of commercial machine learning algorithms. To accomplish the aforementioned objectives, the project team’s work centered around the identification of key event signatures and application of the identified event signatures for event detection and event classification. The industry-validated, semi-supervised machine learning strategy employed for event signature identification involved several major tasks, including data-preprocessing, generation of an overabundance of features, normal data identification, normality modeling, and event signature identification through a methodical, quantitative ranking of features in order of relevance to each studied event type. Throughout the project, data quality issues and mitigation techniques were investigated. In this report, insights are provided regarding the readiness of the provided synchrophasor datasets for application of machine learning and data analytics. The methodologies employed for this technical strategy are summarized in this report. With regards to data preprocessing and feature generation, the provided Training and Test Datasets were ingested into GE’s big data environment. Subsequently, the team applied bad data cleansing and data imputation scripts, event detection scripts, and application programming interfaces (APIs) to the datasets for convenient data access. The project team completed development and validation of dozens of physics-based, statistics-based and transformation-based feature functions used for the extraction of over 60 synchrophasor features. Using a new parallel feature generation technology developed on this project, over 60 features have been rapidly generated for the full two years’ worth of Training and Test Dataset data associated with both the Eastern and Western interconnects. Even accommodating for temporal down-sampling inherent to the feature extraction procedure, this parallel feature generation activity resulted in a massive feature set with a storage requirement approximately equal to that of the raw training dataset itself. With regards to normal data identification and normality modeling, a normality model was built using the feature data extracted from the Training Dataset and iteratively refined subsequent to incremental adjustments and expansions of the Training Dataset feature data. With respect to event characterization and signature identification, an event signature identification pipeline was developed and used in conjunction with the normality model to identify over 15 event signatures for key event categories within the Training Dataset. The identified event signatures were used to characterize hundreds of key events in terms of relative severity, duration, and location of the event. An investigation was undertaken to identify correlated and causal factors involved in transformer events. A separate investigation into temporal trends in ring-down analysis results was undertaken to determine possible associations between system dynamics and various other factors such as loading, season or year. To validate the identified event signatures, additional work was undertaken to develop signature-based anomaly detection and classification tools suitable for convenient application to the synchrophasor datasets. The anomaly detection and classification tools, suitable for online application, were then applied to the entirety of the Eastern Interconnect Training and Test Datasets. Performance of the event detection and classification tools was evaluated upon receipt of the Test Dataset event logs (i.e., the labels for events contained in the Test Dataset), and promising results were obtained despite several challenges (documented herein) associated with application of supervised or semi-supervised machine learning methods to large-scale, anonymized datasets. Finally, the detection and classification tools were used to detect, classify, and characterize thousands of new events not included in the original event logs provided by the DOE within both the Training and Test Datasets.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Lignetics CRADA closeout report

Lignetics is one of the leading domestic producers of wood pellet fuels used for domestic heating. The company also produces pellets for other applications including animal bedding. The primary feedstock used for the production of Lignetics products is woody biomass. To improve the economics of pellet production, Lignetics is evaluating advanced preprocessing technologies developed by researchers at INL. In the proposed project, INL and Lignetics will work collaboratively to assess the feasibility of integrating INL’s advanced preprocessing technologies (fractional milling, high moisture pelleting and low temperature drying) in Lignetics production facilities. The successful completion of the project relies on the integration of the fractional milling and high moisture pelleting (HMP) process and low temperature drying to improve the economics of woody biomass pellet production for use in biofuel and bioproduct applications. In the HMP process, the biomass loses some moisture during compression and extrusion through the pellet die due frictional heat developed. Also, as pellets have a definite size and shape, they can be further dried in low temperature grain or belt dryers. Belt dryers are better suited to take advantage of low-grade and waste heat because they operate at lower temperatures than rotary dryers (Tumuluru et al.) Rotary dryers, for example, typically require inlet temperatures of 260°C, but more optimally operate around 400°C. In contrast, the inlet temperature of a belt dryer, such as a commercially available vacuum dryer, can be as low as 10°C above the ambient temperature, although more typically they operate at higher temperatures, between 90°C and 200°C. Because of their lower temperature operation, fire hazards and emissions to the air are lower for belt dryers in addition to reducing drying and overall pelleting costs, INL’s technology also has potential to reduce emissions at Lignetics facilities.

09 BIOMASS FUELS↗

Assessment of Machine Learning for Ultrasonic Nondestructive Evaluation of Alkali–Silica Reaction in Concrete

Alkali–silica reaction (ASR) is a type of material degradation in concrete structures that leads to concrete cracking and rebar corrosion, thereby reducing the material’s structural integrity and the overall structure’s lifetime and raising safety concerns. Ultrasonic nondestructive evaluation (NDE) has been proven to be a valuable technique for assessing concrete properties and monitoring ASR progression in concrete. However, the deployment and analysis of ultrasonic NDE and its data requires specialized expertise, often relying on the engineer’s subjective interpretation. With the surge in computational power, artificial intelligence (AI) and machine learning (ML) algorithms have become popular in automating NDE data analysis. Various industrial sectors are increasingly adopting ML algorithms for NDE data analysis with a growing emphasis on AI–assisted automation. Regulatory agencies are also preparing for this technological shift, anticipating corresponding revisions in standards. Thus, there is an urgent need to identify the capabilities and limitations of current ML technologies for the evaluation of concrete material properties and damage status. Furthermore, the effects of various factors on ML model performance must be thoroughly investigated. The study summarized herein evaluated the effectiveness of two ML models (i.e., support vector regression (SVR) and deep neural network (DNN)) in predicting concrete material damage induced by ASR based on the long-term ultrasonic monitoring data. Four distinct concrete specimens were cast with artificially induced ASR, and over a period exceeding 500 days, ultrasonic signals and expansion data were continuously collected. For the SVR model, wave velocity and 12 other wave features were extracted from the ultrasonic signals, with 6 out of 13 features selected as input for the model. Different combinations of training and testing datasets were designed to explore factors influencing prediction performance, including the range of data within training and testing sets, in addition to various signal preprocessing methodologies. These findings suggest the importance of using a training dataset with a broader data range compared with testing datasets for improved model performance alongside consistent signal preprocessing across datasets.

36 MATERIALS SCIENCE↗

Woody Feedstock 2022 State of Technology Report

The U.S. Department of Energy promotes production of advanced liquid transportation fuels from lignocellulosic biomass by funding fundamental and applied research that advances the state of technology (SOT). As part of its involvement in this mission, Idaho National Laboratory completes an annual SOT report for n th -plant and 1 st -plant woody biomass feedstock logistics. The purpose of the SOT is to provide the status of feedstock supply system technology development for woody biomass to biofuels relative to technical targets and cost goals from specific design cases, based on data and experimental results. Conventional feedstock supply systems need to be modified to meet the demands of conversion pathways, specifically to have the ability to adjust the quality of the raw biomass materials. Advanced systems incorporate innovative methods of material handling, preprocessing and supply chain configuration. In advanced designs, variability of the raw biomass can be reduced to produce feedstocks of a uniform format, moving toward biomass commoditization. Against this backdrop, the 2022 Woody SOT for low-ash woody feedstocks utilizes feedstock fractionation by incorporating technologies that can separate the biomass into its anatomical fractions (wood, bark, needle, and extrinsic ash) to reduce impurities and attempt to maximize the retention of usable fractions that satisfy downstream quality considerations. By using a series of air classification steps, this strategy can reduce the extrinsic ash in forest residues, separate out a majority of the incoming needles (which can be supplied to alternate markets), and maximize the retention of whitewood in the usable fraction. The fractionated forest residues are then mixed with clean-pine chips in a 50-50 blend to prepare the feedstock for the desired conversion pathway. The n th -plant analysis estimated the delivered cost for the feedstock at $\$$69.23/dry ton (2016$\$$) which represents a $\$$6.64/dry ton decrease compared to the cost estimate of the 2021 Woody SOT supply system for low-ash woody feedstocks. The quality requirements in the 2022 Woody SOT were identical to those of the 2021 Woody SOT at = 1.00 wt % ash and = 50.51 wt% carbon. The cost savings derive primarily from reductions in dry matter losses during air classification. The GHG emissions for the n th -plant analysis were estimated at 178.39 kg CO2e/dry ton compared to 178.71 kg CO 2 e/dry ton in the 2021 Woody SOT, a decrease of 0.32 kg CO2e/dry ton. The small change stems from an increase in emissions attributed to preprocessing and slightly larger savings in emissions from transportation. In the 1 st -plant analysis of the 2022 Woody SOT system, the average throughput was estimated to be approximately 2,128 dry tons/day or 96.51% of the name plate capacity. During the simulation the daily throughput ranged from 1,090 dry tons/day to 2,200 dry tons/day, or 49.43% to 99.75% of the daily nameplate capacity. After the year of operation 722,403 tons of processed feedstock were produced in total without regard to quality considerations (99.64% of the annual nameplate capacity). The variability in throughput was primarily caused by equipment failures in the system. Regular failures, downtime caused by routine maintenance per manufacturer guidelines, contributed to a majority 62.50% of failures and 62.60% of downtime. Failures due to wear were the other cause of disruption within the system, impacting the rotary shear and orbital screen and accounting for 37.50% of the failures and 37.40% of the total downtime. Ultimately the system was on stream for 87.84% during the simulation period, which is only 2.16 percentage points below the nth-plant assumption for on-stream time. The production cost of the system averaged $\$$71.66/dry ton. The costs ranged from a minimum of $\$$71.23/dry ton to a maximum of $\$$2,115.30/dry ton. When dry matter losses (disposed low-quality fractions as well as other losses such as in grinders) were considered the costs increased to an average of $\$$75.11/dry ton with a minimum of $\$$74.69/dry ton and a maximum of $\$$2,136.86/dry ton...

09 BIOMASS FUELS↗

A reproducible study design for the MIMIC-IV in-hospital mortality task

Open, tabular electronic health record (EHR) datasets such as MIMIC-III and MIMIC-IV have become critical resources for developing machine learning (ML) models addressing clinical prediction tasks, including hospital readmission, length of stay, and in-hospital mortality (IHM). While MIMIC-III has benefited from well-established preprocessing pipelines and standardized feature sets, MIMIC-IV remains comparatively challenging to work with because there are no standardized benchmarks to support reproducibility and comparability across studies. To address this limitation, we present a rigorously curated MIMIC-IV custom feature set optimized for IHM prediction, constructed through a reproducible preprocessing pipeline and feature selection strategy.

97 MATHEMATICS AND COMPUTING↗

Upgrading of Raw Coal and Coal Waste for Coal-Derived Graphene Process

Conference presentation at 47th International Technical Conference on Clean Energy (Clearwater Clean Energy Conference), Clearwater, Florida, July 23–27, 2023. The University of North Dakota Energy & Environmental Research Center (EERC) conducted a laboratory-scale coal-derived graphene (CDG) project focused on developing a technological process for making graphite from four U.S. domestic coals and coal wastes. Coal and coal waste preprocessing methods were developed and applied to clean and upgrade the coal precursors prior to graphitization and subsequent conversion to graphene products. Carbonization and graphitization of these preprocessed coals and coal wastes has produced graphite, which was used to make graphene oxide (GO) and reduced graphene oxide (rGO). Graphene quantum dots (GQDs) were also made from the raw and upgraded coal precursors.

01 COAL, LIGNITE, AND PEAT↗

Real-time infrared spectroscopy coupled with blind source separation for nuclear waste process monitoring

On-line infrared absorbance spectroscopy enables rapid measurement of solution-phase molecular species. Many spectra-to-concentration models exist for spectral data, with some models able to handle overlapping spectral bands and nonlinearities. However, model accuracy is limited by the quality of training data used in model fitting. The process spectra of nuclear waste simulants at the Savannah River Site display incongruity between training and process spectra; the glycolate spectral signature in the training data does not match the glycolate signature in Savannah River National Laboratory process data. A novel blind source separation algorithm is proposed that preprocesses spectral data so that process spectra more closely resemble training spectra, thereby improving model quantification accuracy when unexpected sources of variation appear in process spectra. The novel blind source separation preprocessing algorithm is shown to improve nitrate quantification from an R 2 of 0.934 to 0.988 and from 0.267 to 0.978 in two instances analyzing nuclear waste simulants from the Slurry Receipt Adjustment Tank and Slurry Mix Evaporator cycle at the Savannah River Site.

Crouse, Steven H.↗

Environmental Quenching of Low-surface-brightness Galaxies Near Hosts from Large Magellanic Cloud to Milky Way Mass Scales

Low-surface-brightness galaxies (LSBGs) are excellent probes of quenching and other environmental processes near massive galaxies. We study an extensive sample of LSBGs near massive hosts in the local universe that are distributed across a diverse range of environments. The LSBGs with surface-brightness ${\mu }_{\mathrm{eff},{g}}\gt 24.2\,\mathrm{mag}\,{\mathrm{arcsec}}^{-2}$ are drawn from the Dark Energy Survey Year 3 catalog while the hosts with masses $9.0\lt \mathrm{log}({{ \mathcal M }}_{\star }/{M}_{\odot })\lt 11.0$ comparable to the Milky Way and the Large Magellanic Cloud are selected from the z0MGS sample. We study the projected radial density profiles of LSBGs as a function of their color and surface brightness around hosts in both the rich Fornax–Eridanus cluster environment and the low-density field. We detect an overdensity with respect to the background density, out to 2.5 times the virial radius for both hosts in the cluster environment and the isolated field galaxies. When the LSBG sample is split by g − i color or surface brightness μ eff, g , we find the LSBGs closer to their hosts are significantly redder and brighter, like their high-surface-brightness counterparts. The LSBGs form a clear “red sequence” in both the cluster and isolated environments that is visible beyond the virial radius of the hosts. This suggests preprocessing of infalling LSBGs and a quenched backsplash population around both host samples. More so, the relative prominence of the “blue cloud” feature implies that preprocessing is ongoing near the isolated hosts compared to the cluster environment where the LSBGs are already well processed.

79 ASTRONOMY AND ASTROPHYSICS↗

Evaluation of native Earth system model output with ESMValTool v2.6.0

Earth system models (ESMs) are state-of-the-art climate models that allow numerical simulations of the past, present-day, and future climate. To extend our understanding of the Earth system and improve climate change projections, the complexity of ESMs heavily increased over the last decades. As a consequence, the amount and volume of data provided by ESMs has increased considerably. Innovative tools for a comprehensive model evaluation and analysis are required to assess the performance of these increasingly complex ESMs against observations or reanalyses. One of these tools is the Earth System Model Evaluation Tool (ESMValTool), a community diagnostic and performance metrics tool for the evaluation of ESMs. Input data for ESMValTool needs to be formatted according to the CMOR (Climate Model Output Rewriter) standard, a process that is usually referred to as “CMORization”. While this is a quasi-standard for large model intercomparison projects like the Coupled Model Intercomparison Project (CMIP), this complicates the application of ESMValTool to non-CMOR-compliant climate model output. In this paper, we describe an extension of ESMValTool introduced in v2.6.0 that allows seamless reading and processing of “native” climate model output, i.e., operational output produced by running the climate model through the standard workflow of the corresponding modeling institute. This is achieved by an extension of ESMValTool's preprocessing pipeline that performs a CMOR-like reformatting of the native model output during runtime. Thus, the rich collection of diagnostics provided by ESMValTool is now fully available for these models. For models that use unstructured grids, a further preprocessing step required to apply many common diagnostics is regridding to a regular latitude–longitude grid. Extensions to ESMValTool's regridding functions described here allow for more flexible interpolation schemes that can be used on unstructured grids. Currently, ESMValTool supports nearest-neighbor, bilinear, and first-order conservative regridding from unstructured grids to regular grids. Example applications of this new native model support are the evaluation of new model setups against predecessor versions, assessing of the performance of different simulations against observations, CMORization of native model data for contributions to model intercomparison projects, and monitoring of running climate model simulations. For the latter, new general-purpose diagnostics have been added to ESMValTool that are able to plot a wide range of variable types. Currently, five climate models are supported: CESM2 (experimental; at the moment, only surface variables are available), EC-Earth3, EMAC, ICON, and IPSL-CM6. As the framework for the CMOR-like reformatting of native model output described here is implemented in a general way, support for other climate models can be easily added.

58 GEOSCIENCES↗

Explainable Machine Learning for Functional Data

Black-box machine learning models are recognized as useful tools for prediction applications, but the algorithmic complexity of some models causes interpretation challenges. Explainability methods have been proposed to provide insight into these models, but there is little research focused on supervised modeling with functional data inputs. We argue that, especially in applications of high consequence, it is important to explicitly model the functional dependence in a black-box analysis to not obscure or misrepresent patterns in explanations. As such, we propose the V ariable importance E xplainable E lastic S hape A nalysis (VEESA) pipeline for training supervised machine learning models with functional inputs. The pipeline is an analysis process that includes the data preprocessing, modeling, and post-hoc explanations. The preprocessing is done using elastic functional principal components analysis, which accounts for vertical and horizontal variability in functional data and, ultimately, allows for explanations in the original data space that identify the important functional variability without bias due to correlated variables. Here, we demonstrate the pipeline on two high-consequence applications: explosives classification for national security and inkjet printer identification in forensic science. The applications exhibit the VEESA pipeline’s ability to provide an understanding of the characteristics of the functional data useful for prediction. Code for implementing the pipeline is available in the veesa R package (and supplemental python code).

Elastic Shape Analysis↗

Bio-project “derisking” through development of systematic methodologies and frameworks for risk assessment

One of the primary hindrances to producing a viable, sustainable domestic biomass industry for renewable biofuels, bio-products and bio-power is the lack of understanding and quantification of the risks associated with both the biomass supply chain and preprocessing and conversion technologies. Currently a consistent method for assessing, comparing, and quantifying risks in biomass supply chains does not exist, creating a major investment barrier to bioenergy projects in the U.S. The lack of a standardized approach has resulted in bioenergy stakeholders independently using inconsistent approaches and evaluation criteria, leading to unreliable and incomparable assessments of risks and financing barriers to bio-project development. Along with the challenges of inconsistent risk assessment for supply chain risk, technology specific risks based on variability in biomass properties are not fully understood and can pose significant unforeseen challenges for bioenergy projects. In many cases these properties have not yet been identified and the impacts on the proposed technology and products unquantified. This is particularly challenging for emerging preprocessing and conversion technologies. Without a firm understanding of the preprocessing/conversion technology-specific critical properties, the risk of a proposed bio-project cannot be fully evaluated. To address inconsistent risk evaluation in the biomass supply chain supporting project financing, a Biomass Supply Chain Risk Standards (BSCRS) framework was developed. The BSCRS framework includes a comprehensive list of known and perceived risks (Risk Indicators) to the supply chain developed through 100’s of interviews with bioenergy industry experts spanning from feedstock growers and suppliers to representatives from the financial sector. These risks have been organized into a manageable hierarchy of Risk Categories and Risk Factors that can be practically assessed. This BSCRS framework also provides mitigation strategies for multiple Risk Indicators from best available industry practices and research findings. Additionally, a risk quantification methodology for each Risk Factor, Risk Category, and the bio-project as a whole was developed to enable capital markets to assess feedstock risk more efficiently and more accurately. Multiple case studies representing existing bio-projects have been used to evaluate and verify the BSCRS framework and scoring methodology. To address technological risk along with the supply chain risk captured in the developed BSCRS framework, this work also focuses on development of a systematic criticality assessment tool using well-accepted, quantitative risk analysis methods to evaluate bioenergy feedstock critical properties impacting system unit operations. The proposed Failure Mode and Effect Analysis (FMEA) approach uses a team of subject area experts (SAEs) for each targeted unit operation within a system. Collectively, the team will develop and use a quantitative scoring system to assess the material attributes, process parameters, and quality attributes for key unit operations that have already been identified. The FMEA process generates Risk Priority Numbers (RPNs) for the various failures and predominant causes for each material/process unit/product combination resulting in a semi-quantitative, standardized methodology for assessing technological risk and biomass properties contributing to that risk.

09 BIOMASS FUELS↗