Application of Big Data Analytics and Machine Learning to Large-Scale Synchrophasor Datasets: Evaluation of Dataset ‘Machine Learning-Readiness’
Not Available
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Not Available
Abstract not provided
Not provided.
Not provided.
Not provided.
Explore the source record for details and available documents.
Peregrine, a software tool developed at Oak Ridge National Laboratory (ORNL), was used to collect and analyze in-situ monitoring (ISM) data from a Concept Laser M2 (Colibrium Additive) laser powder bed fusion (L-PBF) printer and an ExOne Innovent (Desktop Metal) binder jet printer. Data for four builds (print jobs) were saved to HDF5 (high performance data) files for release. Additionally, process anomalies were annotated by the authors across 37 image stacks (i.e., print layers) and are also provided as HDF5 files.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
2,3-butanediol (2,3-BDO) is an economically important platform chemical that can be used in a variety of chemical feedstocks, liquid fuels, and biosynthetic building blocks. While 2,3-BDO can be efficiently produced by fermentation, the fermentation requires continuous monitoring and control to maximize 2,3-BDO yields and minimize inhibitory coproducts. Because of the time required for sampling and at-line measurement of fermentation samples with high pressure liquid chromatography (HPLC), the ability for operators to perform real-time modification to fermentation conditions is limited. To overcome this challenge, researchers from the National Renewable Energy Laboratory (NREL) have developed a calibration model which can predict the concentration of several analytes in real-time using near-infrared (NIR) spectra of the filtered fermentation broth. While significantly reducing the need for off-line sampling, NREL expects this technology to play a critical role in maximizing 2,3-BDO production. 2,3-butanediol (2,3-BDO) is a useful chemical platform that can be used to create a variety of products. For instance, 2,3-BDO can be (1) dehydrated and converted into methyl ethyl ketone, a liquid fuel additive or (2) deoxydehydrated into 1,3-butadiene for synthetic rubber, which can also be oligomerized in high yields to gasoline, diesel, and jet fuel. In order to maximize 2,3-BDO production, frequent measurement of fermentation samples is needed, as small changes in oxygen concentration can drive the fermentation to undesired products. For example, oxygen-deficient conditions result in glycerol production, while excess oxygen concentrations result in acetoin production. This results in the need for measuring dissolved oxygen, glucose, and xylose concentrations in order to optimize the aeration rate of the fermentation. Traditional monitoring methods occurs off-line and can take up to 30 minutes per sample. With multiple fermenters and high-pressure liquid chromatography (HPLC) injectors, resulting in the need for multiple samples, the sampling process can take hours to complete. NREL’s calibration model can predict the glucose, xylose, 2,3-BDO, acetoin, and glycerol concentrations from NIR spectra of filtered fermentation liquor samples. Using a partial least-squares (PLS) calibration model, NREL’s model can monitor the concentration of these analytes during subsequent fermentations at bench- and pilot-scale, demonstrating the utility of NIR spectroscopy combined with chemometrics for real-time, at-line monitoring of 2,3-BDO fermentations.
DEEPEN stands for DE-risking Exploration of geothermal Plays in magmatic ENvironments. As part of the development of the DEEPEN 3D play fairway analysis (PFA) methodology for magmatic plays (conventional hydrothermal, superhot EGS, and supercritical), weights needed to be developed for use in the weighted sum of the different favorability index models produced from geoscientific exploration datasets. This was done using two different approaches: one based on expert opinions, and one based on statistical learning. This GDR submission includes the datasets used to produce the statistical learning-based weights. While expert opinions allow us to include more nuanced information in the weights, expert opinions are subject to human bias. Data-centric or statistical approaches help to overcome these potential human biases by focusing on and drawing conclusions from the data alone. The drawback is that, to apply these types of approaches, a dataset is needed. Therefore, we attempted to build comprehensive standardized datasets mapping anomalies in each exploration dataset to each component of each play. This data was gathered through a literature review focused on magmatic hydrothermal plays along with well-characterized areas where superhot or supercritical conditions are thought to exist. Datasets were assembled for all three play types, but the hydrothermal dataset is the least complete due to its relatively low priority. For each known or assumed resource, the dataset states what anomaly in each exploration dataset is associated with each component of the system. The data is only a semi-quantitative, where values are either high, medium, or low, relative to background levels. In addition, the dataset has significant gaps, as not every possible exploration dataset has been collected and analyzed at every known or suspected geothermal resource area, in the context of all possible play types. The following training sites were used to assemble this dataset: - Conventional magmatic hydrothermal: Akutan (from AK PFA), Oregon Cascades PFA, Glass Buttes OR, Mauna Kea (from HI PFA), Lanai (from HI PFA), Mt St Helens Shear Zone (from WA PFA), Wind River Valley (From WA PFA), Mount Baker (from WA PFA). - Superhot EGS: Newberry (EGS demonstration project), Coso (EGS demonstration project), Geysers (EGS demonstration project), Eastern Snake River Plain (EGS demonstration project), Utah FORGE, Larderello, Kakkonda, Taupo Volcanic Zone, Acoculco, Krafla. - Supercritical: Coso, Geysers, Salton Sea, Larderello, Los Humeros, Taupo Volcanic Zone, Krafla, Reyjanes, Hengill. **Disclaimer: Treat the supercritical fluid anomalies with skepticism. They are based on assumptions due to the general lack of confirmed supercritical fluid encounters and samples at the sites included in this dataset, at the time of assembling the dataset. The main assumption was that the supercritical fluid in a given geothermal system has shared properties with the hydrothermal fluid, which may not be the case in reality. Once the datasets were assembled, principal component analysis (PCA) was applied to each. PCA is an unsupervised statistical learning technique, meaning that labels are not required on the data, that summarized the directions of variance in the data. This approach was chosen because our labels are not certain, i.e., we do not know with 100% confidence that superhot resources exist at all the assumed positive areas. We also do not have data for any known non-geothermal areas, meaning that it would be challenging to apply a supervised learning technique. In order to generate weights from the PCA, an analysis of the PCA loading values was conducted. PCA loading values represent how much a feature is contributing to each principal component, and therefore the overall variance in the data.
Minirhizotron technology is widely used to study root growth and development. Yet, standard approaches for tracing roots in minirhiztron imagery is extremely tedious and time consuming. Machine learning approaches can help to automate this task. However, lack of enough annotated training data is a major limitation for the application of machine learning methods. Transfer learning is a useful technique to help with training when available datasets are limited. In this paper, we investigated the effect of pre-trained features from the massives-cale, irrelevant ImageNet dataset and a relatively moderate-scale, but relevant peanut root dataset on switchgrass root imagery segmentation applications. We compiled two minirhizotron image datasets to accomplish this study: one with 17,550 peanut root images and another with 28 switchgrass root images. Both datasets were paired with manually labeled ground truth masks. Deep neural networks based on the U-net architecture were used with different pre-trained features as initialization for automated, precise pixel-wise root segmentation in minirhizotron imagery. We observed that features pre-trained on a closely related but relatively moderate size dataset like our peanut dataset were more effective than features pre-trained on the large but unrelated ImageNet dataset. Here, we achieved high quality segmentation on peanut root dataset with 99.04% accuracy at the pixel-level and overcame errors in human-labeled ground truth masks. By applying transfer learning technique on limited switchgrass dataset with features pre-trained on peanut dataset, we obtained 99% segmentation accuracy in switchgrass imagery using only 21 images for training (fine tuning). Furthermore, the peanut pre-trained features can help the model converge faster and have much more stable performance.
Machine learning interatomic potentials (MLIPs) balance high accuracy and lower costs compared to density functional theory calculations, but their performance often depends on the size and diversity of training datasets. Large datasets improve model accuracy and generalization but are computationally expensive to produce and train on, while smaller datasets risk discarding rare but important atomic environments and compromising MLIP accuracy/reliability. Here, we develop an information-theoretical framework to quantify the efficiency of dataset compression methods and propose an algorithm that maximizes this efficiency. By framing atomistic dataset compression as an instance of the minimum set cover (MSC) problem over atom-centered environments, our method identifies the smallest subset of structures that contains as much information as possible from the original dataset while pruning redundant information. The approach is extensively demonstrated on the GAP-20 and TM23 datasets and validated on 64 varied datasets from the ColabFit repository. Across all cases, MSC consistently retains outliers, preserves dataset diversity, and reproduces the long-tail distributions of forces even at high compression rates, outperforming other subsampling methods. Furthermore, MLIPs trained on MSC-compressed datasets exhibit reduced error for out-of-distribution data even in low-data regimes. We explain these results using an outlier analysis and show that such quantitative conclusions could not be achieved with conventional dimensionality reduction methods. The algorithm is implemented in the open-source QUESTS package and can be used for several tasks in atomistic modeling, from data subsampling, outlier detection, and training improved MLIPs at a lower cost.
Abstract. High-resolution gridded datasets of meteorological variables are needed in order to resolve fine-scale hydrological gradients in complex mountainous terrain. Across the United States, the highest available spatial resolution of gridded datasets of daily meteorological records is approximately 800 m. This work presents gridded datasets of daily precipitation and mean temperature for the East–Taylor subbasin (in the western United States) covering a 12-year period (2008–2019) at a high spatial resolution (400 m). The datasets are generated using a downscaling framework that uses data-driven models to learn relationships between climate variables and topography. We observe that downscaled datasets of precipitation and mean temperature exhibit smoother spatial gradients (while preserving the spatial variability) when compared to their coarser counterparts. Additionally, we also observe that when downscaled datasets are upscaled to the original resolution (800 m), the mean residual error is almost zero, ensuring no bias when compared with the original data. Furthermore, the downscaled datasets are observed to be linearly related to elevation, which is consistent with the methodology underlying the original 800 m product. Finally, we validate the spatial patterns exhibited by downscaled datasets via an example use case that models lidar-derived estimates of snowpack. The presented dataset constitutes a valuable resource to resolve fine-scale hydrological gradients in the mountainous terrain of the East–Taylor subbasin, which is an important study area in the context of water security for the southwestern United States and Mexico. The dataset is publicly available at https://doi.org/10.15485/1822259 (Mital et al., 2021).
The Xanthos-Lake v1.0 dataset provides the input data, trained machine-learning models, and simulation outputs needed to characterize lake water balance, snow and ice conditions, and mixing-layer temperature within the Xanthos global hydrological modeling framework. The dataset supports lake representation across a wide range of lake sizes and hydroclimatic conditions by combining xLSIM, a basin-specific machine-learning emulator of lake snow, ice, ice-cover fraction, and mixing-layer temperature, with the Xanthos-Lake water-balance model. The archive contains NetCDF datasets used to train and evaluate xLSIM, trained model weights, processed meteorological and lake-property inputs, and basin- and lake-category-specific simulation outputs. These materials are organized into four primary data groups, described below. Snowice_model_inputs: Contains the NetCDF input data used to train xLSIM. The xLSIM machine-learning framework uses three lake-based datasets. The meteorological forcing dataset provides monthly relative humidity, specific humidity, surface wind speed, maximum and minimum air temperature, downward longwave and shortwave radiation, snowfall, surface air pressure, and total precipitation. Lake surface area is included as an additional static predictor. The target-state dataset provides lake ice thickness, snow depth, snow cover, and lake mixing-layer temperature, while a companion lake-surface dataset provides the lake ice-cover fraction. Before training, ice thickness and snow depth are converted from meters to centimeters, mixing-layer temperature is converted from kelvin to degrees Celsius and constrained to nonnegative values, and ice-cover fraction is converted from a fraction to a percentage. The predictor variables are normalized using statistics calculated across the selected lakes and time steps. Snowice_model_outputs: Contains the NetCDF outputs generated by xLSIM. For each basin, xLSIM produces a file containing observed and predicted lake-state variables for the training, validation, and testing periods. The modeled variables include lake ice thickness, snow depth, snow cover, mixing-layer temperature, and lake ice-cover fraction. For basins without a sufficiently persistent snow-and-ice signal, the emulator predicts only mixing-layer temperature. The outputs also include training and validation loss histories, the selected model configuration, identifiers of the lakes used in training, and SHAP-based feature-importance information at the global, lake, and seasonal-regime levels. The trained machine-learning model weights are provided separately within the dataset archive. Together, these files support model evaluation and subsequent coupling with the Xanthos-Lake water-balance framework. XanthosLAKES: Contains the NetCDF input data used by the Xanthos-Lake framework. Monthly meteorological inputs include relative and specific humidity, downward shortwave and longwave radiation, mean, maximum, and minimum air temperature, wind speed, precipitation, snowfall, and surface air pressure. Static lake-property datasets provide lake identifiers, geographic locations, surface area, volume, mean depth, elevation, drainage area, fetch, outlet-routing information, and associated Xanthos grid-cell attributes. Separate bathymetric datasets provide the coefficients of the area–depth and volume–depth relationships for each aggregated lake unit. GLEV-based records provide observed lake surface area and evaporation data used to initialize lake states, define reference conditions, and calibrate and evaluate the model. Xanthos-Lake Outputs: Contains the basin- and lake-category-specific NetCDF outputs generated by Xanthos-Lake. Monthly variables include lake surface area, storage volume, outlet discharge, evaporation rate, evaporation volume, lake–groundwater exchange, lake inflow, ice thickness, snow depth, snow-cover fraction, ice-cover fraction, and mixing-layer temperature. The files also contain lake-specific calibration and validation statistics, including normalized root-mean-square error, mean absolute error, Nash–Sutcliffe efficiency, Kling–Gupta efficiency, and percent bias. Stored calibrated and derived parameters include the weir discharge coefficient, fractional freeboard, groundwater exchange coefficient, reference water level, corresponding reference surface area and storage volume, weir-width adjustment factor, and the fraction of routed inflow entering the lake. Basin identifiers, lake category, simulation period, calibration and validation periods, and parameter-schema information are retained as NetCDF metadata.
Effective monitoring of global water resources is increasingly critical due to climate change and population growth. Advancements in remote sensing technology, specifically in spatial, spectral, and temporal resolutions, are revolutionizing water resource monitoring, leading to more frequent and high-quality surface water extent maps using various techniques such as traditional image processing and machine learning algorithms. However, satellite imagery datasets contain trade-offs that result in inconsistencies in performance, such as disparities in measurement principles between optical (e.g., Sentinel-2) and radar (e.g., Sentinel-1) sensors and differences in spatial and spectral resolutions among optical sensors. Therefore, developing accurate and robust surface water mapping solutions requires independent validations from multiple datasets to identify potential biases within the imagery and algorithms. However, high-quality validation datasets are expensive to build, and few contain information on water resources. For this purpose, we introduce a globally sampled, high-spatial-resolution dataset labeled using 3 m PlanetScope imagery. Our surface water extent dataset comprises 100 images, each with a size of 1024×1024 pixels, which were sampled using a stratified random sampling strategy covering all 14 biomes. We highlighted urban and rural regions, lakes, and rivers, including braided rivers and coastal regions. We evaluated two surface water extent mapping methods using our dataset – Dynamic World, based on Sentinel-2, and the NASA IMPACT model, based on Sentinel-1. Dynamic World achieved a mean intersection over union (IoU) of 72.16 % and F1 score of 79.70 %, while the NASA IMPACT model had a mean IoU of 57.61 % and F1 score of 65.79 %. Performance varied substantially across biomes, highlighting the importance of evaluating models on diverse landscapes to assess their generalizability and robustness. Our dataset can be used to analyze satellite products and methods, providing insights into their advantages and drawbacks. Our dataset offers a unique tool for analyzing satellite products, aiding the development of more accurate and robust surface water monitoring solutions. The dataset can be accessed via https://doi.org/10.25739/03nt-4f29.
We consider a high-dimensional nonlinear computational model of a dynamical system, parameterized by a vector-valued control parameter, in the presence of uncertainties represented by an uncontrolled parameter modeled by a vector-valued random variable, and possibly with stochastic excitation. The objective is to construct a statistical surrogate model where the input is any deterministic value of the control parameter, and the output is a vector-valued observation of the computational model, which is a random vector whose probability measure is updated using a target dataset. To construct this statistical surrogate model, the stochastic response of the computational model must be built, which is a vector-valued time-discretized stochastic process in high dimension, depending on the control parameter. It is assumed that the computational cost of a single evaluation of the deterministic model is high. For the probabilistic updating, we consider a subset of the components of the observation of the computational model, defined as the “identification observation” of the computational model, for which a small target dataset is available. Therefore, the target dataset is associated with partial observability, corresponding to an incomplete data case. Given a prior probability model of the random control and uncontrolled parameters, a training dataset is constructed, consisting of realizations of the random triplet composed of the stochastic response, the random identification observation, and the random control parameter. Since the computational cost of a single evaluation of the deterministic model is assumed to be large, the training dataset is also of small size. The main challenges in this problem are the high dimensionality, partial observability leading to incomplete data in the target dataset for the identification observation of the computational model (which is not sufficient to identify the computational stochastic responses), and the availability of a small training dataset. To address these challenges, we propose a methodology based on statistical methods for constructing necessary reduced representations, direct probabilistic learning under constraints using probabilistic learning on manifolds (PLoM) constrained by the target dataset, and the use of a weak formulation of the Fourier transform of probability measures. Statistical conditioning is also employed to explore the learned dataset. The constructed predictive statistical surrogate model can be implemented in the context of online computation. Here, we apply this approach to a problem of nonlinear stochastic dynamics in high dimensions within the framework of deformable solids mechanics.