Application of Big Data Analytics and Machine Learning to Large-Scale Synchrophasor Datasets: Evaluation of Dataset ‘Machine Learning-Readiness’
Not Available
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Not Available
Abstract not provided
Not provided.
Not provided.
Not provided.
Explore the source record for details and available documents.
Peregrine, a software tool developed at Oak Ridge National Laboratory (ORNL), was used to collect and analyze in-situ monitoring (ISM) data from a Concept Laser M2 (Colibrium Additive) laser powder bed fusion (L-PBF) printer and an ExOne Innovent (Desktop Metal) binder jet printer. Data for four builds (print jobs) were saved to HDF5 (high performance data) files for release. Additionally, process anomalies were annotated by the authors across 37 image stacks (i.e., print layers) and are also provided as HDF5 files.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
2,3-butanediol (2,3-BDO) is an economically important platform chemical that can be used in a variety of chemical feedstocks, liquid fuels, and biosynthetic building blocks. While 2,3-BDO can be efficiently produced by fermentation, the fermentation requires continuous monitoring and control to maximize 2,3-BDO yields and minimize inhibitory coproducts. Because of the time required for sampling and at-line measurement of fermentation samples with high pressure liquid chromatography (HPLC), the ability for operators to perform real-time modification to fermentation conditions is limited. To overcome this challenge, researchers from the National Renewable Energy Laboratory (NREL) have developed a calibration model which can predict the concentration of several analytes in real-time using near-infrared (NIR) spectra of the filtered fermentation broth. While significantly reducing the need for off-line sampling, NREL expects this technology to play a critical role in maximizing 2,3-BDO production. 2,3-butanediol (2,3-BDO) is a useful chemical platform that can be used to create a variety of products. For instance, 2,3-BDO can be (1) dehydrated and converted into methyl ethyl ketone, a liquid fuel additive or (2) deoxydehydrated into 1,3-butadiene for synthetic rubber, which can also be oligomerized in high yields to gasoline, diesel, and jet fuel. In order to maximize 2,3-BDO production, frequent measurement of fermentation samples is needed, as small changes in oxygen concentration can drive the fermentation to undesired products. For example, oxygen-deficient conditions result in glycerol production, while excess oxygen concentrations result in acetoin production. This results in the need for measuring dissolved oxygen, glucose, and xylose concentrations in order to optimize the aeration rate of the fermentation. Traditional monitoring methods occurs off-line and can take up to 30 minutes per sample. With multiple fermenters and high-pressure liquid chromatography (HPLC) injectors, resulting in the need for multiple samples, the sampling process can take hours to complete. NREL’s calibration model can predict the glucose, xylose, 2,3-BDO, acetoin, and glycerol concentrations from NIR spectra of filtered fermentation liquor samples. Using a partial least-squares (PLS) calibration model, NREL’s model can monitor the concentration of these analytes during subsequent fermentations at bench- and pilot-scale, demonstrating the utility of NIR spectroscopy combined with chemometrics for real-time, at-line monitoring of 2,3-BDO fermentations.
NASA Data Active Archive Centers, or DAACs, ingest, store and distribute data acquired from satellites, ground systems as well as modelling data. These data are organized by the datasets, each presenting collection of files usually associated with the certain mission, instrument, processing level, parameter(s), algorithm and/or model. The number of datasets offered by a single DAAC to the public varies. GES DISC, for example, currently offers for public use approximately ~1,300 datasets. While each publicly offered dataset comes with supporting documentation, it is challenging for novice and even experienced scientists to navigate among the datasets that offer similar parameters to find the datasets for their particular research application. Supplying dataset documentation with the scientific paper citations that refer to that dataset provides means for the dataset users to educate themselves with the application research that dataset is being used in. Collecting citations of the papers that use the datasets for their research yield valuable insights into application areas of those datasets, information about usage of the dataset groups for specific applications and those application topics. It also gives insights into the “deep metrics” of the dataset usage, as opposed to the common metrics of the dataset usage such as number of users who downloaded the dataset files and volumes of downloaded data. Association of a certain scientific paper with the dataset(s) presents a challenge because most of the paper authors do not properly cite the datasets, datasets usually have cryptic names and Digital Object Identifiers (DOIs) that are used for dataset identification were assigned to the datasets only few years ago. Simple Google or online library search do not provide even meaningful fraction of the results when performed by the dataset name or DOI, however they provide too many results when the search is done by more broader terms such as mission and instrument names. Attempts to create an AI system capable to identify dataset in the scientific papers have already been made using neural networks classifiers on the basis of the dataset mission, instrument and variable name. This method was applied to NASA SEDAC, which has 41 datasets in total. In GES DISC there can be as many as ~100 datasets per mission/instrument with some of the datasets consisting of multiple variables so there is a need for more differentiating parameters for dataset identification in the paper. The approach we are currently investigating is creating AI classifiers that are based on multiple dataset features, or keywords, extracted from the NASA Earthdata Common Dataset Repository (CMR). The features are weighted based on how precisely they can identify a dataset. The classifier uses preprocessed paper text as input and searches for the CMR datasets whose feature sets are the closest to the feature sets contained in the paper. The challenges of dataset identification include variety of ways the paper authors describe the datasets in their papers and incomplete tagging of the CMR dataset description (DIFs).
Scientific datasets are increasingly cited in peer-reviewed journal publications, facilitating easy access to research utilizing those datasets. Datasets undergo a life cycle where older versions of datasets are replaced by newer versions often due to improvements in data resolution, algorithms, and other factors. Unlike peer reviewed documents registered with a single Digital Unique Identifier (DOI), datasets can be updated over time and the newer version of the datasets are registered with a new DOI which is not necessarily linked to the previous version of the dataset. It is challenging when publications citing a dataset need to be traced over the entire life cycle of that dataset. We provide an innovative approach to link the dataset versions and publications using a knowledge graph (KG). KG can help to trace the dataset cited in publications over the entire dataset life cycle and shed light into dataset usage in various applied research areas. We fine-tuned the pretrained NASA IMPACTINDUS Large Language Model (LLM) on a set of labeled publications abstracts. Our results showed that 87% of the publications were classified into one of twenty applied research areas, while the remaining 13% were classified into non-applied research areas. By linking datasets to applied research areas through the KG and employing Global Change Master Directory(GCMD), a well-established controlled vocabulary of scientific keywords describing Earth science datasets, we contribute to a transparent and advanced search and discovery mechanism for datasets across the Earth data ecosystem. The integrated KG and LLM approach is now incorporated and operational in dataset publication management at one of NASA’s Earth science data archival centers.
DEEPEN stands for DE-risking Exploration of geothermal Plays in magmatic ENvironments. As part of the development of the DEEPEN 3D play fairway analysis (PFA) methodology for magmatic plays (conventional hydrothermal, superhot EGS, and supercritical), weights needed to be developed for use in the weighted sum of the different favorability index models produced from geoscientific exploration datasets. This was done using two different approaches: one based on expert opinions, and one based on statistical learning. This GDR submission includes the datasets used to produce the statistical learning-based weights. While expert opinions allow us to include more nuanced information in the weights, expert opinions are subject to human bias. Data-centric or statistical approaches help to overcome these potential human biases by focusing on and drawing conclusions from the data alone. The drawback is that, to apply these types of approaches, a dataset is needed. Therefore, we attempted to build comprehensive standardized datasets mapping anomalies in each exploration dataset to each component of each play. This data was gathered through a literature review focused on magmatic hydrothermal plays along with well-characterized areas where superhot or supercritical conditions are thought to exist. Datasets were assembled for all three play types, but the hydrothermal dataset is the least complete due to its relatively low priority. For each known or assumed resource, the dataset states what anomaly in each exploration dataset is associated with each component of the system. The data is only a semi-quantitative, where values are either high, medium, or low, relative to background levels. In addition, the dataset has significant gaps, as not every possible exploration dataset has been collected and analyzed at every known or suspected geothermal resource area, in the context of all possible play types. The following training sites were used to assemble this dataset: - Conventional magmatic hydrothermal: Akutan (from AK PFA), Oregon Cascades PFA, Glass Buttes OR, Mauna Kea (from HI PFA), Lanai (from HI PFA), Mt St Helens Shear Zone (from WA PFA), Wind River Valley (From WA PFA), Mount Baker (from WA PFA). - Superhot EGS: Newberry (EGS demonstration project), Coso (EGS demonstration project), Geysers (EGS demonstration project), Eastern Snake River Plain (EGS demonstration project), Utah FORGE, Larderello, Kakkonda, Taupo Volcanic Zone, Acoculco, Krafla. - Supercritical: Coso, Geysers, Salton Sea, Larderello, Los Humeros, Taupo Volcanic Zone, Krafla, Reyjanes, Hengill. **Disclaimer: Treat the supercritical fluid anomalies with skepticism. They are based on assumptions due to the general lack of confirmed supercritical fluid encounters and samples at the sites included in this dataset, at the time of assembling the dataset. The main assumption was that the supercritical fluid in a given geothermal system has shared properties with the hydrothermal fluid, which may not be the case in reality. Once the datasets were assembled, principal component analysis (PCA) was applied to each. PCA is an unsupervised statistical learning technique, meaning that labels are not required on the data, that summarized the directions of variance in the data. This approach was chosen because our labels are not certain, i.e., we do not know with 100% confidence that superhot resources exist at all the assumed positive areas. We also do not have data for any known non-geothermal areas, meaning that it would be challenging to apply a supervised learning technique. In order to generate weights from the PCA, an analysis of the PCA loading values was conducted. PCA loading values represent how much a feature is contributing to each principal component, and therefore the overall variance in the data.
NASA Data Centers provide the public with thousands of datasets that result in published papers, reports, and conference proceedings. Collecting accurate metrics on usage of these datasets is key to connecting different areas of knowledge and evaluating the datasets’ impact. While most of the datasets have Digital Object Identifiers (DOIs) assigned, most publications do not cite them hampering the automated search of these publications. Instead, articles mention attributes like organization, instrument, mission, variable, or a publication describing the dataset. Often only domain experts can deduce the dataset that was used in the publication text. The lack of a citation slows the spread of information and reduces the research’s impact. With thousands of papers produced each year, an automated means of labeling datasets is critical. This paper explores a hybrid approach of heuristics and a Natural Language Processing (NLP) Named Entity Recognition (NER) model to find and label the datasets used within Earth Science papers. Heuristics are used to produce the labelled sentences and any potential dataset candidates that can be derived from a sentence. The heuristic labels the sentences with the names of mission, instrument, re-analysis models, and science keywords taken from the Global Change Master Directory (GCMD) ontology. Additionally, it uses those labels to generate the dataset citation candidates. If the mission, instrument, and variable are sufficient to create the citation for the dataset the citation and the label the domain expert reviews the output without going through the NLP model. If the extracted label is not sufficient to label the dataset on its own, the sentence and its associated dataset labels will be inputted into the NER model. The model outputs the labeled sentence and the potential dataset candidates with their associated probabilities. The domain expert then reviews the NER model’s output and the correct labels are determined. The newly labelled papers can then be used as additional training data. This creates an iterative process for the approach to continuously improve. Because all the possible mentions are gathered by the model, the domain expert can quickly and easily label the papers resulting in large time savings.
Three independent, quasi-global, gridded datasets of precipitation (a rain gauge-based dataset, the satellite-only component of the NASA Integrated Multi-satellitE Retrievals for Global Precipitation Measurement mission [IMERG] Final Run precipitation product, and precipitation estimates derived from NASA Soil Moisture Active Passive [SMAP] soil moisture retrievals), are objectively combined into a single pentad precipitation dataset at 36-km resolution using a unique approach based on extended triple collocation. The quality of each of the four datasets is then evaluated against independent observations. When a global land surface model at 36-km resolution is integrated four times, once utilizing the merged precipitation forcing and once with each of the three contributing datasets, the near-surface soil moisture variations produced with the merged forcing validate best against independent satellite-based soil moisture fields. In addition, the merged dataset is found to be more consistent, relative to each contributor, with estimates of air temperature variations across the globe. The merged dataset thus appears to draw successfully on the complementary strengths of each contributor: the particularly high quality of the rain gauge-based dataset in areas of high gauge density, the more uniform accuracy across the globe of the IMERG data, and the moderate accuracy, particularly in semi-arid regions, of the soil moisture retrieval-based data. Plain Language Summary Obtaining measurements of precipitation across the globe can be challenging. Rain gauges in some ways provide the most accurate measurements, but gauges are absent in many parts of the world, and even where they exist, they only measure precipitation at the gauge itself and therefore may not provide an accurate large-scale average. Satellite-based estimates of precipitation largely overcome these problems, but such data have their own issues, notably a “snapshot” (rather than a time-average) character of the measurements and difficulty associated with interpreting the measured radiances in the presence of complex land surfaces. In the present paper, we use a novel approach to generate a “merged” dataset, one that optimally combines the gauge precipitation information and the satellite-based precipitation information with a third set of estimates derived from soil moisture retrievals. The merged precipitation dataset and each of the three contributors (aggregated here to 5-day averages at a spatial resolution of about 36-km) are then evaluated for consistency with independent geophysical fields. The merged dataset is found to perform best, a clear indication that it takes proper advantage of the complementary strengths of each contributor and, accordingly, that the presented approach for merging the different contributors is indeed viable.
Minirhizotron technology is widely used to study root growth and development. Yet, standard approaches for tracing roots in minirhiztron imagery is extremely tedious and time consuming. Machine learning approaches can help to automate this task. However, lack of enough annotated training data is a major limitation for the application of machine learning methods. Transfer learning is a useful technique to help with training when available datasets are limited. In this paper, we investigated the effect of pre-trained features from the massives-cale, irrelevant ImageNet dataset and a relatively moderate-scale, but relevant peanut root dataset on switchgrass root imagery segmentation applications. We compiled two minirhizotron image datasets to accomplish this study: one with 17,550 peanut root images and another with 28 switchgrass root images. Both datasets were paired with manually labeled ground truth masks. Deep neural networks based on the U-net architecture were used with different pre-trained features as initialization for automated, precise pixel-wise root segmentation in minirhizotron imagery. We observed that features pre-trained on a closely related but relatively moderate size dataset like our peanut dataset were more effective than features pre-trained on the large but unrelated ImageNet dataset. Here, we achieved high quality segmentation on peanut root dataset with 99.04% accuracy at the pixel-level and overcame errors in human-labeled ground truth masks. By applying transfer learning technique on limited switchgrass dataset with features pre-trained on peanut dataset, we obtained 99% segmentation accuracy in switchgrass imagery using only 21 images for training (fine tuning). Furthermore, the peanut pre-trained features can help the model converge faster and have much more stable performance.
Biological space experiments are often expensive and difficult to conduct. As such, it is critical to maximize the value of the data that is collected during these experiments. One way to do this is to combine multiple–previously separate–datasets. This can increase the number of replicates for the conditions of interest (and hence statistical power), allow new multi-factor questions to be asked, and potentially highlight new patterns that otherwise would not have been identified from single-dataset studies. However, the process of combining datasets introduces noise due to inherent technical variations between experiments. To better understand the insights that can be gained from multi-dataset analyses and the problems that may arise from joining multiple datasets, several mouse muscle RNA-Seq datasets from the Rodent Research-1 mission were first selected. Then, using the R package DESeq2, principal component analysis (PCA) plots and differentially expressed gene (DEG) lists between ground and flight muscle samples were generated for individual datasets and for different pairwise combinations of datasets. Several new DEGs were identified in the combined datasets, and patterns in the PCA plots were affected depending on which datasets were joined. Understanding the results of this work will be critical for future studies that seek to perform multi-dataset analyses.