Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Curation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Nuclear Physics Exascale Requirements Review: An Office of Science Review sponsored jointly by Advanced Scientific Computing Research and Nuclear Physics, June 15 - 17, 2016, Gaithersburg, Maryland

Imagine being able to predict — with unprecedented accuracy and precision — the structure of the proton and neutron, and the forces between them, directly from the dynamics of quarks and gluons, and then using this information in calculations of the structure and reactions of atomic nuclei and of the properties of dense neutron stars (NSs). Also imagine discovering new and exotic states of matter, and new laws of nature, by being able to collect more experimental data than we dream possible today, analyzing it in real time to feed back into an experiment, and curating the data with full tracking capabilities and with fully distributed data mining capabilities. Making this vision a reality would improve basic scientific understanding, enabling us to precisely calculate, for example, the spectrum of gravity waves emitted during NS coalescence, and would have important societal applications in nuclear energy research, stockpile stewardship, and other areas. This review presents the components and characteristics of the exascale computing ecosystems necessary to realize this vision.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Panorama 360 (Final Report)

This is the final technical report for the DOE-funded Panorama 360 project. Panorama 360 provided a resource for the collection, analysis, and sharing of performance data about end-to-end scientific workflows executing on DOE facilities. The work focused on workflows that include experimental data generation at DOE facilities. The main activities of Panorama 360 include the development of: 1. A distributed repository that stores different types of workflow execution data (e.g., point and time series performance traces at fine- and coarse-grained levels); 2. A set of open-source data capture, curation, and publishing tools fully integrated with a state-of-the-art workflow management system that automates data ingestion to the repository and enables users to discover, query, and process data from the repository; 3. A set of analysis algorithms and machine learning based tools to perform analysis and characterization of the gathered data, which can be used to detect anomalous performance or system faults; and 4. Best practices and recommendations for workflow evaluation, analysis, execution, and architectures.

97 MATHEMATICS AND COMPUTING↗

Panorama 360 (Final Report)

This final technical report from the lead institution, USC grant #DE-SC0012636, serves as the final technical report for collaborative institution UNC-CH grant #DE-SC0012390. The goal was to develop a repository and associated capabilities for data collection, ingestion, and analysis for a broad class of DOE applications that span experimental and simulation science workflows. In particular, this work focuses on workflows that include experimental data generation at DOE facilities. The main activities of Panorama 360 include the development of: (1) A distributed repository that stores different types of workflow execution data (e.g., point and time series performance traces at fine- and coarse-grained levels); (2) A set of open-source data capture, curation, and publishing tools fully integrated with a state-of-the-art workflow management system that automates data ingestion to the repository and enables users to discover, query, and process data from the repository; (3) A set of analysis algorithms and machine learning based tools to perform analysis and characterization of the gathered data, which can be used to detect anomalous performance or system faults; and (4) Best practices and recommendations for workflow evaluation, analysis, execution, and architectures.

97 MATHEMATICS AND COMPUTING↗

Contrastive Machine Learning with Gamma Spectroscopy Data Augmentations for Detecting Shielded Radiological Material Transfers

Data analysis techniques can be powerful tools for rapidly analyzing data and extracting information that can be used in a latent space for categorizing observations between classes of data. Machine learning models that exploit learned data relationships can address a variety of nuclear nonproliferation challenges like the detection and tracking of shielded radiological material transfers. The high resource cost of manually labeling radiation spectra is a hindrance to the rapid analysis of data collected from persistent monitoring and to the adoption of supervised machine learning methods that require large volumes of curated training data. Instead, contrastive self-supervised learning on unlabeled spectra can enhance models that are built on limited labeled radiation datasets. This work demonstrates that contrastive machine learning is an effective technique for leveraging unlabeled data in detecting and characterizing nuclear material transfers demonstrated on radiation measurements collected at an Oak Ridge National Laboratory testbed, where sodium iodide detectors measure gamma radiation emitted by material transfers between the High Flux Isotope Reactor and the Radiochemical Engineering Development Center. Label-invariant data augmentations tailored for gamma radiation detection physics are used on unlabeled spectra to contrastively train an encoder, learning a complex, embedded state space with self-supervision. A linear classifier is then trained on a limited set of labeled data to distinguish transfer spectra between byproducts and tracked nuclear material using representations from the contrastively trained encoder. The optimized hyperparameter model achieves a balanced accuracy score of 80.30%. Any given model—that is, a trained encoder and classifier—shows preferential treatment for specific subclasses of transfer types. Regardless of the classifier complexity, a supervised classifier using contrastively trained representations achieves higher accuracy than using spectra when trained and tested on limited labeled data.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Updates to the Alliance of Genome Resources central infrastructure

The Alliance of Genome Resources (Alliance) is an extensible coalition of knowledgebases focused on the genetics and genomics of intensively studied model organisms. The Alliance is organized as individual knowledge centers with strong connections to their research communities and a centralized software infrastructure, discussed here. Model organisms currently represented in the Alliance are budding yeast, Caenorhabditis elegans, Drosophila, zebrafish, frog, laboratory mouse, laboratory rat, and the Gene Ontology Consortium. The project is in a rapid development phase to harmonize knowledge, store it, analyze it, and present it to the community through a web portal, direct downloads, and application programming interfaces (APIs). Here, we focus on developments over the last 2 years. Specifically, we added and enhanced tools for browsing the genome (JBrowse), downloading sequences, mining complex data (AllianceMine), visualizing pathways, full-text searching of the literature (Textpresso), and sequence similarity searching (SequenceServer). We enhanced existing interactive data tables and added an interactive table of paralogs to complement our representation of orthology. To support individual model organism communities, we implemented species-specific “landing pages” and will add disease-specific portals soon; in addition, we support a common community forum implemented in Discourse software. We describe our progress toward a central persistent database to support curation, the data modeling that underpins harmonization, and progress toward a state-of-the-art literature curation system with integrated artificial intelligence and machine learning (AI/ML).

59 BASIC BIOLOGICAL SCIENCES↗

A Comprehensive Northern Hemisphere Particle Microphysics Data Set From the Precipitation Imaging Package

Microphysical observations of precipitating particles are critical data sources for numerical weather prediction models and remote sensing retrieval algorithms. However, obtaining coherent data sets of particle microphysics is challenging as they are often unindexed, distributed across disparate institutions, and have not undergone a uniform quality control process. This work introduces a unified, comprehensive Northern Hemisphere particle microphysical data set from the National Aeronautics and Space Administration precipitation imaging package (PIP), accessible in a standardized data format and stored in a centralized, public repository. Data is collected from 10 measurement sites spanning 34° latitude (37°N–71°N) over 10 years (2014–2023), which comprise a set of 1,070,000 precipitating minutes. The provided data set includes measurements of a suite of microphysical attributes for both rain and snow, including distributions of particle size, vertical velocity, and effective density, along with higher-order products including an approximation of volume-weighted equivalent particle densities, liquid equivalent snowfall, and rainfall rate estimates. The data underwent a rigorous standardization and quality assurance process to filter out erroneous observations to produce a self-describing, scalable, and achievable data set. Case study analyses demonstrate the capabilities of the data set in identifying physical processes like precipitation phase-changes at high temporal resolution. Bulk precipitation characteristics from a multi-site intercomparison also highlight distinct microphysical properties unique to each location. This curated PIP data set is a robust database of high-quality particle microphysical observations for constraining future precipitation retrieval algorithms, and offers new insights toward better understanding regional and seasonal differences in bulk precipitation characteristics.

54 ENVIRONMENTAL SCIENCES↗

Lowering the barrier to access information-rich transient kinetic data for machine learning methods

Transient kinetic data contain a wealth of information about intrinsic features of a catalyst as well as the reaction mechanism. Currently, high volume transient data is underutilized, and data science methods could both increase the value of information that can be extracted from this data, integrate experimental with theoretical data sources, and accelerate the pace of catalyst technology advancement. Transient kinetic characterizations with simple probe molecules exhibiting reversible adsorption, irreversible adsorption and bulk-surface diffusion are presented as training components for similar experiments with more complex surface reactions. In conclusion, by increasing the availability and accessibility of transient kinetic data through details of its structure and acquisition, we aim to decrease the barrier for data scientists to apply machine learning methods to this valuable data source.

Catalysis↗

Using National-Scale Data to Inform Coal Power Plant Redevelopment and Coal Community Revitalization Planning

In the United States, individual coal plants have high variation in their construction and operation. Coal plants commonly have multiple units, with different commissioning ages, operational position, and maintenance conditions – even multiple fuels. These units may be in changing states of utilization, with variation in productivity depending on economic contexts. Coal plants themselves may be temporarily offline, mothballed, retired, or decommissioned due to owner/operator portfolios or market conditions, and the current posture of the plant may be difficult to ascertain without local news sources (Tarekegne et al. 2021). Anticipated dates of plant closures are prospectively represented to regulators and to the Energy Information Administration (EIA), but may also experience uncertainty due to the engineering, environmental, economic and social complexity of closing a very large power plant. As noted in this report, closure of the first unit to the last within a single plant may span more than 20 years. Yet tracking and predicting the position of our nation’s coal fleets on a national scale is deeply important. In developing a case for federal datasets and an approach to building and curating these data, this report identified that data must have high fidelity at the plant-level to be meaningful in aggregate. Plant-scale detail is also a match for its application: national priorities in federal investment range from economic revitalization for coal plant communities to maintaining reliable electric grid operations. This document is organized into a statistical review of coal plant retirements, past and prospective, in the United States; trends in the relationship between closures and existing federal investment programs for community revitalization; and potential methods for accelerating site redevelopment through advanced geospatial analysis based on federal data.

01 COAL, LIGNITE, AND PEAT↗

Deep transfer learning for star cluster classification: I. application to the PHANGS– HST survey

ABSTRACT We present the results of a proof-of-concept experiment that demonstrates that deep learning can successfully be used for production-scale classification of compact star clusters detected in Hubble Space Telescope(HST) ultraviolet-optical imaging of nearby spiral galaxies ($D\lesssim 20\, \textrm{Mpc}$) in the Physics at High Angular Resolution in Nearby GalaxieS (PHANGS)–HST survey. Given the relatively small nature of existing, human-labelled star cluster samples, we transfer the knowledge of state-of-the-art neural network models for real-object recognition to classify star clusters candidates into four morphological classes. We perform a series of experiments to determine the dependence of classification performance on neural network architecture (ResNet18 and VGG19-BN), training data sets curated by either a single expert or three astronomers, and the size of the images used for training. We find that the overall classification accuracies are not significantly affected by these choices. The networks are used to classify star cluster candidates in the PHANGS–HST galaxy NGC 1559, which was not included in the training samples. The resulting prediction accuracies are 70 per cent, 40 per cent, 40–50 per cent, and 50–70 per cent for class 1, 2, 3 star clusters, and class 4 non-clusters, respectively. This performance is competitive with consistency achieved in previously published human and automated quantitative classification of star cluster candidate samples (70–80 per cent, 40–50 per cent, 40–50 per cent, and 60–70 per cent). The methods introduced herein lay the foundations to automate classification for star clusters at scale, and exhibit the need to prepare a standardized data set of human-labelled star cluster classifications, agreed upon by a full range of experts in the field, to further improve the performance of the networks introduced in this study.

Wei, Wei↗

Machine learning framework for predicting uranium enrichments from M400 CZT gamma spectra

A machine learning framework was developed for predicting uranium enrichments from M400 CZT gamma spectra. This framework leverages the availability of a large amount of measured M400 gamma spectra and uses a recently updated version of Gamma Detector Response and Analysis Software (GADRAS) for gamma spectrum analysis and generation. It also leverages the existing machine learning modules in Python for gamma spectrum data processing, curation, model training, benchmarking, and optimization of the deep machine learning models. The framework is used to develop a deep learning model to analyze gamma spectra from a set of U 3 O 8 samples with enrichments ranging from 0.31 to 93.17% and UF 6 cylinders with enrichments ranging from 0.2 to 4.95%, and the model performance is tested using a set of measured spectra and the respective declared enrichment values. Results show that the model can correctly classify 99.35% of the U 3 O 8 sample enrichments, and can predict the samples’ enrichments within an average absolute error of 0.099% (in percentage points of enrichment). For the UF 6 cylinders, the average absolute error was approximately 0.03%, with an accuracy of 98% in classifying discrete enrichment values of UF 6 samples. Finally, the results also show that the model has performed significantly better in terms of predicting enrichments in UF 6 cylinders based on measured gamma spectra than the GEM code, with a standard deviation (of the relative errors) of 2.23% (compared with the 11.51% value for the GEM code) based on results from a set of test data.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Frictionless knowledge injection for few-shot learning

Cutting-edge machine learning methods often require large volumes of curated training data, precluding their use in national security problems with rare events in massive datasets. We present a method for incorporating abstract knowledge into models tailored for sparse data. A subject matter expert defines salient concepts using data examples, which are encoded in the model’s embedding space. Models are then trained to respect these concepts. This method enables knowledge injection, yielding effective models with limited labeled data and the ability to assess model sensitivity for subject matter expertise across the nonproliferation mission space, as demonstrated with Raman spectra analysis.

Stomps, Jordan [ORNL] (ORCID:0000000178114479)↗

A data-centric weak supervised learning for highway traffic incident detection

Using the data from loop detector sensors for near-real-time detection of traffic incidents on highways is crucial to averting major traffic congestion. While recent supervised machine learning methods offer solutions to incident detection by leveraging human-labeled incident data, the false alarm rate is often too high to be used in practice. Specifically, the inconsistency in the human labeling of the incidents significantly affects the performance of supervised learning models. To that end, we focus on a data-centric approach to improve the accuracy and reduce the false alarm rate of traffic incident detection on highways. We develop a weak supervised learning workflow to generate high-quality training labels for the incident data without the ground truth labels, and we use those generated labels in the supervised learning setup for final detection. This approach comprises three stages. First, we introduce a data preprocessing and curation pipeline that processes traffic sensor data to generate high-quality training data through leveraging labeling functions, which can be domain knowledge-related or simple heuristic rules. Second, we evaluate the training data generated by weak supervision using three supervised learning models-random forest, k-nearest neighbors, and a support vector machine ensemble-and long short-term memory classifiers. The results show that the accuracy of all of the models improves significantly after using the training data generated by weak supervision. Third, we develop an online real-time incident detection approach that leverages the model ensemble and the uncertainty quantification while detecting incidents. Finally, we show that our proposed weak supervised learning workflow achieves a high incident detection rate (0.90) and low false alarm rate (0.08).

97 MATHEMATICS AND COMPUTING↗

Mapping use cases and dataset needs for benchmarking buildings data

A perennial challenge in buildings research is the lack of high-quality datasets that can be relied upon for a wide array of tasks, including model calibration and improving energy efficiency and load flexibility. Instrumenting a building for data collection is resource intensive, so it is important to be methodical in the approach and ensure that resulting data are flexible and useful for a broad range of analyses. This study aims to fill the gaps in characterizing potential use cases for buildings datasets and mapping them to dataset needs using a well-defined data infrastructure. Here, we have developed a systematic mapping strategy between buildings dataset needs and use cases to help streamline the processes of efficiently targeting datasets, designing building sensing systems, and determining buildings research use cases. We selected 14 prospective use cases and 11 refined buildings data categories for developing the preliminary dataset-needs-to-use-cases mapping matrix (‘DN-UC mapping matrix’) with generic ‘Tags’—a detailed sub-level of data categories extracted by justifying the needs of an aspect of the datasets to use cases. We present two example applications of the developed mapping matrix to demonstrate use of the mapping matrix and its effectiveness.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Spectroscopic features of dissolved iodine in pristine and gamma-irradiated nitric acid solutions

While iodine speciation is important for a wide range of nuclear activities, understanding the mechanisms of the transformations of iodine between chemical forms and the sensitivity of these transitions to solution conditions and exposure to radiation remains an active area of research. This work curates spectroscopic data from several experimental techniques and establishes their sensitivity and limitations in detecting changes in iodine speciation in both neutral and acidic regimes. The techniques include Raman spectroscopy, Fourier Transform Infrared (FTIR) spectroscopy, 127 I NMR spectroscopy, and ultraviolet-visible (UV-Vis) spectroscopy. Analysis of these data indicates that these commonly accessible spectroscopies often have dynamic ranges of measurable concentrations that do not always overlap between all techniques. The experimental techniques are disparately sensitive to iodide (I - ), molecular iodine (I 2 ), iodate (IO 3 - ), and periodate (IO 4 - ) species. Raman, FTIR, and NMR spectra were subsequentially analyzed using two-dimensional correlation analyses to generate high-resolution autocorrelation spectra. Here, the use of these spectroscopies is then extended to tracking acidification-induced and gamma irradiation-induced transformations of dissolved sodium iodate in deionized water and concentrated nitric acid. Both dissolution into nitric acid and irradiation with a gamma source are demonstrated to perturb the iodine speciation promoting their assembly into molecular iodine (I 2 ) and/or triiodide (I 3 - ). While I 2 and I 3 - species are undetectable with FTIR spectroscopy and 127 I NMR spectroscopy, the species can be detected with UV-Vis spectroscopy, and in some instances, I 3 - can be detected with Raman spectroscopy in the low wavenumber region. Ultimately, the results of this work provide a path to designing optimal combinations of techniques to detect forms of iodine across a wide range of concentrations and conditions.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

Cluster-Graph Fingerprinting: A Framework for Quantitative Analysis of Machine-Learned Interatomic Model Training and Simulation Data

Machine-learned interatomic models represent a significant advancement in simulation methods, extending the predictive ability of first-principles methods to previously inaccessible length and time scales. However, the data-driven nature of these models can lead to difficult-to-detect errors that can compromise prediction accuracy. To address this challenge, we introduce a novel fingerprinting approach based on the Chebyshev Interaction Model for Efficient Simulation (ChIMES) ML-IAM graph-based descriptor. Our strategy enables efficient and statistically rigorous analysis of system configurations used in ML-IAM training and those generated by their application, e.g., in molecular dynamics simulations. We demonstrate that these fingerprints can effectively assess novelty of a configuration relative to an existing data set and determine dissimilarity among individual configurations, which are two key tasks in workflows for active learning-based ML-IAM training, data set curation, and on-the-fly uncertainty quantification.

36 MATERIALS SCIENCE↗

Machine Learning Prediction of the Experimental Transition Temperature of Fe(II) Spin-Crossover Complexes

Spin-crossover (SCO) complexes are materials that exhibit changes in the spin state in response to external stimuli, with potential applications in molecular electronics. It is challenging to know a priori how to design ligands to achieve the delicate balance of entropic and enthalpic contributions needed to tailor a transition temperature close to room temperature. Here, we leverage the SCO complexes from the previously curated SCO-95 data set [Vennelakanti et al. J. Chem. Phys. 159, 024120 (2023)] to train three machine learning (ML) models for transition temperature (T 1/2 ) prediction using graph-based revised autocorrelations as features. We perform feature selection using random forest-ranked recursive feature addition (RF-RFA) to identify the features essential to model transferability. Of the ML models considered, the full feature set RF and recursive feature addition RF models perform best, achieving moderate correlation to experimental T 1/2 values. We then compare ML T 1/2 predictions to those from three previously identified best-performing density functional approximations (DFAs) which accurately predict SCO behavior across SCO-95, finding that the ML models predict T 1/2 more accurately than the best-performing DFAs. In addition, we study ML model predictions for a set of 18 SCO complexes for which only estimated T 1/2 values are available. Upon excluding outliers from this set, the RF-RFA RF model shows a strong correlation to estimated T 1/2 values with a Pearson’s r of 0.82. In contrast, DFA-predicted T 1/2 values have large errors and show no correlation to estimated T 1/2 values over the same set of complexes. Overall, our study demonstrates slightly superior performance of ML models in comparison with some of the best-performing DFAs, and we expect ML models to improve further as larger data sets of SCO complexes are curated and become available for model training.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Changes in Four Decades of Near‐CONUS Tropical Cyclones in an Ensemble of 12 km Thermodynamic Global Warming Simulations

We evaluate tropical cyclones (TCs) in a set of thermodynamic global warming (TGW) simulations over the continental United States (CONUS). A 12 km simulation forced by ERA5 provides a 40‐year historical (1980–2019) control. Four complimentary future scenarios are generated using thermodynamic deltas applied to lateral boundary, interior, and surface forcing. We curate a data set of 4,498 6‐hourly TC snapshots in the control and find a corresponding “twin” in each counterfactual, permitting a paired comparison. Warming results in an increase in mean dynamical TC intensity and moisture‐related quantities, with the latter being more pronounced. TC inner cores contract slightly but outer storm size remains unchanged. The frequency with which TCs become more intense is only moderately consistent, with snapshots having increased hazards ranging from 50% to 80% depending on warming level. The fractions of TCs undergoing rapid intensification and weakening both increase across all warming simulations, suggesting elevated short‐term intensity variability.

54 ENVIRONMENTAL SCIENCES↗

Edge AI-Enhanced Traffic Monitoring and Anomaly Detection Using Multimodal Large Language Models

This paper addresses the challenge of traffic monitoring and incident detection in remote areas, utilizing multimodal large language models (LLMs) deployed on edge AI devices. The key novelty of the LLM is to convert real-time video streams into descriptive texts, enabling low-bandwidth transmissions and reliable detection of anomalies and incidents in environments of intermittent connectivity. The model is developed based on fine-tuning open-source LLMs and extending it with multi-modal capabilities to analyze video frames. Our work also involves deploying this model on edge devices such as Nvidia IGX Orin and is planned to be tested in realistic environments in future work. The methodology includes data set curation, iterative model fine-tuning and compression, and hardware-based optimization. This approach aims to enhance traffic safety and response speed in remote areas, marking a significant advancement in the application of AI for traffic monitoring and safety management.

Peruski, Ryan [University of Tennessee, Knoxville ↗