Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning and data science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Enriching the Twitter Stream Increasing Data Mining Yield and Quality Using Machine Learning

Social media data streams are important sources of real-time and historical global information for science applications. At the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC), we are exploring the Twitter data stream for its potential in augmenting the validation program of NASA Earth science missions, specifically the Global Precipitation Measurement (GPM) mission. We have implemented a tweet processing infrastructure that outputs classified precipitation tweets. Inputs are "passive" tweets, along with a smaller number of tweets from "active" participants, i.e., those knowingly contributing to our effort. The "active" tweets, presumably of higher quality, enrich the Twitter stream. "Active" sources include data scraped from other social media (e.g., public Facebook posts) and data from existing crowdsourcing programs (e.g., mPING reports). In addition, there is likely relevant precipitation information in images and documents that are the end points of links often included in tweets. Information derived from these "active" sources could then be tweeted into the Twitter stream, thus enriching its quality. The objective of our current work is to mine these tweet­ linked images and documents, using neural networks, to increase the information content and quality related to precipitation. For images, we classified them as either precipitation-related or not. For training and validation, we used images obtained via the Google custom search API. We created two models: (1) by training a simple Convolutional Neural Network and (2) by using transfer learning principles to adapt a pre-trained object recognition model. For documents, both those linked to tweets and the tweet contents, we trained Hierarchical Attention Networks to determine precipitation occurrence, type, and intensity. For training and validation, we used a keyword-filtered tweet data set labelled with ground truth data from Dark Sky (an API to retrieve weather-related labels) and the National Severe Storms Laboratory's Multi­ Radar/Multi-Sensor (MRMS) system. Our results demonstrated the efficacy of our machine learning approaches for enriching the Twitter stream, to derive information potentially useful for validation of earth science satellite data.

Albayrak, Arif↗

Uncovering electronic and geometric descriptors of chemical activity for metal alloys and oxides using unsupervised machine learning

Here, we show that unsupervised machine learning (ML) using principal component analysis (PCA) provides a straightforward pathway for developing accurate and interpretable electronic-structure descriptors of the chemical and catalytic properties of materials. We demonstrate the approach by finding chemisorption descriptors for metal alloys and surface oxygens on metals and metal oxides. In both cases, the principal component (PC) descriptors yield ML models that predict the material’s chemical properties with competitive accuracy compared to ML models built using established descriptors. Importantly, interpreting the electronic-structure patterns captured by each PC descriptor via signal reconstruction suggests potential design motifs for future electronic-structure descriptor design and allows us to identify links between a material’s geometric and catalytic properties. Ultimately, we show that the unsupervised ML approach provides a route to find electronic-structure descriptors of the catalytic properties of materials that readily connect to geometric structure and composition.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Learning continuous models for continuous physics

Abstract Dynamical systems that evolve continuously over time are ubiquitous throughout science and engineering. Machine learning (ML) provides data-driven approaches to model and predict the dynamics of such systems. A core issue with this approach is that ML models are typically trained on discrete data, using ML methodologies that are not aware of underlying continuity properties. This results in models that often do not capture any underlying continuous dynamics—either of the system of interest, or indeed of any related system. To address this challenge, we develop a convergence test based on numerical analysis theory. Our test verifies whether a model has learned a function that accurately approximates an underlying continuous dynamics. Models that fail this test fail to capture relevant dynamics, rendering them of limited utility for many scientific prediction tasks; while models that pass this test enable both better interpolation and better extrapolation in multiple ways. Our results illustrate how principled numerical analysis methods can be coupled with existing ML training/testing methodologies to validate models for science and engineering applications.

97 MATHEMATICS AND COMPUTING↗

Evaluating the factors influencing accuracy, interpretability, and reproducibility in the use of machine learning classifiers in biology to enable standardization

The complexity and variability of biological data has promoted the increased use of machine learning methods to understand processes and predict outcomes. These same features complicate reliable, reproducible, interpretable, and responsible use of such methods, resulting in questionable relevance of the derived. outcomes. Here we systematically explore challenges associated with applying machine learning to predict and understand biological processes using a well- characterized in vitro experimental system. We evaluated factors that vary while applying machine learning classifers: (1) type of biochemical signature (transcripts vs. proteins), (2) data curation methods (pre- and post-processing), and (3) choice of machine learning classifier. Using accuracy, generalizability, interpretability, and reproducibility as metrics, we found that the above factors significantly mod- ulate outcomes even within a simple model system. Our results caution against the unregulated use of machine learning methods in the biological sciences, and strongly advocate the need for data standards and validation tool-kits for such studies.

59 BASIC BIOLOGICAL SCIENCES↗

Mapping causal patterns in crystalline solids

The evolution of the atomic structures of the combinatorial library of Sm-substituted thin film BiFeO 3 along the phase transition boundary from the ferroelectric rhombohedral phase to the non-ferroelectric orthorhombic phase is explored using scanning transmission electron microscopy. Localized properties, including polarization, lattice parameter, and chemical composition, are parameterized from atomic-scale imaging, and their causal relationships are reconstructed using a linear non-Gaussian acyclic model. This approach is further extended to explore the spatial variability of the causal coupling using the sliding window transform method, which revealed that new causal relationships emerged at both the expected locations, such as domain walls and interfaces, and at additional regions forming clusters in the vicinity of the walls or spatially distributed features. While the exact physical origins of these relationships are unclear, they likely represent nanophase-separated regions in the morphotropic phase boundaries. Overall, we posit that an in-depth understanding of complex disordered materials away from thermodynamic equilibrium necessitates understanding not only the generative processes that can lead to observed microscopic states but also the causal links between multiple interacting subsystems.

Causal inference↗

Department of Energy’s Atmospheric System Research (ASR) Program’s Workshop on the Future of Atmospheric Large Eddy Simulation (LES): Workshop Report

Large-eddy simulation (LES) is used as a tool to understand physical processes such as turbulence, aerosols, clouds, precipitation, radiation, the interactions among all these, and their interactions with the underlying surface. Over the next 10 years, LES will drive fundamental progress in open scientific questions in these areas as LES is increasingly used to gain understanding of complex interacting physical processes involving atmospheric turbulence. This growth will be driven both by scientific demand and the expansion of computational resources needed to conduct LES, and the form that the growth takes will largely be determined by how computational resources are leveraged for scientific gain. In particular, we suggest that computational resources are likely to be leveraged in two separate but not necessarily distinct ways. On one hand, growth in computational resources will allow LES to be made more routine, that is, performed more frequently, while on the other hand, the computational expense (measured in total floating point operations) afforded to individual LES will expand dramatically, allowing simulations to increase in both domain size and resolution as well as physical detail. Current U.S. Department of Energy (DOE) projects such as LES ARM Symbiotic Simulation and Observation Activity (LASSO) are leading the way in conducting routine LES, building large, public databases that are accessible for data science, sensitivity studies, and training for machine learning. LES will also become more routine as it becomes more accessible for individual researchers to address their scientific questions of interest. Scientific questions addressed by LES over the next 10 years are likely to include cloud organization and aggregation; aerosol cloud interactions and atmospheric chemistry (including geo-engineering); urban-scale LES; atmospheric extreme events, ranging from small-scale severe weather to wildfires; and ocean-wave-atmosphere interactions. Further LES-related research will likely grow significantly in areas related to societal impact studies of air quality and extreme weather events, applications to renewable energy forecasting and resource assessment, and aid in decision-making processes. The growth in the use of LES in atmospheric science research will drive the need for better physical process representations (e.g., cloud aerosol microphysics, radiation, and atmospheric chemistry) at the scales resolved by LES. To date, many of the process representations used by LES have been taken directly from coarser-resolution models. Promising methods for LES process representations include superdroplet and quadrature methods for microphysics, 3D approaches for radiation, and better representation of chemistry and aerosol processes. At LES resolution, land-atmosphere interactions for complex terrains, land cover/types, biogeochemistry, and plant canopy models are needed as an improvement beyond traditional and widely used Monin-Obuhkov similarity theory.

54 ENVIRONMENTAL SCIENCES↗

Department of Energy’s Atmospheric System Research (ASR) Program’s Workshop on the Future of Atmospheric Large Eddy Simulation (LES) (Workshop Report)

Large-eddy simulation (LES) is used as a tool to understand physical processes such as turbulence, aerosols, clouds, precipitation, radiation, the interactions among all these, and their interactions with the underlying surface. Over the next 10 years, LES will drive fundamental progress in open scientific questions in these areas as LES is increasingly used to gain understanding of complex interacting physical processes involving atmospheric turbulence. This growth will be driven both by scientific demand and the expansion of computational resources needed to conduct LES, and the form that the growth takes will largely be determined by how computational resources are leveraged for scientific gain. In particular, we suggest that computational resources are likely to be leveraged in two separate but not necessarily distinct ways. On one hand, growth in computational resources will allow LES to be made more routine, that is, performed more frequently, while on the other hand, the computational expense (measured in total floating point operations) afforded to individual LES will expand dramatically, allowing simulations to increase in both domain size and resolution as well as physical detail. Current U.S. Department of Energy (DOE) projects such as LES ARM Symbiotic Simulation and Observation Activity (LASSO) are leading the way in conducting routine LES, building large, public databases that are accessible for data science, sensitivity studies, and training for machine learning. LES will also become more routine as it becomes more accessible for individual researchers to address their scientific questions of interest. Scientific questions addressed by LES over the next 10 years are likely to include cloud organization and aggregation; aerosol cloud interactions and atmospheric chemistry (including geo-engineering); urban-scale LES; atmospheric extreme events, ranging from small-scale severe weather to wildfires; and ocean-wave-atmosphere interactions. Further LES-related research will likely grow significantly in areas related to societal impact studies of air quality and extreme weather events, applications to renewable energy forecasting and resource assessment, and aid in decision-making processes.

54 ENVIRONMENTAL SCIENCES↗

Advanced Offshore Hazard Forecasting to Enable Resilient Offshore Operations

Paper prepared for the Offshore Technology Conference, 2024. Hazards in the offshore environment can imperil successful energy operations, whether those operations are conventional, renewable, or for decarbonization. The expanding accessibility of data science and the advanced applications of machine learning (ML) models creates an opportunity to assess potential hazards and the infrastructure they impact. We present a use case demonstrating the combined application of published ML tools to U.S. federal waters of the Gulf of Mexico, an actively explored region for offshore energy that is affected by variable metocean conditions and geologic processes contributing to potential hazards.

Mark-Moser, Mackenzie K.↗

Rapid data acquisition and machine learning-assisted composition design of functionally graded alloys via wire arc additive manufacturing

Abstract The lack of high-quality datasets in materials science hinders artificial intelligence (AI)-driven alloy design. To address this challenge, wire arc additive manufacturing (WAAM) was employed to fabricate graded alloys, generating extensive data for machine learning (ML)-assisted property prediction. ML models were developed using high-throughput experiments, computational models, and genetic algorithm to optimize feature selection, successfully predicting hardness and porosity. The ML model demonstrated its efficacy by designing a gradient alloy with enhanced properties. However, scaling up revealed uncertainties in tensile property and porosity due to differences in size and thermal conditions between the designed alloy build and the gradient print used to construct the ML model. This underscores the need for uncertainty quantification and process optimization in WAAM-driven alloy design. Our work advances AI-integrated additive manufacturing, offering a rapid approach to exploring process–structure–property relationships and accelerating materials development.

Wang, Xin↗

Combustion machine learning: Principles, progress and prospects

Progress in combustion science and engineering has led to the generation of large amounts of data from large-scale simulations, high-resolution experiments, and sensors. This corpus of data offers enormous opportunities for extracting new knowledge and insights—if harnessed effectively. Machine learning (ML) techniques have demonstrated remarkable success in data analytics, thus offering a new paradigm for data-intense analyses and scientific investigations through combustion machine learning (CombML). While data-driven methods are utilized in various combustion areas, recent advances in algorithmic developments, the accessibility of open-source software libraries, the availability of computational resources, and the abundance of data have together rendered ML techniques ubiquitous in scientific analysis and engineering. This article examines ML techniques for applications in combustion science and engineering. Starting with a review of sources of data, data-driven techniques, and concepts, we examine supervised, unsupervised, and semi-supervised ML methods. Various combustion examples are considered to illustrate and to evaluate these methods. Next, we review past and recent applications of ML approaches to problems in combustion, spanning fundamental combustion investigations, propulsion and energy-conversion systems, and fire and explosion hazards. Challenges unique to CombML are discussed and further opportunities are identified, focusing on interpretability, uncertainty quantification, robustness, consistency, creation and curation of benchmark data, and the augmentation of ML methods with prior combustion-domain knowledge.

33 ADVANCED PROPULSION SYSTEMS↗

Interpretable Machine Learning for Molecular Biosignatures: a Novel Single-Sample Feature Importance Method That Is Sensitive To Statistical Interactions

Isotope ratio mass spectrometry (IRMS) of volatiles (e.g., CO 2 ) promises to be a powerful tool for potential biosignature detection for future missions to ocean worlds (OW) such as Europa and Enceladus. Machine learning (ML) methods for IRMS data could enable science autonomy by onboard prediction of seawater chemistry and biosignature presence. However, ML models are likely to be complex and involve statistical interactions between features (variables), which can make predictions seem opaque and enigmatic. For ML predictions as significant as extraterrestrial biosignatures, we must place extraordinary confidence in models. It is therefore essential that these models make interpretable predictions (i.e., human-understandable) and include false-prediction diagnostics. We achieve high accuracy and interpretability in ML biosignature and seawater chemistry models for OW through a nearest-neighbors feature selection tool that detects statistical interactions between predictors, constructs interaction networks for visualization of selected features working together to make a prediction, and reports single-sample feature importance scores for false-detection diagnostics. Here we develop a novel single-sample nearest-neighbors projected distance regression(ssNPDR) feature selection method that improves upon existing single-sample algorithms through the inclusion of statistical interactions while providing false-prediction diagnostics for ML models.

geochemistry↗

INCREASING THE TRANSPARENCY AND REPRODUCIBILITY OF SPACE RADIATION SCIENCE: THE RADIATION BIOLOGY ONTOLOGY

Among the primary objectives of the Open/Open-Source Science paradigm are making scientific investigation data transparent and results reproducible [1], objectives shared by the FAIR principles [2]. To accomplish this, the conceptual framework that includes all the investigation objects needs to be accurately captured and communicated to all data consumers. A large part of this requires using metadata standards to annotate data collected. These standards should be readily accessible, informed by scientific community consensus and sufficiently specific to encompass all of the important aspects of the investigation. Starting in 2020 we have been co-leading an open consortium to develop a new metadata standard, the Radiation Biology Ontology (RBO), through the Open Biological and Biomedical Ontologies (OBO) Foundry [3]. We began by transforming many of the terms from the National Council on Radiation Protection and Measurement into concepts that can be formally related to existing OBO Foundry classes or attributes. We then identified and imported into the RBO existing OBO Foundry classes that have obvious relevance for radiation biomedicine (for example, concepts from the Environment Ontology that describe radiative processes, and concepts from the Gene Ontology dealing with molecular and cellular responses to radiation). Finally, we scrutinized datasets from investigations of radiation effects held in NASA GeneLab and LSDA repositories and added additional classes, instances, and attributes into the RBO that should be used to annotate these data. We developed the RBO using the open-source tools of GitHub and publish the RBO periodically through the NIH/NCBI BioPortal website, so systems worldwide can leverage the knowledge it contains [4]. This initial phase of concept modeling has yielded an RBO that at present has more than 300 declared concepts, with more than 3500 additional concepts imported from other OBO Foundry ontologies. While this first phase has focused on concepts for annotating samples, environments, exposures, and measurements, the next phase will center on supporting annotation of results and findings, such as concept models of molecular, cellular and tissue effects. The value of the RBO will be determined in part by our ability to engage the community in its development, and we have established a Radiobiology Informatics Consortium with unrestricted membership as the owner of the RBO in order to encourage investigators, system owners and other to join in this effort. Anyone can report issues or request new concept modeling or other features directly on GitHub. By using the BioPortal application programming interface, systems can pose dynamic queries to the latest version of the RBO for information on individual classes or entire hierarchies; this design eliminates the need for systems to be updated in order to use newer versions of the RBO. We hope to contribute to the advancement of open radiobiological science through the continued, open development of the RBO, that will provide more precise, machine-interpretable descriptions of investigations, as well as support data meta-analysis through machine learning or other artificial intelligence methods. REFERENCES [1] Open science in space. Nature Medicine, 2021. 27(9): p. 1485-1485. [2] Wilkinson, M.D., et al., The FAIR Guiding Principles for scientific data management and stewardship. Sci Data, 2016. 3: p. 160018. [3] Smith, B., et al., The OBO Foundry: coordinated evolution of ontologies to support biomedical data integration. Nat Biotechnol, 2007. 25(11): p. 1251-5. [4] Whetzel, P.L., et al., BioPortal: enhanced functionality via new Web services from the National Center for Biomedical Ontology to access and use ontologies in software applications. Nucleic Acids Res, 2011. 39(Web Server issue): p. W541-5.

informatics↗

2020 ETI Annual Summer School: Data Science and Engineering

The Consortium for Enabling Technologies & Innovation (ETI) was established in 2019 to address emerging technologies within the context of nuclear nonproliferation. ETI creates a research and education environment to support cross-cutting technologies across three core disciplines: 1) computer and engineering science research specifically in a form of machine learning and high performance computing (HPC), 2) advanced manufacturing, and 3) nuclear detection technologies. For outreach and development, ETI hosted the first of three summer schools from August 24-28, 2020 with the theme of “Data Science and Engineering”. The school was hosted in an on-line format and had over 200 participants. The recorded content is available on-line as a resource for students. The summer school had four modules: 1) Fundamentals of data Applications, 2) Computational Machine Learning, 3) Bayesian Modeling and Inference, and 4) Data Science for Safeguards. Modules contained both lectures as well as student exercises. Poll Everywhere was utilized in some modules as an on-line method to engage large groups of students. Upcoming ETI Summer Schools include Novel Instrumentation in 2021 and Advanced Manufacturing in 2022.

Biegalski, Steven R.↗

Sub-pilot-scale Production of High-Value Products from U.S. Coals

Investigators from the University of Utah, University of Wyoming and Marshall University pursued a program to study the conversion of raw coal to high-value products of carbon fiber and silicon carbide. Team members also developed an initial framework for a data portal that can incorporate laboratory data on coal processing and product quality, and also work with tools for machine learning for data analysis, data visualization and economic assessment. Experimental R&D efforts focused on the conversion of raw coal to coal tar and other byproducts, and the resulting tar intermediates were upgraded to form anisotropic and isotropic pitch materials. These pitch materials were produced from coal using both thermal (pyrolysis) and chemical (mild solvolysis liquefaction) decomposition of raw coal. Four different coals were studied: Utah bituminous coal (Sufco), Wyoming PRB coal (Black Thunder), Illinois bituminous coal (Illinois #6), and West Virginia bituminous coal (Flying Eagle). Both metallurgical-grade coking coals and lower-grade steam coals were investigated, and controlled secondary gas-phase reactions were used during a two-stage pyrolysis process to induce cracking and condensation reactions among the pyrolytic tar species. This approach successfully improved the performance of the lower grade coals for yielding pitch materials, with properties more consistent with a commercial-grade pitch that had previously demonstrated success for quality carbon fiber production. The use of waste plastic materials was also studied, to help improve physical and chemical characteristics of the intermediate tars and final pitch product; in particular, for lowering the pitch softening point to an acceptable level for melt spinning carbon fiber. Mild solvolysis liquefaction was also used as a method for producing pitch for carbon fiber production. As expected, significantly higher pitch yields were obtained using this approach, and waste plastic materials were also successfully used to reduce pitch softening point to an acceptable level. The plastic materials were also utilized to create a solvent for the mild solvolysis process, and this plastic-derived solvent was shown to provide results consistent with more expensive commercial chemical solvents, and could thus avoid the need for costly recovery and recycle of a liquefaction solvent. Additional experimental R&D focused on the production of silicon carbide (β-SiC) from the residual char byproduct from pitch production, and also on the production of carbon fiber from the anisotropic pitch. SiC was successfully synthesized using a mixture of residual char and sandstone at a ratio of 1:1. Reaction temperature and residence time were optimized and yielded a product purity of 81%. For carbon fiber production, the most successful pitch samples were obtained from the mild solvolysis liquefaction approach, combined with the use of a plastic (HDPE)-derived solvent. Fiber properties improved over time as laboratory fiber production methodologies improved, and final yields of carbon fiber were obtained with a diameter of 12.14 ± 1.10 um, Modulus of 173.73 ± 15.25 GPa, and Tensile Strength of 1.04 ± 0.10 GPa. A proof-of-concept Modern Community Research Data Portal (MCRDP) was developed and deployed for coal and coal-derived pitch characterization, with the full support of (i) remote web-based access, (ii) distributed analysis, (iii) interactive visualization and exploration, (iv) shared and long-term data access, (v) advanced query capabilities and (vi) real-time collaboration. The Coal to Products Data Portal “coaltoproducts.org” provides researchers with space to store and share data within a project, tools for analyzing and understanding data for scientific investigation, and the ability to publish data to the broader community for reproducibility. The portal leverages the Material Commons 2.0 (MC) platform developed by the Center for PRedictive Integrated Structural Materials Science (PRISMS) of the University of Michigan, to achieve long-term longevity of data collections and, more importantly, collaborative science. A number of data visualization tools were also assessed and implemented for interrogating the experimental and modeling data. The machine learning portion of this project analyzed datasets from two different coal conversion processes performed on a diverse set of coal samples from both the coal pyrolysis experiments and the solvent liquefaction experiments. The work was initiated by exploring standard regression models on the pyrolysis data, aiming to understand the impact of sample characteristics and processing conditions on key product metrics. Over the course of the project, the focus expanded to include a variety of machine learning tools, delving into both supervised and unsupervised learning methods. Models tested on the pyrolysis data included linear, ridge, lasso, elastic-net, Gaussian process, random forest regression, and AutoSklearn, and the approach was continually refined to enhance predictive accuracy and model interpretability. Similar techniques were applied to the liquefaction data with an additional focus on feature engineering. Along with mesophase content, additional outputs of interest were the pitch yield, softening point, and QI content. Insights derived from these analyses are crucial in determining the factors influencing the quality and yield of coal-derived products. As the work progressed, the research evolved from foundational model comparisons to analyses of random forests, decision paths, and feature importance scores. A thorough market analysis was performed to examine the prospects of coal-based carbon fibers. The best opportunities for coal come from its lower and more stable price relative to petroleum, particularly for subbituminous coals, which is the primary advantage that a coal refinery may have over a petroleum refinery. Before a commercial CTP production facility can be modeled, however, several things need to be understood regarding the nature of the would-be coal refinery. These include the technology to be deployed, the size of facility, the volume(s) of co-product(s), and the waste and emissions profile of the plant. The volume of co-products and waste may be substantial and will require separate market analysis to ensure viability. In the near-term, the importance of coal tar pitch, in the form of carbon pitch, to the aluminum and steel industries is likely to overshadow the alternative use of this material as an input for carbon fiber. The importance of steel and aluminum in building materials, and the need for carbon materials in their manufacturing, will ensure that demand for these products remains for the long run. In addition, carbon fiber may also be the best substitute for steel and aluminum well into the future. While society will eventually be able to shift production of much of its electricity needs to renewables, it will not be able to shift away from fossil fuels for production of high-strength construction and vehicular materials. Demand for carbon fiber is expected to increase quickly, but the volume of carbon fiber and the amount of coal that would be needed to produce even a sizeable share of this market may still be relatively small compared to current coal production. Thus, other coal-based products like graphene, graphite, carbon foams, resins, and carbon-based building products will play important roles in sustaining coal production as coal-fired power generation continues to decline.

01 COAL, LIGNITE, AND PEAT↗