Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning tools”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Sandtank-ML: An Educational Tool at the Interface of Hydrology and Machine Learning

Hydrologists and water managers increasingly face challenges associated with extreme climatic events. At the same time, historic datasets for modeling contemporary and future hydrologic conditions are increasingly inadequate. Machine learning is one promising technological tool for navigating the challenges of understanding and managing contemporary hydrological systems. However, in addition to the technical challenges associated with effectively leveraging ML for understanding subsurface hydrological processes, practitioner skepticism and hesitancy surrounding ML presents a significant barrier to adoption of ML technologies among practitioners. In this paper, we discuss an educational application we have developed—Sandtank-ML—to be used as a training and educational tool aimed at building user confidence and supporting adoption of ML technologies among water managers. We argue that supporting the adoption of ML methods and technologies for subsurface hydrological investigations and management requires not only the development of robust technologic tools and approaches, but educational strategies and tools capable of building confidence among diverse users.

54 ENVIRONMENTAL SCIENCES↗

From PINNs to PIKANs: recent advances in physics-informed machine learning

Physics-Informed Neural Networks (PINNs) have emerged as a key tool in Scientific Machine Learning since their introduction in 2017, enabling the efficient solution of ordinary and partial differential equations using sparse measurements. Over the past few years, significant advancements have been made in the training and optimization of PINNs, covering aspects such as network architectures, adaptive refinement, domain decomposition, and the use of adaptive weights and activation functions. A notable recent development is the Physics-Informed Kolmogorov-Arnold Networks (PIKANS), which leverage a representation model originally proposed by Kolmogorov in 1957, offering a promising alternative to traditional PINNs. In this review, we provide a comprehensive overview of the latest advancements in PINNs, focusing on improvements in network design, feature expansion, optimization techniques, uncertainty quantification, and theoretical insights. We also survey key applications across a range of fields, including biomedicine, fluid and solid mechanics, geophysics, dynamical systems, heat transfer, chemical engineering, and beyond. Lastly, we review computational frameworks and software tools developed by both academia and industry to support PINN research and applications.

Kolmogorov-Arnold networks↗

KG-Hub—building and exchanging biological knowledge graphs

Knowledge graphs (KGs) are a powerful approach for integrating heterogeneous data and making inferences in biology and many other domains, but a coherent solution for constructing, exchanging, and facilitating the downstream use of KGs is lacking. Here we present KG-Hub, a platform that enables standardized construction, exchange, and reuse of KGs. Features include a simple, modular extract–transform–load pattern for producing graphs compliant with Biolink Model (a high-level data model for standardizing biological data), easy integration of any OBO (Open Biological and Biomedical Ontologies) ontology, cached downloads of upstream data sources, versioned and automatically updated builds with stable URLs, web-browsable storage of KG artifacts on cloud infrastructure, and easy reuse of transformed subgraphs across projects. Current KG-Hub projects span use cases including COVID-19 research, drug repurposing, microbial–environmental interactions, and rare disease research. KG-Hub is equipped with tooling to easily analyze and manipulate KGs. KG-Hub is also tightly integrated with graph machine learning (ML) tools which allow automated graph ML, including node embeddings and training of models for link prediction and node classification.

59 BASIC BIOLOGICAL SCIENCES↗

Machine Learning Based AFP Inspection: A Tool for Characterization and Integration

Automated Fiber Placement (AFP) has become a standard manufacturing technique in the creation of large scale composite structures due to its high production rates. However, the associated rapid layup that accompanies AFP manufacturing has a tendency to induce defects. We forward an inspection system that utilizes machine learning (ML) algorithms to locate and characterize defects from profilometry scans coupled with a data storage system and a user interface (UI) that allows for informed manufacturing. A Keyence LJ-7080 blue light profilometer is used for fast 2D height profiling. After scans are collected, they are process by ML algorithms, displayed to an operator through the UI, and stored in a database. The overall goal of the inspection system is to add an additional tool for AFP manufacturing. Traditional AFP inspection is done manually adding to manufacturing time and being subject to inspector errors or fatigue. For large parts, the inspection process can be cumbersome. The proposed inspection system has the capability of accelerating this process while still keeping a human inspector integrated and in control. This allows for the rapid capability of the automated inspection software and the robustness of a human checking for defects that the system either missed or misclassified.

Sacco, Christopher↗

Computationally evaluating high-yield metabolites for sustainable aviation fuel (SAF) using machine learning

The computational tool described in this report helps identify promising biological pathways that produce SAF platform molecules (either a drop-in SAF, or a precursor that can be easily converted to a drop-in SAF). The workflow the computational tool follows first identifies possible biological pathways from a user-defined metabolite. These pathways may, or may not lead to a SAF platform molecule, thus the second step involves insilico testing of the end product of each pathway to assess whether it is, or is not, a SAF platform molecule. The identification of biological pathways performed in the first step is facilitated by linking the metabolite to a biological reaction database. Pathways are found by identifying pathways in the reaction database that include the metabolite. The computational tool includes an alternative way to find pathways. The alternative way develops a Flux Balanced Analysis (FBA), and modifying the FBA to include reactions that transform the metabolite. These modifications serve as a basis for understanding, in a semi-quantitative way, if there is an increase in the flux to desirable products. The second step, in silico testing of the end-products, is accomplished by estimating key physical properties relevant to SAF. When good models are available, we have integrated those models into the computational tool. In a few instances, we have developed our own models. In all instances, we have validated the models against available measured data. Finally, we have evaluated the effectiveness of our computational tool by genetically engineering Rhodosporidium toruloides. Validation occurred without the use of a FBA, and further validation is required.

09 BIOMASS FUELS↗

A New Modeling Framework for Geothermal Operational Optimization with Machine Learning (GOOML)

Geothermal power plants are excellent resources for providing low carbon electricity generation with high reliability. However, many geothermal power plants could realize significant improvements in operational efficiency from the application of improved modeling software. Increased integration of digital twins into geothermal operations will not only enable engineers to better understand the complex interplay of components in larger systems but will also enable enhanced exploration of the operational space with the recent advances in artificial intelligence (AI) and machine learning (ML) tools. Such innovations in geothermal operational analysis have been deterred by several challenges, most notably, the challenge in applying idealized thermodynamic models to imperfect as-built systems with constant degradation of nominal performance. This paper presents GOOML: a new framework for Geothermal Operational Optimization with Machine Learning. By taking a hybrid data-driven thermodynamics approach, GOOML is able to accurately model the real-world performance characteristics of as-built geothermal systems. Further, GOOML can be readily integrated into the larger AI and ML ecosystem for true state-of-the-art optimization. This modeling framework has already been applied to several geothermal power plants and has provided reasonably accurate results in all cases. Therefore, we expect that the GOOML framework can be applied to any geothermal power plant around the world.

15 GEOTHERMAL ENERGY↗

Constituent Data Replacement Tool

The purpose of this tool is to estimate key parameters that may be missing in public wastewater composition datasets. The tool can be applied to develop complete treatment and critical mineral extraction profiles for leachate, produced water and other aqueous waste streams. The tool applies machine learning algorithms to replace missing data in a user’s water data set that are adjusted based on user preferences for options including algorithm type, number of features, and classification variables. The tool can use the user’s data alone or combine user data with the NEWTS USGS Produced Water Database for more robust training. This research was funded by the U.S. Department of Energy’s Office Fossil Energy and Carbon Management (FECM) through National Energy Technology Laboratory’s ongoing research under the Water Management for Power System Field Work Proposal, DE-FECM 1022428 and Critical Minerals Field Work Proposal, DE-FECM 1022420.

Aqueous Chemistry↗

Robust Machine Learning

UQ4ML is a code repository for a set of tools for the development of robust machine learning methods, uncertainty quantification and explainability of machine learning methods. The goal of these tools is to develop more robust and statistically rigorous machine learning methods for scientific applications. These tools are developed in Python, a high-level programming language that takes advantage of the Python ecosystem of high-quality open-source packages for machine learning.

Oyen, Diane↗

Predicting Fiber Failure of Plain Weave Fabric with Recursive Multiscale Micromechanics

Recent advances in the development of machine learning (ML) algorithms have enabled the creation of predictive models that can improve decision making, decrease computational cost, and improve efficiency in a variety of fields. As an organization begins to develop and implement such models, the data used in the training, validation, and testing of machine learning models, the model parameters, and the use cases or limitations of the models must be properly stored to ensure models are both fully traceable and used correctly. In the context of predicting material behavior, advances in computationally intense, physics-based, modeling of material behavior at various length scales, and the emergence of Integrated Computational Materials Engineering (ICME) have driven the need for developing data-driven surrogate models of the physics-based simulation tools using machine learning (ML) techniques. Surrogate model development allows for accurate material behavior prediction at a fraction of the cost of its physics-based counterpart, allowing for multiscale simulations of real-world applications, further enabling the ability to design fit-for-purpose materials for a reasonable computational investment. However, training such models requires extensive data, and thus effective data management is necessary to reach the full potential that ML can offer to material design and ICME. This paper proposes a generalized, robust schema that allows organizations to store both real (experimental) and virtual (simulation) data used to train machine learning models and the defining model parameters and architectures. The developed schema allows for various types of data inputs and outputs, including single point values, time-series data, and images that can be used in for various types of machine learning models while following outlined best practices for effective data management. An effective schema for machine learning data and models can help prevent the recreation of virtual/real training data and surrogate models, can help reduce the time to create new models similar to existing ones by offering a starting point in the hyperparameter determination stages, minimize resources devoted to verification and validation (V&V) and certification of models, and ensure that data and surrogate models are not misused due to full traceability of both the data and ML model. It also allows organizations access to models that have already been developed, such that they can be used in the design of new materials, enabling the overall goals of ICME.

Failure↗

Panorama 360 (Final Report)

This is the final technical report for the DOE-funded Panorama 360 project. Panorama 360 provided a resource for the collection, analysis, and sharing of performance data about end-to-end scientific workflows executing on DOE facilities. The work focused on workflows that include experimental data generation at DOE facilities. The main activities of Panorama 360 include the development of: 1. A distributed repository that stores different types of workflow execution data (e.g., point and time series performance traces at fine- and coarse-grained levels); 2. A set of open-source data capture, curation, and publishing tools fully integrated with a state-of-the-art workflow management system that automates data ingestion to the repository and enables users to discover, query, and process data from the repository; 3. A set of analysis algorithms and machine learning based tools to perform analysis and characterization of the gathered data, which can be used to detect anomalous performance or system faults; and 4. Best practices and recommendations for workflow evaluation, analysis, execution, and architectures.

97 MATHEMATICS AND COMPUTING↗

Panorama 360 (Final Report)

This final technical report from the lead institution, USC grant #DE-SC0012636, serves as the final technical report for collaborative institution UNC-CH grant #DE-SC0012390. The goal was to develop a repository and associated capabilities for data collection, ingestion, and analysis for a broad class of DOE applications that span experimental and simulation science workflows. In particular, this work focuses on workflows that include experimental data generation at DOE facilities. The main activities of Panorama 360 include the development of: (1) A distributed repository that stores different types of workflow execution data (e.g., point and time series performance traces at fine- and coarse-grained levels); (2) A set of open-source data capture, curation, and publishing tools fully integrated with a state-of-the-art workflow management system that automates data ingestion to the repository and enables users to discover, query, and process data from the repository; (3) A set of analysis algorithms and machine learning based tools to perform analysis and characterization of the gathered data, which can be used to detect anomalous performance or system faults; and (4) Best practices and recommendations for workflow evaluation, analysis, execution, and architectures.

97 MATHEMATICS AND COMPUTING↗

Towards On-Chip Learning for Low Latency Reasoning with End-to-End Synthesis

The Software Defined Architectures (SODA) Synthesizer is an open-source compiler-based tool able to automatically generate domain-specialized systems targeting Application-Specific Integrated Circuits (ASICs) or Field Programmable Gate Arrays (FPGAs) starting from high-level programming. SODA is composed of a frontend, SODA-OPT, which leverages the multilevel intermediate representation (MLIR) framework to interface with productive programming tools (e.g., machine learning frame-works), identify kernels suitable for acceleration, and perform high-level optimizations, and of a state-of-the-art high-level synthesis backend, Bambu from the PandA framework, to generate custom accelerators. One specific application of the SODA Synthesizer is the generation of accelerators to enable ultra-low latency inference and control on autonomous systems for scientific discovery (e.g., electron microscopes, sensors in particle accelerators, etc.). This paper provides an overview of the flow in the context of the generation of accelerators for edge processing to be integrated in transmission electron microscopy (TEM) devices, focusing on use cases from precision material synthesis. We show the tool in action with an example of design space exploration for inference on reconfigurable devices with a conventional deep neural network model (LeNet). Finally, we discuss the research directions and opportunities enabled by SODA in the area of autonomous control for scientific experimental workflows.

Castellana, Vito G.↗

Detecting Satellite Laser Ranging Station Data and Operational Anomalies with Machine Learning Isolation Forests at NASA's CDDIS

The International Laser Ranging Service (ILRS) is currently composed of 45 active satellite laser ranging (SLR) stations with several more set to join the network over the next several years. Station changes and histories are logged to files, but not always in real time. Sometimes these details are not added until long after changes have been made to the station –on occasion, years later. This in addition to unexpected hardware errors and other system issues that are not immediately detected impact the products generated by analysts. The ILRS Central Bureau (CB) and NASA’s Crustal Dynamics Data Information System (CDDIS) have worked to provide tools for station engineers to use. This includes the creation of station plots which contain temperature and pressure information along with LAser GEOdynamic Satellite (LAGEOS) and LAser RElativity Satellite (LARES) tracking information that enable the monitoring of station performance and todetermine whether the station has undergone any changes. As next steps, the CDDIS is working to enhance these station performance monitoring tools through machine learning. Isolation forest is an unsupervised machine learning algorithm commonly applied to anomaly detection. In this poster, the CDDIS details the steps taken to track anomalies within SLR station performance using isolation forest with LAGEOS and LARES satellite data.

Benjamin P Michael↗

Plasma confinement state classification in fusion power plants: Profile reflectometer and ensemble diagnostics

As Fusion Pilot Plants (FPPs) are increasingly viewed as within reach, many engineering challenges remain. Not many diagnostics are expected to be available in a reactor environment. Survivability, maintainability, and limited port space substantially restrict the number of FPP-relevant diagnostics. One remaining challenge is developing tools and devices to extract plasma state information necessary for controlling an FPP from a limited subset of diagnostics. This work is part of an overarching project to address this challenge. The specific diagnostic subset to be used in FPPs is still under debate. We take the approach of developing machine-learning-based tools for different significant plasma state parameters, using already known FPP-viable diagnostics. Previously we developed a plasma confinement mode classifier utilizing the Electron Cyclotron Emission (ECE) diagnostic. Here, we expand on this by developing a Profile Reflectometer (PR) based classifier with 97% test accuracy, and an ensemble model that combines the ECE and PR models into a single model, achieving 99% test accuracy.

Clark, Randall [Univ. of California, San Diego, CA↗

Flow Redirection and Induction in Steady State (FLORIS) Wind Plant Power Production Data Sets

This dataset contains turbine- and plant-level power outputs for 252,500 cases of diverse wind plant layouts operating under a wide range of yawing and atmospheric conditions. The power outputs were computed using the Gaussian wake model in NREL's FLOw Redirection and Induction in Steady State (FLORIS) model, version 2.3.0. The 252,500 cases include 500 unique wind plants generated randomly by a specialized Plant Layout Generator (PLayGen) that samples randomized realizations of wind plant layouts from one of four canonical configurations: (i) cluster, (ii) single string, (iii) multiple string, (iv) parallel string. Other wind plant layout parameters were also randomly sampled, including the number of turbines (25-200) and the mean turbine spacing (3D-10D, where D denotes the turbine rotor diameter). For each layout, 500 different sets of atmospheric conditions were randomly sampled. These include wind speed in 0-25 m/s, wind direction in 0 deg.-360 deg., and turbulence intensity chosen from low (6%), medium (8%), and high (10%). For each atmospheric inflow scenario, the individual turbine yaw angles were randomly sampled from a one-sided truncated Gaussian on the interval 0 deg.-30 deg. oriented relative to wind inflow direction. This random data is supplemented with a collection of yaw-optimized samples where FLORIS was used to determine turbine yaw angles that maximize power production for the entire plant. To generate this data, a subset of cases were selected (50 atmospheric conditions from 50 layouts each for a total of additional 2,500 cases) for which FLORIS was re-run with wake steering control optimization. The IEA onshore reference turbine, which has a 130 m rotor diameter, a 110 m hub height, and a rated power capacity of 3.4 MW was used as the turbine for all simulations. The simulations were performed using NREL's Eagle high performance computing system in February 2021 as part of the Spatial Analysis for Wind Technology Development project funded by the U.S. Department of Energy Wind Energy Technologies Office. The data was collected, reformatted, and preprocessed for this OEDI submission in May 2023 under the Foundational AI for Wind Energy project funded by the U.S. Department of Energy Wind Energy Technologies Office. This dataset is intended to serve as a benchmark against which new artificial intelligence (AI) or machine learning (ML) tools may be tested. Baseline AI/ML methods for analyzing this dataset have been implemented, and a link to their repository containing those models has been provided. The .h5 data file structure can be found in the GitHub repository under explore_wind_plant_data_h5.ipynb.

AI↗

Development of Solar Flare and Energetic Particle Prediction Portal (SEP 3 )

Solar activity is a primary factor determining the state of the Earth’s space environment, geomagnetic and ionospheric disturbances, and radiation hazards. In the current state of knowledge, machine learning (ML) methods provide essential tools for processing data, investigating relationships among various physical properties and characteristics, uncovering hidden connections, and predicting hazardous solar events. The primary difficulty in developing and applying modern machine-learning tools in heliophysics is that the essential data are scattered among over a hundred data repositories developed by instrument teams of space missions and ground-based observatories. In addition, statistical and ML methods require long time series of homogeneous measurements. To facilitate ML-ready data preparation and access, we have developed an interactive database of solar flares integrating the most essential datasets (https://solarflare.njit.edu/). The database performs an initial data processing and is automatically updated. In addition, we are developing the Solar Energetic Particle Prediction Portal (SEP3, https://sun.njit.edu/SEP3), which hosts web applications that allow users to retrieve the database records. The Portal has a search page for browsing the events from the most widely used catalogs and a dedicated space to share the most recent achievements of the team. The interactive widget can display soft X-ray and proton flux time series from GOES satellites and the flare records. The data portal has been used to evaluate the forecasts of solar proton events and investigate machine-learning approaches to SEP prediction.

SMD↗

Graph-Learning-Assisted State and Event Tracking for Solar-Penetrated Power Grids with Heterogeneous Data Sources

Unlike transmission systems, distribution systems do not typically contain sufficient metering to enable real-time state estimation. The lack of sufficient real-time measurements prohibits accurate and timely monitoring of the state of distribution systems. As a result, control and optimal operation of distribution systems, especially those containing large numbers of renewable generation units are not possible without proper data and information about the current state of the system. The main motivation of this project is to address this shortcoming by developing an approach which provides “predicted” real-time measurements so that they can be used to execute a distribution system state estimator. Thus, the objective of the project is to make the distribution systems fully observable, such that the hosting capacity for solar generation can be accurately estimated, and unnecessary solar curtailments can be avoided. In order to accomplish this goal, the project investigated the use of a grid-model-informed machine learning (ML) tool which integrates heterogeneous data streams obtained from AMI meters, SCADA as well as PMU measurements and created synchronous measurement snapshots for the state estimator (SE); and developed a hybrid robust SE which provides not only accurate state estimates but also real-time feedback for the ML model refinement.

14 SOLAR ENERGY↗

Development of Data-Driven Models for Performance Prediction and Chemical Dosing of a Full-Scale Controlled Phosphorus Precipitation Reactor

This study evaluated the use of data-driven models to improve control of a struvite precipitation reactor that removes phosphorus from wastewater while producing a fertilizer product. The researchers developed predictive models for influent orthophosphate concentration, effluent orthophosphate concentration, and phosphorus removal using operational data from a full-scale MagPrex™ reactor at a water resource recovery facility in Denver, Colorado. Model predictions were used to recommend magnesium chloride dosing adjustments needed to achieve a target effluent phosphorus concentration. Several machine learning approaches were tested, with ridge regression providing the best predictions for influent orthophosphate concentration and phosphorus removal, and XGBoost providing the best predictions for effluent orthophosphate concentration. Simulation results indicated that the decision-support approach could correctly identify dosing adjustments in most cases and reduce chemical use. Full-scale implementation achieved lower accuracy due to changing operating conditions and limited historical data in some operating ranges. Here, the results demonstrate the potential of data-driven tools to support phosphorus recovery process control while also identifying practical limitations that affect deployment in full-scale systems.

42 ENGINEERING↗