Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “statistical learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Advanced Signal Decomposition Analysis and Anomaly Detection in Photovoltaic Systems

With the rapid expansion of large-scale photovoltaic (PV) plants, it is paramount for solar stakeholders to understand the reliability and efficiency of their plants to inform maintenance decisions, increase production, and understand the design factors that impact performance. Diagnosing underperformance in PV plants is challenging due to the relatively few monitoring points with respect to the large geographic footprint of the plant. This work introduces a cutting-edge method that transforms the analysis and management of key factors influencing PV plant performance, including performance loss rate (PLR), recoverable soiling, and major system changes. Identifying these factors is critical for deriving actionable insights. Leveraging advanced analytical techniques such as wavelet transformation, robust regression, and extreme point analysis, this approach provides a nuanced understanding of these factors. This method has been tested across two synthetic datasets and one real dataset, consistently surpassing existing benchmarks by achieving a lower median mean absolute error and reduced error variability across all comparable components.

14 SOLAR ENERGY↗

How Can Probabilistic Solar Power Forecasts Be Used to Lower Costs and Improve Reliability in Power Spot Markets? A Review and Application to Flexiramp Requirements

Net load uncertainty in electricity spot markets is rapidly growing. There are five general approaches by which system operators and market participants can use probabilistic forecasts of wind, solar, and load to help manage this uncertainty. These include operator situation awareness, resource risk hedging, reserves procurement, definition of contingencies, and explicit stochastic optimization. We review these approaches, and then provide a case study in which a method for using probabilistic solar forecasts to define needs for reserves is developed and evaluated. The case study has three parts. First, we describe building blocks for enhancing the Watt-Sun solar forecasting system to produce probabilistic irradiance and power forecasts. Second, relationships between Watt-Sun forecasts for multiple sites in California and the system's need for flexible ramp capability (flexiramp) are defined by machine learning and statistical methods. Third, the performance of present methods to defining flexiramp requirements, which are not conditioned on weather and renewables forecasts, is compared with that of probabilistic solar forecast-based requirements, using a multi-timescale production costing model with an 1820-bus representation of the WECC power system. Significant potential savings in fuel and flexiramp procurement costs from using solar-informed reserve requirements are found.

14 SOLAR ENERGY↗

Anticipating Technical Expertise and Capability Evolution in Research Communities Using Dynamic Graph Transformers

The ability to anticipate global technical expertise and capability evolution trends is essential for national and global security, especially in safety-critical domains such as nuclear nonproliferation (NN) and rapidly emerging fields like artificial intelligence (AI). Here, in this work, we extend traditional statistical relational learning approaches (e.g., link prediction in collaboration networks) and formulate a problem of anticipating technical expertise and capability evolution using dynamic heterogeneous graph representations. We develop novel capabilities to forecast collaboration patterns, authorship behavior, and technical capability evolution at different granularities (e.g., scientist and institution levels) in two distinct research fields. We implement a dynamic graph transformer (DGT) neural architecture, which pushes the state-of-the-art graph neural network models by: 1) forecasting heterogeneous (rather than homogeneous) nodes and edges; and 2) relying on both discrete- and continuous-time inputs. We demonstrate that our DGT models predict collaboration, partnership, and expertise patterns with 0.26, 0.73, and 0.53 mean reciprocal rank values for AI and 0.48, 0.93, and 0.22 for NN domains. DGT model performance exceeds the best-performing static graph baseline models by 30%–80% across AI and NN domains. Our findings demonstrate that DGT models boost inductive task performance when previously unseen nodes appear in the test data for the domains with emerging collaboration patterns (e.g., AI). Specifically, models accurately predict which established scientists will collaborate with early career scientists and vice versa in the AI domain.

97 MATHEMATICS AND COMPUTING↗

Machine learning tools for epigenetics

The software provides machine learning analysis and visualization to detect patterns in epigenetic data, including conventional machine learning and statistical methods, and open-source packages like pyBigWig (https://github.com/deeptools/pyBigWig) for data processing. The software is written in python, it uses some python libraries.

Kim, Anastasiia↗

Portable, heterogeneous ensemble workflows at scale using libEnsemble

libEnsemble is a Python-based toolkit for running dynamic ensembles, developed as part of the DOE Exascale Computing Project. The toolkit utilizes a unique generator–simulator–allocator paradigm, where generators produce input for simulators, simulators evaluate those inputs, and allocators decide whether and when a simulator or generator should be called. The generator steers the ensemble based on simulation results. Generators may, for example, apply methods for numerical optimization, machine learning, or statistical calibration. libEnsemble communicates between a manager and workers. Flexibility is provided through multiple manager–worker communication substrates each of which has different benefits. These include Python’s multiprocessing, mpi4py, and TCP. Multisite ensembles are supported using Balsam or Globus Compute. We overview the unique characteristics of libEnsemble as well as current and potential interoperability with other packages in the workflow ecosystem. We highlight libEnsemble’s dynamic resource features: libEnsemble can detect system resources, such as available nodes, cores, and GPUs, and assign these in a portable way. These features allow users to specify the number of processors and GPUs required for each simulation; and resources will be automatically assigned on a wide range of systems, including Frontier, Aurora, and Perlmutter. Such ensembles can include multiple simulation types, some using GPUs and others using only CPUs, sharing nodes for maximum efficiency. We also describe the benefits of libEnsemble’s generator–simulator coupling, which easily exposes to the user the ability to cancel, and portably kill, running simulations based on models that are updated with intermediate simulation output. We demonstrate libEnsemble’s capabilities, scalability, and scientific impact via a Gaussian process surrogate training problem for the longitudinal density profile at the exit of a plasma accelerator stage. In conclusion, the study uses gpCAM for the surrogate model and employs either Wake-T or WarpX simulations, highlighting efficient use of resources that can easily extend to exascale.

Dynamic ensembles↗

Evaluating Offshore Infrastructure Integrity

Drilling in the offshore environment involves a complex network of infrastructure including pipelines, platforms, rigs, subsea installations, ports, and terminals. Government and industry partners have developed this network over many decades and it remains a critical part of the United States (U.S.) energy portfolio. Many of the major components of this system have been designed with a 20- to 30-year lifespan, yet consistent and growing energy demands support the need to extend the design life of existing infrastructure or repurpose it for secondary needs (i.e. enhanced oil recovery, carbon storage, and new wells). As a result, a growing portion of the offshore infrastructure in the U.S. is approaching or has exceeded its original design life. A critical step in ensuring the continued safe and effective operation of offshore infrastructure is developing a comprehensive understanding of the state of offshore infrastructure and the factors that effect it. The purpose of this project is to assess the current state of existing infrastructure and identify the factors involved in infrastructure degradation through the development and application of big data analytics, machine learning, and advanced spatio-temporal analysis. The project leverages existing data at NETL and combines it with new information on offshore oil and gas structures and the ambient offshore environment in an effort to identify patterns associated with infrastructure integrity. Building on the identified trends and patterns, this project incorporates exploratory analytics and spatial analysis tools in conjunction with machine learning and statistical models to characterize the condition of existing platforms in the offshore environment and predict their risk of failure.

02 PETROLEUM↗

Physics-Informed Learning Machines for Multiscale and Multiphysics Problems (PHILMS) (Technical Report)

The research work at University of California Santa Barbara (UCSB) resulted in several new developments in the areas of scientific machine learning, numerical analysis, and practical methods for data-driven modeling, prediction, reductions, and simulation. Many of the projects were carried out in collaboration with members of the national laboratories at Sandia National Laboratories (SNL), Pacific Northwestern National Laboratories (PNNL), and other institutions. Results included developing new scientific machine learning methods, related theory and mathematical frameworks for analysis and training, data-driven numerical solvers, and related tools and software for scientific computation. During the support period, over 16+ papers were submitted for publication, and 4 open-source software packages were developed and released (available at http://atzberger.org/). In addition, 7+ students and 2 post-docs were mentored in collaboration with the laboratory staff for future careers in academia, government labs, and industry.

97 MATHEMATICS AND COMPUTING↗

Materials Characterization, Prediction and Control Project: Summary Report on Data Analytics Framework

This report summarizes the activities performed under the data analytics Vertex in the Materials Characterization, Prediction and Control Project funded under laboratory directed research and development at Pacific Northwest National Laboratory. The data analytics Vertex developed models for associating global or local process parameters, microstructural features, and performance properties of friction-stir-processed 316L stainless steel plates. Statistical, machine learning, and deep learning models, as well as generative artificial intelligence approaches, were used to develop the associations between the process-structure-property data streams. These associations formed the basis for predicting global properties of parts manufactured under different process envelopes, providing a basis for predicting performance using data driven as well as physics-informed and physics-constrained approaches. Additionally, the associations were used to predict local process parameters and microstructural features of the product, predictive relationships that have the potential to form the basis of a control framework that could eventually modulate a friction-stir process to maintain product quality.

316L stainless steel↗

Optimizing enzymes for plastic upcycling using machine learning design and high throughput experiments

Plastic use is ubiquitous in the modern world, and polyethylene terephthalate (PET) is one of the most abundantly produced plastics (and the most highly produced polyester), with ~65 million metric tons manufactured annually. To the consumer, PET is likely most recognizable as the plastic used to make beverage bottles. Like many plastics, traditional mechanical or chemical means of PET deconstruction and upcycling are costly and inefficient. Because of these challenges, recycled plastic is generally of lower quality and is more expensive to produce than virgin plastic derived from petroleum. Ultimately, this results in most plastic ending up as waste. We view plastic waste as an underutilized resource which, with the development of more efficient and high-quality recycling processes, could (1) generate significant economic value while (2) decreasing petroleum usage and greenhouse gas emissions, as well as (3) minimizing its negative environmental and health impacts. Biocatalytic recycling, or biomanufacturing the basic building blocks of new plastic from plastic waste, is a promising approach to plastic reuse that complements existing recycling technologies. Recently, biological enzymes capable of breaking down PET have garnered significant attention as an attractive means of dealing with the plastic problem. These enzymes are currently undergoing pilot studies for implementation in industrial-scale enzyme-based recycling. However, there are significant limitations to current enzymes, including the need to perform costly pre-processing of the plastic waste before the enzymes are able to work. Further optimization of these enzymes is necessary to make these technologies competitive, and ultimately incentivise industry-wide adoption of this biology-based green recycling technology. n this work we demonstrate a means to design and generate performant biological enzymes, capable of efficiently deconstructing plastic waste. Specifically, we applied recent advances in artificial intelligence, machine learning, and statistical analysis to design new versions and discover natural enzymes capable of breaking down PET. We focused on optimizing key properties that are important for industrial-scale enzymatic recycling such as pH and thermotolerance. Normal testing of enzymatic plastic-deconstruction is extremely labor intensive and so through this work we also developed a robotic-assisted experimental pipeline capable of characterizing thousands of candidate enzymes. The results of this iterative, AI-guided, multi-discipline approach have led to increases in enzymatic breakdown of over 150X over starting enzymes. This work supports the rapidly developing and transformative field of biocatalytic solutions to environmental problems beyond the discovery and predictive understanding of enzymes for polymer recycling, and has wide implications for tackling numerous energy problems such as carbon capture and fixation (e.g., engineering carbon monoxide dehydrogenase and the rubisco-pathway), biomining (e.g., design of lanthanide-binding proteins) and biomanufacturing (e.g., lignin-deconstruction enzymes).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Meeting Global Health Needs via Infectious Disease Forecasting: Development of a Reliable Data-Driven Framework

Infectious diseases (IDs) have a significant detrimental impact on global health. Timely and accurate ID forecasting can result in more informed implementation of control measures and prevention policies. To meet the operational decision-making needs of real-world circumstances, we aimed to build a standardized, reliable, and trustworthy ID forecasting pipeline and visualization dashboard that is generalizable across a wide range of modeling techniques, IDs, and global locations. We forecasted 6 diverse, zoonotic diseases (brucellosis, campylobacteriosis, Middle East respiratory syndrome, Q fever, tick-borne encephalitis, and tularemia) across 4 continents and 8 countries. We included a wide range of statistical, machine learning, and deep learning models (n=9) and trained them on a multitude of features (average n=2326) within the One Health landscape, including demography, landscape, climate, and socioeconomic factors. The pipeline and dashboard were created in consideration of crucial operational metrics—prediction accuracy, computational efficiency, spatiotemporal generalizability, uncertainty quantification, and interpretability—which are essential to strategic data-driven decisions. While no single best model was suitable for all disease, region, and country combinations, our ensemble technique selects the best-performing model for each given scenario to achieve the closest prediction. For new or emerging diseases in a region, the ensemble model can predict how the disease may behave in the new region using a pretrained model from a similar region with a history of that disease. The data visualization dashboard provides a clean interface of important analytical metrics, such as ID temporal patterns, forecasts, prediction uncertainties, and model feature importance across all geographic locations and disease combinations. As the need for real-time, operational ID forecasting capabilities increases, this standardized and automated platform for data collection, analysis, and reporting is a major step forward in enabling evidence-based public health decisions and policies for the prevention and mitigation of future ID outbreaks.

60 APPLIED LIFE SCIENCES↗

Classification of Photovoltaic Failures with Hidden Markov Modeling, an Unsupervised Statistical Approach

Failure detection methods are of significant interest for photovoltaic (PV) site operators to help reduce gaps between expected and observed energy generation. Current approaches for field-based fault detection, however, rely on multiple data inputs and can suffer from interpretability issues. In contrast, this work offers an unsupervised statistical approach that leverages hidden Markov models (HMM) to identify failures occurring at PV sites. Using performance index data from 104 sites across the United States, individual PV-HMM models are trained and evaluated for failure detection and transition probabilities. This analysis indicates that the trained PV-HMM models have the highest probability of remaining in their current state (87.1% to 93.5%), whereas the transition probability from normal to failure (6.5%) is lower than the transition from failure to normal (12.9%) states. A comparison of these patterns using both threshold levels and operations and maintenance (O&M) tickets indicate high precision rates of PV-HMMs (median = 82.4%) across all of the sites. Although additional work is needed to assess sensitivities, the PV-HMM methodology demonstrates significant potential for real-time failure detection as well as extensions into predictive maintenance capabilities for PV.

classification↗

Evaluation of obstacle modelling approaches for resource assessment and small wind turbine siting: case study in the northern Netherlands

Abstract. Growth in adoption of distributed wind turbines for energy generation is significantly impacted by challenges associated with siting and accurate estimation of the wind resource. Small turbines, at hub heights of 40 m or less, are greatly impacted by terrestrial obstacles such as built structures and vegetation that can cause complex wake effects. While some progress in high-fidelity complex fluid dynamics (CFD) models has increased the potential accuracy for modelling the impacts of obstacles on turbulent wind flow, these models are too computationally expensive for practical siting and resource assessment applications. To understand the efficacy of available models in situ, this study evaluates classic and commonly used methods alongside new state-of-the-art lower-order models derived from CFD simulations and machine learning approaches. This evaluation is conducted using a subset of an extensive original dataset of measurements from more than 300 operational wind turbines in the northern Netherlands. The results show that data-driven methods (e.g. machine learning and statistical modelling) are most effective at predicting production at real sites with an average error in annual energy production of 2.5 %. When sufficient data may not be available de novo to support these data-driven approaches, models derived from high-fidelity simulations show promise and reliably outperform classic methods. On average these models have 6.3 %–11.5 % error compared with 26 % for classic methods and 27 % baseline error for reanalysis data without obstacle correction. While more performant on average, these methods are also sensitive to the quality of obstacle descriptions and reanalysis inputs.

17 WIND ENERGY↗

Verification, Validation, and Calibration Through a Causal Lens

While typical validation and verification approaches focus on identifying the associations between data elements using statistical and machine learning methods, the novel methods in this paper focus instead on identifying causal relationships between data elements. Statistical and machine-learning-based approaches are strictly data-driven, meaning that they provide quantitative comparison measures between data sets without explicitly considering the hypotheses behind them. This can lead to the erroneous conclusion that, if two data sets are close enough, the models that generated them are similar. In addition, when experimental and simulated data differ to an extent that fails to meet the acceptance criteria, calibration techniques are used to tweak simulation model parameters to reduce the gap between the two types of data. This produces the false expectation that a simulation model will match reality. The methods presented in this paper move away from these strictly data-driven methods for validation and calibration toward more robust, model-driven methods based on causal inference. Causal inference aims to identify the possible mechanisms that might have generated data. Thus, this analysis targets the prediction of the effects when one (or more) of the identified mechanisms are altered. There are many approaches to identify, quantify, and illustrate causal relationships. For the scope of this paper, directed graphs are employed as causal models. If the directed graph lacks cycles, it is known as a directed acyclic graph. A node in such a graph represents an observed data element while a directed edge connecting two nodes represents a causal relationship between two variables. The developed causal methods are designed to extract causal models from simulation models and experimental data. Causal models capture the causal relationships between data elements (e.g., simulated and experimental data). In this context, validation and verification are performed by comparing causal models. The proposed approach does not only inform system analysts on how a simulation model matches real-world data, but also identifies elements of the simulation model that should be revised when discrepancies between simulation and experimental data are observed. Through these causal methods, analysts can identify the portion of the model equation(s) that are behind an edge connecting two variables. Hence, once the structural differences between causal models have been determined, model calibration can occur by changing only those model parameters that impact the identified causal relationships.

97 MATHEMATICS AND COMPUTING↗

VALIDATION, VERIFICATION, AND CALIBRATION THROUGH A CAUSAL LENS

This paper presents an alternative method based on causal inference to perform validation, verification, and calibration of simulation models. While classical validation and verification approaches focus on the identification of the associations between data elements using statistical and machine learning methods, the novel methods in this paper focus instead on the identification of causal relationships between data elements. Statistical and machine learning-based approaches are strictly data-driven, meaning that they provide quantitative comparison measures between datasets without explicitly considering the hypotheses behind them. This can lead to the erroneous conclusion that, if two data sets are close enough, then the models that generated them are similar. In addition, when experimental and simulated data differ to an extent that fails to meet the acceptance criteria, calibration techniques are used to tweak simulation model parameters to reduce the gap between the two types of data. This produces the false expectation that a simulation model will match reality. The methods presented in this paper move away from these strictly data-driven methods for validation and calibration toward more robust, model-driven methods based on causal inference. Causal inference aims to identify the possible mechanisms that might have generated data. Thus, this analysis targets the prediction of the effects when one (or more) of the identified mechanisms are altered. There are many approaches to identify, quantify and illustrate causal relationships. For the scope of this paper, directed graphs are employed as causal models. If the directed graph lacks cycles it is known as a directed acyclic graph (DAG). A node in such a graph represents an observed data element while a directed edge connecting two nodes represents a causal relationship between two variables. The developed causal methods are designed to extract causal models from simulation models and from experimental data. Causal models capture the causal relationships between data elements (e.g., simulated and experimental data). In this context, validation and verification are performed by comparing causal models. The proposed approach does not only inform system analysts on how a simulation model matches real-world data, but also identifies elements of the simulation model that should be revised when discrepancies between simulation and experimental data are observed. Through these causal methods, analysts have a means to identify the portion of the model equation(s) that are behind an edge connecting two variables. Hence, once the structural differences between causal models have been determined, model calibration can occur by changing only those model parameters that impact the identified causal relationships.

97 MATHEMATICS AND COMPUTING↗

Flood Susceptibility Mapping Using Machine Learning and Geospatial-Sentinel-1 SAR Integration for Enhanced Early Warning Systems

This study presents a comprehensive framework for flood susceptibility mapping by integrating geospatial factors with both statistical and machine learning models. Thirteen Flood-related factors, including DEM, slope, TWI, NDVI, etc., are extracted as features of models, and historical flood data derived from Sentinel-1 SAR from 2018 to 2023 are used as the target variables of the models. These datasets are analyzed using a frequency-based statistical model and three machine learning models, including Random Forest, XGBoost, and CNN, to generate flood susceptibility maps. The performance of each model is evaluated through AUC; and SHAP scores are separately generated for Machine learning (ML) models to explain each feature contribution in the ML model. The generated susceptibility maps are validated by high-flood-risk locations monitored by flood sensors, BLE inundation models, and flood-prone areas suggested by the Local Community Task Force. The results indicate that the XGBoost model outperforms all other models, with an AUC of 0.92 and demonstrates the highest alignment with recommended high-flood-risk locations, while the frequency-based statistical model showed the weakest performance with an AUC of 0.65. SHAP value graphs highlight the elevation, slope, and TWI as the most influential features across all models. The susceptibility maps generated by the machine learning model show strong agreement with the BLE map and high-flood-risk areas identified by the local Community Task Force.

Google Engine↗

Robust Machine Learning

UQ4ML is a code repository for a set of tools for the development of robust machine learning methods, uncertainty quantification and explainability of machine learning methods. The goal of these tools is to develop more robust and statistically rigorous machine learning methods for scientific applications. These tools are developed in Python, a high-level programming language that takes advantage of the Python ecosystem of high-quality open-source packages for machine learning.

Oyen, Diane↗