Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “statistical learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Revisiting trends in the exchange current for hydrogen evolution

Nørskov and collaborators proposed a simple kinetic model to explain the volcano relation for the hydrogen evolution reaction on transition metal surfaces such that j 0 = k 0 f(ΔG H ) where j 0 is the exchange current density, f(ΔG H ) is a function of the hydrogen adsorption free energy ΔG H as computed from density functional theory, and k 0 is a universal rate constant. Herein, focusing on the hydrogen evolution reaction in acidic medium, we revisit the original experimental data and find that the fidelity of this kinetic model can be significantly improved by invoking metal-dependence on k 0 such that the logarithm of k 0 linearly depends on the absolute value of ΔG H . Here, we further confirm this relationship using additional experimental data points obtained from a critical review of the available literature. Our analyses show that the new model decreases the discrepancy between calculated and experimental exchange current density values by up to four orders of magnitude. Furthermore, we show the model can be further improved using machine learning and statistical inference methods that integrate additional material properties.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

BRAKER3: Fully automated genome annotation using RNA-seq and protein evidence with GeneMark-ETP, AUGUSTUS, and TSEBRA

Gene prediction has remained an active area of bioinformatics research for a long time. Still, gene prediction in large eukaryotic genomes presents a challenge that must be addressed by new algorithms. The amount and significance of the evidence available from transcriptomes and proteomes vary across genomes, between genes, and even along a single gene. User-friendly and accurate annotation pipelines that can cope with such data heterogeneity are needed. The previously developed annotation pipelines BRAKER1 and BRAKER2 use RNA-seq or protein data, respectively, but not both. A further significant performance improvement integrating all three data types was made by the recently released GeneMark-ETP. We here present the BRAKER3 pipeline that builds on GeneMark-ETP and AUGUSTUS, and further improves accuracy using the TSEBRA combiner. BRAKER3 annotates protein-coding genes in eukaryotic genomes using both short-read RNA-seq and a large protein database, along with statistical models learned iteratively and specifically for the target genome. We benchmarked the new pipeline on genomes of 11 species under an assumed level of relatedness of the target species proteome to available proteomes. BRAKER3 outperforms BRAKER1 and BRAKER2. The average transcript-level F1-score is increased by about 20 percentage points on average, whereas the difference is most pronounced for species with large and complex genomes. BRAKER3 also outperforms other existing tools, MAKER2, Funannotate, and FINDER. The code of BRAKER3 is available on GitHub and as a ready-to-run Docker container for execution with Docker or Singularity. Overall, BRAKER3 is an accurate, easy-to-use tool for eukaryotic genome annotation.

59 BASIC BIOLOGICAL SCIENCES↗

Mass of 101 Sn and Bayesian extrapolations to the proton drip line

The favorable energy configurations of nuclei at magic numbers of 𝑁 neutrons and 𝑍 protons are fundamental for understanding the evolution of nuclear structure. The 𝑍 = 50 (tin) isotopic chain is a frontier for such studies, with particular interest at and around the doubly magic 100 Sn isotope, for which the mass is a topic of debate. Precise mass values for neutron-deficient isotopes provide necessary anchor points for mass models to test extrapolations near the proton drip line, where experimental studies remain out of reach. In this work, we report a Penning trap mass measurement of 101 Sn . The determined mass excess of −59889.89⁢(96) keV for 101 Sn represents a factor-of-300 improvement over the current precision and indicates that 101 Sn is less bound than previously thought. Mass predictions from a recently developed Bayesian model combination framework employing statistical machine learning and nuclear masses computed within seven global models based on nuclear density functional theory agree within 1⁢𝜎 with experimental masses from the 48 ≤ 𝑍 ≤ 52 isotopic chains. The framework's resilience to new mass data gave confidence in the extrapolation of tin masses down to 𝑁 = 46. Our calculations suggest that 96 Sn is a two-proton drip line nucleus and predict a mass excess of −58090⁢(800) keV for 100 Sn , showing a preference within 1⁢𝜎 for the mass of 100 Sn derived from the 𝛽-delayed 𝑄 value measured at GSI.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Rolling Root Mean Square Based Multimodal Anomaly Detection for Real Time Monitoring of Smart Grid

Reliable real-time monitoring is valuable for maintaining the operational integrity of modern electrical smart grids. Deployment of heterogeneous sensing technologies in substations has enabled high-resolution, multichannel waveform monitoring, but also introduces challenges for anomaly detection due to noise, baseline drift, and modality-dependent signal characteristics. In this work, we present a computationally efficient unsupervised method for multimodal event detection based on Rolling Root Mean Square based Event Detection (RRMSED). The method is developed using in-house, field deployed sensors collecting data at a utility substation. The sensing system comprises voltage and current sensors, triaxial accelerometers, and magnetometers, collectively capturing electrical, vibrational, and magnetic waveform measurements at high temporal resolution. RRMSED operates by extracting rolling RMS energy features and their first-order temporal differences from consecutive waveform segments for each channel and then applying channel-specific statistical thresholds learned from historical data. A persistence-based exceedance logic is employed to robustly identify transient events while suppressing impulsive noise, and to provide precise temporal localization with high resolution. The framework is designed for continuous server-side operation and can be deployed in real time without requiring complex models. Experiments on simulated waveform data with known ground truth demonstrate low false positive (FP) and false negative (FN) rates. Application to real substation data shows RRMSED to identify events that are not captured by conventional monitoring indicators including fast transient detection algorithm currently deployed in the system. These results indicate that rolling RMS based features provide an effective and practical basis for real-time multimodal event detection in smart-grid substations.

Mukherjee, Subrata [ORNL] (ORCID:0000000309930338)↗

Advanced Signal Decomposition Analysis and Anomaly Detection in Photovoltaic Systems

With the rapid expansion of large-scale photovoltaic (PV) plants, it is paramount for solar stakeholders to understand the reliability and efficiency of their plants to inform maintenance decisions, increase production, and understand the design factors that impact performance. Diagnosing underperformance in PV plants is challenging due to the relatively few monitoring points with respect to the large geographic footprint of the plant. This work introduces a cutting-edge method that transforms the analysis and management of key factors influencing PV plant performance, including performance loss rate (PLR), recoverable soiling, and major system changes. Identifying these factors is critical for deriving actionable insights. Leveraging advanced analytical techniques such as wavelet transformation, robust regression, and extreme point analysis, this approach provides a nuanced understanding of these factors. This method has been tested across two synthetic datasets and one real dataset, consistently surpassing existing benchmarks by achieving a lower median mean absolute error and reduced error variability across all comparable components.

14 SOLAR ENERGY↗

How Can Probabilistic Solar Power Forecasts Be Used to Lower Costs and Improve Reliability in Power Spot Markets? A Review and Application to Flexiramp Requirements

Net load uncertainty in electricity spot markets is rapidly growing. There are five general approaches by which system operators and market participants can use probabilistic forecasts of wind, solar, and load to help manage this uncertainty. These include operator situation awareness, resource risk hedging, reserves procurement, definition of contingencies, and explicit stochastic optimization. We review these approaches, and then provide a case study in which a method for using probabilistic solar forecasts to define needs for reserves is developed and evaluated. The case study has three parts. First, we describe building blocks for enhancing the Watt-Sun solar forecasting system to produce probabilistic irradiance and power forecasts. Second, relationships between Watt-Sun forecasts for multiple sites in California and the system's need for flexible ramp capability (flexiramp) are defined by machine learning and statistical methods. Third, the performance of present methods to defining flexiramp requirements, which are not conditioned on weather and renewables forecasts, is compared with that of probabilistic solar forecast-based requirements, using a multi-timescale production costing model with an 1820-bus representation of the WECC power system. Significant potential savings in fuel and flexiramp procurement costs from using solar-informed reserve requirements are found.

14 SOLAR ENERGY↗

Anticipating Technical Expertise and Capability Evolution in Research Communities Using Dynamic Graph Transformers

The ability to anticipate global technical expertise and capability evolution trends is essential for national and global security, especially in safety-critical domains such as nuclear nonproliferation (NN) and rapidly emerging fields like artificial intelligence (AI). Here, in this work, we extend traditional statistical relational learning approaches (e.g., link prediction in collaboration networks) and formulate a problem of anticipating technical expertise and capability evolution using dynamic heterogeneous graph representations. We develop novel capabilities to forecast collaboration patterns, authorship behavior, and technical capability evolution at different granularities (e.g., scientist and institution levels) in two distinct research fields. We implement a dynamic graph transformer (DGT) neural architecture, which pushes the state-of-the-art graph neural network models by: 1) forecasting heterogeneous (rather than homogeneous) nodes and edges; and 2) relying on both discrete- and continuous-time inputs. We demonstrate that our DGT models predict collaboration, partnership, and expertise patterns with 0.26, 0.73, and 0.53 mean reciprocal rank values for AI and 0.48, 0.93, and 0.22 for NN domains. DGT model performance exceeds the best-performing static graph baseline models by 30%–80% across AI and NN domains. Our findings demonstrate that DGT models boost inductive task performance when previously unseen nodes appear in the test data for the domains with emerging collaboration patterns (e.g., AI). Specifically, models accurately predict which established scientists will collaborate with early career scientists and vice versa in the AI domain.

97 MATHEMATICS AND COMPUTING↗

Stochastic Gradient-Based Distributed Bayesian Estimation in Cooperative Sensor Networks

Distributed Bayesian inference provides a full quantification of uncertainty offering numerous advantages over point estimates that autonomous sensor networks are able to exploit. However, fully-decentralized Bayesian inference often requires large communication overheads and low network latency, resources that are not typically available in practical applications. In this paper, we propose a decentralized Bayesian inference approach based on stochastic gradient Langevin dynamics, which produces full posterior distributions at each of the nodes with significantly lower communication overhead. We provide analytical results on convergence of the proposed distributed algorithm to the centralized posterior, under typical network constraints. Finally, we also provide extensive simulation results to demonstrate the validity of the proposed approach.

42 ENGINEERING↗

Machine learning tools for epigenetics

The software provides machine learning analysis and visualization to detect patterns in epigenetic data, including conventional machine learning and statistical methods, and open-source packages like pyBigWig (https://github.com/deeptools/pyBigWig) for data processing. The software is written in python, it uses some python libraries.

Kim, Anastasiia↗

Portable, heterogeneous ensemble workflows at scale using libEnsemble

libEnsemble is a Python-based toolkit for running dynamic ensembles, developed as part of the DOE Exascale Computing Project. The toolkit utilizes a unique generator–simulator–allocator paradigm, where generators produce input for simulators, simulators evaluate those inputs, and allocators decide whether and when a simulator or generator should be called. The generator steers the ensemble based on simulation results. Generators may, for example, apply methods for numerical optimization, machine learning, or statistical calibration. libEnsemble communicates between a manager and workers. Flexibility is provided through multiple manager–worker communication substrates each of which has different benefits. These include Python’s multiprocessing, mpi4py, and TCP. Multisite ensembles are supported using Balsam or Globus Compute. We overview the unique characteristics of libEnsemble as well as current and potential interoperability with other packages in the workflow ecosystem. We highlight libEnsemble’s dynamic resource features: libEnsemble can detect system resources, such as available nodes, cores, and GPUs, and assign these in a portable way. These features allow users to specify the number of processors and GPUs required for each simulation; and resources will be automatically assigned on a wide range of systems, including Frontier, Aurora, and Perlmutter. Such ensembles can include multiple simulation types, some using GPUs and others using only CPUs, sharing nodes for maximum efficiency. We also describe the benefits of libEnsemble’s generator–simulator coupling, which easily exposes to the user the ability to cancel, and portably kill, running simulations based on models that are updated with intermediate simulation output. We demonstrate libEnsemble’s capabilities, scalability, and scientific impact via a Gaussian process surrogate training problem for the longitudinal density profile at the exit of a plasma accelerator stage. In conclusion, the study uses gpCAM for the surrogate model and employs either Wake-T or WarpX simulations, highlighting efficient use of resources that can easily extend to exascale.

Dynamic ensembles↗

Roundup causes embryonic development failure and alters metabolic pathways and gut microbiota functionality in non-target species

Background: Research around the weedkiller Roundup is among the most contentious of the twenty-first century. Scientists have provided inconclusive evidence that the weedkiller causes cancer and other life-threatening diseases, while industry-paid research reports that the weedkiller has no adverse effect on humans or animals. Much of the controversial evidence on Roundup is rooted in the approach used to determine safe use of chemicals, defined by outdated toxicity tests. We apply a system biology approach to the biomedical and ecological model species Daphnia to quantify the impact of glyphosate and of its commercial formula, Roundup, on fitness, genome-wide transcription and gut microbiota, taking full advantage of clonal reproduction in Daphnia. We then apply machine learning-based statistical analysis to identify and prioritize correlations between genome-wide transcriptional and microbiota changes. Results: We demonstrate that chronic exposure to ecologically relevant concentrations of glyphosate and Roundup at the approved regulatory threshold for drinking water in the US induce embryonic developmental failure, induce significant DNA damage (genotoxicity), and interfere with signaling. Furthermore, chronic exposure to the weedkiller alters the gut microbiota functionality and composition interfering with carbon and fat metabolism, as well as homeostasis. Using the “Reactome,” we identify conserved pathways across the Tree of Life, which are potential targets for Roundup in other species, including liver metabolism, inflammation pathways, and collagen degradation, responsible for the repair of wounds and tissue remodeling. Conclusions: Our results show that chronic exposure to concentrations of Roundup and glyphosate at the approved regulatory threshold for drinking water causes embryonic development failure and alteration of key metabolic functions via direct effect on the host molecular processes and indirect effect on the gut microbiota. The ecological model species Daphnia occupies a central position in the food web of aquatic ecosystems, being the preferred food of small vertebrates and invertebrates as well as a grazer of algae and bacteria. The impact of the weedkiller on this keystone species has cascading effects on aquatic food webs, affecting their ability to deliver critical ecosystem services.

59 BASIC BIOLOGICAL SCIENCES↗

Research Needs for Trusted Analytics in National Security Settings

As artificial intelligence, machine learning, and statistical modeling methods become commonplace in national security applications, the drive to create trusted analytics becomes increasingly important. The goal of this report is to identify areas of research that can provide the foundational understanding and technical prerequisites for the development and deployment of trusted analytics in national security settings. Our review of the literature covered several disjoint research communities, including computer science, statistics, human factors, and several branches of psychology and cognitive science, which tend not to interact with one another or cite each other's literatures. As a result, there exists no agreed-upon theoretical framework for understanding how various factors influence trust and no well-established empirical paradigm for studying these effects. This report therefore takes three steps. First, we define several key terms in an effort to provide a unifying language for trusted analytics and to manage the scope of the problem. Second, we outline an empirical perspective that identifies key independent, moderating, and dependent variables in assessing trusted analytics. Though not a substitute for a theoretical framework, the empirical perspective does support research and development of trusted analytics in the national security domain. Finally, we discuss several research gaps relevant to developing trusted analytics for the national security mission space.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Evaluating Offshore Infrastructure Integrity

Drilling in the offshore environment involves a complex network of infrastructure including pipelines, platforms, rigs, subsea installations, ports, and terminals. Government and industry partners have developed this network over many decades and it remains a critical part of the United States (U.S.) energy portfolio. Many of the major components of this system have been designed with a 20- to 30-year lifespan, yet consistent and growing energy demands support the need to extend the design life of existing infrastructure or repurpose it for secondary needs (i.e. enhanced oil recovery, carbon storage, and new wells). As a result, a growing portion of the offshore infrastructure in the U.S. is approaching or has exceeded its original design life. A critical step in ensuring the continued safe and effective operation of offshore infrastructure is developing a comprehensive understanding of the state of offshore infrastructure and the factors that effect it. The purpose of this project is to assess the current state of existing infrastructure and identify the factors involved in infrastructure degradation through the development and application of big data analytics, machine learning, and advanced spatio-temporal analysis. The project leverages existing data at NETL and combines it with new information on offshore oil and gas structures and the ambient offshore environment in an effort to identify patterns associated with infrastructure integrity. Building on the identified trends and patterns, this project incorporates exploratory analytics and spatial analysis tools in conjunction with machine learning and statistical models to characterize the condition of existing platforms in the offshore environment and predict their risk of failure.

02 PETROLEUM↗

Physics-Informed Learning Machines for Multiscale and Multiphysics Problems (PHILMS) (Technical Report)

The research work at University of California Santa Barbara (UCSB) resulted in several new developments in the areas of scientific machine learning, numerical analysis, and practical methods for data-driven modeling, prediction, reductions, and simulation. Many of the projects were carried out in collaboration with members of the national laboratories at Sandia National Laboratories (SNL), Pacific Northwestern National Laboratories (PNNL), and other institutions. Results included developing new scientific machine learning methods, related theory and mathematical frameworks for analysis and training, data-driven numerical solvers, and related tools and software for scientific computation. During the support period, over 16+ papers were submitted for publication, and 4 open-source software packages were developed and released (available at http://atzberger.org/). In addition, 7+ students and 2 post-docs were mentored in collaboration with the laboratory staff for future careers in academia, government labs, and industry.

97 MATHEMATICS AND COMPUTING↗

Materials Characterization, Prediction and Control Project: Summary Report on Data Analytics Framework

This report summarizes the activities performed under the data analytics Vertex in the Materials Characterization, Prediction and Control Project funded under laboratory directed research and development at Pacific Northwest National Laboratory. The data analytics Vertex developed models for associating global or local process parameters, microstructural features, and performance properties of friction-stir-processed 316L stainless steel plates. Statistical, machine learning, and deep learning models, as well as generative artificial intelligence approaches, were used to develop the associations between the process-structure-property data streams. These associations formed the basis for predicting global properties of parts manufactured under different process envelopes, providing a basis for predicting performance using data driven as well as physics-informed and physics-constrained approaches. Additionally, the associations were used to predict local process parameters and microstructural features of the product, predictive relationships that have the potential to form the basis of a control framework that could eventually modulate a friction-stir process to maintain product quality.

316L stainless steel↗

Optimizing enzymes for plastic upcycling using machine learning design and high throughput experiments

Plastic use is ubiquitous in the modern world, and polyethylene terephthalate (PET) is one of the most abundantly produced plastics (and the most highly produced polyester), with ~65 million metric tons manufactured annually. To the consumer, PET is likely most recognizable as the plastic used to make beverage bottles. Like many plastics, traditional mechanical or chemical means of PET deconstruction and upcycling are costly and inefficient. Because of these challenges, recycled plastic is generally of lower quality and is more expensive to produce than virgin plastic derived from petroleum. Ultimately, this results in most plastic ending up as waste. We view plastic waste as an underutilized resource which, with the development of more efficient and high-quality recycling processes, could (1) generate significant economic value while (2) decreasing petroleum usage and greenhouse gas emissions, as well as (3) minimizing its negative environmental and health impacts. Biocatalytic recycling, or biomanufacturing the basic building blocks of new plastic from plastic waste, is a promising approach to plastic reuse that complements existing recycling technologies. Recently, biological enzymes capable of breaking down PET have garnered significant attention as an attractive means of dealing with the plastic problem. These enzymes are currently undergoing pilot studies for implementation in industrial-scale enzyme-based recycling. However, there are significant limitations to current enzymes, including the need to perform costly pre-processing of the plastic waste before the enzymes are able to work. Further optimization of these enzymes is necessary to make these technologies competitive, and ultimately incentivise industry-wide adoption of this biology-based green recycling technology. n this work we demonstrate a means to design and generate performant biological enzymes, capable of efficiently deconstructing plastic waste. Specifically, we applied recent advances in artificial intelligence, machine learning, and statistical analysis to design new versions and discover natural enzymes capable of breaking down PET. We focused on optimizing key properties that are important for industrial-scale enzymatic recycling such as pH and thermotolerance. Normal testing of enzymatic plastic-deconstruction is extremely labor intensive and so through this work we also developed a robotic-assisted experimental pipeline capable of characterizing thousands of candidate enzymes. The results of this iterative, AI-guided, multi-discipline approach have led to increases in enzymatic breakdown of over 150X over starting enzymes. This work supports the rapidly developing and transformative field of biocatalytic solutions to environmental problems beyond the discovery and predictive understanding of enzymes for polymer recycling, and has wide implications for tackling numerous energy problems such as carbon capture and fixation (e.g., engineering carbon monoxide dehydrogenase and the rubisco-pathway), biomining (e.g., design of lanthanide-binding proteins) and biomanufacturing (e.g., lignin-deconstruction enzymes).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Meeting Global Health Needs via Infectious Disease Forecasting: Development of a Reliable Data-Driven Framework

Infectious diseases (IDs) have a significant detrimental impact on global health. Timely and accurate ID forecasting can result in more informed implementation of control measures and prevention policies. To meet the operational decision-making needs of real-world circumstances, we aimed to build a standardized, reliable, and trustworthy ID forecasting pipeline and visualization dashboard that is generalizable across a wide range of modeling techniques, IDs, and global locations. We forecasted 6 diverse, zoonotic diseases (brucellosis, campylobacteriosis, Middle East respiratory syndrome, Q fever, tick-borne encephalitis, and tularemia) across 4 continents and 8 countries. We included a wide range of statistical, machine learning, and deep learning models (n=9) and trained them on a multitude of features (average n=2326) within the One Health landscape, including demography, landscape, climate, and socioeconomic factors. The pipeline and dashboard were created in consideration of crucial operational metrics—prediction accuracy, computational efficiency, spatiotemporal generalizability, uncertainty quantification, and interpretability—which are essential to strategic data-driven decisions. While no single best model was suitable for all disease, region, and country combinations, our ensemble technique selects the best-performing model for each given scenario to achieve the closest prediction. For new or emerging diseases in a region, the ensemble model can predict how the disease may behave in the new region using a pretrained model from a similar region with a history of that disease. The data visualization dashboard provides a clean interface of important analytical metrics, such as ID temporal patterns, forecasts, prediction uncertainties, and model feature importance across all geographic locations and disease combinations. As the need for real-time, operational ID forecasting capabilities increases, this standardized and automated platform for data collection, analysis, and reporting is a major step forward in enabling evidence-based public health decisions and policies for the prevention and mitigation of future ID outbreaks.

60 APPLIED LIFE SCIENCES↗