Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Statistical learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Knowledge graph-aided Bayesian active learning for top- K genetic interaction discovery

In silico methods for predicting the effects of multi-gene perturbations hold great promise for advancing functional genomics, computational drug discovery, and disease modeling. However, the development of these predictive algorithms for mammalian systems has been hampered by limited datasets and high experimental costs. In this study, we present a Bayesian active learning framework designed to discover pairwise host gene knockdowns that effectively inhibit viral proliferation in an in vitro HIV-1 infection model. Our method leverages a biological knowledge graph as side information and employs a computationally efficient batch diversification approach. We evaluated this framework using a dataset of viral load measurements obtained from multi-day dual-gene depletion experiments, encompassing all possible pairwise knockdowns of over 350 host genes associated with HIV infection. We demonstrate that our framework rapidly identifies the most effective gene knockdown pairs for reducing viral load. Furthermore, we show that incorporating side information enhances performance during the early stages of active learning (low data regime), while our batch diversification strategy significantly boosts performance in later stages (high data regime). This framework is general and can be adapted to explore gene interactions in other contexts, such as synthetic lethality prediction and mapping epistatic effects across quantitative trait loci.

Computational biology and bioinformatics↗

Revisiting trends in the exchange current for hydrogen evolution

Nørskov and collaborators proposed a simple kinetic model to explain the volcano relation for the hydrogen evolution reaction on transition metal surfaces such that j 0 = k 0 f(ΔG H ) where j 0 is the exchange current density, f(ΔG H ) is a function of the hydrogen adsorption free energy ΔG H as computed from density functional theory, and k 0 is a universal rate constant. Herein, focusing on the hydrogen evolution reaction in acidic medium, we revisit the original experimental data and find that the fidelity of this kinetic model can be significantly improved by invoking metal-dependence on k 0 such that the logarithm of k 0 linearly depends on the absolute value of ΔG H . Here, we further confirm this relationship using additional experimental data points obtained from a critical review of the available literature. Our analyses show that the new model decreases the discrepancy between calculated and experimental exchange current density values by up to four orders of magnitude. Furthermore, we show the model can be further improved using machine learning and statistical inference methods that integrate additional material properties.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A practical and efficient approach for Bayesian quantum state estimation

Bayesian inference is a powerful paradigm for quantum state tomography, treating uncertainty in meaningful and informative ways. Yet the numerical challenges associated with sampling from complex probability distributions hampers Bayesian tomography in practical settings. In this article, we introduce an improved, self-contained approach for Bayesian quantum state estimation. Leveraging advances in machine learning and statistics, our formulation relies on highly efficient preconditioned Crank–Nicolson sampling and a pseudo-likelihood. We theoretically analyze the computational cost, and provide explicit examples of inference for both actual and simulated datasets, illustrating improved performance with respect to existing approaches.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

BRAKER3: Fully automated genome annotation using RNA-seq and protein evidence with GeneMark-ETP, AUGUSTUS, and TSEBRA

Gene prediction has remained an active area of bioinformatics research for a long time. Still, gene prediction in large eukaryotic genomes presents a challenge that must be addressed by new algorithms. The amount and significance of the evidence available from transcriptomes and proteomes vary across genomes, between genes, and even along a single gene. User-friendly and accurate annotation pipelines that can cope with such data heterogeneity are needed. The previously developed annotation pipelines BRAKER1 and BRAKER2 use RNA-seq or protein data, respectively, but not both. A further significant performance improvement integrating all three data types was made by the recently released GeneMark-ETP. We here present the BRAKER3 pipeline that builds on GeneMark-ETP and AUGUSTUS, and further improves accuracy using the TSEBRA combiner. BRAKER3 annotates protein-coding genes in eukaryotic genomes using both short-read RNA-seq and a large protein database, along with statistical models learned iteratively and specifically for the target genome. We benchmarked the new pipeline on genomes of 11 species under an assumed level of relatedness of the target species proteome to available proteomes. BRAKER3 outperforms BRAKER1 and BRAKER2. The average transcript-level F1-score is increased by about 20 percentage points on average, whereas the difference is most pronounced for species with large and complex genomes. BRAKER3 also outperforms other existing tools, MAKER2, Funannotate, and FINDER. The code of BRAKER3 is available on GitHub and as a ready-to-run Docker container for execution with Docker or Singularity. Overall, BRAKER3 is an accurate, easy-to-use tool for eukaryotic genome annotation.

59 BASIC BIOLOGICAL SCIENCES↗

Mass of 101 Sn and Bayesian extrapolations to the proton drip line

The favorable energy configurations of nuclei at magic numbers of 𝑁 neutrons and 𝑍 protons are fundamental for understanding the evolution of nuclear structure. The 𝑍 = 50 (tin) isotopic chain is a frontier for such studies, with particular interest at and around the doubly magic 100 Sn isotope, for which the mass is a topic of debate. Precise mass values for neutron-deficient isotopes provide necessary anchor points for mass models to test extrapolations near the proton drip line, where experimental studies remain out of reach. In this work, we report a Penning trap mass measurement of 101 Sn . The determined mass excess of −59889.89⁢(96) keV for 101 Sn represents a factor-of-300 improvement over the current precision and indicates that 101 Sn is less bound than previously thought. Mass predictions from a recently developed Bayesian model combination framework employing statistical machine learning and nuclear masses computed within seven global models based on nuclear density functional theory agree within 1⁢𝜎 with experimental masses from the 48 ≤ 𝑍 ≤ 52 isotopic chains. The framework's resilience to new mass data gave confidence in the extrapolation of tin masses down to 𝑁 = 46. Our calculations suggest that 96 Sn is a two-proton drip line nucleus and predict a mass excess of −58090⁢(800) keV for 100 Sn , showing a preference within 1⁢𝜎 for the mass of 100 Sn derived from the 𝛽-delayed 𝑄 value measured at GSI.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Rolling Root Mean Square Based Multimodal Anomaly Detection for Real Time Monitoring of Smart Grid

Reliable real-time monitoring is valuable for maintaining the operational integrity of modern electrical smart grids. Deployment of heterogeneous sensing technologies in substations has enabled high-resolution, multichannel waveform monitoring, but also introduces challenges for anomaly detection due to noise, baseline drift, and modality-dependent signal characteristics. In this work, we present a computationally efficient unsupervised method for multimodal event detection based on Rolling Root Mean Square based Event Detection (RRMSED). The method is developed using in-house, field deployed sensors collecting data at a utility substation. The sensing system comprises voltage and current sensors, triaxial accelerometers, and magnetometers, collectively capturing electrical, vibrational, and magnetic waveform measurements at high temporal resolution. RRMSED operates by extracting rolling RMS energy features and their first-order temporal differences from consecutive waveform segments for each channel and then applying channel-specific statistical thresholds learned from historical data. A persistence-based exceedance logic is employed to robustly identify transient events while suppressing impulsive noise, and to provide precise temporal localization with high resolution. The framework is designed for continuous server-side operation and can be deployed in real time without requiring complex models. Experiments on simulated waveform data with known ground truth demonstrate low false positive (FP) and false negative (FN) rates. Application to real substation data shows RRMSED to identify events that are not captured by conventional monitoring indicators including fast transient detection algorithm currently deployed in the system. These results indicate that rolling RMS based features provide an effective and practical basis for real-time multimodal event detection in smart-grid substations.

Mukherjee, Subrata [ORNL] (ORCID:0000000309930338)↗

Advanced Signal Decomposition Analysis and Anomaly Detection in Photovoltaic Systems

With the rapid expansion of large-scale photovoltaic (PV) plants, it is paramount for solar stakeholders to understand the reliability and efficiency of their plants to inform maintenance decisions, increase production, and understand the design factors that impact performance. Diagnosing underperformance in PV plants is challenging due to the relatively few monitoring points with respect to the large geographic footprint of the plant. This work introduces a cutting-edge method that transforms the analysis and management of key factors influencing PV plant performance, including performance loss rate (PLR), recoverable soiling, and major system changes. Identifying these factors is critical for deriving actionable insights. Leveraging advanced analytical techniques such as wavelet transformation, robust regression, and extreme point analysis, this approach provides a nuanced understanding of these factors. This method has been tested across two synthetic datasets and one real dataset, consistently surpassing existing benchmarks by achieving a lower median mean absolute error and reduced error variability across all comparable components.

14 SOLAR ENERGY↗

How Can Probabilistic Solar Power Forecasts Be Used to Lower Costs and Improve Reliability in Power Spot Markets? A Review and Application to Flexiramp Requirements

Net load uncertainty in electricity spot markets is rapidly growing. There are five general approaches by which system operators and market participants can use probabilistic forecasts of wind, solar, and load to help manage this uncertainty. These include operator situation awareness, resource risk hedging, reserves procurement, definition of contingencies, and explicit stochastic optimization. We review these approaches, and then provide a case study in which a method for using probabilistic solar forecasts to define needs for reserves is developed and evaluated. The case study has three parts. First, we describe building blocks for enhancing the Watt-Sun solar forecasting system to produce probabilistic irradiance and power forecasts. Second, relationships between Watt-Sun forecasts for multiple sites in California and the system's need for flexible ramp capability (flexiramp) are defined by machine learning and statistical methods. Third, the performance of present methods to defining flexiramp requirements, which are not conditioned on weather and renewables forecasts, is compared with that of probabilistic solar forecast-based requirements, using a multi-timescale production costing model with an 1820-bus representation of the WECC power system. Significant potential savings in fuel and flexiramp procurement costs from using solar-informed reserve requirements are found.

14 SOLAR ENERGY↗

Anticipating Technical Expertise and Capability Evolution in Research Communities Using Dynamic Graph Transformers

The ability to anticipate global technical expertise and capability evolution trends is essential for national and global security, especially in safety-critical domains such as nuclear nonproliferation (NN) and rapidly emerging fields like artificial intelligence (AI). Here, in this work, we extend traditional statistical relational learning approaches (e.g., link prediction in collaboration networks) and formulate a problem of anticipating technical expertise and capability evolution using dynamic heterogeneous graph representations. We develop novel capabilities to forecast collaboration patterns, authorship behavior, and technical capability evolution at different granularities (e.g., scientist and institution levels) in two distinct research fields. We implement a dynamic graph transformer (DGT) neural architecture, which pushes the state-of-the-art graph neural network models by: 1) forecasting heterogeneous (rather than homogeneous) nodes and edges; and 2) relying on both discrete- and continuous-time inputs. We demonstrate that our DGT models predict collaboration, partnership, and expertise patterns with 0.26, 0.73, and 0.53 mean reciprocal rank values for AI and 0.48, 0.93, and 0.22 for NN domains. DGT model performance exceeds the best-performing static graph baseline models by 30%–80% across AI and NN domains. Our findings demonstrate that DGT models boost inductive task performance when previously unseen nodes appear in the test data for the domains with emerging collaboration patterns (e.g., AI). Specifically, models accurately predict which established scientists will collaborate with early career scientists and vice versa in the AI domain.

97 MATHEMATICS AND COMPUTING↗

Stochastic Gradient-Based Distributed Bayesian Estimation in Cooperative Sensor Networks

Distributed Bayesian inference provides a full quantification of uncertainty offering numerous advantages over point estimates that autonomous sensor networks are able to exploit. However, fully-decentralized Bayesian inference often requires large communication overheads and low network latency, resources that are not typically available in practical applications. In this paper, we propose a decentralized Bayesian inference approach based on stochastic gradient Langevin dynamics, which produces full posterior distributions at each of the nodes with significantly lower communication overhead. We provide analytical results on convergence of the proposed distributed algorithm to the centralized posterior, under typical network constraints. Finally, we also provide extensive simulation results to demonstrate the validity of the proposed approach.

42 ENGINEERING↗

Machine learning tools for epigenetics

The software provides machine learning analysis and visualization to detect patterns in epigenetic data, including conventional machine learning and statistical methods, and open-source packages like pyBigWig (https://github.com/deeptools/pyBigWig) for data processing. The software is written in python, it uses some python libraries.

Kim, Anastasiia↗

Portable, heterogeneous ensemble workflows at scale using libEnsemble

libEnsemble is a Python-based toolkit for running dynamic ensembles, developed as part of the DOE Exascale Computing Project. The toolkit utilizes a unique generator–simulator–allocator paradigm, where generators produce input for simulators, simulators evaluate those inputs, and allocators decide whether and when a simulator or generator should be called. The generator steers the ensemble based on simulation results. Generators may, for example, apply methods for numerical optimization, machine learning, or statistical calibration. libEnsemble communicates between a manager and workers. Flexibility is provided through multiple manager–worker communication substrates each of which has different benefits. These include Python’s multiprocessing, mpi4py, and TCP. Multisite ensembles are supported using Balsam or Globus Compute. We overview the unique characteristics of libEnsemble as well as current and potential interoperability with other packages in the workflow ecosystem. We highlight libEnsemble’s dynamic resource features: libEnsemble can detect system resources, such as available nodes, cores, and GPUs, and assign these in a portable way. These features allow users to specify the number of processors and GPUs required for each simulation; and resources will be automatically assigned on a wide range of systems, including Frontier, Aurora, and Perlmutter. Such ensembles can include multiple simulation types, some using GPUs and others using only CPUs, sharing nodes for maximum efficiency. We also describe the benefits of libEnsemble’s generator–simulator coupling, which easily exposes to the user the ability to cancel, and portably kill, running simulations based on models that are updated with intermediate simulation output. We demonstrate libEnsemble’s capabilities, scalability, and scientific impact via a Gaussian process surrogate training problem for the longitudinal density profile at the exit of a plasma accelerator stage. In conclusion, the study uses gpCAM for the surrogate model and employs either Wake-T or WarpX simulations, highlighting efficient use of resources that can easily extend to exascale.

Dynamic ensembles↗

Roundup causes embryonic development failure and alters metabolic pathways and gut microbiota functionality in non-target species

Background: Research around the weedkiller Roundup is among the most contentious of the twenty-first century. Scientists have provided inconclusive evidence that the weedkiller causes cancer and other life-threatening diseases, while industry-paid research reports that the weedkiller has no adverse effect on humans or animals. Much of the controversial evidence on Roundup is rooted in the approach used to determine safe use of chemicals, defined by outdated toxicity tests. We apply a system biology approach to the biomedical and ecological model species Daphnia to quantify the impact of glyphosate and of its commercial formula, Roundup, on fitness, genome-wide transcription and gut microbiota, taking full advantage of clonal reproduction in Daphnia. We then apply machine learning-based statistical analysis to identify and prioritize correlations between genome-wide transcriptional and microbiota changes. Results: We demonstrate that chronic exposure to ecologically relevant concentrations of glyphosate and Roundup at the approved regulatory threshold for drinking water in the US induce embryonic developmental failure, induce significant DNA damage (genotoxicity), and interfere with signaling. Furthermore, chronic exposure to the weedkiller alters the gut microbiota functionality and composition interfering with carbon and fat metabolism, as well as homeostasis. Using the “Reactome,” we identify conserved pathways across the Tree of Life, which are potential targets for Roundup in other species, including liver metabolism, inflammation pathways, and collagen degradation, responsible for the repair of wounds and tissue remodeling. Conclusions: Our results show that chronic exposure to concentrations of Roundup and glyphosate at the approved regulatory threshold for drinking water causes embryonic development failure and alteration of key metabolic functions via direct effect on the host molecular processes and indirect effect on the gut microbiota. The ecological model species Daphnia occupies a central position in the food web of aquatic ecosystems, being the preferred food of small vertebrates and invertebrates as well as a grazer of algae and bacteria. The impact of the weedkiller on this keystone species has cascading effects on aquatic food webs, affecting their ability to deliver critical ecosystem services.

59 BASIC BIOLOGICAL SCIENCES↗

Research Needs for Trusted Analytics in National Security Settings

As artificial intelligence, machine learning, and statistical modeling methods become commonplace in national security applications, the drive to create trusted analytics becomes increasingly important. The goal of this report is to identify areas of research that can provide the foundational understanding and technical prerequisites for the development and deployment of trusted analytics in national security settings. Our review of the literature covered several disjoint research communities, including computer science, statistics, human factors, and several branches of psychology and cognitive science, which tend not to interact with one another or cite each other's literatures. As a result, there exists no agreed-upon theoretical framework for understanding how various factors influence trust and no well-established empirical paradigm for studying these effects. This report therefore takes three steps. First, we define several key terms in an effort to provide a unifying language for trusted analytics and to manage the scope of the problem. Second, we outline an empirical perspective that identifies key independent, moderating, and dependent variables in assessing trusted analytics. Though not a substitute for a theoretical framework, the empirical perspective does support research and development of trusted analytics in the national security domain. Finally, we discuss several research gaps relevant to developing trusted analytics for the national security mission space.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Evaluating Offshore Infrastructure Integrity

Drilling in the offshore environment involves a complex network of infrastructure including pipelines, platforms, rigs, subsea installations, ports, and terminals. Government and industry partners have developed this network over many decades and it remains a critical part of the United States (U.S.) energy portfolio. Many of the major components of this system have been designed with a 20- to 30-year lifespan, yet consistent and growing energy demands support the need to extend the design life of existing infrastructure or repurpose it for secondary needs (i.e. enhanced oil recovery, carbon storage, and new wells). As a result, a growing portion of the offshore infrastructure in the U.S. is approaching or has exceeded its original design life. A critical step in ensuring the continued safe and effective operation of offshore infrastructure is developing a comprehensive understanding of the state of offshore infrastructure and the factors that effect it. The purpose of this project is to assess the current state of existing infrastructure and identify the factors involved in infrastructure degradation through the development and application of big data analytics, machine learning, and advanced spatio-temporal analysis. The project leverages existing data at NETL and combines it with new information on offshore oil and gas structures and the ambient offshore environment in an effort to identify patterns associated with infrastructure integrity. Building on the identified trends and patterns, this project incorporates exploratory analytics and spatial analysis tools in conjunction with machine learning and statistical models to characterize the condition of existing platforms in the offshore environment and predict their risk of failure.

02 PETROLEUM↗

Physics-Informed Learning Machines for Multiscale and Multiphysics Problems (PHILMS) (Technical Report)

The research work at University of California Santa Barbara (UCSB) resulted in several new developments in the areas of scientific machine learning, numerical analysis, and practical methods for data-driven modeling, prediction, reductions, and simulation. Many of the projects were carried out in collaboration with members of the national laboratories at Sandia National Laboratories (SNL), Pacific Northwestern National Laboratories (PNNL), and other institutions. Results included developing new scientific machine learning methods, related theory and mathematical frameworks for analysis and training, data-driven numerical solvers, and related tools and software for scientific computation. During the support period, over 16+ papers were submitted for publication, and 4 open-source software packages were developed and released (available at http://atzberger.org/). In addition, 7+ students and 2 post-docs were mentored in collaboration with the laboratory staff for future careers in academia, government labs, and industry.

97 MATHEMATICS AND COMPUTING↗

Materials Characterization, Prediction and Control Project: Summary Report on Data Analytics Framework

This report summarizes the activities performed under the data analytics Vertex in the Materials Characterization, Prediction and Control Project funded under laboratory directed research and development at Pacific Northwest National Laboratory. The data analytics Vertex developed models for associating global or local process parameters, microstructural features, and performance properties of friction-stir-processed 316L stainless steel plates. Statistical, machine learning, and deep learning models, as well as generative artificial intelligence approaches, were used to develop the associations between the process-structure-property data streams. These associations formed the basis for predicting global properties of parts manufactured under different process envelopes, providing a basis for predicting performance using data driven as well as physics-informed and physics-constrained approaches. Additionally, the associations were used to predict local process parameters and microstructural features of the product, predictive relationships that have the potential to form the basis of a control framework that could eventually modulate a friction-stir process to maintain product quality.

316L stainless steel↗