Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “statistical learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

A deep generative model enables automated structure elucidation of novel psychoactive substances

Over the past decade, the illicit drug market has been reshaped by the proliferation of clandestinely produced designer drugs. These agents, referred to as new psychoactive substances (NPSs), are designed to mimic the physiological actions of better-known drugs of abuse while skirting drug control laws. The public health burden of NPS abuse obliges toxicological, police and customs laboratories to screen for them in law enforcement seizures and biological samples. However, the identification of emerging NPSs is challenging due to the chemical diversity of these substances and the fleeting nature of their appearance on the illicit market. Here, in this study, we present DarkNPS, a deep learning-enabled approach to automatically elucidate the structures of unidentified designer drugs using only mass spectrometric data. Our method employs a deep generative model to learn a statistical probability distribution over unobserved structures, which we term the structural prior. We show that the structural prior allows DarkNPS to elucidate the exact chemical structure of an unidentified NPS with an accuracy of 51% and a top-10 accuracy of 86%. Our generative approach has the potential to enable de novo structure elucidation for other types of small molecules that are routinely analysed by mass spectrometry.

Cheminformatics↗

Model orthogonalization and Bayesian forecast mixing via principal component analysis

One can improve predictability in the unknown domain by combining forecasts of imperfect complex computational models using a Bayesian statistical machine learning framework. In many cases, however, the models used in the mixing process are similar. In addition to contaminating the model space, the existence of such similar, or even redundant, models during the multimodeling process can result in misinterpretation of results and deterioration of predictive performance. In this paper we describe a method based on the principal component analysis that eliminates model redundancy. We show that by adding model orthogonalization to the proposed Bayesian model combination framework, one can arrive at better prediction accuracy and reach excellent uncertainty quantification performance.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Neural networks for parameter estimation in intractable models

The goal is to use deep learning models to estimate parameters in statistical models when standard likelihood estimation methods are computationally infeasible. For instance, inference for max-stable processes is exceptionally challenging even with small datasets, but simulation is straightforward. Data from model simulations are used to train deep neural networks and learn statistical parameters from max-stable models. The proposed neural network-based method provides a competitive alternative to current approaches, as demonstrated by considerable accuracy and computational time improvements. Finally, it serves as a proof of concept for deep learning in statistical parameter estimation and can be extended to other estimation problems.

97 MATHEMATICS AND COMPUTING↗

Density estimation via measure transport: Outlook for applications in the biological sciences

Abstract One among several advantages of measure transport methods is that they allow or a unified framework for processing and analysis of data distributed according to a wide class of probability measures. Within this context, we present results from computational studies aimed at assessing the potential of measure transport techniques, specifically, the use of triangular transport maps, as part of a workflow intended to support research in the biological sciences. Scenarios characterized by the availability of limited amount of sample data, which are common in domains such as radiation biology, are of particular interest. We find that when estimating a distribution density function given limited amount of sample data, adaptive transport maps are advantageous. In particular, statistics gathered from computing series of adaptive transport maps, trained on a series of randomly chosen subsets of the set of available data samples, leads to uncovering information hidden in the data. As a result, in the radiation biology application considered here, this approach provides a tool for generating hypotheses about gene relationships and their dynamics under radiation exposure.

gene expression data↗

An adaptive adversarial domain adaptation approach for corn yield prediction

Recently, statistical machine learning and deep learning methods have been widely explored for corn yield prediction. Though successful, machine learning models generated within a specific spatial domain often lose their validity when directly applied to new regions. To address this issue, we designed an unsupervised adaptive domain adversarial neural network (ADANN). Specifically, through domain adversarial training, the ADANN model reduced the impact of domain shift by projecting data from different domains into the same subspace. Also, the ADANN model was designed to be trained in an adaptive way, which guaranteed the model can learn the domain-invariant features and perform accurate yield prediction simultaneously. Informative variables including time-series vegetation indices and sequential weather observations were first collected from multiple data sources and aggregated to the county level. Then, we trained the ADANN model with the extracted features and corresponding reported county-level corn yield from the U.S. Department of Agriculture (USDA). Finally, the trained model was evaluated in four testing years 2016–2019. The U.S. corn belt was used as the study area and counties under study were grouped into two diverse ecological regions. Overall, the experimental results showed that the developed ADANN model had better performance than three other state-of-the-art machine learning models in both local experiments (train and test in the same region) and transfer experiments (train and test in different regions). As the first study using adversarial learning for crop yield prediction, this research demonstrates a novel solution for improving model transferability on crop yield prediction.

59 BASIC BIOLOGICAL SCIENCES↗

Accelerating Markov Chain Monte Carlo sampling with diffusion models

Global fits of physics models require efficient methods for exploring high-dimensional and/or multimodal posterior functions. We introduce a novel method for accelerating Markov Chain Monte Carlo (MCMC) sampling by pairing a Metropolis-Hastings algorithm with a diffusion model that can draw global samples with the aim of approximating the posterior. We briefly review diffusion models in the context of image synthesis before providing a streamlined diffusion model tailored towards low-dimensional data arrays. We then present our adapted Metropolis-Hastings algorithm which combines local proposals with global proposals taken from a diffusion model that is regularly trained on the samples produced during the MCMC run. Our approach leads to a significant reduction in the number of likelihood evaluations required to obtain an accurate representation of the Bayesian posterior across several analytic functions, as well as for a physical example based on a global fit of parton distribution functions. Our method is extensible to other MCMC techniques, and we briefly compare our method to similar approaches based on normalising flows. A code implementation can be found at https://github.com/NickHunt-Smith/MCMC-diffusion.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

An overview of visualization and visual analytics applications in water resources management

Recent advances in information, communication, and environmental monitoring technologies have increased the availability, spatiotemporal resolution, and quality of water-related data, thereby leading to the emergence of many innovative big data applications. Among these applications, visualization and visual analytics, also known as the visual computing techniques, empower the synergy of computational methods (e.g., machine learning and statistical models) with human reasoning to improve the understanding and solution toward complex science and engineering problems. These approaches are frequently integrated with geographic information systems and cyberinfrastructure to provide new opportunities and methods for enhancing water resources management. Here, we present a comprehensive review of recent hydroinformatics applications that employ visual computing techniques to (1) support complex data-driven research problems, and (2) support the communication and decision-makings in the water resources management sector. Then, we conduct a technical review of the state-of-the-art web-based visualization technologies and libraries to share our experiences on developing shareable, adaptive, and interactive visualizations and visual interfaces for water resources management applications. We close with a vision that applies the emerging visual computing technologies and paradigms to develop the next generation of hydroinformatics applications.

54 ENVIRONMENTAL SCIENCES↗

Characterizing soil water content variability across spatial scales from optimized high-resolution distributed temperature sensing technique

Fiber-optic Distributed Temperature Sensing, when combined with the Single-probe Heat-pulse technique can measure soil moisture (θ) across spatial scales. The key limitation of this system is in obtaining the relationship between soil thermal conductivity (λ) and θ for a specific field. Using the Department of Energy Atmospheric Radiation Measurement (ARM) site, this study tested a new methodology to account for the spatial variability in the λ-θ relationship using a Gaussian processes model. The resulting accurate θ measurements (RMSE = 0.03 m 3 m –3 ) were used to characterize the spatial variability of θ across scales and to develop an empirical equation that can correct for the changes in the θ spatial variability observed at different spatial resolutions. In addition, the number of required samples to accurately characterize θ and its variability over scales ranging from 5 m and 350 m were estimated. Finally, these findings provide key information to scale soil moisture from centimeters to hundreds of meters for process understanding.

54 ENVIRONMENTAL SCIENCES↗

Knowledge graph-aided Bayesian active learning for top- K genetic interaction discovery

In silico methods for predicting the effects of multi-gene perturbations hold great promise for advancing functional genomics, computational drug discovery, and disease modeling. However, the development of these predictive algorithms for mammalian systems has been hampered by limited datasets and high experimental costs. In this study, we present a Bayesian active learning framework designed to discover pairwise host gene knockdowns that effectively inhibit viral proliferation in an in vitro HIV-1 infection model. Our method leverages a biological knowledge graph as side information and employs a computationally efficient batch diversification approach. We evaluated this framework using a dataset of viral load measurements obtained from multi-day dual-gene depletion experiments, encompassing all possible pairwise knockdowns of over 350 host genes associated with HIV infection. We demonstrate that our framework rapidly identifies the most effective gene knockdown pairs for reducing viral load. Furthermore, we show that incorporating side information enhances performance during the early stages of active learning (low data regime), while our batch diversification strategy significantly boosts performance in later stages (high data regime). This framework is general and can be adapted to explore gene interactions in other contexts, such as synthetic lethality prediction and mapping epistatic effects across quantitative trait loci.

Computational biology and bioinformatics↗

Revisiting trends in the exchange current for hydrogen evolution

Nørskov and collaborators proposed a simple kinetic model to explain the volcano relation for the hydrogen evolution reaction on transition metal surfaces such that j 0 = k 0 f(ΔG H ) where j 0 is the exchange current density, f(ΔG H ) is a function of the hydrogen adsorption free energy ΔG H as computed from density functional theory, and k 0 is a universal rate constant. Herein, focusing on the hydrogen evolution reaction in acidic medium, we revisit the original experimental data and find that the fidelity of this kinetic model can be significantly improved by invoking metal-dependence on k 0 such that the logarithm of k 0 linearly depends on the absolute value of ΔG H . Here, we further confirm this relationship using additional experimental data points obtained from a critical review of the available literature. Our analyses show that the new model decreases the discrepancy between calculated and experimental exchange current density values by up to four orders of magnitude. Furthermore, we show the model can be further improved using machine learning and statistical inference methods that integrate additional material properties.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

BRAKER3: Fully automated genome annotation using RNA-seq and protein evidence with GeneMark-ETP, AUGUSTUS, and TSEBRA

Gene prediction has remained an active area of bioinformatics research for a long time. Still, gene prediction in large eukaryotic genomes presents a challenge that must be addressed by new algorithms. The amount and significance of the evidence available from transcriptomes and proteomes vary across genomes, between genes, and even along a single gene. User-friendly and accurate annotation pipelines that can cope with such data heterogeneity are needed. The previously developed annotation pipelines BRAKER1 and BRAKER2 use RNA-seq or protein data, respectively, but not both. A further significant performance improvement integrating all three data types was made by the recently released GeneMark-ETP. We here present the BRAKER3 pipeline that builds on GeneMark-ETP and AUGUSTUS, and further improves accuracy using the TSEBRA combiner. BRAKER3 annotates protein-coding genes in eukaryotic genomes using both short-read RNA-seq and a large protein database, along with statistical models learned iteratively and specifically for the target genome. We benchmarked the new pipeline on genomes of 11 species under an assumed level of relatedness of the target species proteome to available proteomes. BRAKER3 outperforms BRAKER1 and BRAKER2. The average transcript-level F1-score is increased by about 20 percentage points on average, whereas the difference is most pronounced for species with large and complex genomes. BRAKER3 also outperforms other existing tools, MAKER2, Funannotate, and FINDER. The code of BRAKER3 is available on GitHub and as a ready-to-run Docker container for execution with Docker or Singularity. Overall, BRAKER3 is an accurate, easy-to-use tool for eukaryotic genome annotation.

59 BASIC BIOLOGICAL SCIENCES↗

Mass of 101 Sn and Bayesian extrapolations to the proton drip line

The favorable energy configurations of nuclei at magic numbers of 𝑁 neutrons and 𝑍 protons are fundamental for understanding the evolution of nuclear structure. The 𝑍 = 50 (tin) isotopic chain is a frontier for such studies, with particular interest at and around the doubly magic 100 Sn isotope, for which the mass is a topic of debate. Precise mass values for neutron-deficient isotopes provide necessary anchor points for mass models to test extrapolations near the proton drip line, where experimental studies remain out of reach. In this work, we report a Penning trap mass measurement of 101 Sn . The determined mass excess of −59889.89⁢(96) keV for 101 Sn represents a factor-of-300 improvement over the current precision and indicates that 101 Sn is less bound than previously thought. Mass predictions from a recently developed Bayesian model combination framework employing statistical machine learning and nuclear masses computed within seven global models based on nuclear density functional theory agree within 1⁢𝜎 with experimental masses from the 48 ≤ 𝑍 ≤ 52 isotopic chains. The framework's resilience to new mass data gave confidence in the extrapolation of tin masses down to 𝑁 = 46. Our calculations suggest that 96 Sn is a two-proton drip line nucleus and predict a mass excess of −58090⁢(800) keV for 100 Sn , showing a preference within 1⁢𝜎 for the mass of 100 Sn derived from the 𝛽-delayed 𝑄 value measured at GSI.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Rolling Root Mean Square Based Multimodal Anomaly Detection for Real Time Monitoring of Smart Grid

Reliable real-time monitoring is valuable for maintaining the operational integrity of modern electrical smart grids. Deployment of heterogeneous sensing technologies in substations has enabled high-resolution, multichannel waveform monitoring, but also introduces challenges for anomaly detection due to noise, baseline drift, and modality-dependent signal characteristics. In this work, we present a computationally efficient unsupervised method for multimodal event detection based on Rolling Root Mean Square based Event Detection (RRMSED). The method is developed using in-house, field deployed sensors collecting data at a utility substation. The sensing system comprises voltage and current sensors, triaxial accelerometers, and magnetometers, collectively capturing electrical, vibrational, and magnetic waveform measurements at high temporal resolution. RRMSED operates by extracting rolling RMS energy features and their first-order temporal differences from consecutive waveform segments for each channel and then applying channel-specific statistical thresholds learned from historical data. A persistence-based exceedance logic is employed to robustly identify transient events while suppressing impulsive noise, and to provide precise temporal localization with high resolution. The framework is designed for continuous server-side operation and can be deployed in real time without requiring complex models. Experiments on simulated waveform data with known ground truth demonstrate low false positive (FP) and false negative (FN) rates. Application to real substation data shows RRMSED to identify events that are not captured by conventional monitoring indicators including fast transient detection algorithm currently deployed in the system. These results indicate that rolling RMS based features provide an effective and practical basis for real-time multimodal event detection in smart-grid substations.

Mukherjee, Subrata [ORNL] (ORCID:0000000309930338)↗

Advanced Signal Decomposition Analysis and Anomaly Detection in Photovoltaic Systems

With the rapid expansion of large-scale photovoltaic (PV) plants, it is paramount for solar stakeholders to understand the reliability and efficiency of their plants to inform maintenance decisions, increase production, and understand the design factors that impact performance. Diagnosing underperformance in PV plants is challenging due to the relatively few monitoring points with respect to the large geographic footprint of the plant. This work introduces a cutting-edge method that transforms the analysis and management of key factors influencing PV plant performance, including performance loss rate (PLR), recoverable soiling, and major system changes. Identifying these factors is critical for deriving actionable insights. Leveraging advanced analytical techniques such as wavelet transformation, robust regression, and extreme point analysis, this approach provides a nuanced understanding of these factors. This method has been tested across two synthetic datasets and one real dataset, consistently surpassing existing benchmarks by achieving a lower median mean absolute error and reduced error variability across all comparable components.

14 SOLAR ENERGY↗

How Can Probabilistic Solar Power Forecasts Be Used to Lower Costs and Improve Reliability in Power Spot Markets? A Review and Application to Flexiramp Requirements

Net load uncertainty in electricity spot markets is rapidly growing. There are five general approaches by which system operators and market participants can use probabilistic forecasts of wind, solar, and load to help manage this uncertainty. These include operator situation awareness, resource risk hedging, reserves procurement, definition of contingencies, and explicit stochastic optimization. We review these approaches, and then provide a case study in which a method for using probabilistic solar forecasts to define needs for reserves is developed and evaluated. The case study has three parts. First, we describe building blocks for enhancing the Watt-Sun solar forecasting system to produce probabilistic irradiance and power forecasts. Second, relationships between Watt-Sun forecasts for multiple sites in California and the system's need for flexible ramp capability (flexiramp) are defined by machine learning and statistical methods. Third, the performance of present methods to defining flexiramp requirements, which are not conditioned on weather and renewables forecasts, is compared with that of probabilistic solar forecast-based requirements, using a multi-timescale production costing model with an 1820-bus representation of the WECC power system. Significant potential savings in fuel and flexiramp procurement costs from using solar-informed reserve requirements are found.

14 SOLAR ENERGY↗

Anticipating Technical Expertise and Capability Evolution in Research Communities Using Dynamic Graph Transformers

The ability to anticipate global technical expertise and capability evolution trends is essential for national and global security, especially in safety-critical domains such as nuclear nonproliferation (NN) and rapidly emerging fields like artificial intelligence (AI). Here, in this work, we extend traditional statistical relational learning approaches (e.g., link prediction in collaboration networks) and formulate a problem of anticipating technical expertise and capability evolution using dynamic heterogeneous graph representations. We develop novel capabilities to forecast collaboration patterns, authorship behavior, and technical capability evolution at different granularities (e.g., scientist and institution levels) in two distinct research fields. We implement a dynamic graph transformer (DGT) neural architecture, which pushes the state-of-the-art graph neural network models by: 1) forecasting heterogeneous (rather than homogeneous) nodes and edges; and 2) relying on both discrete- and continuous-time inputs. We demonstrate that our DGT models predict collaboration, partnership, and expertise patterns with 0.26, 0.73, and 0.53 mean reciprocal rank values for AI and 0.48, 0.93, and 0.22 for NN domains. DGT model performance exceeds the best-performing static graph baseline models by 30%–80% across AI and NN domains. Our findings demonstrate that DGT models boost inductive task performance when previously unseen nodes appear in the test data for the domains with emerging collaboration patterns (e.g., AI). Specifically, models accurately predict which established scientists will collaborate with early career scientists and vice versa in the AI domain.

97 MATHEMATICS AND COMPUTING↗

Machine learning tools for epigenetics

The software provides machine learning analysis and visualization to detect patterns in epigenetic data, including conventional machine learning and statistical methods, and open-source packages like pyBigWig (https://github.com/deeptools/pyBigWig) for data processing. The software is written in python, it uses some python libraries.

Kim, Anastasiia↗