Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning competition”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Optimizing enzymes for plastic upcycling using machine learning design and high throughput experiments

Plastic use is ubiquitous in the modern world, and polyethylene terephthalate (PET) is one of the most abundantly produced plastics (and the most highly produced polyester), with ~65 million metric tons manufactured annually. To the consumer, PET is likely most recognizable as the plastic used to make beverage bottles. Like many plastics, traditional mechanical or chemical means of PET deconstruction and upcycling are costly and inefficient. Because of these challenges, recycled plastic is generally of lower quality and is more expensive to produce than virgin plastic derived from petroleum. Ultimately, this results in most plastic ending up as waste. We view plastic waste as an underutilized resource which, with the development of more efficient and high-quality recycling processes, could (1) generate significant economic value while (2) decreasing petroleum usage and greenhouse gas emissions, as well as (3) minimizing its negative environmental and health impacts. Biocatalytic recycling, or biomanufacturing the basic building blocks of new plastic from plastic waste, is a promising approach to plastic reuse that complements existing recycling technologies. Recently, biological enzymes capable of breaking down PET have garnered significant attention as an attractive means of dealing with the plastic problem. These enzymes are currently undergoing pilot studies for implementation in industrial-scale enzyme-based recycling. However, there are significant limitations to current enzymes, including the need to perform costly pre-processing of the plastic waste before the enzymes are able to work. Further optimization of these enzymes is necessary to make these technologies competitive, and ultimately incentivise industry-wide adoption of this biology-based green recycling technology. n this work we demonstrate a means to design and generate performant biological enzymes, capable of efficiently deconstructing plastic waste. Specifically, we applied recent advances in artificial intelligence, machine learning, and statistical analysis to design new versions and discover natural enzymes capable of breaking down PET. We focused on optimizing key properties that are important for industrial-scale enzymatic recycling such as pH and thermotolerance. Normal testing of enzymatic plastic-deconstruction is extremely labor intensive and so through this work we also developed a robotic-assisted experimental pipeline capable of characterizing thousands of candidate enzymes. The results of this iterative, AI-guided, multi-discipline approach have led to increases in enzymatic breakdown of over 150X over starting enzymes. This work supports the rapidly developing and transformative field of biocatalytic solutions to environmental problems beyond the discovery and predictive understanding of enzymes for polymer recycling, and has wide implications for tackling numerous energy problems such as carbon capture and fixation (e.g., engineering carbon monoxide dehydrogenase and the rubisco-pathway), biomining (e.g., design of lanthanide-binding proteins) and biomanufacturing (e.g., lignin-deconstruction enzymes).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

A Generalized Transformer-Based Pulse Detection Algorithm

Pulse-like signals are ubiquitous in the field of single molecule analysis, e.g., electrical or optical pulses caused by analyte translocations in nanopores. The primary challenge in processing pulse-like signals is to capture the pulses in noisy backgrounds, but current methods are subjectively based on a user-defined threshold for pulse recognition. Here, we propose a generalized machine-learning based method, named pulse detection transformer (PETR), for pulse detection. PETR determines the start and end time points of individual pulses, thereby singling out pulse segments in a time-sequential trace. It is objective without needing to specify any threshold. It provides a generalized interface for downstream algorithms for specific application scenarios. PETR is validated using both simulated and experimental nanopore translocation data. It returns a competitive performance in detecting pulses through assessing them with several standard metrics. Finally, the generalization nature of the PETR output is demonstrated using two representative algorithms for feature extraction.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Robust Power System Stability Assessment Against Adversarial Machine Learning-Based Cyberattacks via Online Purification

The increasing complexity associated with renewable generation brings more challenges to power system stability assessment (SA). Data-driven approaches based on machine learning (ML) techniques for stability assessment have received significant research interest and shown their promising performance. However, ML-based models are recognized to be vulnerable to adversarial disturbances, where a slight perturbation to power system measurements could lead to unacceptable errors. To address this issue, this paper develops a novel lightweight mitigation strategy, i.e., robust online stability assessment (ROSA), to enhance the ML-based assessment model against both white-box and the black-box adversarial disturbances (i.e., purification) in the online implementation. The ROSA involves a supervised learning-based module for the primary stability assessment and a self-supervised learning-based module. Further, the two modules are trained jointly with different objective (loss) functions and implemented in sequence. A suitable purification objective and various time-series data augmentation methods are designed for SA applications to tackle adversarial disturbances adaptively. Case studies are performed, and the comparative results have clearly illustrated the competitive, robust accuracy against various adversarial scenarios and verified the effectiveness of the proposed online purification strategy.

24 POWER TRANSMISSION AND DISTRIBUTION↗

A Better Method to Calculate Fuel Burnup in Pebble Bed Reactors Using Machine Learning

Burnup measurement is an important step in material control and accountancy (MC&A) at nuclear reactors, and may be done by examining gamma spectra of fuel samples. Traditional approaches rely on known correlations to specific photopeaks (e.g. 137 Cs) and operate via a standard linear regression method. However, the quality of these regression methods is limited even in the best case, and is significantly poorer at short fuel cool-down times, due to the elevated radiation background by short life-time isotopes, and self-shielding effect of the fuel. For practical operation of pebble bed reactors (PBRs), quick measurements (in minutes) and short cooling times (in hours) are required from a safety and security perspective. We investigated the efficacy and performance of machine learning (ML) methods to predict the burnup of the pebble fuel from full gamma spectra (rather than specific discrete photopeaks) and found a full-spectrum ML approach to far outperform baseline regression predictions in all measurement and cooling conditions - including in operational-like measurement conditions. We also performed model and data ablation experiments to determine the relative performance impact of our ML methods' capacity to model data nonlinearities and the inherent additional information in full spectra. Applying our ML methods, we found a number of surprising results, including improved accuracy at shorter fuel cooling times (the opposite of the norm), remarkable robustness to spectrum compression (via rebinning), and competitive burnup predictions even when using background signal only (i.e. explicitly omitting known isotope photopeaks).

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Learning protein fitness models from evolutionary and assay-labeled data

Machine learning-based models of protein fitness typically learn from either unlabeled, evolutionarily related sequences or variant sequences with experimentally measured labels. For regimes where only limited experimental data are available, recent work has suggested methods for combining both sources of information. Toward that goal, we propose a simple combination approach that is competitive with, and on average outperforms more sophisticated methods. Our approach uses ridge regression on site-specific amino acid features combined with one probability density feature from modeling the evolutionary data. Within this approach, we find that a variational autoencoder-based probability density model showed the best overall performance, although any evolutionary density model can be used. Moreover, our analysis highlights the importance of systematic evaluations and sufficient baselines.

59 BASIC BIOLOGICAL SCIENCES↗

GASNet-EX RMA Communication Performance on Recent Supercomputing Systems

Partitioned Global Address Space (PGAS) programming models, typified by systems such as Unified Parallel C (UPC) and Fortran coarrays, expose one-sided Remote Memory Access (RMA) communication as a key building block for High Performance Computing (HPC) applications. Architectural trends in supercomputing make such programming models increasingly attractive, and newer, more sophisticated models such as UPC++, Legion and Chapel that rely upon similar communication paradigms are gaining popularity. GASNet-EX is a portable, open-source, high-performance communication library designed to efficiently support the networking requirements of PGAS runtime systems and other alternative models in emerging exascale machines. The library is an evolution of the popular GASNet communication system, building upon 20 years of lessons learned. We present microbenchmark results which demonstrate the RMA performance of GASNet-EX is competitive with MPI implementations on four recent, high-impact, production HPC systems. These results are an update relative to previously published results on older systems. The networks measured here are representative of hardware currently used in six of the top ten fastest supercomputers in the world, and all of the exascale systems on the U.S. DOE road map.

Hargrove, Paul H↗

Optical Control of Adaptive Nanoscale Domain Networks

Adaptive networks can sense and adjust to dynamic environments to optimize their performance. Understanding their nanoscale responses to external stimuli is essential for applications in nanodevices and neuromorphic computing. However, it is challenging to image such responses on the nanoscale with crystallographic sensitivity. Here, the evolution of nanodomain networks in (PbTiO 3 ) n /(SrTiO 3 ) n superlattices (SLs) is directly visualized in real space as the system adapts to ultrafast repetitive optical excitations that emulate controlled neural inputs. The adaptive response allows the system to explore a wealth of metastable states that are previously inaccessible. Their reconfiguration and competition are quantitatively measured by scanning x-ray nanodiffraction as a function of the number of applied pulses, in which crystallographic characteristics are quantitatively assessed by assorted diffraction patterns using unsupervised machine-learning methods. The corresponding domain boundaries and their connectivity are drastically altered by light, holding promise for light-programable nanocircuits in analogy to neuroplasticity. Phase-field simulations elucidate that the reconfiguration of the domain networks is a result of the interplay between photocarriers and transient lattice temperature. The demonstrated optical control scheme and the uncovered nanoscopic insights open opportunities for the remote control of adaptive nanoscale domain networks.

36 MATERIALS SCIENCE↗

The LHC Olympics 2020 a community challenge for anomaly detection in high energy physics

A new paradigm for data-driven, model-agnostic new physics searches at colliders is emerging, and aims to leverage recent breakthroughs in anomaly detection and machine learning. In order to develop and benchmark new anomaly detection methods within this framework, it is essential to have standard datasets. To this end, we have created the LHC Olympics 2020, a community challenge accompanied by a set of simulated collider events. Participants in these Olympics have developed their methods using an R&D dataset and then tested them on black boxes: datasets with an unknown anomaly (or not). Furthermore, methods made use of modern machine learning tools and were based on unsupervised learning (autoencoders, generative adversarial networks, normalizing flows), weakly supervised learning, and semi-supervised learning. This paper will review the LHC Olympics 2020 challenge, including an overview of the competition, a description of methods deployed in the competition, lessons learned from the experience, and implications for data analyses with future datasets as well as future colliders.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Combining resonant and tail-based anomaly detection

In many well-motivated models of the electroweak scale, cascade decays of new particles can result in highly boosted hadronic resonances (e.g., Z / W / h ). This can make these models rich and promising targets for recently developed resonant anomaly detection methods powered by modern machine learning. We demonstrate this using the state-of-the-art classifying anomalies through outer density estimation () method applied to supersymmetry scenarios with gluino pair production. We show that , despite being model agnostic, is nevertheless competitive with dedicated cut-based searches, while simultaneously covering a much wider region of parameter space. The gluino events also populate the tails of the missing energy and H T distributions, making this a novel combination of resonant and tail-based anomaly detection. Published by the American Physical Society 2024

Astronomy & Astrophysics↗

Introduction to Fuzzy Set Theory

An introduction to fuzzy set theory is described. Topics covered include: neural networks and fuzzy systems; the dynamical systems approach to machine intelligence; intelligent behavior as adaptive model-free estimation; fuzziness versus probability; fuzzy sets; the entropy-subsethood theorem; adaptive fuzzy systems for backing up a truck-and-trailer; product-space clustering with differential competitive learning; and adaptive fuzzy system for target tracking.

Kosko, Bart↗

smol_model_benchmark

Machine learning models for protein–small molecule binding prediction using graph (GAT), image-based (ResNet), and transformer-based SMILES approaches (ChemBERTa and MIST). The code includes training pipelines and experiments on the dataset from the BELKA Kaggle competition.

Gibson, Kaetlyn [Los Alamos National Lab]↗

Uncovering electronic and geometric descriptors of chemical activity for metal alloys and oxides using unsupervised machine learning

Here, we show that unsupervised machine learning (ML) using principal component analysis (PCA) provides a straightforward pathway for developing accurate and interpretable electronic-structure descriptors of the chemical and catalytic properties of materials. We demonstrate the approach by finding chemisorption descriptors for metal alloys and surface oxygens on metals and metal oxides. In both cases, the principal component (PC) descriptors yield ML models that predict the material’s chemical properties with competitive accuracy compared to ML models built using established descriptors. Importantly, interpreting the electronic-structure patterns captured by each PC descriptor via signal reconstruction suggests potential design motifs for future electronic-structure descriptor design and allows us to identify links between a material’s geometric and catalytic properties. Ultimately, we show that the unsupervised ML approach provides a route to find electronic-structure descriptors of the catalytic properties of materials that readily connect to geometric structure and composition.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Comparing Automated Posterior Estimation Techniques for Modeling Strong Lenses In Ground-based Survey Data

Current and future ground-based cosmological surveys, such as the Dark Energy Survey (DES), and the Vera Rubin Observatory Legacy Survey of Space and Time (LSST), are predicted to discover thousands to tens of thousands of strong gravitational lenses. The large number of strong lenses discoverable in future surveys will make strong lensing a highly competitive and complementary cosmic probe. However, conventional lens modeling techniques are unable to scale up to the sheer number of lenses that will be discovered through upcoming surveys. Therefore, the use of automated lens analysis techniques is necessary. We demonstrate that machine learning methods can be used to automate the inference of informative model posteriors of strong lensing systems in ground-based surveys with credible uncertainty estimation. We present two Simulation-Based Inference (SBI) approaches for lens parameter estimation of galaxy-galaxy lenses. We demonstrate applications of Neural Posteriors Estima tors (NPEs) and Bayesian Neural Network (BNNs) to automate the inference of a 12-parameter lensing system for DES-like ground-based imaging data. We apply a suite of diagnostics (e.g., posterior coverage and SBC) to validate the performance of our methods. We find that NPEs outperform the BNN, producing posterior distributions that are for the most part both more accurate and more precise; in particular, several source-light model parameters are systematically biased in the BNN implementation.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Adaptive Quantum Generative Training using an Unbounded Loss Function

We propose a generative quantum learning algorithm using the Adaptive Derivative-Assembled Problem Tailored ansatz (ADAPT) framework in which the loss function to be minimized is the maximal quantum Rényi divergence of order two, an unbounded function that mitigates barren plateaus which inhibit training variational circuits. We benchmark this method against other state-of-the-art adaptive algorithms by learning random two-local thermal states. We perform numerical experiments of up to 12 qubits comparing our method learning algorithms that use linear objective functions and show that Rényi-ADAPT is capable of constructing shallow quantum circuits competitive with existing methods, while the gradients remain favorable resulting from the maximal Rényi divergence loss function.

quantum algorithms, quantum machine learning, quan↗

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY↗

Machine Learning for Correlated Intelligence. LDRD SAND Report

The Machine Learning for Correlated Intelligence Laboratory Directed Research & Development (LDRD) Project explored competing a variety of machine learning (ML) classification techniques against a known, open source dataset through the use of a rapid and automated algorithm research & development (RD) infrastructure. This approach relied heavily on creating an infrastructure in which to provide a pipeline for automatic target recognition (ATR) ML algorithm competition. Results are presented for nine ML classifiers against a primary dataset using the pipeline infrastructure developed for this project. New approaches to feature set extraction are presented and discussed as well.

97 MATHEMATICS AND COMPUTING↗

Modeling Protein–Protein and Protein–Ligand Interactions by the ClusPro Team in CASP16

ABSTRACT In the CASP16 experiment, our team employed hybrid computational strategies to predict both protein–protein and protein–ligand complex structures. For protein–protein docking, we combined physics‐based sampling—using ClusPro FFT docking and molecular dynamics—with AlphaFold (AF)‐based sampling, followed by AF‐based refinement. Our method produced numerous high‐accuracy complex models, including cases where AF alone failed, underscoring the critical role of physics‐based sampling alongside deep learning‐based refinement. For protein–ligand docking, we integrated the ClusPro LigTBM template‐based approach with a machine learning‐based confidence model for rescoring. The method preserves conserved interaction fragments derived from homologous complexes, followed by local resampling using physics‐based sampling and a diffusion model. Our template‐based strategy achieved a mean lDDT‐PLI of 0.69 across 233 targets, which was highly competitive. These results demonstrate that combining physics‐based modeling with AI‐driven refinement can significantly enhance the accuracy of both protein–protein and protein–ligand structure predictions.

Ashizawa, Ryota [Department of Applied Mathematics↗

An adaptive sampling augmented Lagrangian method for stochastic optimization with deterministic constraints

The primary goal of this paper is to provide an efficient solution algorithm based on the augmented Lagrangian framework for optimization problems with a stochastic objective function and deterministic constraints. Our main contribution is combining the augmented Lagrangian framework with adaptive sampling, resulting in an efficient optimization methodology validated with practical examples. To achieve the presented efficiency, here we consider inexact solutions for the augmented Lagrangian subproblems, and through an adaptive sampling mechanism, we control the variance in the gradient estimates. Furthermore, we analyze the theoretical performance of the proposed scheme by showing equivalence to a gradient descent algorithm on a Moreau envelope function, and we prove sublinear convergence for convex objectives and linear convergence for strongly convex objectives with affine equality constraints. The worst-case sample complexity of the resulting algorithm, for an arbitrary choice of penalty parameter in the augmented Lagrangian function, is $\mathscr{O}$(ϵ -3-δ ) , where ϵ > 0 is the expected error of the solution and δ > 0 is a user-defined parameter. If the penalty parameter is chosen to be $\mathscr{O}$(ϵ -1 ), we demonstrate that the result can be improved to $\mathscr{O}$(ϵ -2 ) , which is competitive with the other methods employed in the literature. Moreover, if the objective function is strongly convex with affine equality constraints, we obtain $\mathscr{O}$(ϵ -1 log(1/ϵ)) complexity. Finally, we empirically verify the performance of our adaptive sampling augmented Lagrangian framework in machine learning optimization and engineering design problems, including topology optimization of a heat sink with environmental uncertainty.

97 MATHEMATICS AND COMPUTING↗