Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning competition”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Importance of kernel bandwidth in quantum machine learning

Quantum kernel methods are considered a promising avenue for applying quantum computers to machine learning problems. Identifying hyperparameters controlling the inductive bias of quantum machine learning models is expected to be crucial given the central role hyperparameters play in determining the performance of classical machine learning methods. In this work we introduce the hyperparameter controlling the bandwidth of a quantum kernel and show that it controls the expressivity of the resulting model. We use extensive numerical experiments with multiple quantum kernels and classical data sets to show consistent change in the model behavior from underfitting (bandwidth too large) to overfitting (bandwidth too small), with optimal generalization in between. We draw a connection between the bandwidth of classical and quantum kernels and show analogous behavior in both cases. Furthermore, we show that optimizing the bandwidth can help mitigate the exponential decay of kernel values with qubit count, which is the cause behind recent observations that the performance of quantum kernel methods decreases with qubit count. Here, we reproduce these negative results and show that if the kernel bandwidth is optimized, the performance instead improves with growing qubit count and becomes competitive with the best classical methods.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Sub-microsecond Transformers for Jet Tagging on FPGAs

We present the first sub-microsecond transformer implementation on an FPGA achieving competitive performance for state-of-the-art high-energy physics benchmarks. Transformers have shown exceptional performance on multiple tasks in modern machine learning applications, including jet tagging at the CERN Large Hadron Collider (LHC). However, their computational complexity prohibits use in real-time applications, such as the hardware trigger system of the collider experiments up until now. In this work, we demonstrate the first application of transformers for jet tagging on FPGAs, achieving $\mathcal{O}(100)$ nanosecond latency with superior performance compared to alternative baseline models. We leverage high-granularity quantization and distributed arithmetic optimization to fit the entire transformer model on a single FPGA, achieving the required throughput and latency. Furthermore, we add multi-head attention and linear attention support to hls4ml, making our work accessible to the broader fast machine learning community. This work advances the next-generation trigger systems for the High Luminosity LHC, enabling the use of transformers for real-time applications in high-energy physics and beyond.

Laatu, Lauri [Imperial Coll., London]↗

Machine learning-assisted ultrafast flash sintering of high-performance and flexible silver–selenide thermoelectric devices

Flexible thermoelectric generators (TEGs) have shown immense potential for serving as a power source for wearable electronics and the Internet of Things. A key challenge preventing large-scale application of TEGs lies in the lack of a high-throughput processing method, which can sinter thermoelectric (TE) materials rapidly while maintaining their high thermoelectric properties. Herein, we integrate high-throughput experimentation and Bayesian optimization (BO) to accelerate the discovery of the optimum sintering conditions of silver–selenide TE films using an ultrafast intense pulsed light (flash) sintering technique. Due to the nature of the high-dimensional optimization problem of flash sintering processes, a Gaussian process regression (GPR) machine learning model is established to rapidly recommend the optimum flash sintering variables based on Bayesian expected improvement. For the first time, an ultrahigh-power factor flexible TE film (a power factor of 2205 μW m -1 K -2 with a zT of 1.1 at 300 K) is demonstrated with a sintering time less than 1.0 second, which is several orders of magnitude shorter than that of conventional thermal sintering techniques. Further, the films also show excellent flexibility with 92% retention of the power factor (PF) after 10 3 bending cycles with a 5 mm bending radius. In addition, a wearable thermoelectric generator based on the flash-sintered films generates a very competitive power density of 0.5 mW cm -2 at a temperature difference of 10 K. This work not only shows the tremendous potential of high-performance and flexible silver–selenide TEGs but also demonstrates a machine learning-assisted flash sintering strategy that could be used for ultrafast, high-throughput and scalable processing of functional materials for a broad range of energy and electronic applications.

25 ENERGY STORAGE↗

HD-Bind: Encoding of Molecular Structure with Low Precision, Hyperdimensional Binary Representations

Publicly available collections of drug-like molecules have grown to comprise tens of billions of compounds due to advances in combinatorial chemistry. Traditional methods for identifying "hit" molecules from a large collection of potential drug-like candidates have relied on biophysical theory to compute approximations to the Gibbs free energy of the binding interaction between the drug and its protein target. These approaches have the major drawback that they require exceptional computing capabilities for even relatively small collections of molecules. Hyperdimensional Computing (HDC) is a recently-proposed learning paradigm that represents data with high-dimension binary vectors; this allows the use of low-precision binary vector arithmetic to create models of the data that can be learned without the need for the gradient-based optimization required in many conventional machine learning and deep learning methods. This algorithmic simplicity allows for acceleration in hardware that has been previously demonstrated in a range of application areas. We consider existing HDC approaches for molecular property classification and introduce two novel encodings of a commonly-used molecular representation, the extended connectivity fingerprint (ECFP). We show that HDC-based inference methods are as much as 91 times more efficient than traditional machine learning methods, and achieve an acceleration of nearly nine orders of magnitude compared to molecular docking. Our results show that HDC accelerated methods retain competitive accuracy on a number of well-studied tasks such as molecular property predictions using the MoleculeNet dataset, and bind/no-bind activity classification using the DUD-E and LIT-PCBA datasets. Our work thus motivates further investigation into molecular representation learning to develop ultraefficient pre-screening tools.

Jones, William↗

Prediction of Solute Segregation at Metal/Oxide Interfaces Using Machine Learning Approaches

The atomic structure and chemistry at metal/oxide interfaces play a crucial role in determining their properties. However, studying semi-coherent metal/oxide interfaces that include misfit dislocations through density functional theory (DFT) is often computationally expensive due to the large number of atoms involved, ranging from hundreds to thousands. In this study, we explore solute segregation behavior at the Fe/Y 2 O 3 interface—an important model interface for cladding applications in nuclear fission reactors—by combining DFT calculations with a machine learning (ML) approach. ML models are trained using DFT-calculated segregation energies (𝐸 𝑆𝑒𝑔 ) to identify the key chemical and geometric factors influencing solute segregation at metal/oxide interfaces, revealing the competition between these features in determining 𝐸 𝑆𝑒𝑔 . Moreover, the segregation behavior at a specific Fe/Y 2 O 3 interface is predicted with high accuracy using ML models trained on data from this interface. Furthermore, it is found that the ML models could also predict solute segregation at a different Fe/Y 2 O 3 interface with a new orientation relationship (OR), at a computational cost of less than 1/45 of that required for similar DFT calculations.

36 - MATERIALS SCIENCE↗

On the Solution of ℓ 0 -Constrained Sparse Inverse Covariance Estimation Problems

The sparse inverse covariance matrix is used to model conditional dependencies between variables in a graphical model to fit a multivariate Gaussian distribution. Estimating the matrix from data are well known to be computationally expensive for large-scale problems. Sparsity is employed to handle noise in the data and to promote interpretability of a learning model. Although the use of a convex ℓ 1 regularizer to encourage sparsity is common practice, the combinatorial ℓ 0 penalty often has more favorable statistical properties. In this paper, we directly constrain sparsity by specifying a maximally allowable number of nonzeros, in other words, by imposing an ℓ 0 constraint. Here, we introduce an efficient approximate Newton algorithm using warm starts for solving the nonconvex ℓ 0 -constrained inverse covariance learning problem. Numerical experiments on standard data sets show that the performance of the proposed algorithm is competitive with state-of-the-art methods.

$\ell_0$-Constrained↗

Bingo: A Customizable Framework for Symbolic Regression with Genetic Programming

In this paper, we introduce Bingo, a flexible and customizable yet performant Python framework for symbolic regression with genetic programming. Bingo maintains a modular code structure for simple abstraction and easily swappable components. Fitness functions, selection methods, and constant optimization methods allow for easy problem-specific customization. Bingo also maintains several features for increased efficiency such as parallelism, equation simplification, and a C++ backend. We compare Bingo’s performance to other genetic programming for symbolic regression (GPSR) methods to show that it is both competitive and flexible.

machine learning↗

Protein model quality assessment using rotation–equivariant transformations on point clouds

Machine learning research concerning protein structure has seen a surge in popularity over the last years with promising advances for basic science and drug discovery. Working with macromolecular structure in a machine learning context requires an adequate numerical representation, and researchers have extensively studied representations such as graphs, discretized 3D grids, and distance maps. As part of CASP14, we explored a new and conceptually simple representation in a blind experiment: atoms as points in 3D, each with associated features. These features—initially just the basic element type of each atom—are updated through a series of neural network layers featuring rotation-equivariant convolutions. Starting from all atoms, we further aggregate information at the level of alpha carbons before making a prediction at the level of the entire protein structure. We find that this approach yields competitive results in protein model quality assessment despite its simplicity and despite the fact that it incorporates minimal prior information and is trained on relatively little data. As a result, its performance and generality are particularly noteworthy in an era where highly complex, customized machine learning methods such as AlphaFold 2 have come to dominate protein structure prediction.

59 BASIC BIOLOGICAL SCIENCES↗

Ion correlations explain kinetic selectivity in diffusion-limited solid-state synthesis reactions

Establishing viable solid-state synthesis pathways for novel inorganic materials remains a major challenge in materials science. Previous pathway design methods using pairwise reaction approaches have navigated the thermodynamic landscape with first-principles data but lack kinetic information, limiting their effectiveness. This gap leads to suboptimal precursor selection and predictions, especially for reactions forming competing phases with similar formation energies, where ion diffusion is a critical influence. Here we demonstrate an inorganic synthesis framework by incorporating machine learning-derived transport properties through ‘liquid-like’ product layers into a thermodynamic cellular reaction model. In the Ba–Ti–O system, known for its competitive polymorphism, we obtain accurate predictions of phase formation with varying BaO:TiO2 ratios as a function of time and temperature. We find that diffusion–thermodynamics interplay governs phase compositions, with cross-ion transport coefficients critical for predicting diffusion-limited selectivity. This work bridges length scales and timescales by integrating solid-state reaction kinetics with first-principles thermodynamics and spatial reactivity.

Atomistic models↗

Results of the Photometric LSST Astronomical Time-series Classification Challenge (PLAsTiCC)

Abstract Next-generation surveys like the Legacy Survey of Space and Time (LSST) on the Vera C. Rubin Observatory (Rubin) will generate orders of magnitude more discoveries of transients and variable stars than previous surveys. To prepare for this data deluge, we developed the Photometric LSST Astronomical Time-series Classification Challenge (PLAsTiCC), a competition that aimed to catalyze the development of robust classifiers under LSST-like conditions of a nonrepresentative training set for a large photometric test set of imbalanced classes. Over 1000 teams participated in PLAsTiCC, which was hosted in the Kaggle data science competition platform between 2018 September 28 and 2018 December 17, ultimately identifying three winners in 2019 February. Participants produced classifiers employing a diverse set of machine-learning techniques including hybrid combinations and ensemble averages of a range of approaches, among them boosted decision trees, neural networks, and multilayer perceptrons. The strong performance of the top three classifiers on Type Ia supernovae and kilonovae represent a major improvement over the current state of the art within astronomy. This paper summarizes the most promising methods and evaluates their results in detail, highlighting future directions both for classifier development and simulation needs for a next-generation PLAsTiCC data set.

79 ASTRONOMY AND ASTROPHYSICS↗

A non-cooperative meta-modeling game for automated third-party calibrating, validating and falsifying constitutive laws with parallelized adversarial attacks

The evaluation of constitutive models, especially for high-risk and high-regret engineering applications, requires efficient and rigorous third-party calibration, validation and falsification. While there are numerous efforts to develop paradigms and standard procedures to validate models, difficulties may arise due to the sequential, manual, and often biased nature of the commonly adopted calibration and validation processes, thus slowing down data collections, hampering the progress towards discovering new physics, increasing expenses and possibly leading to misinterpretations of the credibility and application ranges of proposed models. This work attempts to introduce concepts from game theory and machine learning techniques to overcome many of these existing difficulties. Here, we introduce an automated meta-modeling game where two competing AI agents systematically generate experimental data to calibrate a given constitutive model and to explore its weakness such that the experiment design and model robustness can be improved through competitions. The two agents automatically search for the Nash equilibrium of the meta-modeling game in an adversarial reinforcement learning framework without human intervention. In particular, a protagonist agent seeks to find the more effective ways to generate data for model calibrations, while an adversary agent tries to find the most devastating test scenarios that expose the weaknesses of the constitutive model calibrated by the protagonist. By capturing all possible design options of the laboratory experiments into a single decision tree, we recast the design of experiments as a game of combinatorial moves that can be resolved through deep reinforcement learning by the two competing players. Our adversarial framework emulates idealized scientific collaborations and competitions among researchers to achieve a better understanding of the application range of the learned material laws and prevent misinterpretations caused by conventional AI-based third-party validation. Numerical examples are given to demonstrate the wide applicability of the proposed meta-modeling game with adversarial attacks on both human-crafted constitutive models and machine learning models.

97 MATHEMATICS AND COMPUTING↗

An Entropy-Maximization Approach to Automated Training Set Generation for Interatomic Potentials

Machine learning-based interatomic potentials are currently garnering a lot of attention as they strive to achieve the accuracy of electronic structure methods at the computational cost of empirical potentials. Given their generic functional forms, the transferability of these potentials is highly dependent on the quality of the training set, the generation of which can be highly labor-intensive. Good training sets should at once contain a very diverse set of configurations while avoiding redundancies that incur cost without providing benefits. We formalize these requirements in a local entropy-maximization framework and propose an automated sampling scheme to sample from this objective function. We show that this approach generates much more diverse training sets than unbiased sampling and is competitive with hand-crafted training sets.

74 ATOMIC AND MOLECULAR PHYSICS↗

Multifidelity Ensemble Kalman Filtering Using Surrogate Models Defined by Theory-Guided Autoencoders

Data assimilation is a Bayesian inference process that obtains an enhanced understanding of a physical system of interest by fusing information from an inexact physics-based model, and from noisy sparse observations of reality. The multifidelity ensemble Kalman filter (MFEnKF) recently developed by the authors combines a full-order physical model and a hierarchy of reduced order surrogate models in order to increase the computational efficiency of data assimilation. The standard MFEnKF uses linear couplings between models, and is statistically optimal in case of Gaussian probability densities. This work extends the MFEnKF into to make use of a broader class of surrogate model such as those based on machine learning methods such as autoencoders non-linear couplings in between the model hierarchies. We identify the right-invertibility property for autoencoders as being a key predictor of success in the forecasting power of autoencoder-based reduced order models. We propose a methodology that allows us to construct reduced order surrogate models that are more accurate than the ones obtained via conventional linear methods. Numerical experiments with the canonical Lorenz'96 model illustrate that nonlinear surrogates perform better than linear projection-based ones in the context of multifidelity ensemble Kalman filtering. We additionality show a large-scale proof-of-concept result with the quasi-geostrophic equations, showing the competitiveness of the method with a traditional reduced order model-based MFEnKF.

97 MATHEMATICS AND COMPUTING↗

When ancient numerical demons meet physics-informed machine learning: adjoint-based gradients for implicit differentiable modeling

Recent advances in differentiable modeling, a genre of physics-informed machine learning that trains neural networks (NNs) together with process-based equations, have shown promise in enhancing hydrological models' accuracy, interpretability, and knowledge-discovery potential. Current differentiable models are efficient for NN-based parameter regionalization, but the simple explicit numerical schemes paired with sequential calculations (operator splitting) can incur numerical errors whose impacts on models' representation power and learned parameters are not clear. Implicit schemes, however, cannot rely on automatic differentiation to calculate gradients due to potential issues of gradient vanishing and memory demand. Here we propose a “discretize-then-optimize” adjoint method to enable differentiable implicit numerical schemes for the first time for large-scale hydrological modeling. The adjoint model demonstrates comprehensively improved performance, with Kling–Gupta efficiency coefficients, peak-flow and low-flow metrics, and evapotranspiration that moderately surpass the already-competitive explicit model. Therefore, the previous sequential-calculation approach had a detrimental impact on the model's ability to represent hydrological dynamics. Furthermore, with a structural update that describes capillary rise, the adjoint model can better describe baseflow in arid regions and also produce low flows that outperform even pure machine learning methods such as long short-term memory networks. The adjoint model rectified some parameter distortions but did not alter spatial parameter distributions, demonstrating the robustness of regionalized parameterization. Despite higher computational expenses and modest improvements, the adjoint model's success removes the barrier for complex implicit schemes to enrich differentiable modeling in hydrology.

58 GEOSCIENCES↗

Structure of the divergent human astrovirus MLB capsid spike

Despite their worldwide prevalence and association with human disease, the molecular bases of human astrovirus (HAstV) infection and evolution remain poorly characterized. Here, we report the structure of the capsid protein spike of the divergent HAstV MLB clade (HAstV MLB). While the structure shares a similar folding topology with that of classical-clade HAstV spikes, it is otherwise strikingly different. We find no evidence of a conserved receptor-binding site between the MLB and classical HAstV spikes, suggesting that MLB and classical HAstVs utilize different receptors for host-cell attachment. We provide evidence for this hypothesis using a novel HAstV infection competition assay. Comparisons of the HAstV MLB spike structure with structures predicted from its sequence reveal poor matches, but template-based predictions were surprisingly accurate relative to machine-learning-based predictions. Our data provide a foundation for understanding the mechanisms of infection by diverse HAstVs and can support structure determination in similarly unstudied systems.

59 BASIC BIOLOGICAL SCIENCES↗

Optimizing enzymes for plastic upcycling using machine learning design and high throughput experiments

Plastic use is ubiquitous in the modern world, and polyethylene terephthalate (PET) is one of the most abundantly produced plastics (and the most highly produced polyester), with ~65 million metric tons manufactured annually. To the consumer, PET is likely most recognizable as the plastic used to make beverage bottles. Like many plastics, traditional mechanical or chemical means of PET deconstruction and upcycling are costly and inefficient. Because of these challenges, recycled plastic is generally of lower quality and is more expensive to produce than virgin plastic derived from petroleum. Ultimately, this results in most plastic ending up as waste. We view plastic waste as an underutilized resource which, with the development of more efficient and high-quality recycling processes, could (1) generate significant economic value while (2) decreasing petroleum usage and greenhouse gas emissions, as well as (3) minimizing its negative environmental and health impacts. Biocatalytic recycling, or biomanufacturing the basic building blocks of new plastic from plastic waste, is a promising approach to plastic reuse that complements existing recycling technologies. Recently, biological enzymes capable of breaking down PET have garnered significant attention as an attractive means of dealing with the plastic problem. These enzymes are currently undergoing pilot studies for implementation in industrial-scale enzyme-based recycling. However, there are significant limitations to current enzymes, including the need to perform costly pre-processing of the plastic waste before the enzymes are able to work. Further optimization of these enzymes is necessary to make these technologies competitive, and ultimately incentivise industry-wide adoption of this biology-based green recycling technology. n this work we demonstrate a means to design and generate performant biological enzymes, capable of efficiently deconstructing plastic waste. Specifically, we applied recent advances in artificial intelligence, machine learning, and statistical analysis to design new versions and discover natural enzymes capable of breaking down PET. We focused on optimizing key properties that are important for industrial-scale enzymatic recycling such as pH and thermotolerance. Normal testing of enzymatic plastic-deconstruction is extremely labor intensive and so through this work we also developed a robotic-assisted experimental pipeline capable of characterizing thousands of candidate enzymes. The results of this iterative, AI-guided, multi-discipline approach have led to increases in enzymatic breakdown of over 150X over starting enzymes. This work supports the rapidly developing and transformative field of biocatalytic solutions to environmental problems beyond the discovery and predictive understanding of enzymes for polymer recycling, and has wide implications for tackling numerous energy problems such as carbon capture and fixation (e.g., engineering carbon monoxide dehydrogenase and the rubisco-pathway), biomining (e.g., design of lanthanide-binding proteins) and biomanufacturing (e.g., lignin-deconstruction enzymes).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

A Generalized Transformer-Based Pulse Detection Algorithm

Pulse-like signals are ubiquitous in the field of single molecule analysis, e.g., electrical or optical pulses caused by analyte translocations in nanopores. The primary challenge in processing pulse-like signals is to capture the pulses in noisy backgrounds, but current methods are subjectively based on a user-defined threshold for pulse recognition. Here, we propose a generalized machine-learning based method, named pulse detection transformer (PETR), for pulse detection. PETR determines the start and end time points of individual pulses, thereby singling out pulse segments in a time-sequential trace. It is objective without needing to specify any threshold. It provides a generalized interface for downstream algorithms for specific application scenarios. PETR is validated using both simulated and experimental nanopore translocation data. It returns a competitive performance in detecting pulses through assessing them with several standard metrics. Finally, the generalization nature of the PETR output is demonstrated using two representative algorithms for feature extraction.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Robust Power System Stability Assessment Against Adversarial Machine Learning-Based Cyberattacks via Online Purification

The increasing complexity associated with renewable generation brings more challenges to power system stability assessment (SA). Data-driven approaches based on machine learning (ML) techniques for stability assessment have received significant research interest and shown their promising performance. However, ML-based models are recognized to be vulnerable to adversarial disturbances, where a slight perturbation to power system measurements could lead to unacceptable errors. To address this issue, this paper develops a novel lightweight mitigation strategy, i.e., robust online stability assessment (ROSA), to enhance the ML-based assessment model against both white-box and the black-box adversarial disturbances (i.e., purification) in the online implementation. The ROSA involves a supervised learning-based module for the primary stability assessment and a self-supervised learning-based module. Further, the two modules are trained jointly with different objective (loss) functions and implemented in sequence. A suitable purification objective and various time-series data augmentation methods are designed for SA applications to tackle adversarial disturbances adaptively. Case studies are performed, and the comparative results have clearly illustrated the competitive, robust accuracy against various adversarial scenarios and verified the effectiveness of the proposed online purification strategy.

24 POWER TRANSMISSION AND DISTRIBUTION↗