Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Generative Adversarial Network”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Learning functional priors and posteriors from data and physics

In this work, we develop a new Bayesian framework based on deep neural networks to be able to extrapolate in space-time using historical data and to quantify uncertainties arising from both noisy and gappy data in physical problems. Specifically, the proposed approach has two stages: (1) prior learning and (2) posterior estimation. At the first stage, we employ the physics-informed Generative Adversarial Networks (PI-GAN) to learn a functional prior either from a prescribed function distribution, e.g., Gaussian process, or from historical data and physics. At the second stage, we employ the Hamiltonian Monte Carlo (HMC) method to estimate the posterior in the latent space of PI-GANs. In addition, we use two different approaches to encode the physics: (1) automatic differentiation, used in the physicsinformed neural networks (PINNs) for scenarios with explicitly known partial differential equations (PDEs), and (2) operator regression using the deep operator network (DeepONet) for PDE-agnostic scenarios. We then test the proposed method for (1) meta-learning for one-dimensional regression, and forward/inverse PDE problems (combined with PINNs); (2) PDE-agnostic physical problems (combined with DeepONet), e.g., fractional diffusion as well as saturated stochastic (100-dimensional) flows in heterogeneous porous media; and (3) spatial-temporal regression problems, i.e., inference of a marine riser displacement field using experimental data from the Norwegian Deepwater Programme (NDP). The results demonstrate that the proposed approach can provide accurate predictions as well as uncertainty quantification given very limited scattered and noisy data, since historical data could be available to provide informative priors. In summary, the proposed method is capable of learning flexible functional priors, e.g., both Gaussian and non-Gaussian process, and can be readily extended to big data problems by enabling mini-batch training using stochastic HMC or normalizing flows since the latent space is generally characterized as low dimensional.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Symplectic machine learning model for fast simulation of space-charge effects

Symplectic simulation of space-charge effects is crucial for the design and operation of high-intensity particle accelerators. Traditional methods for simulating these effects are often computationally expensive, resulting in significant overhead. In this work, we introduce a generative model based on a U-Net architecture within a generative adversarial network framework to efficiently simulate space-charge effects. The model is trained to predict the transverse multiparticle space-charge Hamiltonian, which can be physically computed using a gridless spectral method. The one-step symplectic transverse transfer map for the particles is then obtained by differentiating the predicted Hamiltonian. Benchmarking results demonstrate that this generative model achieves an order of magnitude higher computational efficiency compared to the spectral method, providing a highly efficient alternative for simulating space-charge effects with a large number of particles. By maintaining symplecticity, the model effectively preserves the phase-space structure and mitigates nonphysical errors in long-term simulations. This model has been integrated into jutrack, a novel autodifferentiable accelerator modeling code developed in the julia programming language.

Beam code development & simulation techniques↗

UVCGAN: UNet Vision Transformer cycle-consistent GAN for unpaired image-to-image translation

Unpaired image-to-image translation has broad applications in art, design, and scientific simulations. One early breakthrough was CycleGAN that emphasizes one-to-one mappings between two unpaired image domains via generative-adversarial networks (GAN) coupled with the cycle-consistency constraint, while more recent works promote one-to-many mapping to boost diversity of the translated images. Motivated by scientific simulation and one-to-one needs, this work revisits the classic CycleGAN framework and boosts its performance to outperform more contemporary models without relaxing the cycle-consistency constraint. To achieve this, we equip the generator with a Vision Transformer (ViT) and employ necessary training and regularization techniques. Compared to previous best-performing models, our model performs better and retains a strong correlation between the original and translated image. An accompanying ablation study shows that both the gradient penalty and self-supervised pre-training are crucial to the improvement. To promote reproducibility and open science, the source code, hyperparameter configurations, and pre-trained model are available at https: //github.com/LS4GAN/uvcgan.

97 MATHEMATICS AND COMPUTING↗

Fitting a deep generative hadronization model

Hadronization is a critical step in the simulation of high-energy particle and nuclear physics experiments. As there is no first principles understanding of this process, physically-inspired hadronization models have a large number of parameters that are fit to data. Deep generative models are a natural replacement for classical techniques, since they are more flexible and may be able to improve the overall precision. Proof of principle studies have shown how to use neural networks to emulate specific hadronization when trained using the inputs and outputs of classical methods. However, these approaches will not work with data, where we do not have a matching between observed hadrons and partons. In this paper, we develop a protocol for fitting a deep generative hadronization model in a realistic setting, where we only have access to a set of hadrons in data. Our approach uses a variation of a Generative Adversarial Network with a permutation invariant discriminator. We find that this setup is able to match the hadronization model in Herwig with multiple sets of parameters. This work represents a significant step forward in a longer term program to develop, train, and integrate machine learning-based hadronization models into parton shower Monte Carlo programs.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

ResNet and CycleGAN for pulse shape discrimination of He-4 detector pulses: Recovering pulses conventional algorithms fail to label unanimously

Pulse shape discrimination (PSD) capable detectors, such as He-4, that respond to neutron and gamma-ray 7 interactions have a threshold deposited energy value below which n/γ discrimination vanishes when using 8 conventional PSD algorithms. Recent attempts in applying supervised learning based artificial neural 9 networks for PSD use the pulses in the separated regions to train the networks so they can be used to classify 10 another set of separated pulses. In doing so, pulses previously indistinguishable are not recovered for 11 classification, which would have increased the number of neutron and gamma-ray pulses that could be used 12 for further analysis. Assuming the reason why conventional PSD algorithms have unseparated regions is 13 because the parameter space of the algorithms fail to capture the intrinsic (but subtle) distinguishing 14 behavior of some of the neutron and gamma-ray pulses, a cycle-consistent generative adversarial network 15 (CycleGAN) was trained to amplify those differences and extract well separated neutron and gamma-ray 16 clusters. Results show that, once the network is trained with pulses from separated and unseparated regions, 17 it was able to transform the pulses in the unseparated region to improve the PSD. Subsequent n/γ 18 classification was performed using deep residual network (ResNet) that takes pulses with 512 data points 19 as an input. Two different ResNets were explored – simple ResNet and modified ResNet which takes 20 segmented pulse inputs in the first layer and the corresponding time axis values in the last hidden layer. 21 The later approach enables the network to extract time correlated pulse features to enhance its ability to 22 capture the pulse behaviors relevant for PSD. Although it achieves slightly lower accuracy, 99.41% versus 23 99.89%, based on simply counting the number of correct n/γ labels assigned, compared to the simple 24 ResNet, the modified ResNets architecture was able to decreases the cross-entropy loss function by half, 25 which implies that the correct n/γ labels assigned are less likely to be accidental. PSD parameter 26 distributions based on n/γ classification by ResNet before and after transforming unseparated pulses using 27 CycleGAN show that by enhancing the separation between neutrons and gamma-rays, the transformation 28 helps improve the performance of classifier networks that are trained using labeled dataset. The 29 enhancement of neutron and gamma-ray separation by the CycleGAN increased the PSD figure of merit 30 (FOM) by up to 70% in some regions. Here, the results show that, if a given detector achieves clear separation 31 between neutron and gamma-ray pulses in any energy region, such neural network approaches can help 32 lower the energy threshold for the separation and increasing the number of neutron and gamma-ray pulses 33 that can be used for further analysis.

4He↗

Effectiveness of denoising diffusion probabilistic models for fast and high-fidelity whole-event simulation in high-energy heavy-ion experiments

Artificial intelligence (AI) generative models, such as generative adversarial networks (GANs), variational autoencoders, and normalizing flows, have been widely used and studied as efficient alternatives for traditional scientific simulations. However, they have several drawbacks, including training instability and inability to cover the entire data distribution, especially for regions where data are rare. This is particularly challenging for whole-event, full-detector simulations in high-energy heavy-ion experiments, such as sPHENIX at the Relativistic Heavy Ion Collider and Large Hadron Collider experiments, where thousands of particles are produced per event and interact with the detector. This work investigates the effectiveness of denoising diffusion probabilistic models (DDPMs) as an AI-based generative surrogate model for the sPHENIX experiment that includes the heavy-ion event generation and response of the entire calorimeter stack. DDPM performance in sPHENIX simulation data is compared with a popular rival, GANs. Results show that both DDPMs and GANs can reproduce the data distribution where the examples are abundant (low-to-medium calorimeter energies). Nonetheless, DDPMs significantly outperform GANs, especially in high-energy regions where data are rare. Additionally, DDPMs exhibit superior stability compared to GANs. The results are consistent between both central and peripheral centrality heavy-ion collision events. Moreover, DDPMs offer a substantial speedup of approximately a factor of 100 compared to the traditional Geant4 simulation method.

42 ENGINEERING↗

Leveraging generative artificial intelligence to bridge domain gaps in wind turbine research

A central challenge in wind turbine health monitoring is the scarcity of real-world data due to limited instrumentation, leading researchers to rely on simulation models that often suffer from reduced fidelity. However, even within simulation environments, discrepancies arise because of modeling assumptions, and configuration fidelities, creating domain gaps that limit the transferability of learned representations. Here, to investigate domain translation under controlled conditions, this project explores the use of generative artificial intelligence, specifically cycle-consistent generative adversarial networks (CGANs), to bridge the gap between OpenFAST simulation models representing 1.5 MW and 5 MW wind turbines. A physics-informed CGAN architecture is introduced, where a simplified turbine tower dynamics model is incorporated into the training loss to ensure physically consistent outputs. Quantitative results showed moderate to high agreement in frequency-domain features. Incorporating the physics-informed loss function improved the R 2 values by 30%, reduced the RMSE from 1.39 to 1.1 m/s 2 , and reduced training time by 82%. Furthermore, under increased turbulence intensity (IEC Category A), the RMSE remained stable at approximately 1.1 m/s 2 . While the present study is entirely simulation-based, it establishes a pipeline for evaluating physics-informed generative domain translation, which may serve as a foundation for future simulation-to-reality validation studies.

17 WIND ENERGY↗

Analysis of Defects in Metal Additive Manufacturing with Augmented Data Generation

Laser powder bed fusion (LPBF) is a method of additive manufacturing (AM) that selectively melts and fuses together microscopic metallic powder. LPBF offers the benefit of producing custom structures out of high strength metals that can be difficult to fabricate with conventional methods. The challenge of LPBF is that 3D printed structures often have internal pores due to process flaws. Pulsed thermal tomography (PTT) is a method for reconstructing the depth profile of materials, allowing the visualization internal voids in solids. In prior work, we developed a convolutional neural network (CNN) which, having been trained on simulated 2D PTT images of subsurface elliptical defects, was able to classify the semi-major radii, semi-minor radii, and angular orientation of the best-fit ellipses in previously unseen PTT images. The unseen PTT images contained subsurface irregular defect shapes imported from scanning electron microscopy (SEM) images of metallic LPBF-printed specimens. Training the CNN on irregular defect shapes instead of on elliptical shapes would make the resulting classifications more descriptive of actual defect shapes. However, this requires a much higher volume of SEM images of material defects, which are difficult to obtain because of random occurrence of defects in LPBF. To address this challenge, we developed a generative adversarial network (GAN) to augment the existing dataset of SEM defect images. The GAN model is demonstrated to create novel yet realistic defect shapes that can be used as input for simulated PTT images to train CNN.

36 MATERIALS SCIENCE↗

Pulsed Thermal Tomography Nondestructive Examination of Additively Manufactured Reactor Materials and Components (Final Technical Report)

Metal Additive Manufacturing (AM) is a promising method for cost-efficient fabrication of complex shape structures for applications in harsh environment, such as in a nuclear reactor. However, internal defects (pores) occur in high-strength AM alloys, which are manufactured with Laser Powder Bed Fusion (LPBF) AM method. Pulsed Infrared Thermography (PIT) is an efficient nondestructive evaluation (NDE) method to examine actual structures, because this method offers one-sided non-contact measurements, and fast processing of large sample areas. However, imaging of material defects, particularly defects with sizes at microscopic level, is challenging. In this report, we benchmark the performance of several Unsupervised Learning (UL) algorithms designed to enhance imaging of microscopic defects in metals with PIT. UL aims to learn the latent principal patterns (dictionaries) in PIT data to detect defects with minimal human supervision. Performance of Independent Component Analysis (ICA), Sparse Coding (SC), Principal Component Analysis (PCA) and Exploratory Factor Analysis (EFA) was compared using F-score, UL model training time and defects reconstruction time. We obtained the average F-score of 0.75, and a highest F-score of 0.89 for the EFA algorithm. Overall, EFA outperforms other UL algorithms considered in this study. In another approach, we investigate Thermal Tomography (TT), which is a computational method for reconstruction of depth profile of internal material defects from PIT nondestructive evaluation (NDE). TT algorithm obtains depth reconstructions of thermal effusivity, which has been shown to provide visualization of subsurface internals defects in metals. In many applications, one needs to determine the defect shape and orientation from reconstructed effusivity images. Interpretation of TT images is non-trivial because of blurring, which increases with depth due to heat diffusion-based nature of image formation. We have developed a deep learning convolutional neural network (CNN) to classify size and orientation of subsurface material defects in TT images. CNN was trained with TT images produced with computer simulations of 2D metallic structures (thin plates) containing elliptical subsurface voids. Performance of CNN was investigated using test TT images developed with computer simulations of plates containing elliptical defects, and defects with shape imported from scanning electron microscopy (SEM) images. CNN demonstrated the ability to classify radii and angular orientation of elliptical defects in previously unseen test TT images. We have also demonstrated that CNN trained on TT images of elliptical defects is capable of classifying shape and orientation of irregular defects. Training the CNN on irregular defect shapes instead of on elliptical shapes would make the resulting classifications more descriptive of actual defect shapes. However, this requires a much higher volume of SEM images of material defects, which are difficult to obtain because of random occurrence of defects in LPBF. To address this challenge, we developed a generative adversarial network (GAN) to augment the existing dataset of SEM defect images. The GAN model is demonstrated to create novel yet realistic defect shapes that can be used as input for simulated PTT images to train CNN. We also investigate several approaches based on Gaussian Random Circle and Bezier Curves for constructing parametric models of irregular-shape defects.

36 MATERIALS SCIENCE↗

Development of Multiresolution Capabilities for the Holistic Energy Resource Optimization Network (HERON) tool A progress update

INL researchers work on technoeconomic analyses for integrated energy systems (IES) using the Framework for Optimization of ResourCes and Economics (FORCE). Within FORCE, researchers use the Holistic Energy Resource Optimization Network (HERON) tool to conduct optimization of grid portfolios under uncertain market conditions. These optimizations determine optimal capacities for all IES components and strategies for resource dispatch which maximize some economic metric (e.g., net present value). Resource dispatch occurs on finer timescales (typically hours) and thus are asked to respond to a given time series (e.g. hourly load demand profiles for a grid, or pre-determined electricity prices). Volatile and complex bidding dynamics as well as poorly forecasted weather events within deregulated markets add uncertainty to the time series; FORCE can address this uncertainty by training a reduced order model on historical time series and generate unique synthetic time series which represent individual scenarios or realizations of the market. The IES configuration can be simulated under these different sampled realizations and a stochastic optimization is conducted which optimizes the expected value of the desired economic metric. The training of a synthetic time series generator is limited by the chosen time resolution; dynamics can occur on different time scales. Seasonal demand trends can dominate faster dynamical events (such as power outages from certain sectors or severe weather events) which might not get captured correctly by the trained model. In this report, we investigate different ways of addressing the training and generation of time series on multiple time scales using three main algorithms: wavelet decomposition, dynamic mode decomposition, and generative adversarial networks for time series. We demonstrate a time series analysis that yields information on not just the frequency space but also temporal space: where a fast Fourier transform can provide what frequencies dominate, the new algorithms can provide when the frequencies dominate as well. These analyses can help improve IES optimization by allowing researchers to couple simulations at different timescales when it is most needed - seasonal, day-ahead, and real time optimization - with greater computational efficiency. Future work will include implementation of a subset of the proposed algorithms into the FORCE toolset and application of these analyses into multiple timescale optimization.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Report on the AAPM grand challenge on deep generative modeling for learning medical image statistics

Abstract Background The findings of the 2023 AAPM Grand Challenge on Deep Generative Modeling for Learning Medical Image Statistics are reported in this Special Report. Purpose The goal of this challenge was to promote the development of deep generative models for medical imaging and to emphasize the need for their domain‐relevant assessments via the analysis of relevant image statistics. Methods As part of this Grand Challenge, a common training dataset and an evaluation procedure was developed for benchmarking deep generative models for medical image synthesis. To create the training dataset, an established 3D virtual breast phantom was adapted. The resulting dataset comprised about 108 000 images of size 512 512. For the evaluation of submissions to the Challenge, an ensemble of 10 000 DGM‐generated images from each submission was employed. The evaluation procedure consisted of two stages. In the first stage, a preliminary check for memorization and image quality (via the Fréchet Inception Distance [FID]) was performed. Submissions that passed the first stage were then evaluated for the reproducibility of image statistics corresponding to several feature families including texture, morphology, image moments, fractal statistics, and skeleton statistics. A summary measure in this feature space was employed to rank the submissions. Additional analyses of submissions was performed to assess DGM performance specific to individual feature families, the four classes in the training data, and also to identify various artifacts. Results Fifty‐eight submissions from 12 unique users were received for this Challenge. Out of these 12 submissions, 9 submissions passed the first stage of evaluation and were eligible for ranking. The top‐ranked submission employed a conditional latent diffusion model, whereas the joint runners‐up employed a generative adversarial network, followed by another network for image superresolution. In general, we observed that the overall ranking of the top 9 submissions according to our evaluation method (i) did not match the FID‐based ranking, and (ii) differed with respect to individual feature families. Another important finding from our additional analyses was that different DGMs demonstrated similar kinds of artifacts. Conclusions This Grand Challenge highlighted the need for domain‐specific evaluation to further DGM design as well as deployment. It also demonstrated that the specification of a DGM may differ depending on its intended use.

Radiology, Nuclear Medicine & Medical Imaging↗

Adoption of image-driven machine learning for microstructure characterization and materials design: A Perspective

Microstructure characterization enables the development of structure-processing-property relationships critical to several research areas within the broad field of materials science, from alloy design to the assessment of corrosion resistance, and failure analysis. Conventional approaches to material characterization have relied on either qualitative inference by the human ex-pert or software applications that can extract high-level features from images, such as boundary segmentation, average grain diameter, etc. Such approaches rely heavily on subject matter expert user intervention and knowledge of what phases or more generally, what microstructural features, are of interest. The recent surge in the adoption of machine learning techniques to address problems in materials engineering has brought with it an increased interest and application of Image Driven Machine Learning (IDML) approaches. In this work, we review the applications of IDML to the field of materials characterization. A canonical hierarchy of stages is defined, which when put sequentially together completes an IDML study: problem definition, dataset building, model selection and training, model evaluation, and integration with existing instrumentation or simulation workflow. The studies reviewed in this work are analyzed from the perspective of each of these stages. Such a review permits agranular assessment of the field, for example the impact of IDML on materials characterization at the nanoscale, the size of a typical dataset required to train a semantic segmentation model on electron microscopy images, ubiquitousness of transfer learning in the domain, etc. Finally, we discuss the importance of interpretability and explainability in the field of IDML for materials characterization, and provide an overview of two emerging techniques in the field: semantic segmentation and generative adversarial networks.

Baskaran, Arun↗

Improving microstructures segmentation via pretraining with synthetic data

Image analysis of material microstructures through microscopy is an integral capability in the field of materials science. The topological and chemical information obtained through microscopy allow us to draw vital connections between material microstructures, properties, and processing. While scanning electron microscopy (SEM) is able to yield a considerable wealth of information interpretable by the intuition of experts, there has been considerable interest in using machine learning, convolutional neural networks (CNNs) in particular, for such image analysis task. Training CNNs for an image analysis task requires a large annotated dataset. However, in many materials science applications, obtaining a large annotated dataset is cost and labor intensive. In this work, we study the use of synthetic data to enlarge the available annotated experimental data of uranium oxide. We utilize a modified Potts model to simulate uranium oxide particles with morphologies similar to those observed experimentally. We then leverage an image-to-image translation model to synthesize the simulated particles as if they are acquired with SEM. Through this process, we obtain pairs of particle images and their corresponding SEM representations, which corresponds to pairs of annotations and images. Unlike previous works, we leverage synthetic data for pretraining a CNN model prior, and finetune that model further with experimental data. We experimentally demonstrate that using synthetic data as incremental learning process benefits the overall performance compared to training a model on combined synthetic and experimental data.

36 MATERIALS SCIENCE↗

Inverse design of hypoeutectoid pearlite steel microstructures using a deep learning and genetic algorithm optimization framework

Goal-oriented microstructure design in metallic materials is a challenging task due to complex structure-property relationships. Traditional experimental and computational approaches are time-intensive and economically inefficient, limiting their applicability for large-scale design space exploration. Here, in this work, we propose an end-to-end framework that integrates deep learning models with genetic optimization to design microstructures with targeted mechanical properties. Deep learning models enable accurate forward design, while their integration with genetic optimization enables efficient inverse design within a few hours, compared to days or weeks using conventional finite element simulations. The framework combines experimental characterization and finite element modeling to analyze the influence of microstructural features on the mechanical behavior of hypoeutectoid steels. Data from both experiments and simulations are used to train the deep learning models. To demonstrate its effectiveness, we apply the framework to 0.63% carbon steel with proeutectoid ferrite and pearlite phases, commonly used in industrial applications. In this study, 2D microstructures were used for modeling, selected primarily for computational efficiency and to establish proof of concept. The framework successfully optimizes microstructures for targeted yield strength, ultimate strength, and stress concentration factors while significantly reducing computational time. Beyond hypoeutectoid steels, this scalable framework can be extended to other material systems and integrated with additive manufacturing, offering an efficient approach for accelerating microstructure design for specific engineering applications.

ConvLSTM↗

Machine learning-based microstructure prediction during laser sintering of alumina

Abstract Predicting material’s microstructure under new processing conditions is essential in advanced manufacturing and materials science. This is because the material’s microstructure hugely influences the material’s properties. We demonstrate an elegant machine learning algorithm that faithfully predicts the microstructure under new conditions, without the need of knowing the governing laws. We name this algorithm, RCWGAN-GP, which is regression-based conditional generative adversarial networks with Wasserstein loss function and gradient penalty. This algorithm was trained with experimental SEM micrographs from laser-sintered alumina under various laser powers. The RCWGAN-GP realistically regenerates the SEM micrographs under the trained laser powers. Impressively, it also faithfully predicts the alumina’s microstructure under unexplored laser powers. The predicted microstructure features, including the morphology of the sintered particles and the pores, match the experimental SEM micrographs very well. We further quantitatively examined the prediction accuracy of the RCWGAN-GP. We trained the algorithm with computer-created micrograph datasets of secondary-phase growth governed by the well-known Johnson–Mehl–Avrami (JMA) equation. The RCWGAN-GP accurately regenerates the micrographs at the trained time series, in terms of the grains’ shapes, sizes, and spatial distributions. More importantly, the predicted secondary phase fraction accurately follows the JMA curve.

08 HYDROGEN↗

Materials representation and transfer learning for multi-property prediction

The adoption of machine learning in materials science has rapidly transformed materials property prediction. Hurdles limiting full capitalization of recent advancements in machine learning include the limited development of methods to learn the underlying interactions of multiple elements as well as the relationships among multiple properties to facilitate property prediction in new composition spaces. To address these issues, we introduce the Hierarchical Correlation Learning for Multi-property Prediction (H-CLMP) framework that seamlessly integrates: (i) prediction using only a material's composition, (ii) learning and exploitation of correlations among target properties in multi-target regression, and (iii) leveraging training data from tangential domains via generative transfer learning. The model is demonstrated for prediction of spectral optical absorption of complex metal oxides spanning 69 three-cation metal oxide composition spaces. H-CLMP accurately predicts non-linear composition-property relationships in composition spaces for which no training data are available, which broadens the purview of machine learning to the discovery of materials with exceptional properties. This achievement results from the principled integration of latent embedding learning, property correlation learning, generative transfer learning, and attention models. The best performance is obtained using H-CLMP with transfer learning [H-CLMP(T)] wherein a generative adversarial network is trained on computational density of states data and deployed in the target domain to augment prediction of optical absorption from composition. H-CLMP(T) aggregates multiple knowledge sources with a framework that is well suited for multi-target regression across the physical sciences.

36 MATERIALS SCIENCE↗

The LHC Olympics 2020 a community challenge for anomaly detection in high energy physics

A new paradigm for data-driven, model-agnostic new physics searches at colliders is emerging, and aims to leverage recent breakthroughs in anomaly detection and machine learning. In order to develop and benchmark new anomaly detection methods within this framework, it is essential to have standard datasets. To this end, we have created the LHC Olympics 2020, a community challenge accompanied by a set of simulated collider events. Participants in these Olympics have developed their methods using an R&D dataset and then tested them on black boxes: datasets with an unknown anomaly (or not). Furthermore, methods made use of modern machine learning tools and were based on unsupervised learning (autoencoders, generative adversarial networks, normalizing flows), weakly supervised learning, and semi-supervised learning. This paper will review the LHC Olympics 2020 challenge, including an overview of the competition, a description of methods deployed in the competition, lessons learned from the experience, and implications for data analyses with future datasets as well as future colliders.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Breaking the barrier of human-annotated training data for machine learning-aided plant research using aerial imagery

Machine learning (ML) can accelerate biological research. However, the adoption of such tools to facilitate phenotyping based on sensor data has been limited by (i) the need for a large amount of human-annotated training data for each context in which the tool is used and (ii) phenotypes varying across contexts defined in terms of genetics and environment. This is a major bottleneck because acquiring training data is generally costly and time-consuming. This study demonstrates how a ML approach can address these challenges by minimizing the amount of human supervision needed for tool building. A case study was performed to compare ML approaches that examine images collected by an uncrewed aerial vehicle to determine the presence/absence of panicles (i.e. “heading”) across thousands of field plots containing genetically diverse breeding populations of 2 Miscanthus species. Automated analysis of aerial imagery enabled the identification of heading approximately 9 times faster than in-field visual inspection by humans. Leveraging an Efficiently Supervised Generative Adversarial Network (ESGAN) learning strategy reduced the requirement for human-annotated data by 1 to 2 orders of magnitude compared to traditional, fully supervised learning approaches. The ESGAN model learned the salient features of the data set by using thousands of unlabeled images to inform the discriminative ability of a classifier so that it required minimal human-labeled training data. This method can accelerate the phenotyping of heading date as a measure of flowering time in Miscanthus across diverse contexts (e.g. in multistate trials) and opens avenues to promote the broad adoption of ML tools.

59 BASIC BIOLOGICAL SCIENCES↗