Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Adversarial Machine Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Fitting a deep generative hadronization model

Hadronization is a critical step in the simulation of high-energy particle and nuclear physics experiments. As there is no first principles understanding of this process, physically-inspired hadronization models have a large number of parameters that are fit to data. Deep generative models are a natural replacement for classical techniques, since they are more flexible and may be able to improve the overall precision. Proof of principle studies have shown how to use neural networks to emulate specific hadronization when trained using the inputs and outputs of classical methods. However, these approaches will not work with data, where we do not have a matching between observed hadrons and partons. In this paper, we develop a protocol for fitting a deep generative hadronization model in a realistic setting, where we only have access to a set of hadrons in data. Our approach uses a variation of a Generative Adversarial Network with a permutation invariant discriminator. We find that this setup is able to match the hadronization model in Herwig with multiple sets of parameters. This work represents a significant step forward in a longer term program to develop, train, and integrate machine learning-based hadronization models into parton shower Monte Carlo programs.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Improve Learning from Crowds via Generative Augmentation

Crowdsourcing provides an efficient label collection schema for supervised machine learning. However, to control annotation cost, each instance in the crowdsourced data is typically annotated by a small number of annotators. This creates a sparsity issue and limits the quality of machine learning models trained on such data. In this paper, we study how to handle sparsity in crowdsourced data using data augmentation. Specifically, we propose to directly learn a classifier by augmenting the raw sparse annotations. We implement two principles of high-quality augmentation using Generative Adversarial Networks: 1) the generated annotations should follow the distribution of authentic ones, which is measured by a discriminator; 2) the generated annotations should have high mutual information with the ground-truth labels, which is measured by an auxiliary network. Extensive experiments and comparisons against an array of state-of-the-art learning from crowds methods on three real-world datasets proved the effectiveness of our data augmentation framework. It shows the potential of our algorithm for low-budget crowdsourcing in general.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Systematic Evaluation of Backdoor Data Poisoning Attacks on Image Classifiers

Backdoor data poisoning attacks have recently been demonstrated in computer vision research as a potential safety risk for machine learning (ML) systems. Traditional data poisoning attacks manipulate training data to induce unreliability of an ML model, whereas backdoor data poisoning attacks maintain system performance unless the MLmodel is presented with an input containing an embedded“trigger” that provides a predetermined response advantageous to the adversary. Our work builds upon prior back-door data-poisoning research for ML image classifiers and systematically assesses different experimental conditions including types of trigger patterns, persistence of trigger patterns during retraining, poisoning strategies, architectures (ResNet-50, NasNet, NasNet-Mobile), datasets (Flowers, CIFAR-10), and potential defensive regularization techniques (Contrastive Loss, Logit Squeezing, Manifold Mixup,Soft-Nearest-Neighbors Loss). Experiments yield four key findings. First, the success rate of backdoor poisoning at-tacks varies widely, depending on several factors, including model architecture, trigger pattern and regularization technique. Second, we find that poisoned models are hard to detect through performance inspection alone. Third, regularization typically reduces backdoor success rate, although it can have no effect or even slightly increase it, depending on the form of regularization. Finally, backdoors inserted through data poisoning can be rendered ineffective after just a few epochs of additional training on a small set of clean data without affecting the model’s performance.

Truong, Loc T.↗

Explainable machine learning of the underlying physics of high-energy particle collisions

We present an implementation of an explainable and physics-aware machine learning model capable of inferring the underlying physics of high-energy particle collisions using the information encoded in the energy-momentum four-vectors of the final state particles. We demonstrate the proof-of-concept of our White Box AI approach using a Generative Adversarial Network (GAN) which learns from a DGLAP-based parton shower Monte Carlo event generator. The constrained generator network architecture mimics the structure of a parton shower exhibiting similarities with Recurrent Neural Networks (RNNs). We show, for the first time, that our approach leads to a network that is able to learn not only the final distribution of particles, but also the underlying parton branching mechanism, i.e. the Altarelli-Parisi splitting function, the ordering variable of the shower, and the scaling behavior. While the current work is focused on perturbative physics of the parton shower, we foresee a broad range of applications of our framework to areas that are currently difficult to address from first principles in QCD. Examples include nonperturbative and collective effects, factorization breaking and the modification of the parton shower in heavy-ion, and electron-nucleus collisions.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Modeling design and control problems involving neural network surrogates

Here, we consider nonlinear optimization problems that involve surrogate models represented by neural networks. We demonstrate first how to directly embed neural network evaluation into optimization models, highlight a difficulty with this approach that can prevent convergence, and then characterize stationarity of such models. We then present two alternative formulations of these problems in the specific case of feedforward neural networks with ReLU activation: as a mixed-integer optimization problem and as a mathematical program with complementarity constraints. For the latter formulation we prove that stationarity at a point for this problem corresponds to stationarity of the embedded formulation. Each of these formulations may be solved with state-of-the-art optimization methods, and we show how to obtain good initial feasible solutions for these methods. We compare our formulations on three practical applications arising in the design and control of combustion engines, in the generation of adversarial attacks on classifier networks, and in the determination of optimal flows in an oil well network.

97 MATHEMATICS AND COMPUTING↗

Decoding structure-spectrum relationships with physically organized latent spaces

Here, a semisupervised machine learning method for the discovery of structure-spectrum relationships is developed and then demonstrated using the specific example of interpreting x-ray absorption near-edge structure (XANES) spectra. This method constructs a one-to-one mapping between individual structure descriptors and spectral trends. Specifically, an adversarial autoencoder is augmented with a rank constraint (RankAAE). The RankAAE methodology produces a continuous and interpretable latent space, where each dimension can track an individual structure descriptor. As a part of this process, the model provides a robust and quantitative measure of the structure-spectrum relationship by decoupling intertwined spectral contributions from multiple structural characteristics. This makes it ideal for spectral interpretation and the discovery of descriptors. The capability of this procedure is showcased by considering five local structure descriptors and a database of >50 000 simulated XANES spectra across eight first-row transition metal oxide families. The resulting structure-spectrum relationships not only reproduce known trends in the literature but also reveal unintuitive ones that are visually indiscernible in large datasets. The results suggest that the RankAAE methodology has great potential to assist researchers in interpreting complex scientific data, testing physical hypotheses, and revealing patterns that extend scientific insight.

36 MATERIALS SCIENCE↗

Cross-Layered Cyber-Physical Power System State Estimation towards a Secure Grid Operation

In the Smart Grid paradigm, this critical infrastructure operation is increasingly exposed to cyber-threats due to the increased dependency on communication networks. An adversary can launch an attack on a power grid operation through False Data Injection into system measurements and/or through attacks on the communication network, such as flooding the communication channels with unnecessary data or intercepting messages. A cross-layered strategy that combines power grid data, communication grid monitoring and Machine Learning based processing is a promising solution for detecting cyberthreats. In this paper, an implementation of an integrated solution of a cross-layer framework is presented. The advantage of such a framework is the augmentation of valuable data that enhances the detection of anomalies in the operation of power grid. IEEE 118-bus system is built in Simulink to provide a power grid testing environment and communication network data is emulated using SimComponents. The performance of the framework is investigated under various FDI and communication attacks.

cyber security, network security, cyber-physical s↗

Predictive Data-driven Platform for Subsurface Energy Production

Subsurface energy activities such as unconventional resource recovery, enhanced geothermal energy systems, and geologic carbon storage require fast and reliable methods to account for complex, multiphysical processes in heterogeneous fractured and porous media. Although reservoir simulation is considered the industry standard for simulating these subsurface systems with injection and/or extraction operations, reservoir simulation requires spatio-temporal “Big Data” into the simulation model, which is typically a major challenge during model development and computational phase. In this work, we developed and applied various deep neural network-based approaches to (1) process multiscale image segmentation, (2) generate ensemble members of drainage networks, flow channels, and porous media using deep convolutional generative adversarial network, (3) construct multiple hybrid neural networks such as convolutional LSTM and convolutional neural network-LSTM to develop fast and accurate reduced order models for shale gas extraction, and (4) physics-informed neural network and deep Q-learning for flow and energy production. We hypothesized that physicsbased machine learning/deep learning can overcome the shortcomings of traditional machine learning methods where data-driven models have faltered beyond the data and physical conditions used for training and validation. We improved and developed novel approaches to demonstrate that physics-based ML can allow us to incorporate physical constraints (e.g., scientific domain knowledge) into ML framework. Outcomes of this project will be readily applicable for many energy and national security problems that are particularly defined by multiscale features and network systems.

58 GEOSCIENCES↗

Adversarial autoencoder ensemble for fast and probabilistic reconstructions of few-shot photon correlation functions for solid-state quantum emitters

Second-order photon correlation measurements [g (2) (τ) functions] are widely used to classify single-photon emission purity in quantum emitters or to measure the multiexciton quantum yield of emitters that can simultaneously host multiple excitations – such as quantum dots – by evaluating the value of g (2) (τ = 0). Accumulating enough photons to accurately calculate this value is time consuming and could be accelerated by fitting of few-shot photon correlations. Here, we develop an uncertainty-aware, deep adversarial autoencoder ensemble (AAE) that reconstructs noise-free g (2) (τ) functions from noise-dominated, few-shot inputs. The model is trained with simulated g (2) (τ) functions that are facilely generated by Poisson sampling time bins. The AAE reconstructions are performed orders-of-magnitude faster, with reconstruction errors and estimates of g (2) (τ = 0) that are lower in variance and similar in accuracy compared to Maximum likelihood estimation and Levenberg-Marquardt least-squares fitting approaches, for simulated and experimentally measured few-shot g (2) (τ) functions (~100 two-photon events) of InP/ZnS/ZnSe and CdS/CdSe/CdS quantum dots. The deep-ensemble model comprises eight individual autoencoders, allowing for probabilistic reconstructions of noise-free g (2) (τ) functions, and we show that the predicted variance scales inversely with number of shots, with comparable uncertainties to computationally intensive Markov chain Monte Carlo sampling. Furthermore, this work demonstrates the advantage of machine learning models to perform uncertainty-aware, fast, and accurate reconstructions of simple Poisson-distributed photon correlation functions, allowing for on-the-fly reconstructions and accelerated materials characterization of solid-state quantum emitters.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

DER Cybersecurity Detection and Response Suite

SAND2024-08475O The Distributed Energy Resource (DER) Cybersecurity Detection and Response Suite is a solution for distributed energy resource (DER) systems. The DER Security Orchestration, Automation, and Response (SOAR) solution that uses alerts from signature- and behavior-based Intrusion Detection Systems are intended to be deployed as bump-in-the-wire (BITW) devices in front of DER equipment. The fielded application would use multiple intrusion detection systems that report data to SOAR to respond to cyberattacks. The suite consists of two software components: • The proactive intrusion detection and mitigation system (PIDMS) secures grid-edge photovoltaic smart inverters and other equipment in distributed energy resource systems. It is a distributed BITW solution; cyber and physical data are automatically processed using network inspection tools and custom machine learning algorithms to detect abnormal events and correlate cyber-physical events. • The Security Orchestration, Automation, and Response for Distributed Energy Resources (SOAR4DER) application ingests data from several intrusion detection systems to quickly block attacks and revert DER systems to good states. Using a collection of intrusion detection system technologies on a BITW device, it incorporates physical and cyber data to detect abnormal and potential malicious behaviors. Multiple SOAR playbooks then use the intrusion detection system data streams to automatically defend the system. SOAR4DER system testing showed detection and response times under 30 seconds for all adversary reconnaissance, denial-of-service attacks, malicious Modbus commands, brute-force logins, and machine-in-the-middle attacks. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Johnson, Jay↗

Transforming the $v$ World: A New Multivariate Transformer Energy Estimator for NOvA

The NOvA Transformer Energy Estimator (Transformer_EE) is a universal machine learning tool currently used to infer the incoming beam neutrino energy and the outgoing lepton energy in both near andfar detectors. It uses a unique, highly flexible framework for simultaneous multivariate prediction that supports many possible loss functions. A spectral reweighting and flattening scheme lessens training bias. A feature noising subroutine enables adversarial-like training, mitigating sensitivities to certain systematic effects at marginal resolution loss at inference time. The state of the Transformer_EE will be reviewed, and its robustness with respect to several NOvA Near and Far Detector systematics highlighted.

Tong, Leon [Minnesota U.] (ORCID:0000000231625965)↗

Complete Evaluation on Advanced Reactor Machine Learning Subversion Attacks (Final)

Navigating through the world of Artificial Intelligence (AI) in nuclear reactors and their Instrumentation and Control (I&C) systems demands a careful, deliberate journey. AI’s capability to manage massive datasets and streamline control systems has indeed carved out a significant role in various sectors, including nuclear energy. However, while AI, and particularly Large Language Models (LLMs), bring a lot to the table in terms of operational efficiency and anomaly detection, they also expose the sector to a new breed of cybersecurity threats, like Inference Attacks, Adversarial Attacks, and Trojan Attacks. This guide is designed to be a straightforward manual, diving deep into the intertwining worlds of AI and cybersecurity within nuclear reactors, and tailoring insights for three crucial audiences: I&C Vendors/Developers, Nuclear Regulators, and Nuclear Reactor Operators and Cyber Defense Teams. (1) Section 2, directed at I&C Vendors/Developers, will provide a clear and focused look at several cybersecurity attacks, offering practical recommendations and detailed scenarios related to AI cybersecurity. This section isn’t just about identifying problems but also about giving solid, usable solutions. (2) Section 3, meant for Nuclear Regulators, gets straight to the point about regulations, policy suggestions, and guidelines that are needed to lay down a robust, secure, and ethical foundation for the application of AI in nuclear operations. The focus is on making sure that everything adheres to international standards and laws while being practicable and clear-cut. (3) Section 4, aimed at Nuclear Reactor Operators and Cyber Defense Teams, offers an exhaustive exploration and technical reports, with clear recommendations and scenario analyses vital to protect operational environments and guarantee the secure application of AI in nuclear reactor operations. The goal is simple: as we step into an era where AI becomes a fundamental element of our technological and energy infrastructures, this guide is here to act as a clear, direct handbook, ensuring that AI is implemented within the nuclear sector in a manner that is secure, responsible, and practical. It’s about striking a balance – optimizing the undeniable benefits offered by AI while securing and shielding against potential cyber threats as we move through this new and complex landscape.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Adaptive Cognitive Mechanisms to Maintain Calibrated Trust and Reliance in Automation

Trust calibration for a human–machine team is the process by which a human adjusts their expectations of the automation’s reliability and trustworthiness; adaptive support for trust calibration is needed to engender appropriate reliance on automation. Herein, we leverage an instance-based learning ACT-R cognitive model of decisions to obtain and rely on an automated assistant for visual search in a UAV interface. This cognitive model matches well with the human predictive power statistics measuring reliance decisions; we obtain from the model an internal estimate of automation reliability that mirrors human subjective ratings. The model is able to predict the effect of various potential disruptions, such as environmental changes or particular classes of adversarial intrusions on human trust in automation. Finally, we consider the use of model predictions to improve automation transparency that account for human cognitive biases in order to optimize the bidirectional interaction between human and machine through supporting trust calibration. The implications of our findings for the design of reliable and trustworthy automation are discussed.

60 APPLIED LIFE SCIENCES↗

Leveraging Natural Language Processing and Generative Models in Molecular Chemistry: Property Prediction and Novel Compound Generation

The accurate prediction of molecular properties is important for the rational design and the advancement of green chemistry and sustainable materials research. However, the predictive power of traditional computational chemistry methods is limited due to computational restrictions. Here, in this study, we examine an alternative approach to the accurate prediction of properties of organic compounds: natural language processing (NLP)-based molecular embedding. Using viscosity, partition coefficient (log P), and enthalpy of vaporization as test properties through a survey of comprehensive datasets comprising 5695 data points for viscosity, 25 870 data points for log P, and 2296 data points for enthalpy of vaporization. These are important properties for the design of greener, safer, and sustainable chemical processes. Models were trained using NLP methods such as Mol2vec and fine-tuned ChemBERTa, and results were compared with traditional input featurization techniques such as Morgan fingerprints and quantum chemistry derived sigma profiles and DFT features. Among the various machine learning models, Mol2vec demonstrated superior predictive capabilities, achieving the highest correlation coefficient (R 2 = 0.945) and lowest RMSE (0.106 mPa s) for viscosity, as well as high accuracy for log P and enthalpy of vaporization predictions. These findings establish the Mol2vec featurization technique, graph-convolutional neural networks (GCNN), and fine-tuned ChemBERTa model as powerful tools for predictive modeling of organic compounds properties, offering a significant improvement over previously used featurization techniques and opening up strategies for very-high-throughput computational screening. Finally, we integrated ML models with hybrid language-model-based generative adversarial networks (LM-GAN) to generate novel molecular sequences with desirable properties for different research applications. The ability to computationally design solvents with lower viscosity, lower log P, and lower enthalpy of vaporization offers a data-driven route to accelerating the discovery of sustainable alternatives to traditionally toxic solvents.

ChemBERTa↗

TRIM: AI Guided Random Number Generation for Resource-Constrained IoT Systems

Random numbers often serve as the backbone for many security solutions in diverse domains such as cryptography, side channel leakage prevention, and moving target defense. However, generating true random numbers requires a physical source of entropy (e.g. hardware, quantum, environmental phenomenon) making it difficult to realize at a large scale and at a low cost. On the flip side, pseudorandom number generators (easy to implement) following a specific distribution (e.g. Gaussian) can be easily compromised given a sufficient amount of traces. In this work, we have developed a machine learning-guided generative approach that can be used to create portable, resource-efficient, and cost-effective random number generators with high throughput and true randomness characteristics. We implement the proposed approach as a highly parameterized framework and perform extensive evaluation for different settings. The framework was able to learn from true random sources such as irrational numbers and environmental audio noise and imitate those sources towards generating new good quality random numbers on demand. We have generated more than 1 billion bits and observed robust performance in terms of true randomness metrics obtained from NIST SP 800-22 and FIPS 140-1 randomness test suites achieving a throughput of up to 142.85 Mbps. Compared to the state-of-the-art (SOTA) technique, the iso-cost setup of our framework can achieve more than 500 Mbps in a distributed setting. We have evaluated the efficacy of running the true randomness imitation AI models on target edge devices such as Raspberry Pi 4 (Model B), Nvidia Jetson Nano, Nvidia Jetson Orin Nano and Nvidia Jetson Xavier. We have also looked at the security of the TRIM framework itself against different adversarial threat models.

Cybersecurity↗

Device Feasibility Analysis of Multi-level FeFETs for Neuromorphic Computing

As an emerging non-volatile memory device technology, Ferroelectric Field-Effect Transistors (FeFETs) can enable low-power, adaptive intelligent system design. However, device dimension and operating voltage dependent reliability issues of scaled FeFETs can ultimately lead to degraded performance in solving machine learning tasks. In this article, detailed experimental characterization of FeFET devices of different dimensions have been carried out to explicitly evaluate the non-ideal behavior in device conductance programming properties like number of programming states, cycle-to-cycle (C2C) variations, device-to-device (D2D) variations, and state retention. A hardware-aware software simulation approach has been adopted to capture the adversarial effects of the non-idealities on recognition accuracy through algorithm-level performance assessment by including them in NeuroSim, a popular neural network hardware simulator, to execute a neural network model considering all other hardware constraints. With the added non-idealities, significant accuracy degradation has been observed compared to the ideal scenarios where D2D variations play the most critical role. Thereafter, feasibility of a variation-aware training method has been evaluated to tackle the accuracy drop.

42 ENGINEERING↗

Physics-driven learning of Wasserstein GAN for density reconstruction in dynamic tomography

Object density reconstruction from projections containing scattered radiation and noise is of critical importance in many applications. Existing scatter correction and density reconstruction methods may not provide the high accuracy needed in many applications and can break down in the presence of unmodeled or anomalous scatter and other experimental artifacts. Incorporating machine-learning models could prove beneficial for accurate density reconstruction, particularly in dynamic imaging, where the time evolution of the density fields could be captured by partial differential equations or by learning from hydrodynamics simulations. In this work, we demonstrate the ability of learned deep neural networks to perform artifact removal in noisy density reconstructions, where the noise is imperfectly characterized. Here, we use a Wasserstein generative adversarial network (WGAN), where the generator serves as a denoiser that removes artifacts in densities obtained from traditional reconstruction algorithms. We train the networks from large density time-series datasets, with noise simulated according to parametric random distributions that may mimic noise in experiments. The WGAN is trained with noisy density frames as generator inputs, to match the generator outputs to the distribution of clean densities (time series) from simulations. A supervised loss is also included in the training, which leads to an improved density restoration performance. In addition, we employ physics-based constraints such as mass conservation during the network training and application to further enable highly accurate density reconstructions. Our preliminary numerical results show that the models trained in our frameworks can remove significant portions of unknown noise in density time-series data.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Making Corgis Important for Honeycomb Classification: Adversarial Attacks on Concept-based Explainability Tools

Methods for model explainability have become increasingly critical for testing the fairness and soundness of deep learning. Concept-based interpretability techniques, which use a small set of human-interpretable concept exemplars in order to measure the influence of a concept on a model's internal representation of input, are an important thread in this line of research. In this work we show that these explainability methods can suffer the same vulnerability to adversarial attacks as the models they are meant to analyze. We demonstrate this phenomenon on two well-known concept-based interpretability methods: TCAV and faceted feature visualization. We show that by leveraging the geometry of the problem and carefully perturbing the examples of the concept that is being investigated, we can radically change the output of the interpretability method. The attacks that we propose can either induce positive interpretations (polka dots are an important concept for a model when classifying zebras) or negative interpretations (stripes are not an important factor in identifying images of a zebra). Our work highlights the fact that in safety-critical applications, there is need for security around not only the machine learning pipeline but also the model interpretation process.

Brown, Davis R.↗