Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “adversarial evaluation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Towards Informing an Intuitive Mission Planning Interface for Autonomous Multi-Asset Teams via Image Descriptions

Establishing a basis for certification of autonomous systems using trust and trustworthiness is the focus of Autonomy Teaming and TRAjectories for Complex Trusted Operational Reliability (ATTRACTOR). The Human-Machine Interface (HMI) team is working to capture and utilize the multitude of ways in which humans are already comfortable communicating mission goals and translate that into an intuitive mission planning interface. Several input/output modalities (speech/audio, typing/text, touch, and gesture) are being considered and investigated in the context human-machine teaming for the ATTRACTOR design reference mission (DRM) of Search and Rescue or (more generally) intelligence, surveillance, and reconnaissance (ISR). The first of these investigations, the Human Informed Natural-language GANs Evaluation (HINGE) data collection effort, is aimed at building an image description database to train a Generative Adversarial Network (GAN). In addition to building an image description database, the HMI team was interested if, and how, modality (spoken vs. written) affects different aspects of the image description given. The results will be analyzed to better inform the designing of an interface for mission planning.

Generative Adversarial Network (GAN)↗

Assessing Cybersecurity Resilience of Distributed Ledger Technology in Energy Sector Using the MITRE ATT&CK® ICS Framework

Digitization in the power industry enables wide connectivity among multiple new entrants such as DERs, prosumers, and P2P counterparts within or outside the Distributed Ledger Technology (DLT). The use of DLT to improve resilience in the power grid has growing support, but new technology provides new opportunities for adversaries to cause harm. This work completed by the Cybersecurity- focused task force of IEEE SA P2418.5 evaluates the potential risks by applying the MITRE ATT&CK® ICS matrix to the DLT Engineering and Cybersecurity Stack designed for power systems applications

Gourisetti, Sri Nikhil Gupta↗

Optimal Mitigation Planning For Adversarial Scenarios

We propose a generalized framework which performs an optimal partitioning of a limited budget into various organizational sectors in order to improve the cybersecurity of a smart device or component in the Cyber Physical Energy System (CPS). The framework identifies the adversarial threats and possible attack sequences which can be performed to exploit cyber vulnerabilities of the component. Thereafter, we formulate an Mixed Integer Linear Programming (MILP) optimization problem which aims to evaluate the optimal budget partitions in order to minimize the number of highly likely attack sequences. Though we provide results for using the framework in CPES, the proposed methodology can be extended for multiple domains with a set of known adversarial and mitigation actions.

Purohit, Sumit [Pacific Northwest National Laborat↗

The Double-edged Sword of Data-driven Super-Resolution: Adversarial Super-resolution Models

Data-driven super-resolution (SR) methods are often integrated into imaging pipelines as preprocessing steps to improve downstream tasks such as classification and detection. However, these SR models introduce a previously unexplored attack surface into imaging pipelines. In this paper, we present AdvSR, a framework demonstrating that adversarial behavior can be embedded directly into SR model weights during training, requiring no access to inputs at inference time. Unlike prior attacks that perturb inputs or rely on backdoor triggers, AdvSR operates entirely at the model level. By jointly optimizing for reconstruction quality and targeted adversarial outcomes, AdvSR produces models that appear benign under standard image quality metrics while inducing downstream misclassification. We evaluate AdvSR on three SR architectures (SRCNN, EDSR, SwinIR) paired with a YOLOv11 classifier and demonstrate that AdvSR models can achieve high attack success rates with minimal quality degradation. These findings highlight a new model-level threat for imaging pipelines, with implications for how practitioners source and validate models in safety-critical applications.

Sullivan, Haley [ORNL] (ORCID:0000000274069217)↗

Adversarial autoencoder ensemble for fast and probabilistic reconstructions of few-shot photon correlation functions for solid-state quantum emitters

Second-order photon correlation measurements [g (2) (τ) functions] are widely used to classify single-photon emission purity in quantum emitters or to measure the multiexciton quantum yield of emitters that can simultaneously host multiple excitations – such as quantum dots – by evaluating the value of g (2) (τ = 0). Accumulating enough photons to accurately calculate this value is time consuming and could be accelerated by fitting of few-shot photon correlations. Here, we develop an uncertainty-aware, deep adversarial autoencoder ensemble (AAE) that reconstructs noise-free g (2) (τ) functions from noise-dominated, few-shot inputs. The model is trained with simulated g (2) (τ) functions that are facilely generated by Poisson sampling time bins. The AAE reconstructions are performed orders-of-magnitude faster, with reconstruction errors and estimates of g (2) (τ = 0) that are lower in variance and similar in accuracy compared to Maximum likelihood estimation and Levenberg-Marquardt least-squares fitting approaches, for simulated and experimentally measured few-shot g (2) (τ) functions (~100 two-photon events) of InP/ZnS/ZnSe and CdS/CdSe/CdS quantum dots. The deep-ensemble model comprises eight individual autoencoders, allowing for probabilistic reconstructions of noise-free g (2) (τ) functions, and we show that the predicted variance scales inversely with number of shots, with comparable uncertainties to computationally intensive Markov chain Monte Carlo sampling. Furthermore, this work demonstrates the advantage of machine learning models to perform uncertainty-aware, fast, and accurate reconstructions of simple Poisson-distributed photon correlation functions, allowing for on-the-fly reconstructions and accelerated materials characterization of solid-state quantum emitters.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Security Evaluation of Smart Cards and Secure Tokens: Benefits and Drawbacks for Reducing Supply Chain Risks of Nuclear Power Plants

The supply chain attack pathway is being increasingly used by adversaries to bypass security controls and gain unauthorized access to sensitive networks and equipment (e.g., Critical Digital Assets). Cyber-attacks targeting supply chain generally aim to compromise the environments, products, or services of vendors and suppliers to inject, add, or substitute authentic software and hardware with malicious elements. These malicious elements are deemed to be authentic as they arise from the vendor or supplier (i.e., the supply chain). This research aims to leverage findings and assumptions made from the previous report to determine the security benefits and drawbacks of a smart card- based hardware root of trust. Smart cards can provide devices inside Nuclear Power Plants (NPP) with a secure environment to store keys in and perform sensitive operations such as digital signature generation. These abilities can be leveraged to increase supply chain cybersecurity by autonomously providing NPP Licensees with reports on device integrity, authenticity and measurements of executable and non-executable data.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Diagnosing autonomous vehicle driving criteria with an adversarial evolutionary algorithm

We repurposed an adversarial evolutionary algorithm, Gremlin, from finding driving scenarios where a model of an autonomous vehicle drove poorly to troubleshooting driving quality evaluation criteria. We evaluated the driving performance of a "perfect driver" robot in a virtual town environment using the same fitness criteria intended for a deep learner (DL) trained driver. We found that the fitness evaluation criteria poorly handled turns, and used Gremlin to iteratively improve that criteria. We were confident that the same criteria could then be applied to the DL-based models as originally intended, and that this approach could be used as a general means of troubleshooting autonomous vehicle driving criteria.

Coletti, Mark↗

Adaptive Stress Testing of Trajectory Predictions in Flight Management Systems

To find failure events and their likelihoods in flight-critical systems, we investigate the use of an advanced black-box stress testing approach called adaptive stress testing. We analyze a trajectory predictor from a developmental commercial flight management system which takes as input a collection of lateral waypoints and en-route environmental conditions. Our aim is to search for failure events relating to inconsistencies in the predicted lateral trajectories. The intention of this work is to find likely failures and report them back to the developers so they can address and potentially resolve shortcomings of the system before deployment. To improve search performance, this work extends the adaptive stress testing formulation to be applied more generally to sequential decision-making problems with episodic reward by collecting the state transitions during the search and evaluating at the end of the simulated rollout. We use a modified Monte Carlo tree search algorithm with progressive widening as our adversarial reinforcement learner. The performance is compared to direct Monte Carlo simulations and to the cross-entropy method as an alternative importance sampling baseline. The goal is to find potential problems otherwise not found by traditional requirements-based testing. Results indicate that our adaptive stress testing approach finds more failures and finds failures with higher likelihood relative to the baseline approaches.

adaptive stress testing↗

Model Assumptions and Data Characteristics: Impacts on Domain Adaptation in Building Segmentation

Studies on domain adaptation (DA) for remote sensing (RS) imagery analysis lack consistency in selection and description of evaluation scenarios. Without properly characterizing datasets, model assumptions, and evaluation scenarios, it is difficult to objectively compare DA methods and reach conclusions about their suitability across different applications. With this motivation, this work seeks to empirically assess to which extent the interaction between data characteristics and model assumptions influences the effectiveness of DA methods. Using the widely explored task of building footprint segmentation as a case study, we perform a large-scale study across over 200 DA scenarios that include variations across view angles, areas observed, and sensors used for data acquisition. Rather than adopting different model architectures or optimization criteria, we contrast the performances of two DA methods based on adversarial learning that differ only in their assumptions about source and target domains. Informed by metadata and data characteristics unveiled using traditional computer vision (CV) techniques as well as pretrained deep models, we provide a detailed meta-analysis of experiments highlighting the importance of accurately considering data assumptions for DA in RS segmentation tasks. As demonstrated by a “cherry-picking” exercise, different claims regarding which model is best could be made by selecting different subsets of evaluation scenarios. While well-calibrated assumptions can be beneficial, mismatching assumptions can lead to negative biases in DA applications. Furthermore, this study intends to motivate the community toward more consistent evaluation protocols while providing recommendations and insights toward creating novel benchmark datasets, documenting data characteristics, application-specific knowledge, and model assumptions.

42 ENGINEERING↗

RanCompute: Computational Security in Embedded Devices via Random Input and Output Encodings

An embedded device in an insecure environment is subject to additional security risk through capture and reverse-engineering by a capable adversary. If this device contains a microchip performing sensitive computations, capture of the chip may leak functionality to an adversary. In this paper we propose a novel method in which we randomly encode the input operands and the outputs of a computation, thus not revealing the arithmetic operations being performed. The operations are sequenced in a graph representing the overall application. Once the initialization values are overwritten and lost, the results of these computations are indecipherable by the device performing the calculations as well as by any adversary. The result is transmitted back to a secure server which has stored the initialization values and so can decode the results which appear random to the adversary.

Embedded computing↗

SoK: What does it Mean to Benchmark Database Forensics?

Relational Database Management Systems are the backbone of modern enterprises and public-sector services, and are thus frequent targets of security incidents, insider threats, and thorough regulatory audits. Consequently, databases have become key sources of digital evidence, requiring investigators to reconstruct past activity from audit logs, transaction logs, and backups. Although benchmarking frameworks such as those developed by the Transaction Processing Performance Council (TPC) are widely used to evaluate database performance, they do not capture forensic requirements such as evidentiary completeness, tamper-evidence, chain of custody, or regulatory compliance under GDPR and CCPA. This survey examines the emerging domain of forensic database benchmarking. We gathered prior research on database forensics, secure logging, and tamper-evident data structures; we analyze modern forensic-ready features in commercial and open-source systems (SQL Server Ledger, Oracle Blockchain Tables, PostgreSQL pgAudit, Db2 Audit, Aurora Database Activity Streams, Oracle Real Application Security and IBM Guardium) and assess why existing benchmarks are insufficient. We propose forensic workloads, metrics, and methodologies that incorporate adversarial stressors, deleted-record recovery, and backup analysis. We also identify open research problems and call for a community-driven forensic benchmark suite. The result is an idea for evaluating not only database performance but also forensic soundness, bridging the gap between system engineering, compliance, and digital investigations.

Lenard, Ben↗

ARM-IRL: Adaptive Resilience Metric Quantification Using Inverse Reinforcement Learning

The resilience of safety-critical systems is gaining importance due to the rise in cyber and physical threats, especially within critical infrastructure. Traditional static resilience metrics may not capture dynamic system states, leading to inaccurate assessments and ineffective responses to cyber threats. This work aims to develop a data-driven, adaptive method for resilience metric learning. We propose a data-driven approach using inverse reinforcement learning (IRL) to learn a single, adaptive resilience metric. The method infers a reward function from expert control actions. Unlike previous approaches using static weights or fuzzy logic, this work applies adversarial inverse reinforcement learning (AIRL), training a generator and discriminator in parallel to learn the reward structure and derive an optimal policy. The proposed approach is evaluated on multiple scenarios: optimal communication network rerouting, power distribution network reconfiguration, and cyber–physical restoration of critical loads using the IEEE 123-bus system. The adaptive, learned resilience metric enables faster critical load restoration in comparison to conventional RL approaches.

97 MATHEMATICS AND COMPUTING↗

Evaluating Direct and Indirect Influence on EV Charging Stations Across the US

The adoption of new technology for electric vehicles (EV) and mobility applications can bring underappreciated vulnerabilities to the power grid. One area of potential fraud and adversarial influence is through the business ecosystem of startups that own and deploy EV technology. Yet, there are no models or analyses that map the network of organizations and people that have direct and indirect influence over technologies currently deployed in the grid. To fill this gap, we develop a multilayer network model to measure direct and indirect influence on EV charging stations. First, we create and adversarial socio-technical network (ASTN) model via a data fusion pipeline for different US regions of interest (ROI). Then, we develop an integrated ASTN for Chicago, Los Angeles, New York, and Philadelphia. We rank EV charging companies direct influence within each geographic region as well as indirect influence via social network analysis. While some companies have strong direct and indirect influence (i.e., ChargePoint) others show a mismatch between their influence over charging stations and their position within the social network. For example, Tesla has strong direct influence on stations and weak indirect influence over competitors. In contrast, 7Charge has weak direct influence over stations, but strong indirect influence over competitors.

29 - ENERGY PLANNING, POLICY AND ECONOMY↗

Understanding Impacts of Data Integrity Attacks on Transactive Control Systems

With the rapid growth of internet-connected smart devices capable of exchanging energy price information and adaptively controlling energy consumption of connected loads, the role of transactive control is expected to become more prominent in the modern grid. Transactive control systems integrate the wholesale and retail energy markets, and enable active participation of end users, thereby playing a key role in managing the rising number of distributed assets in the smart grid. The use of internet for the communication of data between the controllers at the building, distribution, and transmission levels makes the system susceptible to cyber-attacks. A skilled adversary can potentially manipulate the exchanged data with the intention of inflicting damage to the system. In this paper, four categories of metrics are put forth for evaluating the performance of the system when under data integrity attacks. A model for scaling-based data integrity attack has been described, and co-simulations have been performed on a 240 bus WECC transmission system model with detailed modeling of distribution systems at specific buses. Our results show that scaling-based data integrity attacks can have non-trivial impacts on the operational, financial, and comfort aspects of the transactive control system.

Smart grid, transactive control system, Buildings-↗

Synthetic Scientific Image Generation with VAE, GAN, and Diffusion Model Architectures

Generative AI (genAI) has emerged as a powerful tool for synthesizing diverse and complex image data, offering new possibilities for scientific imaging applications. This review presents a comprehensive comparative analysis of leading generative architectures, ranging from Variational Autoencoders (VAEs) to Generative Adversarial Networks (GANs) on through to Diffusion Models, in the context of scientific image synthesis. We examine each model's foundational principles, recent architectural advancements, and practical trade-offs. Our evaluation, conducted on domain-specific datasets including microCT scans of rocks and composite fibers, as well as high-resolution images of plant roots, integrates both quantitative metrics (SSIM, LPIPS, FID, CLIPScore) and expert-driven qualitative assessments. Results show that GANs, particularly StyleGAN, produce images with high perceptual quality and structural coherence. Diffusion-based models for inpainting and image variation, such as DALL-E 2, delivered high realism and semantic alignment but generally struggled in balancing visual fidelity with scientific accuracy. Importantly, our findings reveal limitations of standard quantitative metrics in capturing scientific relevance, underscoring the need for domain-expert validation. We conclude by discussing key challenges such as model interpretability, computational cost, and verification protocols, and discuss future directions where generative AI can drive innovation in data augmentation, simulation, and hypothesis generation in scientific research.

Generative Adversarial Networks↗

Device Feasibility Analysis of Multi-level FeFETs for Neuromorphic Computing

As an emerging non-volatile memory device technology, Ferroelectric Field-Effect Transistors (FeFETs) can enable low-power, adaptive intelligent system design. However, device dimension and operating voltage dependent reliability issues of scaled FeFETs can ultimately lead to degraded performance in solving machine learning tasks. In this article, detailed experimental characterization of FeFET devices of different dimensions have been carried out to explicitly evaluate the non-ideal behavior in device conductance programming properties like number of programming states, cycle-to-cycle (C2C) variations, device-to-device (D2D) variations, and state retention. A hardware-aware software simulation approach has been adopted to capture the adversarial effects of the non-idealities on recognition accuracy through algorithm-level performance assessment by including them in NeuroSim, a popular neural network hardware simulator, to execute a neural network model considering all other hardware constraints. With the added non-idealities, significant accuracy degradation has been observed compared to the ideal scenarios where D2D variations play the most critical role. Thereafter, feasibility of a variation-aware training method has been evaluated to tackle the accuracy drop.

42 ENGINEERING↗

ESM data downscaling: a comparison of super-resolution deep learning models

Abstract Climate projections at fine spatial resolutions are required to conduct accurate risk assessment for critical infrastructure and design adaptation planning. Generating these projections using advanced Earth system models (ESM) requires significant computational resources. To address this issue, various statistical downscaling techniques have been introduced to generate fine-resolution data from coarse-resolution simulations. In this study, we evaluate and compare five deep learning-based downscaling techniques, namely, super-resolution convolutional neural networks, fast super-resolution convolutional neural network ESM, efficient sub-pixel convolutional neural network, enhanced deep residual network (EDRN), and super-resolution generative adversarial network (SRGAN). These techniques are applied to a dataset generated by the Energy Exascale Earth System Model (E3SM), focusing on key surface variables such as surface temperature, shortwave heat flux, and longwave heat flux. Models are trained and validated using paired fine-resolution (0.25 $$^{\circ }$$ ∘ ) and coarse-resolution (1 $$^{\circ }$$ ∘ ) monthly data obtained from a 9-year simulation. Next, blind testing is performed using monthly data obtained from two different years outside of the training and validation set. To evaluate the efficiency of each technique, different statistical metrics are used, including mean squared error (MSE), peak signal-to-noise ratio (PSNR), structural similarity index measure (SSIM), and learned perceptual image patch similarity (LPIPS). The results show that EDRN outperforms other algorithms in terms of PSNR, SSIM, and MSE, but struggles to capture fine-scale features in the data. In contrast, SRGAN, a generative model that uses perceptual loss, excels in capturing fine details at boundaries and internal structures, resulting in lower LPIPS than other methods.

Pawar, Nikhil M. (ORCID:0000000211613289)↗

Advanced Reactor Designs Security Analysis, Risk, and Recommendations: Risks, Consequences, and Possible by-Design Mitigation Approaches Associated with Select Advanced Reactors

Next-generation advanced reactors (ARs) incorporate enhanced safety systems, have smaller source terms, and feature compact modular designs, which should lessen their collective risk profiles. However, to fully evaluate risk, security needs to be a part of the equation. Without taking security into consideration, safety systems and components in the new ARs may be vulnerable to sabotage. These base attributes, coupled with enhanced security features specific to AR design through sound engineering and security-by-design (SeBD), should provide developers and operators with lower inherent security risk profiles. Building security early into the AR design may remove or passively secure potential critical targets from an adversary’s reach , thereby increasing overall safety and security. An integrated approach and diverse design team that includes engineering, operations, and security experts are fundamental to building security into the design without sacrificing fundamental operational efficiencies and principles. The objective of this project was to evaluate the security and safety interfaces for five classes of reactors, identify potential security vulnerabilities of structures, systems, and components (SSC), and underscore the need to consider security alongside safety in the design o f these concepts. The five reactor classes evaluated in this project and presented in this report are molten-salt reactors (MSR), high temperature gas reactors (HTGR), sodium-fast reactors (SFR), advanced light-water reactors (ALWR), and microreactors. These designs were selected because they reflect the concepts that are closest to market deployment and have received significant resource investments from the public and private sector. This project assesses the inherent security risks posed by common classes of ARs, provides a methodology and framework to assess security along with safety, and offers an analysis of potential mitigation strategies that could be incorporated. For each AR technology, the SSCs that relate to radionuclide source safety functions are discussed to understand the SSC contribution to safety and relative importance in the protective strategy for the design. The assumptions that went into evaluating each reactor concept originated from generic publicly available nonproprietary information and should not directly be used to qualify an absolute risk profile nor to rank specific AR designs. Instead, the purpose of the analysis is to understand and compare the generic inherent security risks of different AR technologies.

98 - NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL↗