Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Secure data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Explainable Machine Learning for Functional Data

Black-box machine learning models are recognized as useful tools for prediction applications, but the algorithmic complexity of some models causes interpretation challenges. Explainability methods have been proposed to provide insight into these models, but there is little research focused on supervised modeling with functional data inputs. We argue that, especially in applications of high consequence, it is important to explicitly model the functional dependence in a black-box analysis to not obscure or misrepresent patterns in explanations. As such, we propose the V ariable importance E xplainable E lastic S hape A nalysis (VEESA) pipeline for training supervised machine learning models with functional inputs. The pipeline is an analysis process that includes the data preprocessing, modeling, and post-hoc explanations. The preprocessing is done using elastic functional principal components analysis, which accounts for vertical and horizontal variability in functional data and, ultimately, allows for explanations in the original data space that identify the important functional variability without bias due to correlated variables. Here, we demonstrate the pipeline on two high-consequence applications: explosives classification for national security and inkjet printer identification in forensic science. The applications exhibit the VEESA pipeline’s ability to provide an understanding of the characteristics of the functional data useful for prediction. Code for implementing the pipeline is available in the veesa R package (and supplemental python code).

Elastic Shape Analysis

Genesis Mission-Enabled Secure AI to Fortify Energy Process Safety (Genesis-SAFE)

Argonne National Laboratory is supporting the U.S. Department of Transportation’s (USDOT’s) Bureau of Transportation Statistics (BTS) with collaborative research on development and application of privacy preserving AI frameworks that leverage unmatched AI expertise and secure computing resources made available through the U.S. Genesis Mission1 . This research advances U.S. energy security goals by supporting a safe offshore energy industry with secure, domain-specific AI tools to analyze confidential industry datasets collected by BTS to rapidly improve identification of hazards, precursors, and systemic safety risks in high-risk operational environments. The staged, security-first approach begins with development and testing of Argonne’s Genesis Mission-enabled Secure AI to Fortify Energy Process Safety (Genesis-SAFE) framework within Argonne’s accredited secure computing enclave (ABLE) leveraging Argonne’s AI scientific assistant substrate (AISAC). Methods to build synthetic datasets were developed together with BTS for use in preparing synthetic datasets that can be used to validate data containment, governance, and security controls in the ABLE environment. Future research directions would focus on applying the Genesis-SAFE framework to CIPSEA-protected datasets entirely within ABLE to support confidentiality-preserving analysis of safety risks, trends, and contributing factors.

Kim, Hyekyung [Argonne National Laboratory (ANL),

LTE Electrolyzer Data Collection

The goal for NREL is to collect, develop and publish performance metrics relative to low temperature electrolyzer installations. This will be done through the development of: Secure storage solution to house the collection of data from multiple projects Standardization of data to be collected and analyzed. This will be done using data templates developed with the help of partners involved with electrolyzer installations. Analysis that produces metrics of interest for all stakeholders Aggregation of results from multiple projects to view industry progress as a whole Publication of aggregated results in the form of composite data products (CDPs) Collaboration with Idaho National Lab and their work with high temperature electrolyzer installations will enable efficient use of storage and analysis tools.

data

Lightfall v0.0.1

Lightfall is a desktop application for synchrotron beamline instrument control, data acquisition, and live analysis at the Advanced Light Source (ALS). Built on Python and Qt, it provides a native graphical interface for operating beamline hardware, configuring and executing experimental scans, and visualizing results in real time. Key features include direct integration with EPICS control systems, a built-in electronic logbook, remote beamline access over secure tunnels, and an interprocess communication (IPC) architecture that coordinates with external analysis applications via ZMQ and EPICS process variables. This IPC approach allows Lightfall to orchestrate specialized analysis tools—including GPU-accelerated streaming correlators—without embedding them, avoiding the dependency conflicts common in monolithic scientific software platforms. Compared to prior approaches such as Xi-CAM's plugin-based architecture, Lightfall's design cleanly separates instrument control from domain-specific analysis, enabling feedback-driven acquisition where live analysis results can adjust scan parameters during an experiment. Its native Qt interface provides responsive performance for real-time data visualization that web-based alternatives struggle to match. Lightfall is designed for use by beamline scientists and staff operating synchrotron instruments at national user facilities.

Pandolfi, Ronald [Lawrence Berkeley National Labor

Scalable Data Center Capacity for DOE's AI Prototype: A Rapidly Available Gigawatt Data Center for DOE

The multilaboratory Gigawatt Data Center working group was commissioned to identify approaches to rapidly establish federal data centers with scalable capacities up to 1,000 MW. These state-of-the-art facilities will serve as hubs for interdisciplinary collaboration, industry partnerships, and transformative applications of artificial intelligence. The proposed strategic shift includes facilitating multilaboratory collaboration, prioritizing operational efficiency, expanding public–private partnerships, optimizing investments, ensuring long-term contractual flexibility, supporting open science and secure data enclaves, and exploiting high-speed national networks. Owing to their extensive experience and best practices, the US Department of Energy national laboratories are uniquely positioned to lead this initiative. We recommend conducting a feasibility analysis to rapidly identify the optimal sites for this initiative, and the effort will likely involve private industry for design, construction, financing, and operational integration. We also propose establishing multiple geographically diverse sites to ensure energy resilience, high operational reliability, and a diverse user base, thereby effectively addressing the nation’s critical needs.

42 ENGINEERING

Artificial Intelligence for Enhancing Multiscale Analysis: Buildings Focus

This project aims to develop multi-scale building energy data, potentially improving the representation of the U.S. buildings sector in GCAM-USA, an U.S.-focused human-energy-Earth systems model. Existing building energy datasets are typically limited to national or regional levels, which constrains the ability of models to capture fine-scale human-energy-Earth systems interactions and reduces their relevance for decision-making on issues such as energy security, resilience, and energy planning. By leveraging AI and advanced data integration methods, this work fuses multiple existing datasets to enhance the physical and geographic representation of both residential and commercial building energy use. So far, progress includes processing residential building data, designing the data structure for commercial buildings, and testing AI approaches for integrating datasets and addressing spatial-temporal gaps. This effort can not only advances GCAM-USA’s capability in modeling the buildings sector but also supports broader DOE missions, such as developing digital testbeds, enhancing grid resilience analysis, and improving building–energy system modeling at decision-relevant scales.

29 ENERGY PLANNING, POLICY, AND ECONOMY

Security Analysis of a Class of Spread Spectrum Systems Presentation

A method of adding physical layer security to a class of spread spectrum systems has been recently proposed. In this paper, we look into the rate at which an eavesdropper may gain information about the system to decipher the data symbols. The Shannon mutual information is used to measure the rate of information that may be gained by an eavesdropper. The k-nearest neighbors (k-NN) method is used to obtain estimates of relevant entropy values, which will then be used to quantify the rate of information recovery as more data is transmitted. It turns out that such information recovery requires the adoption of special methods that avoid any destructive bias in the estimates. Details of these methods are also presented.

97 - MATHEMATICS AND COMPUTING

Exponential Backoff and Its Security Implications for Safety-Critical OT Protocols over TCP/IP Networks

The convergence of Operational Technology (OT) and Information Technology (IT) networks has become increasingly prevalent with the growth of Industrial Internet of Things (IIoT) applications. This shift, while enabling enhanced automation, remote monitoring, and data sharing, also introduces new challenges related to communication latency and cybersecurity. Oftentimes, legacy OT protocols were adapted to the TCP/IP stack without an extensive review of the ramifications to their robustness, performance, or safety objectives. To further accommodate the IT/OT convergence, protocol gateways were introduced to facilitate the migration from serial protocols to TCP/IP protocol stacks within modern IT/OT infrastructure. However, they often introduce additional vulnerabilities by exposing traditionally isolated protocols to external threats. This study investigates the security and reliability implications of migrating serial protocols to TCP/IP stacks and the impact of protocol gateways, utilizing two widely used OT protocols: Modbus TCP and DNP3. Our protocol analysis finds a significant safety-critical vulnerability resulting from this migration, and our subsequent tests clearly demonstrate its presence and impact. A multi-tiered testbed, consisting of both physical and emulated components, is used to evaluate protocol performance and the effects of device-specific implementation flaws. Through this analysis of specifications and behaviors during communication interruptions, we identify critical differences in fault handling and the impact on time-sensitive data delivery. The findings highlight how reliance on lower-level IT protocols can undermine OT system resilience, and they inform the development of mitigation strategies to enhance the robustness of industrial communication networks.

DNP3

Improving Cyber Situational Understanding

Effective cybersecurity operations require the ability to analyze large amounts of information to assess security risks and formulate defensive strategies against adversaries. This has become more complex in recent years as the sprawl and interconnectivity of devices grows through implementation of virtualization, cloud computing, and Internet of Things (IoT). The amount of data and analysis required for effective cybersecurity command and control decisions far exceeds humans’ capacity to perform manually. We characterize the analysis problem as cyber situational understanding. The research presented to improve cyber situational understanding focuses on vulnerability analysis and threat intelligence. Regarding vulnerabilities, entities must analyze and plan work for between thousands and tens of thousands of software vulnerabilities annually. Entities heavily use network firewalls to limit vulnerability exposure. As a result, some of these vulnerabilities permit exposure to adversarial exploitation, whereas others are inaccessible and therefore present negligible risk of exploitation. Distinguishing between high and low risk software vulnerabilities requires a deep understanding of the vulnerability, network firewall protection, and characteristics of the targeted device. This problem is solved by extracting network service features from vulnerability data features using both machine-learning and natural language processing. Then, the network firewall topology is parsed to determine which vulnerabilities are reachable by adversaries. Ultimately, a state-based safety analysis ascertains which vulnerabilities are unsafe. A related vulnerability analysis problem occurs in cybersecurity operations when associating an entity’s hardware and software assets to public vulnerability databases. Assets often reveal hardware and software through installation artifacts and network service identification, and entities store these artifacts in inventory databases. However, software and hardware vendors apply a standard Common Platform Enumeration (CPE) naming convention when publicly reporting vulnerabilities. Associating these two datasets often requires many hours to days of manual inspection. The proposed solution automates the mapping approach of human analysts using fuzzy matching techniques, natural language processing, and, ultimately, machine learning to present a small set of recommendations for mapping the two datasets. The result significantly reduces human analysis time and reduces the occurrence of false positives in vulnerability notifications. Finally, cyber threat intelligence (CTI) requires associating cyber observable artifacts, such as IP addresses, URIs, and file hashes, with cyber threat tactics, techniques, and procedures. Unfortunately, most CTI data is compartmentalized across multiple organizations and cannot be shared due to the legal and reputational risk with cyber threat being associated with the entity. The approach to solving this problem inovlves using a distributed ledger with anonymous token spending and authentication. This allows a consortium of semi-trusted entities to share the workload of curating CTI for a threat sharing community’s cooperative benefit.

Huff, Philip

Data Centers and Digital Assurance Introduction to Supply Chain and Cybersecurity for Data Centers, Session 1

The first session of the TADA (Technical Assistance for Digital Assurance) Data Centers Cohort Workshop, held on October 30, 2025, introduced foundational concepts of Digital Assurance in the context of data center and grid integration. Sponsored by the U.S. Department of Energy, the workshop brought together utilities, data center operators, developers, and vendors to address cybersecurity and supply chain vulnerabilities. The session emphasized the growing criticality of data centers within the electric grid and the need for secure, real-time, bidirectional communication. Participants explored the principles of Digital Assurance, including cybersecurity, cyber-informed engineering (CIE), and lifecycle security, and applied a threat-vulnerability-consequence framework to identify and mitigate risks at the data center–grid interface. Discussions covered a range of threats such as spoofed dispatch signals and insider threats, architectural vulnerabilities like SCADA interfaces and insecure protocols, and potential consequences including cascading grid failures. The session also raised strategic questions about business value, vendor assurance, and defining cyber boundaries and responsibilities. This foundational workshop set the stage for deeper technical analysis and the development of actionable frameworks in subsequent sessions. Session 1 of 3.

24 - POWER TRANSMISSION AND DISTRIBUTION

Enabling Seamless Transitions from Experimental to Production HPC for Interactive Workflows

The evolving landscape of scientific computing requires seamless transitions from experimental to production HPC environments for interactive workflows. This paper presents a structured transition pathway developed at OLCF that bridges the gap between development testbeds and production systems. We address both technological and policy challenges, introducing frameworks for data streaming architectures, secure service interfaces, and adaptive resource scheduling for time-sensitive workloads and improved HPC interactivity. Our approach transforms traditional batch-oriented HPC into a more dynamic ecosystem capable of supporting modern scientific workflows that require near real-time data analysis, experimental steering, and cross-facility integration.

Etz, Brian [ORNL] (ORCID:0000000208554863)

Scalable and Secure Power Outage Data Reporting: A Hexagonal Geospatial Approach

Power outages disrupt critical infrastructure and cause billions of dollars in economic losses annually in the United States. Accurate and granular outage reporting is vital for effective restoration and mitigation. This paper examines the integration of the Hexagonal Hierarchical Geospatial Indexing System (H3) to enhance power outage reporting, leveraging its uniform grid structure, scalable resolutions, and support for privacy-preserving analysis. Using high-resolution LandScan Global population data and K-anonymization techniques, this work achieves a balance between data granularity and privacy. Results show that lower privacy thresholds (e.g., K-anonymity = 2) enable higher resolution, while stricter thresholds (e.g., >15 people per hex) reduce granularity, potentially affecting localized responses. State-and county-level resolution case studies demonstrate H3’s adaptability and the trade-offs between precision and privacy. The proposed H3-based framework offers a scalable and efficient solution for geospatial data integration within the energy sector, such as outage data, aiding utilities and regulators in improving resilience and response efforts, particularly in disaster-prone regions.

Ahmad, Nasir [ORNL] (ORCID:0000000150677368)

Machine Learning for Anomaly Detection in Neural Network Security and SRF Cavities

This dissertation explores the development and deployment of machine learning approaches to address critical challenges in anomaly detection across two distinct domains: neural network security in federated learning settings and cavity behavior analysis in particle accelerator operations at Jefferson Lab in Newport News, Virginia. Anomaly detection identifies deviations from expected patterns, safeguarding systems in cybersecurity, industry, and research against malicious activities and failures. This dissertation demonstrates how our machine learning approaches enhance detection accuracy and efficiency in both neural network security and industrial applications. First, we investigate vulnerabilities in deep neural networks deployed in federated learning. Although federated learning preserves user privacy by training models locally, it remains vulnerable to backdoor attacks, in which malicious participants embed hidden triggers that induce targeted misbehavior. We propose a self-supervised contrastive learning framework to detect and mitigate such backdoor attacks. In our experiments, this method achieves higher detection accuracy and lower false positive rates than existing defenses, while operating without access to local model updates or original training data and thus preserving the privacy guarantees of the federated setting. Second, we address the operational reliability of superconducting radio-frequency (SRF) cavities at the Continuous Electron Beam Accelerator Facility (CEBAF). Our research leverages an unsupervised learning approach, combined with Principal Component Analysis (PCA) and k-means clustering, to identify anomalous behaviors in SRF cavities. Our method detects subtle anomalous behavior by analyzing SRF signal data. This knowledge allows for the early detection and resolution of potential faults, significantly improving the efficiency and reliability of operations. Third, we extend these insights to time-series anomaly detection more broadly. We design a contrastive-learning based model tailored to increasingly dynamic environments and academic research. This model improves detection accuracy in settings that require real-time monitoring and predictive maintenance. Our research underscores the broader applicability and impact of advanced machine learning techniques in anomaly detection. By extracting meaningful patterns from complex data, machine learning can significantly enhance security in distributed neural networks and improve the efficiency of particle accelerator operations. This dissertation serves as a stepping stone for future investigations into the vast possibilities of anomaly detection, inspiring further exploration and development of machine learning techniques in this field.

Ferguson, Hal [Old Dominion University]

Elucidating the geometric and electronic structure of a fully sulfided analog of an Anderson polyoxomolybdate cluster

The catalytic activity of transition metal sulfide (TMS) clusters in small molecule activation, redox transformations, and charge transfer has inspired the design of novel TMS-based materials for energy-related catalysis and chemical applications. Polyoxometalates (POMs), known for their structural diversity, can in principle be transformed into TMS clusters; however, fully sulfided analogs are rarely isolated, likely due to the strong tendency of uncapped TMS clusters to agglomerate. Here, we report the geometric and electronic structure of a capping ligand-free fully sulfided analog of heptamolybdate Anderson POM [Mo VI 7 O 24 ] 6− , synthesized through the sulfidation of a nanoconfined POM secured within a porous Zr-metal organic framework (NU-1000). A combined computational and experimental analysis indicates that the sulfided counterpart of the Anderson POM is geometrically and electronically more sophisticated than the parent POM. Comparison of experimental pair distribution function (PDF) data with computational simulations confirms that, unlike the oxygen-only [Mo VI 7 O 24 ] 6− cluster, the [Mo IV 7 (μ 3 -S) 6 (μ 2 -SH) 6 (S 2 ) 6 ] 2− polythiometalate (PTM) exhibits diverse sulfur anions (S 2− , HS − , S 2 2− ). DFT calculations indicate that H 2 S acts as a reducing agent, and together with terminal disulfide (S 2 2− ) ligands in the PTM structure, facilitates the complete reduction of all seven Mo VI centers in the parent POM to Mo IV . These findings are supported by X-ray photoelectron spectroscopy (XPS), which confirms exclusive Mo IV , and elemental analysis, which shows quantitative sulfur incorporation. Difference envelope density (DED) mapping further reveals that the PTM clusters are spatially confined within the MOF pores, preventing agglomeration and preserving molecular integrity.

Rabbani, S. M. Gulam [The Ohio State University, C

HPC and Cloud Convergence Beyond Technical Boundaries: Strategies for Economic Sustainability, Standardization, and Data Accessibility

At the IEEE/ACM International Conference for High-Performance Computing, Networking, Storage, and Analysis (SC23), held in Denver, experts discussed the convergence of high-performance computing and cloud computing. Experts explored how this integration could address current scientific computing limitations, enhance computational capabilities, and foster global collaboration while focusing on economic, security, technical, and community challenges and opportunities.

97 MATHEMATICS AND COMPUTING

Generator Frequency Response Droop Monitoring Tool

Monitoring and analyzing the frequency response performance of power generation units is essential for maintaining reliable and secure power system operation. To address this need, an automation tool has been developed to provide a pipeline for processing historical power plant generation data, including large-scale SCADA archives. The tool performs end-to-end processing, including event detection, frequency response (FR) analysis in accordance with NERC standards, and estimation of speed governor droop characteristics. The tool is designed with a modular architecture, allowing individual components of the workflow to be extended, customized, or deployed independently. In addition, the tool provides an API that enables seamless integration with other production systems and operational analytics platforms.

Etingov, PavelV [Pacific Northwest National Labora

ELG spectroscopic systematics analysis of the DESI Data Release 1

Dark Energy Spectroscopic Instrument (DESI) uses more than 2.4 million Emission Line Galaxies (ELGs) for 3D large-scale structure (LSS) analyses in its Data Release 1 (DR1). Such large statistics enable thorough research on systematic uncertainties. In this study, we focus on spectroscopic systematics of ELGs. The redshift success rate (f goodz ) is the relative fraction of secure redshifts among all measurements. It depends on observing conditions, thus introduces non-cosmological variations to the LSS. We, therefore, develop the redshift failure weight (w zfail ) and a per-fibre correction (η zfail ) to mitigate these dependences. They have minor influences on the galaxy clustering. For ELGs with a secure redshift, there are two subtypes of systematics: 1) catastrophics (large) that only occur in a few samples; 2) redshift uncertainty (small) that exists for all samples. The catastrophics represent 0.26% of the total DR1 ELGs, composed of the confusion between [O ii] and sky residuals, double objects, total catastrophics and others. We simulate the realistic 0.26% catastrophics of DR1 ELGs, the hypothetical 1% catastrophics, and the truncation of the contaminated 1.31 < z < 1.33 in the A BACUS S UMMIT ELG mocks. Their P ℓ show non-negligible bias from the uncontaminated mocks. But their influences on the redshift space distortions (RSD) parameters are smaller than 0.2σ. The redshift uncertainty of DR1 ELGs is 8.5km s -1 with a Lorentzian profile. The code for implementing the catastrophics and redshift uncertainty on mocks can be found in https://github.com/Jiaxi-Yu/modelling_spectro_sys.

79 ASTRONOMY AND ASTROPHYSICS

Nonparametric Multiparticle Set Methods for Interpreting Environmental Samples

Collection and analysis of environmental samples is commonly used by a range of stakeholders in nuclear safeguards and security contexts. While the ubiquity of samples and their transport in the environment allow regular collection, developing and demonstrating methods for analyzing these samples is difficult. In this work, an environmental sample consists of a set of one or more individual particles. Recent advances in reactor simulation have allowed us to generate data that are more representative of real-world environmental samples, enabling statistically defensible method development and testing. The most notable of these advances is a drastic increase in the number of material depletion regions, which allows our simulations to capture the variation in isotopic composition seen at length scales consistent with environmental samples. Traditional approaches for handling multiparticle samples treat each particle in the sample individually, estimating the quantity of interest (e.g., core-average burnup) resulting from measurement and analysis of signatures (e.g., nuclide assays) from each individual particle. Individual estimates are then averaged to generate a single estimate of the quantity of interest over the entire sample. In this presentation, we introduce two novel approaches for interpreting environmental samples that comprise of multiple particles: (1) the Quantile-Quantile Comparator, which uses a multivariate generalization of quantile-quantile plots for comparing unknown statistical distributions, and (2) the Set Transformer, an attention-based neural network module designed to model interactions among elements (particles) in the input set (sample). Statistically representative sampling cannot be guaranteed as samples are passively collected and are beholden to what particles are available in the environment. These new analysis methods for set-input problems are expected to be more robust than traditional approaches to issues of sampling bias where particles are not uniformly distributed throughout regions of interest, as well as generally outperform traditional approaches by jointly considering all elements in the set. We will present results comparing the performance of traditional single particle approaches and the novel Quantile-Quantile Comparator and Set Transformer for interpretation of simulated environmental samples.

Phathanapirom, Birdy