Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Secure data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Tight Practical Bounds for Subgraph Densities in Ego-centric Networks

SAND2025-11782O Tight Practical Bounds for Subgraph Densities in Ego-centric Networks is a software tool for calculating the “subgraph spread ratio” for social network analysis. This value is useful in network analysis for determining the amount of exogenous and endogenous pressure on a graph. It can distinguish between networks coming from different sources, e.g. distinguishing a graph of Facebook data versus a graph of Wikipedia data. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Mattes, Connor↗

Automation of Vulnerability and Patch Management: Information Extraction, Association, and Optimization

Vulnerability and patch management is an integral part of a robust cybersecurity program, yet it grows increasingly complex due to the sheer amount of data that must be analyzed. Particularly in Operational Technology (OT) environments, analysis must be done manually because of the lack of automated solutions. Additionally, there are many steps in this process, from the initial discovery of the vulnerability to the implementation of its remediation, and each step in the process requires different data in order to be performed effectively. In this work, we provide approaches and strategies to assist operators in industrial or OT environments throughout the vulnerability management cycle. Security advisories provide key information about mitigation strategies, or actions that can be taken when a patch is unavailable or cannot be installed. Details of these strategies are not shared in public vulnerability databases and must be found manually. We approach this problem by designing a solution to automatically identify that information within vendor security advisories and retrieve it for operator use. We start with an approach that requires domain-specific knowledge of certain frequently-seen reference websites. Next, an approach that can work on an arbitrary website but relies on certain keywords. Finally, an approach that uses Natural Language Processing (NLP) methods and does not require specific knowledge or keywords. Each of these approaches is more general than its predecessor; we demonstrate high accuracy for all approaches Advisories also often contain details of affected products in non-standard or natural language formats. While this information can be easily understood when read by an operator, the non-standard format acts as a barrier to effective automation. We provide an approach for the first step in this process: identifying vendors in security advisories and mapping them to a standard framework for representing digital assets and software products. We evaluate five established string similarity algorithms, plus one of our own design that combines string similarity and information theory, on the task of mapping vendors to their corresponding entries in the Common Platform Enumeration (CPE) repository. Our results show that our proposed metric outperforms all others. Due to the constraints on time, finances, and personnel for organizations, Large Language Models (LLMs) may seem like attractive opportunities for security operators to speed up information gathering; however, it is still not clear whether LLMs can handle vulnerability management tasks well. To answer this question, we perform an empirical study of LLMs’ ability to provide consistent, accurate information about vulnerabilities in order to guide organizations in their adoption of LLMs. We observe poor performance for all models tested, suggesting that these models are not well-suited to the consistent retrieval of accurate vulnerability information. Finally, once vulnerabilities have been identified and any additional information has been obtained, operators must decide which remediation actions to implement based on their available resources. This already-complex problem becomes even more so when we consider that a vulnerability may have multiple avenues for remediation. We formulate this scenario as two knapsack problems and provide solutions, which we then compare against several existing strategies for vulnerability prioritization seen in real operational environments.

McClanahan, Kylie↗

Scalable edge clustering of dynamic graphs via weighted line graphs

Timestamped relational datasets consisting of records (or connections) between pairs of entities are ubiquitous in network science. For applications like peer-to-peer communication, email, various social network interactions, and computer network security, it is useful to organize these records into groups based on how and when they are occurring. Weighted line graphs offer a natural way to model how records are related in such datasets but for large real-world graph topologies, building and utilizing the line graph is prohibitively expensive. Here, we present the framework to cluster the edges of a dynamic graph via the associated line graph that contains two major contributions. The first is a method to work with the line graph implicitly and the second is a distributed scale implementation of an agglomerative hierarchical graph clustering algorithm. We outline a novel hierarchical dynamic graph edge clustering approach that efficiently breaks massive relational datasets into small sets of edges containing events at various timescales. This is in stark contrast to traditional graph clustering algorithms that prioritize highly connected (clique-like) community structures. Our approach relies on constructing a sufficient subgraph of a weighted line graph and applying a hierarchical agglomerative clustering. This approach is related to scalable techniques from spatial clustering, nonlinear-dimension reduction, topological data analysis, and draws particular inspiration from HDBSCAN. As an edge clustering, this method yields an overlapping node clustering. Our algorithm is parallelizable and we demonstrate efficient clustering of a billion-scale, real-world dynamic graph into small edge sets that correlate in topology and time. The entire clustering process for a graph with tens of billions of edges takes just a few minutes of run time on 256 nodes of a distributed compute environment. We argue how the output of the edge clustering is useful for a multitude of data visualization and powerful machine learning tasks, both involving the original massive dynamic graph data and metadata associated with the nodes and edges. Finally, we describe how this approach can be extended to dynamic hypergraphs and dynamic graphs/hypergraphs with unstructured data living on vertices and edges.

Data Analysis↗

Mapping Critical Vulnerabilities in Natural Gas Pipeline Systems through Network Centrality and GIS Analytics

Natural gas plays a central role in the US energy landscape, providing 43% of electricity generation in 2023. Its exclusive recovery ability on pipelines for transmission underscores the importance of understanding the disruption recovery ability of this infrastructure. This study employs a network-based analytical framework integrating geographic information systems (GIS) with multiple centrality measures—betweenness, closeness, degree, and eigenvector—to pinpoint key segments and evaluate the structural robustness of the national pipeline network. Pipelines are grouped by System ID and Operator ID to capture variations across organizational and physical structures. The analysis reveals uneven patterns of network influence, where certain pipelines function as critical connectors or dominant hubs. Spatial mapping highlights geographic dependencies and potential chokepoints, offering a clear view of where targeted risk prevention measures would be most effective. The findings provide practical guidance for prioritizing maintenance, enhancing system robustness, and mitigating risks to ensure a stable and secure energy supply. Future research will expand the framework to incorporate dynamic operational data and real-time network behavior.

Peterson, Steven [ORNL] (ORCID:0000000287672998)↗

Summary of Pilot Project State Technical Assistance on Multi-Sector Analysis for Electric and Petroleum Fuels

The Oregon Energy Security Plan, (ODOE 2024) published in September 2024, builds a strong case for the state to give acute attention to the fuel supply chain. In December 2024, Pacific Northwest National Laboratory (PNNL) in partnership with Oregon Department of Energy (ODOE), announced a pilot project to conduct an analysis that synthesizes current and projected transportation fuel dynamics, supply chain risks, and risk comparators with relevant sectors, such as transportation electrification, sponsored by the Department of Energy’s (DOE) Office of Cybersecurity, Energy Security, and Emergency Response (CESER). The study, is intended to leverage existing modeling and frameworks from a recent 2024 sector coupling analysis supported by the DOEs Office of Electricity (OE) (B. Mitra, S. Pal, et al., Coupling of the Electricity and Transportation Sectors - Part I: Sector Overviews 2024) (B. Mitra, S. Pal and J. Reeve, et al. 2024). While the PNNL team set out to conduct a quantitative risk analysis driven by detailed data that synthesizes current and projected transportation fuel dynamics, supply chain risks, and risk comparators with relevant sectors. The intention was to provide an approach that could be extendable to other parts of the country. They encountered data limitations and adjusted their approach accordingly. This report summarizes PNNL's original plan for executing the study, including limitations for obtaining data requirements for fuel flows and interim products, as well as a risk matrix that can be used to identify supply chain risks.

02 PETROLEUM↗

CEC Quest: Long Duration Energy Storage Impact Analysis Tool

SAND2025-14389O CEC Quest is a Python tool with a user interface designed to analyze the greenhouse gas impacts of long-duration energy storage projects in California. The tool automates data collection from public sources and uses an Application Programming Interface (API) to enable users to download photovoltaic resource availability, marginal operating emissions rate, and utility rate data. It guides users in inputting parameters for a battery energy storage model and uploading site electrical load data, while also prompting for relevant analysis parameters like timestep and grid limits. CEC Quest performs monthly optimization of one year of data to assess impacts on the site’s electrical bill and the grid’s greenhouse gas emissions. Finally, it conducts a lifecycle analysis to evaluate changes over a defined quantification period, with results aggregated through automated report generation. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Rosewater, David [Sandia National Lab. (SNL-CA), L↗

From Bricks to Clicks: Mapping the White Space in Building Innovation

It is a critical national imperative to transform the buildings sector, yet innovation is impeded by deployment failures that leave promising technologies stranded. Conventional market reports and techno-economic analysis provide an insufficient understanding of markets and resource allocation for emerging building technologies. They omit crucial commercialization factors such as ecosystem maturity and adoption friction, where the coordinated participation of a network of suppliers, contractors, financiers, regulators, and integrators is required to scale solutions. This study addresses these gaps by introducing an evaluation framework grounded in front-line data from six years of the DOE's IMPEL incubator, comprising experience from 300 building-sector innovators and the adjacent, complex ecosystem. Our methodology synthesizes top-down market analysis with bottom-up, practitioner-level data across five megatrends: (M1) Affordable materials and industrialized construction; (M2) Healthy and efficient mechanical systems; (M3) Intelligent building operations; (M4) Buildings as grid assets; and (M5) High-density power and cooling for data centers and therein identify twelve "white space" technology opportunities. Next, we develop a multi-criteria scoring rubric to rank these opportunities based on parameters, i.e., Affordability, Quality of Life, Reliability, and Security, yielding composite ‘Demand’ and ‘Maturity’ indices. Our results indicate that the most significant white spaces may not be incremental products but a new class of ‘Ecosystem Enablers’, such as logistics platforms, orchestration layers, and automated compliance software that solve structural deployment gaps. This paper summarizes this transparent, evidence-based, practitioner-informed evaluation framework for policymakers and investors to re-evaluate policy and resource allocation and unlock scalable market transformation.

Singh, Reshma↗

AI Design Assistant

SAND2025-01930O The AI Design Assistant uses ChatGPT to provide a natural language interface to airfoil analysis tools (XFOIL). Most of the code base is glue code, connecting XFOIL (a tool for analyzing airfoils) to the OpenAI interface. Among the more novel features are an airfoil geometry class and methods on how to extract detailed data from XFOIL. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Karcher, Cody↗

Fracture Network Prediction Using Physics-based Machine Learning Algorithms

In recent years, systematic CO2 injection into geological reservoirs across the U.S. has gained traction as a strategy to mitigate greenhouse gas emissions. This approach necessitates precise monitoring to ensure secure containment, minimize risks, and optimize storage management. Our study leverages machine learning (ML) techniques to advance the understanding of CO2 injection processes, focusing on the Illinois Basin. Over a three-year injection period, we analyzed microseismic data, identifying 19 temporal intervals with significant bottom-hole pressure changes. By partitioning microseismic events into these intervals and estimating b-values, we revealed over 100 clusters of events related to fracture initiation or reactivation. Advanced spatial analysis highlighted horizontally-oriented fractures along the NNW-SSE axis. This quantification of fracture networks informs dynamic injection scheduling, work-over strategies, and risk assessments, enhancing carbon capture, utilization, and storage (CCUS) operations. Additionally, our methodology offers valuable insights for oil and gas operations and geothermal development, supporting fracture-based monitoring and risk mitigation.

Kumar, Abhash↗

Radiochemical transport analysis of gamma spectroscopic data to support estimation of molten salt reactor off-gas inventories

This work introduces a novel application of radiochronometry to estimate nuclide inventories in molten salt reactor off-gas systems based on gamma spectroscopic data from the Molten Salt Reactor Experiment. By analyzing isotopic, isobaric, and isomeric activity ratios, key depletion model parameters related to species transport within the reactor system could be inferred. The findings demonstrate the potential of leveraging a limited subset of gamma spectroscopy measurements to accurately estimate nuclide inventories throughout the off-gas system. The approach can be useful in reactor design activities and support analyses relevant to operations, safety, security, and safeguards.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

TalkPipe

SAND2025-11168O TalkPipe is a software tool to help users create and manage complex data analysis tasks involving Large Language Models. Its easy-to-use interface allows users to combine different analytical processes. TalkPipe includes a Python library, a scripting language, and can be run in a Docker container, making it simple to customize and extend. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Bauer, Travis [Sandia National Lab. (SNL-CA), Live↗

Learning from Arctic Microgrids: Cost and Resiliency Projections for Renewable Energy Expansion with Hydrogen and Battery Storage

Electricity in rural Alaska is provided by more than 200 standalone microgrid systems powered predominantly by diesel generators. Incorporating renewable energy generation and storage to these systems can reduce their reliance on costly imported fuel and improve sustainability; however, uncertainty remains about optimal grid architectures to minimize cost, including how and when to incorporate long-duration energy storage. This study implements a novel, multi-pronged approach to assess the techno-economic feasibility of future energy pathways in the community of Kotzebue, which has already successfully deployed solar photovoltaics, wind turbines, and battery storage systems. Using real community load, resource, and generation data, we develop a series of comparison models using the HOMER Pro software tool to evaluate microgrid architectures to meet over 90% of the annual community electricity demand with renewable generation, considering both battery and hydrogen energy storage. We find that near-term planned capacity expansions in the community could enable over 50% renewable generation and reduce the total cost of energy. Additional build-outs to reach 75% renewable generation are shown to be competitive with current costs, but further capacity expansion is not currently economical. We additionally include a cost sensitivity analysis and a storage capacity sizing assessment that suggest hydrogen storage may be economically viable if battery costs increase, but large-scale seasonal storage via hydrogen is currently unlikely to be cost-effective nor practical for the region considered. While these findings are based on data and community priorities in Kotzebue, we expect this approach to be relevant to many communities in the Arctic and Sub-Arctic regions working to improve energy reliability, sustainability, and security.

25 ENERGY STORAGE↗

East Tennessee Technology Park Biological Monitoring and Abatement Program 2024 Calendar Year Report

The East Tennessee Technology Park (ETTP) Biological Monitoring and Abatement Program (BMAP) consists of three tasks that reflect different but complementary approaches to evaluating the ecological integrity of waters near ETTP. These tasks include (1) bioaccumulation monitoring of fish and clams, (2) benthic macroinvertebrate species richness and density monitoring, and (3) fish community monitoring. The sampling and analysis requirements for the ETTP BMAP in calendar year 2024, covering in part both FY 2024 and FY 2025, are outlined in the respective FY sampling and analysis plans (UCOR 2023, 2024). Sampled water bodies and locations for the ETTP BMAP are shown in Figures 1 and 2. This ETTP BMAP report presents the CY 2024 results and provides context with results from previous years. The report also includes Oak Ridge National Laboratory (ORNL)–generated biological monitoring data collected for other US Department of Energy programs, including the UCOR Water Resources Restoration Program (WRRP) off-site fish bioaccumulation data (UCOR 2023) and select Y-12 National Security Complex (Y-12) BMAP fish bioaccumulation data. Historical data collected for the ETTP BMAP and other programs in the nearby Poplar Creek and Clinch River are provided where appropriate. This progress report provides an update on the biological monitoring activities supporting the ETTP UCOR Environmental Compliance organization, which sponsors the ETTP BMAP. In addition to this internal reporting, ETTP BMAP results are provided in the annual remediation effectiveness reports and the annual site environmental reports, both of which are publicly available. BMAP data are also available to the public via the Oak Ridge Environmental Information System (https://ucor.com/oak-ridge-environmental-information-system-oreis/).

54 ENVIRONMENTAL SCIENCES↗

Frictionless knowledge injection for few-shot learning

Cutting-edge machine learning methods often require large volumes of curated training data, precluding their use in national security problems with rare events in massive datasets. We present a method for incorporating abstract knowledge into models tailored for sparse data. A subject matter expert defines salient concepts using data examples, which are encoded in the model’s embedding space. Models are then trained to respect these concepts. This method enables knowledge injection, yielding effective models with limited labeled data and the ability to assess model sensitivity for subject matter expertise across the nonproliferation mission space, as demonstrated with Raman spectra analysis.

Stomps, Jordan [ORNL] (ORCID:0000000178114479)↗

Visual Analytics of Performance of Quantum Computing Systems and Circuit Optimization

Driven by potential exponential speedups in business, security, and scientific scenarios, interest in quantum computing is surging. This interest feeds the development of quantum computing hardware, but several challenges arise in optimizing application performance for hardware metrics (e.g., qubit coherence and gate fidelity). In this work, we describe a visual analytics approach for analyzing the performance properties of quantum devices and quantum circuit optimization. Our approach allows users to explore spatial and temporal patterns in quantum device performance data and it computes similarities and variances in key performance metrics. Detailed analysis of the error properties characterizing individual qubits is also supported. We also describe a method for visualizing the optimization of quantum circuits. The resulting visualization tool allows researchers to design more efficient quantum algorithms and applications by increasing the interpretability of quantum computations.

Chae, Junghoon↗

Deployment and Evaluation of SciStream on OLCF's Advanced Computing Ecosystem (ACE)

The growing demand for real-time analysis, experimental steering, and decision-making in scientific workflows has created a need for tightly coupled integrations between experimental facilities and high-performance computing (HPC) systems. The Department of Energy’s Integrated Research Infrastructure (IRI) initiative highlights data streaming as a key capability for enabling memory-to-memory data transfers, bypassing the limitations of traditional store-and-forward models. SciStream is a toolkit developed by researchers at Argonne National Laboratory (ANL) to support such streaming by addressing cross-domain security, delegated authentication, and application transparency. We deployed and evaluated SciStream on the Oak Ridge Leadership Computing Facility’s (OLCF) Advanced Computing Ecosystem (ACE) infrastructure, leveraging the Olivine OpenShift cluster and its high-bandwidth Data Streaming Nodes (DSNs) as gateway nodes. Our evaluation included synthetic streaming workloads derived from IRI science workflows, a streaming simulator, and integration with RabbitMQ to handle low-level messaging. This report documents the deployment process, performance evaluation, and challenges encountered, along with opportunities for future improvements.

97 MATHEMATICS AND COMPUTING↗

Correlation Between Weather Alerts and Grid Component Failures for Grid Alert

Weather events cause most grid failures. Often, we even get notifications on our phones to take cover or be prepared for an imminent event. If electric grid utilities had a similar warning that also included probable scenarios and the equipment involved, they could prepare and minimize the effects. Recent research at Idaho National Laboratory into electric grid risk analysis methods resulted in a tool that allows for the development of the most likely scenarios given failure probabilities of grid components. INL has a project with the U.S. Department of Energy’s Cybersecurity, Energy Security, and Emergency Response (CESER) program to develop a Grid Alert application that receives messages from the existing emergency alert system, filters and determines components possibly affected by the emergency event, calculates probable scenarios uses MASTERRI and then notifies the utility if there is significant risk. Historical failure data of elements that comprise the U.S. electric grid have been compiled by utilities and organizations such as the international regulatory body North American Electric Reliability Corporation (NERC). Nominal failure rates are obtained from this data. To make this tool possible, estimated failure rates are needed for different component types given the alert type, severity, and location. Historic weather-related grid element failures are correlated with historic weather events from Integrated Public Alert & Warning System (IPAWS). These correlated events and failures are used along with Bayesian updates from the historical norms to provide a modified failure rate for grid elements in the alert areas and calculate probable scenarios. This discusses the Grid Alert project plan but focuses on the data gathered and process used in determining failure rates for possible grid failure scenarios.

24 - POWER TRANSMISSION AND DISTRIBUTION↗

Rapid Monitoring and Defense Approach for Resilience Improvement of Grid Cyber Security

Cyber-physical systems and electric utilities significantly depend on the reliability and efficiency of information and operational technology. However, false data injection attacks based on synchrophasor measurement data pose a serious threat to the safe and reliable operation of modern power systems. Here, to mitigate this problem, a rapid monitoring and defense approach is proposed to defend against cyber attacks. Initially, the Time and Frequency based Convolutional neural Network (TFCN) is proposed to detect different types of attacks. Within the TFCN, the advances are that both time and frequency domain information can be fused without extra spectrum analysis methods, and can save detection time to speed the calculation efficiency using the developed time-frequency block. Next, a comprehensive defense strategy is developed for multiple cyber attacks to ensure the stability and resilience of the power system according to the feedback detection results. The advances of this strategy are that different control strategies can be automatically selected to recover the stability to the greatest extent according to the detected attacks. To verify the effectiveness of the proposed approach, the high-speed frequency measurements collected from the wide-area monitoring system are used. The results demonstrate that the cyber attack detection performance can reach 95.57% accuracy, outperforming both traditional and some advanced neural networks. Importantly, the defense strategy is conducted and verified in a modified IEEE 39 bus system as well, which illustrates profound performance in faster stability restoration.

Comprehensive defense strategy↗